DOCX → HTML
Word document to HTML.
Without the Word junk.
Save As Web Page produces thousands of lines of mso- styles. This produces markup you'd write by hand.
Headings become <h2>, tables become <table>, and nothing Microsoft-specific survives.
Drop a file here.
Or pick one with the button. Or just paste with Ctrl+V. Dozens at a time is fine.
.docx .doc / 25 MB per file / runs in your browser, nothing uploaded
Other formats: .md .markdown .txt .mdown MD → HTML · .txt .text Text → HTML · .txt .csv .tsv CSV → table · .html .htm Google Docs → HTML · .xlsx Excel → table · .pdf docstomd.com
How it works
- 01Drop a .docx onto the box above, or click to pick one. Dozens at a time is fine.
- 02Mammoth reads the document's structure — real heading levels, real tables, real lists — and writes HTML from it. That HTML then goes through DOMPurify, because Mammoth's own docs are explicit that its output isn't sanitised.
- 03Choose what happens to images, then copy the HTML, download the file, or take a zip with the images beside it.
What's supported
- Every .docx Word has written since 2007, plus Word Online and Word for Mac
- Old .doc from Word 97–2003, detected from the file header rather than the extension
- Heading levels as <h1>–<h6>, taken from Word's styles rather than guessed from font size
- Tables as <table>, lists as <ul>/<ol> with real nesting, links, bold, italic, strikethrough
- Blockquotes, code-styled paragraphs as <pre><code>, and captions as <p class="caption">
- Embedded images: inlined as base64, extracted into an images/ folder for the zip, or dropped
What it won't do
- Word's visual formatting — fonts, colours, exact spacing — is deliberately not carried over
- Merged table cells lose their colspan and rowspan; each cell becomes its own <td>
- Track changes and comments are dropped — you get the final text, not the editing history
- Text boxes, SmartArt and charts don't survive; only their text, if any
- Images don't come out of old .doc files — that format hides them where a browser can't reach
- Encrypted documents are refused rather than half-read
What gets thrown away
Word and Google Docs both wrap their content in a layer of junk that only means something inside their own editor. It goes.
Structure stays: headings are headings, tables are tables, lists are lists. It's the decoration that's removed.
- <script> tags
- onclick and friends
- Inline mso- styles
- Dead c1 / c17 classes
- Link tracking parameters
- Office-only tags
- Kept: semantic structure
- Escaped: & < > and quotes
Questions people ask
Going the other way?
DocsToMD is this same tool pointed in reverse — Word, PDF, Excel and HTML into Markdown. Same approach, same privacy model, opposite output.