Guide · Google Docs → HTML
Getting clean HTML out of a Google Doc
Copy a few paragraphs out of a Doc, paste them into a CMS, and you get text wrapped in class names like c1 and c17 that point at a stylesheet you don't have.
Three specific things are wrong with that markup. Once you know what they are, the fix is one paste.
Copy the content, not the link
This trips people up first: Download → Web Page from the Doc's menu gives you a zip with a full HTML document and an embedded stylesheet — the whole page, styles and all.
What you want instead is the clipboard. Selecting content in a Doc and copying puts a rich-text HTML version on it, and that version carries the structure without the stylesheet.
- 01Open the Doc and select the content you want. Ctrl+A works if you want all of it.
- 02Copy with Ctrl+C — not "Copy link", which only gives you a URL to the document.
- 03Paste into the box on the converter page, or press Ctrl+V anywhere on that page. There's no sign-in and no Drive permission involved; your browser hands over the clipboard, nothing else.
Problem one: the class soup
Google's clipboard HTML puts a class on almost every element — c0, c1, c17, and lst-kix_ names on list items. They're generated per document and refer to CSS that stays behind in the Doc.
So they're not just noise. They're dead references: they do nothing on your page, they collide with class names of your own, and they make the markup unreadable when you go to edit it.
They're removed by pattern — c followed by digits, lst-kix_ anything, docs-internal-guid ids. Classes of your own that happen to be in the paste are left alone.
<p class="c3"><span class="c1">A sentence.</span></p><p>A sentence.</p>Problem two: the bold wrapper that isn't bold
Docs wraps copied content in <b style="font-weight:normal">. It's a <b> tag that explicitly turns bold off — a quirk of how the editor tracks formatting internally.
Paste that anywhere the style attribute gets stripped, which is most CMSes, and the whole block turns bold. The tag is unwrapped here rather than kept, so the content comes out at the nesting level it should have been at all along.
The same pass removes empty <span> elements left behind after their class names go. Those are what make a two-paragraph paste twelve lines long.
Problem three: every link is a redirect
Links in a Doc come out pointing at google.com/url?q=https://example.com/… rather than at the destination. Google uses it for click tracking inside the editor.
Published on a page, it means every outbound link on your site routes through Google, shows the wrong URL in the status bar, and breaks the day that redirector changes.
The wrapper is unwound back to the real destination, repeatedly if it's nested. Tracking parameters go too — utm_source and friends, gclid, fbclid, and a handful of others.
<a href="https://www.google.com/url?q=https://example.com/a%3Futm_source%3Ddoc">link</a><a href="https://example.com/a">link</a>What's left, and what's still on you
Headings, paragraphs, lists, tables, links, bold and italic — the document, in other words. Scripts, event handlers and inline styles are gone, because clipboard HTML can come from any page on the web and gets treated as untrusted input.
Two things the cleaner leaves for you. Google's empty spacer paragraphs come through as <p></p>, so delete the ones you don't want. And Docs' images are hosted on Google's servers with URLs that expire — download them and host them yourself, or they'll vanish from your page later.
Select, copy, paste. The clipboard is read in your browser and nothing is sent anywhere.
Google Docs → HTML