Somewhere on most small-business websites there is a PDF doing nothing. A menu. A price list. A product catalogue. A whitepaper, a parts guide, a class timetable. It was uploaded because sending a document was easier than building a page, and it has been sitting there ever since receiving approximately no visitors.
That is a shame, because a PDF is not invisible to search engines. It is a URL like any other: it can be crawled, it can be indexed, and it can appear in results with your domain next to it. Most site PDFs are not found because of three specific, fixable problems — and because their owners never treated them as pages worth discovering.
The key takeaway: search engines can index PDFs, but a PDF has to clear the same pipeline as any other URL — discovered, crawled, evaluated, indexed — and PDFs fail at the discovery and evaluation stages more often than HTML pages do. Submitted is not indexed, and indexed is not ranked.
A PDF is a URL, and that is the whole mental model
Everything follows from one fact: yoursite.com/menu.pdf is an address that returns a file, exactly as yoursite.com/menu is an address that returns HTML. The same discovery pipeline applies. A crawler has to discover the URL, which normally means finding a link to it; fetch it, which means it must be reachable and unblocked; extract meaning from it, which means there must be text to read; and only then does the search engine decide whether to index it, with ranking a separate question after that.
Where PDFs differ is in the details of each stage — and in the fact that nobody bothers with them. Nobody writes a title for a PDF, links to it from a relevant page, or checks whether it is in the sitemap. The result is a file that is technically published and practically undiscoverable.
Problem 1: nothing links to it
This is the most common failure by a wide margin. The PDF was emailed to someone, uploaded to the media library, and the only link lives in a newsletter from two years ago.
An unlinked file is orphaned. Crawlers find URLs primarily by following links, so an orphaned PDF may simply never be discovered — and if it was discovered once, nothing signals that it is worth returning to.
- Link to it from a relevant HTML page, with descriptive anchor text. Not "click here" or "download PDF", but "download the full spring parts catalogue (PDF)". That anchor text is one of the strongest hints a crawler gets about what the file contains.
- Link from more than one place if it matters — the relevant service page, a resources index, the footer if it is genuinely a primary document.
- Give it a home. A one-paragraph HTML page introducing the document, with the download link on it, gives the file an internal link, a place in your navigation, and somewhere to send people who want context first.
- Include it in your XML sitemap. Sitemaps can list PDF URLs alongside HTML ones. It is a discovery hint rather than a guarantee and no substitute for linking, but it is close to free — see our walkthrough on building and validating a sitemap.
Problem 2: there is no text to read
A scanned PDF is a stack of images. Select a paragraph and nothing highlights, because there are no characters — just pictures of characters. A crawler extracting text finds nothing to work with, and a file with no extractable text has almost nothing to be indexed for.
This catches people out constantly, because the document looks completely normal on screen. The test takes five seconds: open it and try to select a sentence. If you can highlight it, there is a text layer.
If there is not, you have three options. Republish from the source — if the document started in a word processor or design tool, export a fresh PDF rather than scanning a printout; a born-digital export carries real text automatically and will be a fraction of the size. Run OCR, which converts pictures of words into selectable text; it is a separate step from converting or compressing a file, and its accuracy depends on scan quality, so treat the output as needing a proofread rather than as automatically correct. Or put the content on an HTML page and keep the PDF as an optional download.
While you are in the file, set its internal document title property. Search engines commonly use a PDF's title metadata to generate the title shown in results, and the default is frequently something like "Microsoft Word - final_v3_REVISED.docx". That string is what a searcher would see.
Problem 3: it is enormous
A heavy PDF causes two problems, and neither is a penalty — they are practical. Crawling: very large files are slower and more expensive to fetch, and crawlers make trade-offs about what is worth downloading and how often to return. Humans: someone clicks your catalogue on a phone, on mobile data, and waits. Most do not wait, and a result that gets clicked and abandoned is not doing the job you published it for.
To get the weight down, in the order that preserves the most quality:
- Export from the source with sensible image settings. Publishing a print-resolution file for online reading is the single biggest cause of bloat; most layout tools offer a screen or web preset.
- Downsample images before assembly if you control the source document. Photos placed at full camera resolution dominate file size.
- Split a very large document into logical parts. A 200-page catalogue is usually better as several section files linked from an HTML index — smaller downloads, more specific titles, and each part can be found for its own subject instead of one file trying to cover everything.
- Compress the finished PDF if the source is gone. A free browser toolset such as EyePDF handles compression, splitting, and merging without an install, which is enough to fix a legacy file inherited from a previous web designer.
After any compression pass, open the file and check that small type and diagrams are still legible. An unreadable file that loads fast is not an improvement.
Naming, URLs, and the small stuff that adds up
PDFs are usually uploaded with whatever filename they had on someone's desktop, and that filename becomes the URL.
- Use descriptive, hyphenated, lowercase filenames.
/downloads/spring-parts-catalogue.pdftells a searcher what the file is;/wp-content/uploads/2019/11/FINAL_v3.pdftells them nothing. Avoid spaces and special characters, which percent-encode into an unreadable URL. - Do not rename casually. If you change a published PDF's URL, add a redirect — otherwise every existing link to it, including any that helped it get discovered, now points at a 404.
- Consider a canonical hint. If the same content exists as both an HTML page and a PDF, you can serve a
Link: rel="canonical"HTTP header on the PDF pointing at the HTML version so the two are not competing. Most site owners will not need this, but it is the correct tool when the duplicate is deliberate. - Check you have not blocked it. A
Disallow:line covering the downloads folder in robots.txt, or a login wall in front of it, keeps the file out of the index regardless of everything else here.
When the right answer is an HTML page instead
This is the honest part. For a lot of content, the PDF should never have been the delivery format.
Use HTML when the content is meant to be read rather than printed or filed: a menu, a price list, a service description, FAQs, an event schedule. HTML pages are easier to crawl, adapt to phone screens, update in seconds, carry proper internal links, and give you real analytics.
Keep the PDF when the document genuinely needs to be a document — something to print, sign, file, or hand over. Forms, technical specifications, warranty terms, application packs, anything with a required layout.
Do both when it is worth it: publish the content as an HTML page and offer the PDF as a download from it. You get the discoverability of a page and the utility of a document, and the download link gives the PDF the internal link it needs to be discovered too. If the document started as a PDF, converting it to an editable file is usually the fastest way to move the content — a PDF-to-Word conversion gets the text out so you can rebuild it as HTML rather than retyping it. Expect to redo the layout, because converters reconstruct an approximation and tables tend to arrive in poor shape.
Checking whether any of this worked
Verify rather than assume, and keep the three stages distinct.
- Index status: run your search console's URL inspection tool on the PDF's exact URL. It reports whether the URL is known, when it was last crawled, and whether it is indexed. A
site:search restricted to your domain also shows which files appear. - Crawl evidence: server logs or your host's analytics show crawler requests for the file. That is proof of fetching, which is not the same as indexing.
- Human evidence: most analytics tools do not record PDF views by default, because there is no page to run a script on. Track clicks on the download link from the HTML page instead — that is the number that tells you whether the document is doing anything for the business.
- Timelines: discovery and indexing take anywhere from days to several weeks, and plenty of URLs are crawled and never indexed. Give a change a few weeks before calling it a failure, and remember that indexed is not ranked. Our guide on submitting URLs to search engines covers what submission does and does not accomplish.
A short working checklist
- [ ] Every published PDF is linked from at least one relevant HTML page, with descriptive anchor text
- [ ] Its text is selectable — if not, republish from source, run OCR, or move the content to HTML
- [ ] The internal document title is set to something a searcher would want to click
- [ ] The filename is lowercase, hyphenated, and descriptive
- [ ] File size is reasonable for its purpose; oversized legacy files are compressed or split
- [ ] The URL appears in your XML sitemap, and robots.txt does not block it
- [ ] Content meant to be read, not filed, has an HTML version
- [ ] Download clicks are tracked from the linking page
FAQ
Can Google index PDF files?
Yes. Major search engines crawl and index PDFs and can show them in results, often with a file-type indicator. The practical constraints are that the file must be discoverable through a link, fetchable, and contain extractable text. A scanned, image-only PDF has nothing to index; an orphaned PDF may never be discovered in the first place.
Should I put PDFs in my XML sitemap?
You can, and it is a reasonable low-effort discovery hint for documents you want found. It is not a substitute for linking to them — a sitemap entry tells a crawler a URL exists, while an internal link tells it the URL matters and gives it context. Do both for documents that count.
Why does my PDF show a strange title in search results?
Because search engines often use the PDF's internal title metadata rather than its filename, and that field frequently retains the source document's name. Open the file's document properties, set a clear title, re-upload, and wait for it to be re-crawled.
Is it bad to have the same content as both a PDF and a web page?
It is normal and generally fine. If you want to be explicit about which version is primary, serve a canonical link header on the PDF pointing at the HTML page. In most cases the simpler approach works: publish the readable version as HTML and offer the PDF as a download from it, so there is no ambiguity about which one you are promoting.
Will compressing a PDF hurt how it ranks?
There is no ranking penalty for file size in either direction. Smaller files are easier to crawl and much better for the person who clicks the result on a phone. The real risk is compressing so hard that small text or diagrams become unreadable — so check the file afterwards, and prefer re-exporting from the original source when you still have it.
Next step
Take fifteen minutes and inventory the PDFs on your site. For each one, answer three questions: is anything linking to it, can you select its text, and is it a sensible size? That inventory alone usually explains why a document nobody has ever found is a document nobody has ever found — and most of the fixes are a link, a re-export, and a sitemap entry.
Then give the pages hosting those documents an extra discovery path: submit your URL to the directory and get a real listing page pointing at them, so crawlers have somewhere besides your own internal links to find your site from.