Fossick

How Researchers Can Search Thousands of PDFs by Concept

· 5 min read

Every researcher eventually hits the same wall: you know you read a study on a specific idea — a method, a finding, a turn of phrase — but you cannot remember which of your 3,000 PDFs it lived in. The filename is a DOI string. The keyword you'd search for isn't in the abstract. And your reference manager only searches metadata you bothered to tag.

This article is about a better way to work: document search for researchers that finds files by meaning, so you can type what you remember about an idea and get the paper back — without uploading a single confidential draft, dataset, or unpublished manuscript to the cloud.

Fossick search panel ranking documents on this Mac by meaning for a plain-language query
Fossick in action: type what you remember, get the document — entirely on your machine. Get the app.

Why keyword search fails a literature library

A serious research corpus is messy. It grows over years, spans dozens of subfields, and arrives in inconsistent formats: publisher PDFs, preprints, scanned book chapters, your own annotated notes, and exported reference records.

Keyword search assumes you remember the exact words on the page. But months later you rarely do. You remember the idea: "that paper arguing small sample sizes inflate effect estimates," or "the review on gut microbiome and mood." The authors may have written "statistical power," "publication bias," or "gut–brain axis" — none of which match your query literally.

That mismatch is exactly what semantic search solves. Instead of matching characters, it matches concepts, so a description in your own words surfaces the document even when the vocabulary differs. For a deeper comparison of the two approaches on a folder of files, see full-text search vs semantic search.

Search by what you remember, not what you filed

Reference managers are excellent at what they do — storing citations, formatting bibliographies, deduplicating entries. But their search operates on metadata: title, author, year, and whatever tags you manually applied.

The problem is that the thing you remember is almost never in the metadata. It's a sentence in the discussion section, a definition in the methods, a caveat in a footnote. To find it, you need search that reads the full text of the document.

Fossick, a desktop app that searches your local documents by meaning, entirely offline, indexes the actual contents of every file — not just the filename or the tags. You describe the passage you're chasing ("the objection about reverse causation in observational studies") and it ranks the documents whose meaning is closest, even if your exact words never appear. It's the practical answer to finding a document when you can't remember its name.

Keeping unpublished work private

Researchers handle material that must not leak: manuscripts under embargo, confidential peer reviews, grant applications, interview transcripts, participant data governed by ethics approvals, and industry-sponsored results under NDA.

Many "AI search" tools work by uploading your files to a server to index them. For a research group, that can be an outright compliance problem — and at minimum it's a risk you shouldn't take with someone else's data.

Fossick runs entirely on your own machine. Indexing, embedding, and every search happen locally, and your documents are never uploaded. The simplest proof is the one you can test yourself:

Unplug the Wi-Fi. Everything still works — because nothing was ever going anywhere.

If privacy is the deciding factor for you, these two pieces go deeper: private document search and searching PDFs without uploading them.

Scanned articles, book chapters, and figures

A lot of the literature that matters is locked inside images. Older journal articles are scanned page images. Book chapters arrive as photocopies. Historical sources, archival material, and even handwriting-adjacent documents hold no selectable text at all — so keyword search returns nothing.

Fossick runs OCR (optical character recognition) on-device, automatically, for scanned PDFs and image files like PNG and JPG. That means a photographed page or a scanned chapter becomes searchable by concept alongside your born-digital PDFs, without sending the scan to a cloud service. The mechanics are covered in searching scanned PDFs and images locally with OCR.

Working across thousands of files at once

A single research project can accumulate thousands of documents; a career accumulates far more. The point of good document search for researchers is to treat that entire archive as one searchable space instead of a folder tree you navigate by memory.

A workable literature workflow looks like this:

  1. Point the tool at your reference library, download folders, and note directories — wherever your PDFs and documents actually live.
  2. Let it index the full text (and OCR the scans) once, locally.
  3. Search by concept whenever a question comes up during writing or reading, and jump straight to the source.

Because the index is local, searches stay fast even as the corpus grows into the tens of thousands. There's more on this pattern in searching across thousands of files at once.

Search, not a chatbot — and why that matters for citations

It's worth being precise about what this tool is and is not. Fossick is search, not chat. There is no generative AI writing summaries, no model answering your questions in prose, and no cloud account holding your files.

For research, that distinction is a feature, not a limitation. A chatbot that paraphrases a paper gives you text you then have to trace back and verify — and that can hallucinate a claim or a citation that was never in the source. Semantic search hands you the actual document, at the passage that matches, so you cite the primary source directly. You stay the interpreter of your own literature.

If you're weighing on-device search against AI assistants, is it safe to search documents with AI walks through the trade-offs. Fossick runs on Windows and Mac; you can download it and index your own library to test it against real searches, and it's free to try during the beta. Full details are on the pricing page when you're ready.

Frequently asked questions

How is this different from searching in Zotero or Mendeley?

Reference managers search metadata — title, author, year, and the tags you added by hand. They don't reliably search the full text of every PDF by meaning. Fossick indexes the actual contents of your files and finds them by concept, so it complements a reference manager rather than replacing it.

Can it search scanned journal articles and book chapters?

Yes. Fossick runs OCR on scanned PDFs and image files automatically and on-device, turning page images into searchable text. That lets you find older or photocopied literature by concept, right alongside your born-digital PDFs, without uploading anything.

Does it summarise papers or answer research questions like ChatGPT?

No. Fossick is semantic search, not a chatbot — there's no generative AI and nothing that writes prose. It points you to the exact document and passage that matches your query, so you read and cite the primary source yourself instead of verifying an AI-generated answer.

Is my unpublished work actually kept private?

Yes. Indexing, embedding, and search all run locally on your computer, and your documents are never uploaded. You can confirm it by disconnecting from the internet — search keeps working, because nothing was ever sent anywhere. Only licensing touches the network.

How many documents can it handle?

It's built to search across thousands of files at once, and the local index keeps searches fast as your library grows into the tens of thousands. You point it at the folders where your PDFs, notes, and documents already live and search the whole archive as one space.