BookTrace — Search Internet Archive Genealogy Books

Find where surnames and places appear in 15,442 catalogued books and 39.3M extracted entities. A finding-aid, not a full-text reader.

New to BookTrace? Read the BookTrace Research Guide for a full explanation of how the index was built, how to interpret every part of a result, and how to document your searches to BCG/GPS standards.

Type a surname, press Enter to add. Add multiple names for FAN-club or family-group searches. (You can also just click Search — anything still typed in the box gets added automatically.)

Search tips

Use ? to match one unknown letter (e.g., H?nks) or * to match any run of letters (e.g., Han*). Wildcards bypass variant expansion.

Narrows results — but given names in the index are incomplete and error-prone, so this filter will miss valid records. Search surname-only first.

Clear all fields

Ready to search

BookTrace searches 15,442 catalogued Internet Archive books, of which 7,179 have their full text indexed for names and places — 39.3 million entity extractions total.

Start with a surname above. Add places, dates, or a given name to narrow results.

How BookTrace works →    What the index does not cover →

About BookTrace

Who built this tool, why, and how to cite a BookTrace search in your research notes.

Read more →

How the index was built

Full disclosure of what's indexed and what isn't — including the 53.5% text-searchable gap, OCR limitations, and what a negative result does and doesn't mean.

Read the transparency report →

Frequently Asked Questions

Basics

What is BookTrace?

BookTrace is a free finding aid that searches 15,442 genealogy-relevant books catalogued from the Internet Archive. It tells you which books mention a surname or place you’re researching, with approximate page numbers, and links you directly to each book on archive.org.

Does BookTrace host the books?

No. BookTrace stores no book content, only an index of extracted names, places, and catalogue metadata. Every result links to the original item on archive.org, where the Internet Archive serves the page images.

Is BookTrace free?

Yes. Like all Evidence Toolbox tools, BookTrace is free with no account, registration, or usage limits.

Is BookTrace affiliated with the Internet Archive?

No. BookTrace is an independent research tool built by Glenside Digital LLC. It is not affiliated with or endorsed by the Internet Archive.

What kinds of books are in the index?

County and local histories, published family genealogies, city directories, vital records abstracts, biographies, military records, passenger lists, and genealogical periodicals, drawn from the Internet Archive’s genealogy and americana collections. Contributors include the Allen County Public Library Genealogy Center, Harvard, the Library of Congress, the New York Public Library, and the University of Michigan.

Does BookTrace search the entire Internet Archive?

No, and this matters for interpreting results. BookTrace covers only the 15,442 books that scored as genealogy-relevant during index construction, not the Archive’s 40+ million items. A zero-result search in BookTrace is a statement about this index, not about the Internet Archive as a whole.

Does the Internet Archive already have full-text search? Why do I need BookTrace?

Yes, and its per-book “Search inside” feature is genuinely useful when you already know which book to search. What IA does not have is the genealogical search model this work requires: variant expansion for OCR-damaged surnames, filtering by record type (county history, published genealogy, city directory), multi-surname proximity queries across the corpus, coverage disclosure per book, or citation-ready documentation of what was searched. BookTrace fills those specific gaps. The two tools complement rather than compete.

When should I use BookTrace vs. IA’s own search?

Use IA when you already know the book and want the passage: “Search inside” returns exact page hits with the sentence around each match. Use BookTrace when you don’t yet know which books to open, when you need to search across many books at once with a genealogical model, or when you need documented negative searches for GPS-standard research. A common workflow uses BookTrace to identify candidate books, then jumps to IA to read each one.

What does BookTrace do that IA search cannot?

Six things: surname variant expansion (Hanks also finds Hankes, Hanx, Henks), given-name abbreviation expansion (William also finds Wm. and Wm), state-name equivalents (Kentucky and KY are treated as the same place), record-type filtering (only show county histories, or only published genealogies), multi-surname proximity searches across the corpus (find books where two or three surnames appear on the same page), and citation-ready documentation of every search, including specific negative-search language for zero-result queries.

Searching

How do I do a basic search?

Type a surname, press Enter to add it, and click Search. Start surname-only with default settings, review the results, then narrow with places, dates, or a source type if the result set is large.

Can I search multiple surnames at once?

Yes, up to twelve. This supports FAN-club searches (Friends, Associates, Neighbors) and family-group searches. By default a book matches if it contains any of your surnames; switch the toggle to “All (AND)” to require every surname in the same book.

What is variant expansion?

Each surname you enter is automatically expanded to include likely spelling and OCR variants; searching Hanks also matches spellings like Hankes and Henks. The exact variants applied are always listed in your search’s citation summary, so your research log records what was actually searched.

Can I use wildcards?

Yes: ? matches one unknown letter (H?nks) and * matches any run of letters (Han*). Wildcard terms bypass variant expansion; they’re searched exactly as written. Terms must begin with at least one fixed letter, and a term containing ? (or a mid-word *) searches book text only, since catalogue metadata can’t be pattern-searched that way.

Why does the given-name field warn me it will miss records?

The index splits each extracted name positionally on the last space, which breaks on inverted names (“Lake, James”), titles (“Mrs. John Lake”), and suffixes (“John Lake Jr.”). The given-name filter therefore drops books where the surname matched but the given name was recorded imperfectly. Search surname-only first; add a given name only to trim an unmanageably large result set.

Are name abbreviations like “Wm.” handled?

Yes, bidirectionally. Searching William also matches Wm. and Wm; searching Jno. also matches John. The table covers standard nineteenth-century written abbreviations (Thos., Chas., Geo., Robt., Margt., Eliz., and so on). Nicknames (Polly, Peggy, Sally) are not included in the current version; that’s a planned future enhancement.

Do state abbreviations work in the Places field?

Yes. KY expands to Kentucky and Ky., and vice versa. You can also enter counties, cities, and other place names.

What does the “Keywords in catalogue” field search?

Only the Internet Archive’s metadata about each book (title, subject, and description), never the book’s text. Use it for terms that describe a book rather than appear in it, like “county history” or “muster rolls.”

What does the proximity setting do?

When you search multiple terms, it restricts text matches to the same page, within roughly ten pages, or anywhere in the book (the default). Because page numbers in the index are estimates, proximity is approximate by nature.

Results

Why are results split into three groups?

Because the three groups carry different evidentiary weight. Text matches (Group 1): every term was found in the book’s extracted text. Split matches (Group 2): some terms matched in text, others only in catalogue metadata. Catalogue matches (Group 3): terms matched only in the metadata, meaning the book’s text was not searched, so these are leads to investigate rather than confirmed text hits.

Why do page numbers have a “~” in front of them?

Because they’re estimates, not recorded page numbers. Each is computed from where the name sits in the book’s OCR text, accurate to roughly ±3 to 10 real pages. Open the book at the estimated page and scroll a few pages in each direction, or use the book’s own index if it has one.

Can BookTrace show me the sentence where my ancestor is mentioned?

BookTrace does not display snippets. It is designed as a finding aid, pointing you to the books and pages where terms appear; reading the passage happens on the Internet Archive, where the images and text live. For a specific book, IA’s own “Search inside” feature returns the sentence around each match. Snippet display within BookTrace is a named future enhancement.

If a surname and a place appear on the same page, are they connected?

Not necessarily. Same-page co-occurrence means both terms were independently extracted from the same region of text, within a few dozen lines of each other. It does not mean they appeared in the same sentence or that the place describes the person. Only reading the passage can establish that.

What does “Contributor” mean on a result card?

The institution that digitized the book, not its publisher, author, or subject.

Why does a result say “~340,000 names extracted from this volume”?

A few items in the corpus, such as annual Catholic Directory volumes, are extraordinarily name-dense and surface constantly for common surnames. BookTrace states the fact and lets you decide; it never down-ranks or hides these books. If they clutter a search, use the source-type filter to exclude directories.

My common surname says results were truncated. What happened?

Very common surnames like Smith can match thousands of books, so each term’s matches are capped. When that happens, BookTrace tells you plainly in a banner and in the citation summary, and suggests narrowing with places, dates, or a source type. Truncation is always disclosed.

What does “Default order” mean in the sort options?

It’s the initial ordering, which surfaces books with stronger, more concentrated term matches first. It isn’t labeled “relevance” because BookTrace doesn’t present an internal scoring formula as an explainable ranking. The other sorts (date, title, contributor) are plainly factual.

Coverage and Limitations

Why is only 46.5% of the index full-text searchable?

Of 15,442 catalogued books, 7,179 have their text indexed. The remaining 8,263 either have no OCR text layer (6,298 books, mostly microfilm scans and handwritten material), had OCR too degraded to index reliably (1,177), or had unreachable text files (788). This is a fact of the corpus, not a pipeline failure. Genealogy sources are heavily handwritten or microfilmed. BookTrace still shows these books’ catalogue entries so their existence is never hidden.

What exactly is an “extracted entity”?

A text span that a statistical language model classified as a person or place in the book’s OCR text: 39.3 million of them across the index. It’s statistical classification, not verified fact, and both false positives and false negatives occur. Read every result as “this string was detected as a name at approximately this page,” and treat it as a signal to open the book and read.

Why can’t I find a woman I know is in one of these books?

Likely the married-women gap. Sources that record a woman only as “Mrs. John Smith” contain her husband’s name, not hers; her own given name isn’t in the text to be extracted. Search under her husband’s name and her maiden surname, and treat a null result for a woman’s given name as weak evidence at best. This is a property of the sources, not the tool, but it is systematic and severe in pre-twentieth-century material.

Why do some books show dates like 2024 or 2025?

Some Internet Archive records carry the scan date rather than the publication date, which is why the field is labeled “Publication/scan year (as recorded by IA).” Filter by date with that imprecision in mind. Also note that 1,544 books (10.0%) have no recorded date at all. The “include unknown date” toggle defaults to on so undated books are not excluded from view without your knowledge.

I got zero results. What does that actually mean?

It depends which of three states you’re in, and BookTrace tells you. A true negative (no matches in text or catalogue anywhere in the index) is the strongest negative signal the tool can produce. A partial negative (not text-searchable) means catalogue matches exist among books whose text isn’t indexed; your term may appear in their pages, and those books are your follow-up list. A partial negative (narrow scope) means your own filters may have excluded matching books. And even a true negative applies only to this 15,442-book index, not the Internet Archive as a whole.

How do I cite a BookTrace search in my research notes?

Every search generates a citation-ready summary stating the terms as entered, the variant expansions actually applied, the filters in effect, and the index scope, suitable for direct paste into a research log. For zero-result searches it includes the negative-search language BCG-standard documentation calls for. For example:

BookTrace (evidencetoolbox.com/tools/booktrace), search for surname “Hanks” and place “Kentucky,” accessed [DATE]; searched both book text and catalogue metadata across 15,442 catalogued Internet Archive books, of which 7,179 have full text indexed.

Does a BookTrace search satisfy the “reasonably exhaustive search” requirement of the GPS?

It contributes to one; it doesn’t complete one. BookTrace makes tens of thousands of county histories, published genealogies, and directories searchable at once and documents exactly what was covered, but it reads only 46.5% of its own catalogue, covers only a curated slice of the Internet Archive, and points to passages it cannot show you. Use it to find and document; verify every finding in the source itself.

Index figures (15,442 books catalogued; 7,179 full-text indexed; 39.3 million extracted entities) reflect the index as of July 2026.

Stay Updated

New tools ship regularly. Subscribe for occasional updates on new genealogy research tools — no more than once or twice a month.

Subscribe

Evidence Toolbox is free and independent. Support it on Ko-Fi.