Evaluating Book UX Usability Heuristics: Physical Books

Why a Book Needed a Usability Evaluation
A client came to HAR with a fair question: why wasn't their book selling the way they had expected. That question sounds simple until you try to answer it, since a dozen different explanations could each account for the same frustrating numbers, and guessing at the first one risks solving the wrong problem. These three things immediately came to mind, what is reasonable to expect for that book’s market; where is the book in its life cycle and what’s been done for its marketing; and whether readers found some shortcoming either by content or by design of the structure or design of the book. This blog post focuses on one facet of the last issue, an exploration of a novel adaptation for evaluating whether there are non-user friendly aspects of physical book design. This approach offers a cost-effective approach to catch some of the easiest things to fix upstream in the project of writing and publishing a book.
The book as a physical object needed its own separate, consistent scrutiny, a way to check whether its binding, its layout, its size, or the way it organized its content was creating friction before a single word of reader feedback is written. This calls for a usability solution, rather than a marketing solution, and usability work already has a name for the kind of systematic pass that answers it: heuristic evaluation. The trouble was that nothing sat ready-made for a bound, printed object. Book UX usability heuristics, the kind built for something held and turned rather than clicked and scrolled, simply didn't exist yet.
Any author or publisher watching sales numbers sit flatter than expected will recognize the instinct to reach for marketing first, more ads, a different retailer push, a new cover reveal. Those are well tread approaches to reversing the trend for a time, but they are the wrong first stop if the book itself is, unbeknownst to publisher or author, quietly working against its own reader. Ruling the object out, or in, before spending further on marketing is a cheaper and more honest place to start. Now, this wasn’t my client’s issue, but this issue of physical book design is one I’d like to share about today.
Adapting Usability Heuristics for a Physical Object
Jakob Nielsen’s ten usability heuristics (Nielsen 1993) assume an interface that responds: a system that shows its status, gives feedback after an action, lets a user undo a wrong click. A website changes when you touch it. A book does not. That difference is the first thing I pondered on before applying the method to something you hold and flip through instead of click and scroll. What follows extends Nielsen's original ten, along with the effectiveness, efficiency, and satisfaction framing set out in ISO 9241, into an object that has never had to answer to either. Each of the ten original heuristics maps to one physical-object adaptation, no gaps, no redundancies. It is a novel application of established usability frameworks, not a claim to have invented usability thinking itself, and that lineage matters, since it is what makes the extension checkable rather than improvised. This is also a work in progress, and I am still musing on how to best create UX heuristic evaluation for books, but it’s in a form worth sharing and beginning discussion around.
Heuristic evaluation is genuinely rare outside software. Nielsen's heuristics were developed in the 1990s specifically for graphic interfaces, and concepts like visibility of system status or the ability to undo and redo map so directly onto a screen that the standard framework feels foreign the moment it is asked to describe a physical object. Physical and spatial designers have their own established toolkit instead, one built on human factors, ergonomics, and anthropometrics, relying on physical prototyping and direct observation of use rather than a heuristic walkthrough. Physical work also carries a higher cost of being wrong. A digital usability problem flagged in review usually gets fixed by editing code or adjusting a layout. A flaw in a physical product means new tooling, new materials, or a structural change, which is exactly why physical designers lean so heavily on testing before anything goes into production rather than reviewing after the fact. Digital interfaces share a vocabulary that physical objects never will, buttons, menus, screens, patterns repeated across nearly every piece of software ever built. A blender, a highway interchange, and a printed book share almost nothing physically, which makes a single standardized heuristic set for physical objects a much harder thing to build than it was for software.
None of this means physical and spatial work has no heuristic tradition at all. It shows up under other names. Industrial design carries a version of it through affordance and signifiers, the vocabulary Don Norman built in The Design of Everyday Things (Norman 1988) for how an object signals its own use without needing an instruction manual. Interestingly enough, the concepts of affordances and signifiers were already familiar to me in my graduate studies in anthropology, through Entanglement Theory, Neomaterialism, object agency, and Structural Linguistics. Other fields, such as architecture and spatial design carry a version of it through building codes, wayfinding standards, and universal design principles, each one functioning as a kind of heuristic evaluation for the built environment even if no one on site would call it that. Product researchers have also adapted heuristic sets for hardware and for IoT devices, objects that pair a physical form with a screen, but those adaptations tend to be built case by case rather than standing as a shared reference the way Nielsen's list stands for software. A printed, bound object with no screen at all sits outside even those efforts.
The temptation, in the absence of more physical-object adaptations of heuristic evaluation, is to conclude it simply does not transfer to a physical object. That conclusion comes too fast. Nielsen's list was never really about screens. It was about how people locate themselves inside a system, how they recover when they take a wrong turn, how much they have to remember versus how much the system shows them. Those questions belong to any designed object a person has to navigate, a printed field guide as much as a checkout flow. The screen was the occasion for the heuristics, and so it is possible to generalize past it.
Take a case closer to the question that started this: a nonfiction paperback that pairs step-by-step instructional content, a recipe, a repair procedure, a workbook exercise, with narrative or reference material that explains the reasoning behind it. It carries two kinds of content at once, application-oriented material that tells the reader what to do next and interpretive material that explains why any of it matters. On a screen, those two registers can live on separate views, one tap apart. On paper, they share the same page, competing for the same eye.
The Ten Heuristics, Translated for Print
Mapped heuristic by heuristic, against all ten of Nielsen's original categories, the translation runs like this:
Flexibility and Efficiency of Use → Portability & Binding Quality. A responsive website adapts to whatever device carries it. A printed object has one fixed size, one weight, one answer to whether someone will actually bring it into the field or leave it in the car, which makes flexibility a property built in at the printing stage rather than adjusted after the fact.
Visibility of System Status → Page Layout Clarity. On a screen that means the interface tells you where you are. In a physical book it becomes whether a running header or a consistent numbering scheme tells the reader where they stand, or whether they have to flip back to the table of contents every time they lose their place.
Recognition Rather Than Recall → Typographic Hierarchy. Becomes a question of whether the reader can tell instructional text from background material at a glance, rather than holding the distinction in memory.
Error Prevention → Sequencing. On a screen this might mean a confirmation dialog. On the page it becomes whether the book tells the reader what they need to know before asking them to act, rather than asking for action first and explaining afterward.
Consistency and Standards → Consistency of Visual Language. Asks whether the same visual language is used the same way throughout, on a screen or a page alike.
Match Between System and Real World → Application vs. Interpretive Balance. Do task-oriented content and background content sit legibly together without one crowding out the other? Dedicated checklists or blank recordation forms are the clearest version of this heuristic, fill-in content sitting right beside the interpretive material that explains it.
User Control and Freedom → Gutter Margin & Page Openness. Can the reader reach every word without fighting the binding to keep the book open, or does text near the fold require cracking the spine to read?
Help and Documentation → Back Matter Usability. Can the reader relocate a specific fact after they have already read the book once? A missing or shallow index is one of the most reliably cited complaints in reviews of nonfiction, the kind readers notice before they notice much else.
Aesthetic and Minimalist Design → Print & Paper Fidelity. Does the physical production stay out of the content's way, or does it call attention to itself: paper that shows text through from the other side, print that reads crisp instead of gray?
Help Users Recognize, Diagnose, and Recover from Errors → Typesetting Legibility. Does the line spacing and type choice let a reader's eye track a line without words merging together, apart from how well the page already distinguishes one kind of content from another?
Next, I’ll apply my use of this evaluative method to three different kinds of physical non-fiction books, which already are highly regarded content by readers, but surfacing ways the physical design of specific non-fiction books have different user needs which could be refined in a future edition in more user-friendly ways. The three types of non-fictions is not intended as a list of all non-fiction types of books, it’s just three books I highly value in my library. So again, this evaluative method is not about content quality, it’s about the physical structure and design of objects which affect user experience. The books I chose for evaluation were: The Prehistory of Texas (Perttula 2004), as of today on Amazon has 4.9 stars, and it provides a much overdue and needed synthesis of Texas’ regional archaeologies. The Field Manual for the Archaeology of Ritual, Religion, and Magic (Augé 2022) is a sui generis, long‑sought text for systematically recognizing material signatures of religio‑magio behavior in the archaeological record, with 5 stars in Amazon. Lastly, there is an herbalist’s touchstone I evaluate, The Modern Herbal Dispensatory (Easley and Horne 2016), which is currently rated 4.8 stars in Amazon.
This full evaluative instrument runs to ten heuristics, one for each of Nielsen's original ten, with no heuristic left to stand for two different physical concerns and none of the physical categories treated as unprecedented. I approached each book category by category, and its sub-categories on which it is evaluated (including particularities to that non-fiction’s needs). Each also carries its own considerations across the different kinds of nonfiction reference books a library holds, since what counts as a failure in a field manual meant to survive a pack and bad light is not always what counts as a failure in a handbook meant to sit on a shelf. Each sub-category (or item) is scored on a severity scale from zero, no issue found, to four, a critical usability failure, with the score anchored to a written description of what that severity looks like on the page.

How the Book UX Usability Heuristics Scoring Works
Every score in this piece was assessed by one reviewer working alone; however, where it is feasible, it’s best practice to have three trained UX researchers scoring independently and then averaging their results is the better practice, since a single evaluator's judgment calls, particularly on the more interpretive items, carry more weight unchecked than they would against two other trained eyes. These scores also measure exactly one thing, how easily and comfortably a reader can physically handle, navigate, and use the object itself. Again, they say nothing about the quality, value, or originality of what is written inside it. A book can score poorly on this instrument and still be an excellent, important, or genuinely unique piece of work in its field. Usability and merit are separate questions, and this instrument only ever answers the first one.
Each book's usability index below is calculated two ways: once treating every individual item as equally important, and once treating every heuristic category as equally important regardless of how many items sit inside it. The two methods usually land close together. When they diverge, that gap itself is worth reporting, since it shows which categories are carrying more or less weight in the final number, and readers deserve to see both figures rather than a single score that hides its own arithmetic.
Each heuristic is scored on the same five-point severity scale before it is anchored to a written description specific to that heuristic:
Score | Meaning |
0 | No Issue. Nothing to flag against this heuristic. |
1 | Cosmetic Issue. Worth fixing eventually, but does not affect use. |
2 | Minor Usability Issue. Slows or mildly frustrates the reader. |
3 | Major Usability Issue. The reader is meaningfully hindered. |
4 | Critical Issue. The reader is blocked or actively misled. |
A single severity anchor makes the scale concrete.
For binding quality and construction, the scale reads as follows.
Severity | What It Looks Like |
0 | Sewn or reinforced binding rated for repeated field use |
2 | Perfect binding adequate for shelf use, showing early wear under handling |
4 | Binding failure within normal use, pages detaching or spine cracking on first read |
What follows applies the instrument to three books, one from each subtype discussed above, each chosen not as a convenient example but because the book itself is doing real, gap-filling work in its market. It’s not a book review and it’s not full disclosure of all the heuristic violations, but meant to propose an approach to evaluating books. Though, I’m happy to share more details with those who ask.
Each section below includes the category-level severity scores, the resulting usability index calculated both ways, and the corresponding spidergram and clustered bar chart. The two chart types answer different questions: averaging the score within a category plots cleanly on the spidergram, showing which heuristics are cumulatively weakest across a title at a glance. In contrast, the clustered bar chart counts how many scores fall at each severity level within a category, which keeps a single critical issue from disappearing into an average alongside several cosmetic ones.
Three Real Books, Scored
The Prehistory of Texas
Representing the edited-volume and sub-specialty handbook subtype of non-fiction, a reference built for a specialist reader who already knows roughly what they are looking for.
In addition to visually analyzing via a spidergram and clustered bar graph, every individual item scored was averaged, to provide an even more zoomed out approach for comparing a book’s usability if more like comparisons were made of the same type of book on the market for benchmarking purposes. In this case, this book returns a usability index of about 53 percent. Even if averaged category by category, the figure lands in nearly the same place, in this case also around 53 percent, so the two methods agree here. The lowest-scoring categories, typographic hierarchy, sequencing, consistency of visual language, and application versus interpretive balance, are worth reading together, since they point to the same underlying issue, a lack of visual and structural differentiation between the book's various kinds of content. This book has considerable editing issues when read across and back and forth, there are some regions mentioned in initial cultural history and archaeological culture maps that never make another appearance, making many holes apparent to the close reader. Between chapters the only consistencies involves a discussion of the setting of the environment and a chronological handling of a given archaeological region, but coverage and depth is quite uneven and many times incomparable to the student or newcomer to Texas archaeology. The practitioner won't find a text easy to scan for what they are looking for, a tidy comparison of site types, idiosyncrasies, issues, and research questions at a glance.

The Field Manual for the Archaeology of Ritual, Religion, and Magic
Representing the fieldwork reference manual subtype of non-fiction, built to travel, to be consulted under time pressure, and to survive repeated handling outside a controlled environment.
Averaged across every individual item scored, this manual returns a usability index of about 84 percent. Averaged category by category, the figure lands in almost the same place, also around 84 percent. Sequencing is the outlier here, the category of whether the information is provided in the natural order of needing to know something before the next thing (especially when expected to act a certain way or to do what is next required), driven by the need for advising non-specialists in etiquette, cautions, or safety with visual cues. Consider how plant identification guides have warning icons around handling certain plants (e.g. Rocky Mountain maple seeds (samaras) have fine, stiff, translucent hairs that have sharp irritating tiny glass needles if touched. While advice on differential behavior of sacred places or sites of magical or religious behavior (e.g., no taking pictures of human remains, or taboos on certain ritual purity, monastic roles, or clan-based or gender rules around touching certain objects), the visitor may miss more seemingly mundane objects which may go missed (e.g., small object in a doorway like a mezuzah post's former site marked with appropriately space nail holes) or untested (e.g. singing stones written off as a metate) or undocumented. For example, a doorway icon next to mezuzah's listing in the Judaism would reinforce the spatial placement of said listed object and allow a reader in an instant to what to pivot towards and what to expect there. So an icon for each item in the two models of identification being utilized throughout sections in devices, site types, ethnicity, and religion. In the same vein as sequencing, the chapter order that does not yet match the order in how a reader encounters an unfamiliar site type in the field, which can have intentioned bouncing around. As a field manual, it benefits from pages that lay flat and can be managed hands-free or with one hand (margins and page openness), and from a durable cover that can weather a backpack (portability).

The Modern Herbal Dispensatory
Representing the step-by-step desk reference subtype of non-fiction, pairing instructional material, recipes and preparations, with the explanatory material behind it.
Averaged across every individual item scored, this guide returns a usability index of about 89 percent. Averaged category by category, the figure comes out lower, around 87 percent, the clearest instance among the three books where the two weighting methods diverge. That gap exists because this book's items are spread unevenly, back matter usability alone carries six scored items while several other categories carry only four, so an item-weighted average leans more heavily on whichever categories happen to have the most items scored inside them, while a category-weighted average does not. That will differ between the type of book you are evaluating. Similar to the field manual’s needs, this step-by-step guide where herbs, tinctures, and infusions are on the same table would benefit from pages which lay flat, hands-free (margins and page openness), and a more resilient cover material from contact with substances on the worktable or just from heavy use (portability). A back matter usability issue of note is the missing glossary, but is surprising light for back batter contents in general, and focuses the reader to flip back and forth for finding some things or search externally for what is meant by this term. In the context in clinical herbalism, it's important to clear up the different historical and multiple contemporaneous terminologies that sound the same but mean different things even in something as simple as energetics categories. Furnishing more comparison tables of content mined from Matthew Wood's deep historical herbalist knowledge would be useful here.

The evaluation, done this way, is closer to a translation than a checklist borrowed wholesale. It carries over the questions that still apply, retires the ones that do not, and builds new ones where the physical object raises problems a screen never has to solve. A website can be updated overnight if users report confusion. A misprinted run of books is a fixed cost that cannot be quietly patched. That asymmetry by itself changes how much weight an evaluator should put on clarity before anything goes to press, and it is exactly what made a physical-object walkthrough worth building for that client's question in the first place, a way to say with confidence whether the book itself was part of the problem before looking anywhere else.
A non-fiction author wondering whether their own book might be working against itself does not need the full instrument to start. Three checks will cover a lot of mileage without it: Open the book flat on a table and see whether it holds itself open without a hand pressing the spine, since a book that will not lay flat is already asking more of a reader who wants to follow instructions or take notes. Turn to the back and see whether a term the body text assumes the reader already knows can actually be found again, since a glossary or index that does not reach the vocabulary the front of the book relies on quietly works against its own explanations. Quality indexing is a must-have for non-fiction, and it is one of the essential backmatter heuristic evaluation items. If you find yourself needing an expert indexer, firms such as Courage to Enlighten who can help. Then flip through five pages in a row and see whether the eye can immediately tell instruction from explanation, since a page where every kind of content looks the same forces a reader to read everything closely just to find the part that matters – that in user experience terms is called “friction,” and is something that can surface in negative reader reviews.
This is the kind of adaptation HAR takes on regularly, not because physical objects are rare in client work, but because usability does not stop being usability when it leaves the screen. Whether the object in question is a book whose sales numbers do not add up, a printed insert packaged inside a consumer product, or a hybrid item that pairs something physical with a companion app, the same underlying questions apply. The tools built for interfaces still have something to say about them, whether you are an author trying to rule out your own book before spending further on marketing, a UX researcher building a version of this framework for your own case, or a product team wondering why a packaged instruction sheet keeps getting ignored. They just need to be asked correctly, and sometimes they need a few new questions added before they can be asked at all.

References
Augé, C. Riley. 2022. Field Manual for the Archaeology of Ritual, Religion, and Magic. New York: Berghahn Books.
Easley, Thomas, and Steven Horne. 2016. The Modern Herbal Dispensatory: A Medicine-Making Guide. Berkeley, CA: North Atlantic Books.
Nielsen, Jakob. 1993. Usability Engineering. Boston: Academic Press.
Norman, Don. 1988. The Design of Everyday Things. New York: Basic Books.
Perttula, Timothy K., ed. 2004. The Prehistory of Texas. College Station: Texas A&M University Press.

Comments