A Biomarker by Any Other Name Would Smell as Sweet: Building an AI-Powered Internal Concept Library

Part two of a two-part series on the idea of Internal Concept Libraries. Read part one.

Building an AI-powered internal concept library

The promise of artificial intelligence in pathology is often framed around image analysis and diagnostic augmentation. But one of the most immediate opportunities may lie elsewhere: transforming decades of unstructured pathology reports into computable, searchable clinical data.

A recent pathology informatics project by Raj Singh, Professor of Pathology at UPenn and PathPresenter co-founder, and Alexander Goel, COO at PhenoML, explored this challenge through the development of an AI-driven extraction pipeline designed around a central principle: preserve all clinically meaningful observations, even when standardized terminology systems fail to represent them completely.

The project focused on constructing an Internal Concept Library capable of bridging the gap between narrative pathology reporting and structured research databases.

The dataset consisted of 1,155 free-text pathology reports drawn from The Cancer Genome Atlas (TCGA), including breast, lung, and prostate cancer cases spanning 1978 to 2013. These reports largely predated modern synoptic reporting practices and therefore represented the type of highly variable narrative text still common throughout historical pathology archives.

The objective was not simply to extract data fields. It was to evaluate whether a concept-first architecture could preserve more clinically useful information than traditional terminology-dependent pipelines.

The system architecture was intentionally designed to separate clinical meaning from terminology standardization.

The pipeline began with ingestion of raw pathology reports. Cases were first routed through disease-specific classification using external terminology APIs and machine learning classifiers. Each report was then processed using large language model prompts tailored to individual disease sites.

Rather than allowing unconstrained extraction, the prompts used a closed-world design philosophy. The model was explicitly instructed that the field list was exhaustive, preventing the AI from hallucinating expected findings that were not actually documented in the report.

Each extracted observation was assigned an internal concept identifier before any attempt at external terminology mapping occurred. This architectural decision represented the core innovation of the project.

In traditional pipelines, extraction is often tightly coupled to successful SNOMED encoding. If no suitable standard code is found, the observation may be excluded from downstream datasets entirely.

In the concept-library model, however, the internal concept persists regardless of SNOMED availability.

Only after extraction were findings submitted to terminology matching systems to identify corresponding SNOMED CT or OMOP standard concepts. When no standard code existed, custom OMOP concept IDs were assigned instead, preserving full queryability within the OMOP ecosystem. The results highlighted the scale of the terminology gap.

Study results
Results of extraction and mapping

Across 1,155 reports, the system extracted 13,323 clinical observations. Of these, approximately 59.6% successfully mapped to SNOMED CT concepts. Roughly 39.1% could not be mapped to standard terminology despite representing clinically meaningful findings. An additional 1.3% were flagged as uncertain due to ambiguity or conflicting evidence within the report text.

Under conventional code-first architectures, a substantial portion of these observations would likely have been excluded from structured datasets entirely. Instead, the Internal Concept Library preserved all extracted findings.

The project also evaluated extraction completeness against the College of American Pathologists electronic Cancer Protocols (CAP eCPs), using the CAP templates as a formal benchmark for information coverage. This approach itself represented a novel methodological contribution, as CAP templates have rarely been used as explicit completeness standards in published large language model pathology extraction research.

Coverage varied by disease site. Breast reports achieved approximately 32.9% CAP field extraction, prostate reports 36.5%, and lung reports 21.8%, with an overall completeness rate of 29.2%.

At first glance, those percentages may appear modest. But the context matters enormously.

The TCGA reports originated from an era before structured synoptic reporting became widespread. Many CAP-defined fields simply did not exist within the source narratives. In other words, lower completeness often reflected absence of documentation rather than extraction failure.

The researchers hypothesized that applying the same pipeline to contemporary synoptic or semi-structured pathology reports would likely produce substantially higher completeness rates.

Several important design lessons emerged from the project.

  • First, closed-world prompting proved essential. Explicitly constraining the extraction schema reduced over-extraction and prevented the model from inferring clinically plausible but undocumented findings.
  • Second, the team found that self-reported LLM confidence scores were poorly calibrated and unreliable for quality assurance. Instead, better performance indicators emerged from structured assertion logic (e.g. forcing every extracted finding into explicit states such as present, absent, or uncertain rather than allowing vague narrative interpretation), explicit review of terminology concordance (verifying that extracted findings align consistently with accepted clinical vocabularies), and cross-field consistency checks (evaluating whether extracted observations remain clinically coherent when considered together).
  • Third, the researchers intentionally avoided creating separate post-extraction normalization stages. Additional normalization layers often compound error propagation while duplicating functionality already present in terminology APIs.

Most importantly, the study demonstrated that pathology AI systems should prioritize preservation of meaning over immediate standardization.

This distinction has major implications for clinical trial recruitment and translational research. Using concept-level querying, researchers could identify highly specific patient cohorts using combinations of biomarkers, grading systems, and morphologic findings regardless of whether standardized terminology mappings existed for every observation.

Queries such as:

  • Triple-negative breast cancers with Ki-67 greater than 50%
  • Gleason score ≥8 with perineural invasion

could operate across the full extracted dataset, not merely the subset successfully encoded in SNOMED.

The project also pointed toward future directions in ontology management itself. Singh and Goel are considering multi-agent AI frameworks capable of automatically reviewing new observations, searching multiple ontologies, drafting candidate mappings, and escalating only genuinely ambiguous cases for human expert review. In such systems, pathologists and informaticians would spend less time performing repetitive terminology lookups and more time adjudicating clinically meaningful ambiguity.

Internal Concept Library multi agent framework

Ultimately, the work reinforces a broader shift occurring across healthcare AI. The goal is no longer simply extracting structured data from pathology reports. The larger challenge is preserving clinical meaning at scale while enabling interoperability, research, and machine reasoning.

Pathologists have always documented nuanced clinical interpretation for human readers. Building systems that preserve that same nuance for computational systems may become one of the defining informatics challenges of precision medicine.

The algorithm is ready. The question is whether our departments are.

As a pathologist, this feels to me like one of the most important moments our specialty has seen in decades.

Not because AI in pathology is new. Most academic departments have been thinking about computational pathology, biomarker quantification, and digital workflows for the last few years. What feels different now is that these tools are no longer being discussed primarily as research infrastructure or future possibilities. They are beginning to shape real treatment decisions.

Roche’s acquisition of PathAI is a significant signal that computational pathology has moved into the center of precision oncology strategy. The implications are larger than a single acquisition.

For years, pathologists have understood something that the broader healthcare ecosystem is only now fully appreciating: the biopsy is not simply supporting information for oncology care. Increasingly, it is where therapeutic eligibility is determined. As AI-enabled companion diagnostics become regulatory-grade clinical tools, pathology workflows are becoming directly connected to treatment access.

The emerging generation of ADCs and biomarker-driven therapies illustrates this clearly. Many of these treatments depend on increasingly nuanced tissue-based interpretation — quantitative scoring, spatial context, low-expression biomarkers, and patterns difficult to assess reproducibly through conventional microscopy alone. The pathologist remains central to interpretation, but the infrastructure around that interpretation is changing rapidly.

The Real Challenge: Deployment Over Development

An algorithm may be developed in a research environment, but it ultimately it has to function inside a clinical department. It must integrate into the daily workflow of pathologists, connect to existing laboratory systems, support regulatory and reporting requirements, and operate within institutions that have already made long-term infrastructure decisions. That deployment challenge is likely to become one of the defining questions of computational pathology over the next several years.

The academic medical centers where many of these patients are diagnosed are extraordinarily complex environments. They are simultaneously delivering clinical care, running trials, training residents and fellows, conducting translational research, supporting tumor boards, and managing large-scale image archives. Any AI companion diagnostic that hopes to achieve broad adoption will need to fit naturally into that ecosystem.

This is why workflow and institutional trust matter as much as algorithmic performance.

Pathology departments do not adopt infrastructure lightly. The systems that succeed over time are the ones that integrate cleanly into clinical operations, support a wide range of departmental needs, and earn confidence through years of daily use. In practice, the operational realities of pathology often determine whether even highly sophisticated technologies can meaningfully reach patients.

None of this diminishes what companies like Roche and PathAI have accomplished. The field needs validated, clinically deployable AI systems, and the progress being made is genuinely exciting for pathology and oncology alike.

But the next phase of the field may be defined less by whether these algorithms can be built, and more by how effectively they can be deployed within the institutions where precision oncology is actually practiced.

About the Author

Dr. Rajendra Singh is a Professor of Pathology at the University of Pennsylvania and co-founder of PathPresenter. He serves as a member of the Digital and Computational Pathology Committee of the CAP, Editorial Board of the WHO for Classification of tumors, 5th Edition and the Board of Digital Pathology Association.

Why a Vendor-Agnostic IMS Is Key to a Future-Proof AI Strategy

by Patrick Myles, MBA
CEO, PathPresenter

As AI adoption accelerates in digital pathology, many institutions are facing a strategic decision that will shape their technology stack for years to come: how tightly should AI be coupled to the core imaging and workflow platform?

At PathPresenter, we’ve taken a very deliberate position. On the AI side, we are intentionally AI-agnostic at the IMS and workflow level. That decision is rooted not in theory, but in hard-earned lessons from the field.

The Hidden Cost of Closed AI Ecosystems

Over the years, we’ve seen multiple institutions make large investments in closed scanning and workflow ecosystems, often driven by early regulatory approvals or the promise of an “end-to-end” solution. In the short term, those systems worked. But over time, they frequently became bottlenecks.

When institutions later wanted to onboard additional scanners, integrate third-party tools, or adopt stronger algorithms from competing vendors, those closed platforms made change slow, expensive, or in some cases impossible. The result was delayed deployments, integration failures, and ultimately costly replacement projects that disrupted operations far more than anyone anticipated.

That model simply doesn’t hold up in today’s AI landscape.

AI Is Changing Too Quickly for Locked-In Platforms

AI is moving at a pace we’ve never seen before. Many algorithms deployed in production today are CNN-based, while the next wave of capability is clearly shifting toward foundation-model approaches. As performance improves and new vendors emerge, the “best” solution will continue to change.

In that environment, locking the core IMS and workflow layer to a specific AI vendor is a strategic risk. Our view is simple: the IMS should remain independent of the AI layer.

A Future-Proof Strategy that Keeps the Platform Stable While AI Evolves

To make that separation practical, we’ve invested heavily in building an AI middleware layer designed to be future-proof. This middleware can ingest, normalize, and orchestrate the clinical and operational data required to deploy AI at scale, including integrations built on HL7, FHIR, and DICOMweb. By abstracting these complexities away from the core platform, institutions can adopt new AI models and workflows over time without needing to rebuild integrations or re-platform each time the ecosystem evolves.

Just as importantly, this approach is informed by real-world experience. We’ve worked with a wide range of institutions and organizations and have seen multiple generations of digital pathology and AI vendors come and go. Some early companies were highly innovative, but struggled to scale or keep pace with market shifts.

Those experiences have consistently reinforced the same conclusion: workflow platforms must remain stable, while the AI layer must remain flexible.

In an AI landscape defined by constant change, future-proofing isn’t about predicting which algorithm will win. It’s about designing systems that can evolve, without disruption, as the technology does.

About the Author

Patrick Myles is CEO of PathPresenter. Previously, he was CEO of Huron Digital Pathology, and Vice President of Business Development for Teledyne DALSA. He served as a board member of the Digital Pathology Association.

PathPresenter and Aisencia partner for AI-Powered Workflows

Calling all AAD2025 attendees! See the future of dermatopathology workflows at the PathPresenter booth (number 816). We have partnered with AI vendor Aisencia to expedite how dermatopathologists sign out cases. Meet with PathPresenter’s Cory Batenchuk, and Catherine Cline and see how to optimize your digital derm workflow for speed, efficiency, and precision.

By combining cutting-edge AI technology with decades of dermatopathology expertise, we have created a unique workflow optimized for speed, efficiency, and precision.

Why Choose PathPresenter + Aisencia?

Our founder, Dr. Rajendra Singh, brings over 20 years of dermatopathology experience to this innovative solution. His insights have been instrumental in designing a system that delivers maximum efficiency and accuracy for dermatopathologists. Aisencia is an AI company specifically focused on dermatopathology and digital workflows, with a first-of-its-kind model that already detects and classifies 50+ types of skin lesions.

Key Features

Smart Case Prioritization

  • AI ranks cases from suspicious to non-suspicious, ensuring high-priority cases are reviewed first.

AI-Powered Macro-Based Reporting

  • Reports are pre-populated by AI for faster sign-out.
  • Upcoming voice command integration allows seamless updates and corrections.

DermPath Assist

  • AI Chatbot trained specifically for dermatopathology 
  • Functions as a co-pathologist, offering diagnostic suggestions based on case features.

Flexible Case Sign-Out

  • Supports TC/PC models for signing out cases by institution or external sites.
  • Dermatologists can push cases back to the main organization when needed.

LIS Integration

  • Fully deployed with CoPath, NovoPath, and Epic Beaker for a seamless workflow experience.

Powered by Aisencia

With AI-based case prioritization and macro-reporting that condenses the results of 50+ precision AI models for dermal lesions, our solution is set to make a significant impact on dermatopathology workflows.

Let’s Connect!

We’re excited to share how AI is reshaping digital pathology. Visit us at AAD2025 (booth 816) and discover how PathPresenter and Aisencia can transform your institution’s dermatopathology practices.

Let’s drive innovation together!