A Biomarker by Any Other Name Would Smell as Sweet: Building an AI-Powered Internal Concept Library

Part two of a two-part series on the idea of Internal Concept Libraries. Read part one.

Building an AI-powered internal concept library

The promise of artificial intelligence in pathology is often framed around image analysis and diagnostic augmentation. But one of the most immediate opportunities may lie elsewhere: transforming decades of unstructured pathology reports into computable, searchable clinical data.

A recent pathology informatics project by Raj Singh, Professor of Pathology at UPenn and PathPresenter co-founder, and Alexander Goel, COO at PhenoML, explored this challenge through the development of an AI-driven extraction pipeline designed around a central principle: preserve all clinically meaningful observations, even when standardized terminology systems fail to represent them completely.

The project focused on constructing an Internal Concept Library capable of bridging the gap between narrative pathology reporting and structured research databases.

The dataset consisted of 1,155 free-text pathology reports drawn from The Cancer Genome Atlas (TCGA), including breast, lung, and prostate cancer cases spanning 1978 to 2013. These reports largely predated modern synoptic reporting practices and therefore represented the type of highly variable narrative text still common throughout historical pathology archives.

The objective was not simply to extract data fields. It was to evaluate whether a concept-first architecture could preserve more clinically useful information than traditional terminology-dependent pipelines.

The system architecture was intentionally designed to separate clinical meaning from terminology standardization.

The pipeline began with ingestion of raw pathology reports. Cases were first routed through disease-specific classification using external terminology APIs and machine learning classifiers. Each report was then processed using large language model prompts tailored to individual disease sites.

Rather than allowing unconstrained extraction, the prompts used a closed-world design philosophy. The model was explicitly instructed that the field list was exhaustive, preventing the AI from hallucinating expected findings that were not actually documented in the report.

Each extracted observation was assigned an internal concept identifier before any attempt at external terminology mapping occurred. This architectural decision represented the core innovation of the project.

In traditional pipelines, extraction is often tightly coupled to successful SNOMED encoding. If no suitable standard code is found, the observation may be excluded from downstream datasets entirely.

In the concept-library model, however, the internal concept persists regardless of SNOMED availability.

Only after extraction were findings submitted to terminology matching systems to identify corresponding SNOMED CT or OMOP standard concepts. When no standard code existed, custom OMOP concept IDs were assigned instead, preserving full queryability within the OMOP ecosystem. The results highlighted the scale of the terminology gap.

Study results
Results of extraction and mapping

Across 1,155 reports, the system extracted 13,323 clinical observations. Of these, approximately 59.6% successfully mapped to SNOMED CT concepts. Roughly 39.1% could not be mapped to standard terminology despite representing clinically meaningful findings. An additional 1.3% were flagged as uncertain due to ambiguity or conflicting evidence within the report text.

Under conventional code-first architectures, a substantial portion of these observations would likely have been excluded from structured datasets entirely. Instead, the Internal Concept Library preserved all extracted findings.

The project also evaluated extraction completeness against the College of American Pathologists electronic Cancer Protocols (CAP eCPs), using the CAP templates as a formal benchmark for information coverage. This approach itself represented a novel methodological contribution, as CAP templates have rarely been used as explicit completeness standards in published large language model pathology extraction research.

Coverage varied by disease site. Breast reports achieved approximately 32.9% CAP field extraction, prostate reports 36.5%, and lung reports 21.8%, with an overall completeness rate of 29.2%.

At first glance, those percentages may appear modest. But the context matters enormously.

The TCGA reports originated from an era before structured synoptic reporting became widespread. Many CAP-defined fields simply did not exist within the source narratives. In other words, lower completeness often reflected absence of documentation rather than extraction failure.

The researchers hypothesized that applying the same pipeline to contemporary synoptic or semi-structured pathology reports would likely produce substantially higher completeness rates.

Several important design lessons emerged from the project.

  • First, closed-world prompting proved essential. Explicitly constraining the extraction schema reduced over-extraction and prevented the model from inferring clinically plausible but undocumented findings.
  • Second, the team found that self-reported LLM confidence scores were poorly calibrated and unreliable for quality assurance. Instead, better performance indicators emerged from structured assertion logic (e.g. forcing every extracted finding into explicit states such as present, absent, or uncertain rather than allowing vague narrative interpretation), explicit review of terminology concordance (verifying that extracted findings align consistently with accepted clinical vocabularies), and cross-field consistency checks (evaluating whether extracted observations remain clinically coherent when considered together).
  • Third, the researchers intentionally avoided creating separate post-extraction normalization stages. Additional normalization layers often compound error propagation while duplicating functionality already present in terminology APIs.

Most importantly, the study demonstrated that pathology AI systems should prioritize preservation of meaning over immediate standardization.

This distinction has major implications for clinical trial recruitment and translational research. Using concept-level querying, researchers could identify highly specific patient cohorts using combinations of biomarkers, grading systems, and morphologic findings regardless of whether standardized terminology mappings existed for every observation.

Queries such as:

  • Triple-negative breast cancers with Ki-67 greater than 50%
  • Gleason score ≥8 with perineural invasion

could operate across the full extracted dataset, not merely the subset successfully encoded in SNOMED.

The project also pointed toward future directions in ontology management itself. Singh and Goel are considering multi-agent AI frameworks capable of automatically reviewing new observations, searching multiple ontologies, drafting candidate mappings, and escalating only genuinely ambiguous cases for human expert review. In such systems, pathologists and informaticians would spend less time performing repetitive terminology lookups and more time adjudicating clinically meaningful ambiguity.

Internal Concept Library multi agent framework

Ultimately, the work reinforces a broader shift occurring across healthcare AI. The goal is no longer simply extracting structured data from pathology reports. The larger challenge is preserving clinical meaning at scale while enabling interoperability, research, and machine reasoning.

Pathologists have always documented nuanced clinical interpretation for human readers. Building systems that preserve that same nuance for computational systems may become one of the defining informatics challenges of precision medicine.

Why Seamless IMS-LIS Integration Is Essential for Digital Pathology

Unified digital workflows connecting IMS to LIS unlock efficiency, compliance, and the full potential of AI for modern pathology

Digital pathology has moved decisively from experimentation into everyday practice. For pathology organizations, it now underpins clinical workflows, operational efficiency, and long-term strategy. To enable remote sign-out and AI-driven analysis, organizations have turned to Image Management Systems (IMS) that can organize, track, and contextualize whole slide images across the diagnostic workflow.

However, the data needed to interpret those images extends well beyond what an IMS typically manages. Patient demographics, clinical context, specimen metadata, test orders, results, finalized reports and more all live in a separate but equally critical system: the Laboratory Information System (LIS). For groups accustomed to thinking primarily in terms of image workflows, the LIS can feel like a parallel universe, yet it is the system of record for nearly all non-image pathology data. When IMS and LIS platforms operate in isolation, inefficiencies emerge, context is lost, and the promise of digital pathology remains only partially realized.

As digital pathology programs mature, it is becoming evident that success depends not only on performance, scale, and regulatory readiness, but also on seamless interoperability. In this article, we explore why deep, context-aware integration between an IMS and a modern LIS is essential to unlocking the full value of digital pathology. We’ll examine the practical challenges this integration resolves and the strategic benefits it enables for pathology groups advancing toward a fully digital, AI-enabled future.

Image Management Systems (IMS) in Digital Pathology

A digital pathology image management system (IMS) goes far beyond mere image viewing. At its best, it is an enterprise-grade platform designed to manage the entire lifecycle of digital pathology images across clinical, research, and educational workflows, and serving as the backbone of digital pathology operations, supporting everything from secure storage to AI integration and multi-institution collaboration.

Core Capabilities of a Modern Digital Pathology IMS

A robust IMS typically includes:

  • Advanced Slide Viewing: High-performance visualization across multiple formats (SVS, NDPI, SCN, DICOM, BigTIFF, and more), synchronized multi-slide viewing, annotations, overlays, and fast rendering.
  • Centralized Data Management: Secure, scalable image storage, on-premises, cloud, or hybrid, with governance controls and long-term data stewardship.
  • Security and Compliance: HIPAA-grade protections, encryption, audit trails, role-based access, and regulatory readiness for clinical use.
  • Workflow Integration: Case management, task tracking, and seamless interoperability with LIS, EHR, PACS, and other enterprise laboratory software systems.
  • User Roles and Access Control: Granular permissions ensure that only authorized users access sensitive data, while maintaining accountability.
  • Collaboration Tools: Support for remote consults, tumor boards, education, and peer review through shared environments.
  • AI and Image Analysis Enablement: Integration with AI models for detection, quantification, and pattern recognition, embedded directly into diagnostic workflows.
  • Regulatory Readiness: Many IMS platforms are designed to meet FDA, HIPAA, GDPR, IVDR, and other regulatory standards. Some, such as PathPresenter’s clinical viewer, are FDA 510(k) cleared and EU-IVDR certified for primary diagnosis when paired with approved scanners.

In short, an IMS provides a comprehensive infrastructure for digital pathology. It can be highly capable on its own, handling slide storage, visualization, annotations, and even AI outputs. But when it operates independently from the laboratory’s core operational system, its impact is inherently limited. Rather than dissolving legacy barriers, a disconnected IMS can unintentionally recreate them in digital form.

True digital pathology maturity is reached only when images and laboratory data function as a single, coherent workflow. Without that alignment, pathology groups often find themselves managing parallel systems that never quite converge.

The Operational Cost of Keeping IMS and LIS Separate

When image workflows and laboratory workflows are not connected at a foundational level, pathology teams encounter familiar but amplified challenges, including:

  • Reliance on manual reconciliation of cases and slides
  • Repeated entry of patient and specimen data
  • Frequent toggling between systems during case review
  • Greater exposure to mismatches, omissions, and reporting errors
  • Disjointed compliance records spanning multiple platforms
  • Collaboration slowdowns across sites and subspecialties

In isolation, each issue may seem manageable. At scale however, particularly in regulated, high-throughput environments, they compound rapidly, eroding both efficiency and confidence in the digital workflow.

Rethinking the LIS: More Than a Back-End Database

For many digital pathology teams, the LIS is perceived primarily as a static repository: the place where orders originate and reports ultimately land. While that historical role remains essential, it no longer reflects how modern laboratories operate.

From Passive Record-Keeping to Active Workflow Orchestration

Today’s laboratories increasingly require systems that do more than store information. A modern LIS such as LigoLab functions as an operational engine: driving workflow logic, enforcing rules, automating handoffs, and presenting the right context at the right moment.

When an IMS is natively integrated into this environment, digital pathology is no longer an external tool that users “visit.” Instead, image review becomes a natural extension of the laboratory workflow, governed by the same logic, permissions, and traceability as the rest of the case lifecycle.

What Integrated LIS–IMS Workflows Enable Day to Day

In a genuinely unified digital pathology environment:

  • Digital slides are opened directly from the laboratory case, not searched for separately
  • Patient, specimen, and order context flows automatically into image review
  • Movement between data, images, and reports feels continuous rather than segmented
  • Auditability covers both diagnostic decisions and image interactions end to end
  • AI outputs appear within established workflows rather than as side-channel tools

This level of integration changes how work is performed: not by adding features, but by removing friction.

LIS-IMS integration

Core Challenges Solved by Deep Integration

1. Smoother, Faster Diagnostic Workflows

When image access is embedded within the laboratory system, pathologists spend less time navigating software and more time interpreting cases. Reduced friction translates directly into faster turnaround and lower cognitive load.

2. Scalable Support for Remote Practice

Integrated platforms make location largely irrelevant. Secure remote sign-out, real-time collaboration, and distributed subspecialty review become standard operations rather than exceptions.

3. Better Decisions Through Complete Context

Images rarely tell the full story on their own. Tight coupling between LIS and IMS ensures that diagnostic interpretation always occurs alongside the complete clinical and specimen context, reducing the risk of incomplete assessment.

4. Stronger Governance and Compliance

When images and laboratory data share a unified audit framework, compliance with regulatory and accreditation requirements becomes simpler and more defensible—without added administrative burden.

5. A Practical Path to Enterprise AI

AI delivers value only when it operates inside real workflows. Integrated LIS–IMS environments provide the governance, data access, and workflow hooks required to deploy AI responsibly and at scale.

Integration in the Real World: From Concept to Practice

A practical example of this approach can be seen in the collaboration between PathPresenter and LigoLab. Rather than treating image management as an external system, PathPresenter’s clinical viewer is embedded directly within LigoLab’s enterprise laboratory platform.

The result is a unified diagnostic experience where pathologists interact with slides in the same environment where cases are managed, orders are tracked, and reports are finalized. This tight coupling reduces navigation overhead, accelerates review cycles, and establishes a solid foundation for future AI-driven enhancements.

Related: LigoLab and PathPresenter Announce Strategic Partnership to Deliver Seamless Digital Pathology Workflows

Building Toward a Unified Digital Pathology Future

At its core, LIS–IMS integration is not about software convenience. It is about redefining how pathology work flows from accessioning through diagnosis and reporting.

When laboratory data and images operate as one system:

  • Operations gain visibility and automation
  • Financial teams benefit from cleaner alignment with billing and revenue workflows
  • Clinical staff gain modern tools that reduce manual effort
  • Patients benefit from faster, more consistent diagnostic outcomes

As digital pathology, AI, and remote diagnostics continue to evolve, this alignment moves from “nice to have” to mission critical. Pathology groups that invest early in deep, contextual integration position themselves not just to digitize slides, but to fundamentally modernize how pathology is practiced.

More Information

Looking for a new (or improved) LIS-IMS integration? Connect with our team to assess your workflow needs and explore the best solution.

Choosing the Right Digital Pathology Workflow Partner: So Many Options

Digital pathology is rapidly evolving into a central component of modern diagnostics, education, and research. The adoption of whole slide imaging, artificial intelligence (AI), and cloud-based platforms is reshaping workflows and collaboration. But technology alone does not guarantee success—choosing the right workflow partner is the critical factor that determines whether digital transformation delivers meaningful impact.

Core Elements of a Strong Workflow Partnership

1. Domain Knowledge and Expertise of the Vendor

Every digital pathology journey is a journey, so choosing a vendor with deep domain knowledge and expertise is critical. To consider:

  • Does the vendor have extensive experience in pathology workflows?
  • Are they able to identify the gaps in pathology workflows?
  • Do they have the experience to serve as a guide to a successful implementation?

Domain expertise and experience contribute directly to a successful implementation.

2. Interoperability

A digital pathology platform must integrate seamlessly with existing systems. This includes:

  • LIS/EMR compatibility through HL7, DICOM, and other standards.
  • Support for diverse file formats (WSI, JPEG, DICOM).
  • Ability to integrate scanners, AI tools, and storage infrastructure.

Interoperability ensures that digital pathology becomes part of the clinical workflow, not an isolated add-on.

3. Scalability and Flexibility

Institutions vary in size and needs, but a platform should be able to grow with them. Key factors include:

  • Handling both small-scale implementations and large archives of millions of slides.
  • Flexible deployment models (cloud, hybrid, on-premise).
  • Configurable workflows that adapt to clinical, educational, and research use cases.

A scalable partner protects investments and avoids costly migrations later.

4. Security and Compliance

Patient data and digital slides are sensitive assets. A trusted partner must provide:

  • HIPAA-compliant infrastructure.
  • Strong data encryption and access controls.
  • Transparent chain-of-custody management.
  • Regular security audits and compliance certifications.

This foundation ensures safety, trust, and regulatory alignment.

5. User Experience and Adoption

Even the most advanced technology fails if users resist it. A partner should prioritize:

  • Intuitive user interfaces that mimic the glass slide experience.
  • Tools that support pathologists’ natural workflows (annotation, conferencing, consults).
  • Minimal technical barriers for both in-house and remote users.

Familiarity breeds confidence, leading to higher adoption and satisfaction.

6. AI Enablement and Future-Proofing

Digital pathology is closely tied to AI innovation. An ideal partner should:

  • Support annotation tools for AI training and validation.
  • Offer seamless deployment of algorithms within the workflow.
  • Provide APIs and middleware for integration with multiple vendors.

Future-proofing with AI ensures the platform remains relevant as the field advances.

7. ROI and Value Creation

Institutions must justify investment with clear returns. Value comes from:

  • Reduced storage and logistics costs.
  • Faster turnaround times for consultations and sign-outs.
  • Enhanced collaboration and education opportunities.
  • Ability to monetize research data through partnerships.

The right partner helps institutions build a compelling business case.

8. Collaboration and Cultural Fit

Technology is only half of the equation; the other half is people. The partner should demonstrate:

  • Willingness to co-develop solutions.
  • Strong track record of customer engagement and support.
  • Transparent communication and long-term commitment.

A true partner invests in mutual success rather than a transactional relationship.

Future Directions

Digital pathology partnerships will increasingly be shaped by:

  • Standardization through DICOM adoption and interoperability frameworks.
  • Enterprise imaging integration, aligning pathology with radiology and other disciplines.
  • Practical AI applications, moving from hype to validated, clinically useful tools.

These trends will demand partners who are both innovative and realistic, balancing future readiness with present-day reliability.

Conclusion

Choosing the right digital pathology workflow partner is about creating bridges—between systems, between people, and between present workflows and future possibilities. The decision should be guided by interoperability, scalability, security, usability, AI readiness, ROI, and cultural fit.

When these elements align, institutions unlock the full potential of digital pathology: efficient operations, powerful collaboration, enriched education, and most importantly, improved patient care.

More Information

PathPresenter has guided a wide range of organizations to success on their digital pathology journey. Interested in discussing your own challenges? Contact us.