A search-based pathological section AI diagnosis auxiliary system and method

CN122817499APending Publication Date: 2026-09-25HANGZHOU XINO INTELLIGENT MEDICINE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610929626.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

该类方案通常存在以下问题:其一,不同医院来源的病理切片在染色方式、扫描设备及制片协议上存在明显差异,导致不同来源切片的特征分布不一致,跨机构病例检索效果不稳定;其二,现有系统多仅返回相似病例列表,缺乏对“为什么相似”的可视化解释,医生难以据此快速验证AI推荐依据;其三,现有系统对相似病例结果缺乏进一步的标签统计分析和反馈优化机制,难以形成持续改进的辅助诊断闭环

Benefits of technology

本发明以病例检索替代传统分类输出,将待诊断病理切片与历史病例在统一特征空间中进行相似度匹配,使诊断参考来源于真实病例而非预设类别边界,从而降低对大规模标注数据的依赖,在小样本或新机构场景下仍能够提供稳定参考,解决现有方法在数据分布变化条件下适用性不足的问题,体现出以“检索驱动诊断”的系统性技术路径创新。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817499A_ABST
    Figure CN122817499A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of pathological diagnosis assistance, and particularly discloses an AI diagnosis assistance system and method for pathological sections based on retrieval, which comprises a data warehousing module, a unified feature coding calling module, a case retrieval library construction module, a retrieval index construction module, a real-time retrieval module, an explainability display module, a diagnosis reference generation module and a self-evolution optimization module; a pre-constructed cross-domain pathological image unified representation model is called to perform unified feature coding on pathological sections of different sources, a case retrieval library and an approximate nearest neighbor retrieval index are constructed, candidate case retrieval and accurate reordering are performed on pathological sections to be diagnosed, a similar case list is obtained, a key similar area retrieval heat map and a diagnosis reference abstract are generated, and case retrieval library updating and retrieval strategy optimization are performed based on doctor feedback. The application realizes comparable retrieval and explainable auxiliary diagnosis of pathological sections of different institutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pathological diagnostic assistance technology, specifically to an AI-assisted system and method for pathological slide diagnosis based on retrieval. Background Technology

[0002] Currently, there are two main technical approaches to AI-assisted pathology diagnosis. The first is the classification model approach, which trains a deep neural network based on a large number of labeled pathology slides and directly outputs the diagnostic category or confidence level for the input slide. The second is the retrieval-assisted approach, which retrieves similar cases from a historical case database for the pathology slide to be diagnosed and provides diagnostic references to doctors. The former relies heavily on large-scale labeled data, and its deployment cost is high in new hospitals, rare diseases, or small sample scenarios. Moreover, the model output usually lacks intuitive and verifiable references. While the latter can provide doctors with similar case references, if the feature representation used is not compatible with pathology slides from different hospitals, with different equipment, and under different slide preparation conditions, it can easily lead to incomparable features between cases, affecting the accuracy of retrieval.

[0003] Existing retrieval-based pathology auxiliary diagnostic systems typically construct case databases using pre-trained features from natural images or simple feature encoding methods, and directly perform similarity matching after the slide to be diagnosed is input. This approach generally suffers from the following problems: First, pathology slides from different hospitals exhibit significant differences in staining methods, scanning equipment, and slide preparation protocols, leading to inconsistent feature distributions and unstable cross-institutional case retrieval results. Second, existing systems mostly only return a list of similar cases, lacking a visual explanation of "why they are similar," making it difficult for doctors to quickly verify the basis of AI recommendations. Third, existing systems lack further label statistical analysis and feedback optimization mechanisms for similar case results, making it difficult to form a continuously improving closed loop for auxiliary diagnosis.

[0004] Therefore, there is an urgent need for a retrieval-based AI-assisted diagnostic system and method for pathological slides, which, based on the unified feature representation capabilities, can achieve cross-institutional case retrieval, interpretable display of retrieval results, generation of diagnostic reference summaries, and feedback-driven optimization, thereby improving its clinical application value. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of existing technologies, this invention provides a retrieval-based AI-assisted diagnostic system and method for pathological slides. This system systematically integrates various processing steps in pathological slide-assisted diagnosis, including case organization, feature invocation, retrieval index construction, real-time retrieval, result interpretation, diagnostic reference generation, and feedback optimization. For pathological slides from different sources, structured data entry and unified feature encoding are first performed, followed by the construction of a case retrieval library and an approximate nearest neighbor retrieval index. For pathological slides to be diagnosed, a pre-constructed cross-domain pathological image unified representation model is first invoked to perform unified feature encoding, followed by rapid retrieval of candidate cases and precise similarity calculation and reordering. Furthermore, a retrieval heatmap of key similar regions and a diagnostic reference summary are generated. Simultaneously, the case retrieval library and retrieval strategy are continuously updated based on physician feedback, thereby achieving transparent and scalable assisted diagnosis for clinical applications, thus addressing the problems raised in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A retrieval-based AI-assisted diagnosis system for pathological slides includes a data entry module, a unified feature encoding module, a case retrieval library construction module, a retrieval index construction module, a real-time retrieval module, an interpretability display module, a diagnostic reference generation module, and a self-evolutionary optimization module. The data entry module receives and structures pathological slide data and corresponding case information from different medical institutions. The unified feature encoding module invokes a pre-built cross-domain unified representation model for pathological images to perform unified feature encoding on the entered pathological slides and the pathological slides to be diagnosed, obtaining slide-level feature vectors. The case retrieval library construction module associates and stores the slide-level feature vectors with corresponding case information to construct a case retrieval library. The retrieval index construction module constructs an approximate nearest neighbor retrieval index based on the slide-level feature vectors. The real-time retrieval module performs candidate case retrieval and precise re-ranking based on query feature vectors to obtain a list of similar cases. The interpretability display module generates a retrieval heatmap of key similar regions. The diagnostic reference generation module generates a diagnostic reference summary. The self-evolutionary optimization module updates the case retrieval library and optimizes the retrieval strategy based on doctor feedback. The above technical solution is based on unified representation invocation, supported by a case retrieval database and a near nearest neighbor index, and centered on two-stage retrieval, heatmap interpretation, diagnostic reference, and feedback optimization, so that the system forms a complete closed loop.

[0007] As a further aspect of this invention, when the data entry module receives pathological slide data from different medical institutions, it records the unique case identifier, slide identifier, diagnostic label, hospital source information, and clinical metadata related to the retrieval in a unified data format. It can also perform sampling and organization of historical pathological slides according to disease type distribution using a preset script, generating a structured retrieval data file. For new case slides continuously accessed by collaborating hospitals, the data entry module receives their original slide images and related diagnostic information through a standard interface and incorporates them into the processing flow according to the same fields. Through this method, the case retrieval library can simultaneously cover historical cases from the source domain and new cases from the target domain, providing a foundation for cross-institutional case reference retrieval.

[0008] As a further aspect of this invention, the unified feature encoding invocation module applies the same encoding rules to both the imported pathological slides and the pathological slides to be diagnosed, and outputs fixed-dimensional slide-level feature vectors. This allows pathological slides from different sources to be mapped to the same feature space, enabling direct similarity comparison between pathological slides from different institutions. It should be noted that the unified representation model for cross-domain pathological images can be pre-trained using an independent underlying representation learning method. In this invention, only the unified representation model is invoked to complete the unified encoding function, without specifically limiting the internal training process of the unified representation model. This ensures that the system solution has a unified metric foundation in cross-institutional scenarios, while maintaining a clear protection boundary between this invention and the underlying cross-domain representation learning algorithm.

[0009] As a further aspect of this invention, the retrieval index building module constructs an approximate nearest neighbor retrieval index based on slice-level feature vectors in the case retrieval database to support efficient queries in large-scale case retrieval scenarios. The approximate nearest neighbor retrieval index can employ a hierarchical graph index, hash index, or other index structures suitable for fast retrieval of high-dimensional features. After the case retrieval database is constructed, the retrieval index building module can synchronously update the index, enabling the system to maintain a high online response capability even as the case scale continues to expand. Compared to calculating the full similarity for each case individually, the approximate nearest neighbor retrieval index significantly reduces the computational load, providing efficiency assurance for real-time clinical auxiliary diagnosis.

[0010] As a further aspect of the present invention, the real-time retrieval module performs a two-stage retrieval. The first stage utilizes the approximate nearest neighbor retrieval index to quickly retrieve a preset number of candidate cases from the case retrieval database. The second stage calculates the similarity between the query feature vector and the corresponding slice-level feature vector of each candidate case and sorts them by similarity to obtain a list of similar cases. In a preferred embodiment, after precise similarity sorting, the real-time retrieval module can further perform a secondary re-sorting of candidate cases based on preset clinical rules, and output at least one of the following in the list of similar cases: diagnostic label, similarity score, key pathological image region, and clinical metadata. Through the combination of approximate retrieval and precise re-sorting, the system ensures both response speed under a large-scale case database and reference quality of the returned similar cases.

[0011] As a further aspect of the present invention, the interpretability display module performs similarity analysis on the regional correspondence between the pathological slide to be diagnosed and similar cases, identifies key diagnostic similarities, and overlays these key similarities as heatmaps onto the pathological images of the pathological slide to be diagnosed and similar cases. Through this technical solution, doctors can not only obtain a list of similar cases but also see which regions the pathological slide to be diagnosed most closely matches those of similar cases, thereby understanding the pathological basis of the search results. This module elevates search-based assisted diagnosis from simply "returning similar cases" to "returning similar cases and explaining the reasons for the similarity," enhancing the system's transparency and clinical credibility.

[0012] As a further aspect of the present invention, the diagnostic reference generation module performs statistical analysis on the diagnostic label distribution of a preset number of similar cases ranked at the top of the similar case list, and performs weighted evaluation, diagnostic consistency analysis, and confidence assessment based on the similarity scores of each similar case to generate a structured diagnostic reference summary. Through this diagnostic reference summary, the system can transform discrete similar case results into quantified diagnostic support information, enabling doctors to obtain statistical information such as label distribution, support level, and consistency level while viewing individual case references, thereby improving the usability of the auxiliary diagnostic results.

[0013] As a further aspect of the present invention, when expanding the case retrieval database, the self-evolutionary optimization module adds new cases confirmed by doctors and their diagnostic tags to the database. When optimizing the retrieval strategy, it records the reference behavior of similar cases ultimately adopted by doctors and adjusts the retrieval parameters, ranking strategy, or case quality weights based on this reference behavior. Preferably, the self-evolutionary optimization module may also include a built-in evaluation submodule to periodically analyze the system's accuracy, false positive distribution, and error sources in cross-domain retrieval scenarios, and guide the updating of the case retrieval database and adjustment of the retrieval strategy based on the analysis results. Through this feedback-driven self-evolutionary mechanism, the system can continuously accumulate new cases, optimize ranking logic, and improve long-term performance with actual use.

[0014] This invention also provides a retrieval-based AI-assisted diagnosis method for pathological slides, applied to the aforementioned system, comprising the following steps: receiving and structuring pathological slide data and corresponding case information from different medical institutions; calling a pre-constructed cross-domain pathological image unified representation model to perform unified feature encoding on the input pathological slides, obtaining slide-level feature vectors, and storing the slide-level feature vectors in association with corresponding case information to construct a case retrieval library; constructing an approximate nearest neighbor retrieval index based on the slide-level feature vectors in the case retrieval library; calling the same cross-domain pathological image unified representation model as described above to perform unified feature encoding on the pathological slide to be diagnosed, obtaining a query feature vector; performing candidate case retrieval and precise reordering based on the query feature vectors to obtain a list of similar cases; performing similarity analysis on the regional correspondence between the pathological slide to be diagnosed and similar cases to generate a retrieval heatmap of key similar regions; generating a diagnostic reference summary based on the diagnostic tag distribution, similarity score, and case quality information in the list of similar cases; and updating the case retrieval library and optimizing the retrieval strategy based on doctor feedback. The method and the system work together to form a complete retrieval-based pathological auxiliary diagnosis scheme.

[0015] Compared with existing technologies, the beneficial effects of the retrieval-based AI diagnostic assistance system and method for pathological slides of the present invention are as follows: This invention replaces traditional classification output with case retrieval, matching the pathological slides to be diagnosed with historical cases in a unified feature space. This allows diagnostic references to come from real cases rather than preset category boundaries, thereby reducing dependence on large-scale labeled data. It can still provide stable references in small sample or new institution scenarios, solving the problem of insufficient applicability of existing methods under conditions of changing data distribution. It embodies the innovative systematic technical path of "retrieval-driven diagnosis".

[0016] This invention utilizes a unified cross-domain pathological image representation model to map pathological slides from different medical institutions to the same feature space. This enables feature comparability across devices, staining methods, and sources without relying on unified acquisition conditions. Furthermore, by combining a case retrieval database construction and indexing mechanism, cross-domain data can directly participate in the retrieval of similar cases, avoiding the mismatch problem of traditional natural image features in pathological scenarios and forming a unified retrieval foundation that can be transferred across institutions.

[0017] This invention, based on the output of a list of similar cases, introduces a key similarity region identification and retrieval heatmap generation mechanism to establish a correspondence between the retrieval results and specific pathological regions, enabling doctors to intuitively understand the source of similarity. It also generates a structured diagnostic reference summary by combining the distribution of diagnostic tags, realizing the transformation from "returning cases" to "interpretive basis + statistical support". This helps to improve the transparency and verifiability of the auxiliary diagnostic process and enhance the acceptability of the system in clinical applications.

[0018] This invention constructs a closed-loop update mechanism based on doctor feedback through a self-evolutionary optimization module. During system operation, it continuously introduces new case data and records diagnostic reference adoption behavior, adaptively adjusting retrieval parameters and sorting strategies. This allows the system performance to be gradually optimized as it is used, adapting to the data characteristics and usage habits of different hospitals. Compared with static models, it has stronger long-term adaptability and scalability, demonstrating continuous evolution characteristics for real-world application scenarios. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the structure of a pathological slide AI diagnostic assistance system based on retrieval, according to the present invention.

[0020] Figure 2 This is a flowchart illustrating an AI-assisted diagnostic method for pathological slides based on retrieval, as described in this invention. Detailed Implementation

[0021] The technical solutions of this embodiment will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0022] This invention provides a retrieval-based AI-assisted diagnostic system and method for pathological slides. The core idea is as follows: For pathological slide data from different medical institutions, structured data entry is first performed, followed by the generation of slide-level feature vectors within the same feature space using a pre-built cross-domain unified pathological image representation model. Based on this, a case retrieval library and an approximate nearest neighbor retrieval index are constructed. For the pathological slide to be diagnosed, rapid retrieval and precise reordering of candidate cases are performed, and a retrieval heatmap of key similar regions and diagnostic reference summaries are further generated. Simultaneously, the case retrieval library and retrieval strategy are continuously updated based on physician feedback, thereby achieving transparent and scalable assisted diagnosis for clinical applications.

[0023] In this invention, the following terms are used consistently: data entry module, unified feature encoding invocation module, case retrieval database construction module, retrieval index construction module, real-time retrieval module, interpretability display module, diagnostic reference generation module, and self-evolutionary optimization module. The module names are consistent throughout the specification, claims, and drawings. Example: Figure 1 and Figure 2 As shown, this embodiment uses cross-institutional pathological slides from four different diagnostic subtypes of the same cancer as an application scenario to fully illustrate a retrieval-based AI-assisted diagnostic system and method for pathological slides. This implementation scenario includes a historical case source domain and an actual target domain for diagnosis. The source domain uses publicly available pathological slide data to form a basic case database, while the target domain uses newly generated pathological slides from a collaborating hospital. The source and target domains differ in scanning equipment, staining conditions, and data acquisition processes. Directly using pre-trained features from natural images for case retrieval can easily lead to problems such as incomparable features, mismatches in similar cases, and insufficient reliability of diagnostic references. This embodiment focuses on four aspects: cross-institutional feature unification, efficient two-stage retrieval, visualization of key similar regions, and feedback-driven optimization. This enables the slides to obtain verifiable, interpretable, and continuously optimizeable assisted diagnostic results without relying on a large number of labeled samples from the target domain.

[0024] In this embodiment, the system first receives historical pathological slide data from the source domain and performs structured processing through the data entry module. The source domain data selects 1000 pathological slides from four different subtypes of the same cancer as basic case data, while the target domain selects 300 pathological slides from the same four different subtypes of the same cancer as search data. During data entry, structured search data files are formed according to case granularity. Each record includes at least a unique case identifier, slide identifier, diagnostic label, hospital source information, and the corresponding feature file path. When necessary, scanning equipment information, staining batch information, and necessary clinical metadata are also recorded simultaneously. For source domain data, the data entry module performs sampling and organization according to disease subtype distribution using a preset script, generating data records in a unified format. For new slides continuously accessed by collaborating hospitals, the data entry module receives the original slide images and their confirmed diagnostic information through a standard interface and writes them into the processing pipeline according to the same fields. In this way, the system establishes a multi-source case management foundation covering different medical institutions, different equipment, and different staining conditions, providing stable input for subsequent unified coding and cross-institutional retrieval.

[0025] After the data is successfully imported, the unified feature encoding module begins to perform unified feature encoding on the imported pathological slides. This module calls a pre-built cross-domain pathological image unified representation model to output a fixed-dimensional slide-level feature vector for each pathological slide. In this embodiment, the unified feature encoding result is set to a 128-dimensional slide-level feature vector. Both source domain and target domain pathological slides use completely consistent encoding rules, ensuring that pathological slides from different sources are mapped to the same feature space. After the unified feature encoding module completes the encoding, it not only outputs the feature vector but also establishes an association between it and the corresponding case's diagnostic label, hospital source information, and patient-related metadata. This association is then passed to the case retrieval database construction module for unified writing into the case retrieval database. The result of this process is that the case retrieval database stores an integrated case knowledge unit consisting of "slide-level feature vector—case label—source information—auxiliary metadata." Since both the source and target domains are encoded using the same unified representation model, newly imported pathological slides to be diagnosed by collaborating hospitals can be directly compared with historical cases in the same metric space, thus making cross-institutional case retrieval feasible.

[0026] After the case retrieval database is constructed, the retrieval index construction module synchronously constructs an approximate nearest neighbor retrieval index based on the 128-dimensional slice-level feature vectors in the case retrieval database. In this embodiment, the approximate nearest neighbor retrieval index used is the HNSW index. The HNSW index organizes high-dimensional features into a hierarchical graph structure, allowing the system to quickly locate neighboring regions in the feature space and return a limited number of candidate cases before performing full cosine similarity calculations on all cases in the case retrieval database during online retrieval, instead of performing full cosine similarity calculations on all cases in the database. As the case retrieval database expands, the role of the HNSW index becomes particularly evident; even when the case retrieval database is continuously supplemented to tens of thousands of case records, the system can still maintain a short online response time, meeting the real-time requirements for clinical auxiliary diagnosis. In this way, the case retrieval database and the index structure are coupled during the construction phase: the case retrieval database ensures the integrity of case information, and the approximate nearest neighbor index ensures online retrieval efficiency. Together, they form the foundational support for cross-institutional retrieval-based auxiliary diagnosis in this embodiment.

[0027] After a doctor submits a new pathology slide for diagnosis, the system enters the real-time retrieval phase. The pathology slide is first processed by the unified feature encoding module, which executes a unified encoding process identical to that of the slides in the database, generating a 128-dimensional query feature vector. Since this query feature vector resides in the same feature space as the slide-level feature vectors in the case retrieval database, it can be directly used for cross-institutional case retrieval. The real-time retrieval module processes this query feature vector using a two-stage retrieval method: the first stage uses the HNSW approximate nearest neighbor index to quickly retrieve K=100 candidate cases from the case retrieval database. In this embodiment, the number of candidate cases K is set to 100; this step trades a small amount of precision for millisecond-level response speed. The second stage calculates the precise cosine similarity between the query feature vector and the corresponding slide-level feature vector for each of these 100 candidate cases, and sorts them from high to low similarity scores to form an initial list of similar cases. Subsequently, the system further performs a secondary re-sorting based on preset clinical rules, such as prioritizing cases with clear diagnostic labels, complete clinical metadata, and high image quality, while downgrading cases with missing labels or incomplete information, ultimately outputting a list of similar cases for doctors' reference. In the output, each similar case is accompanied by at least a diagnostic label and a similarity score. When necessary, key pathological image regions and clinical metadata are also included, so that doctors can not only see "what kind of case it is", but also "what diagnosis the case is, how high the degree of matching is, and whether it has a complete reference background".

[0028] After obtaining the sorted list of similar cases, the system further performs an interpretability display of the search results. The interpretability display module does not simply stop at the "return case list" level; instead, it performs similarity analysis on the regional correspondence between the pathological slide to be diagnosed and the top-ranked similar cases, identifies key similar regions for diagnosis, and overlays these key similar regions as a heatmap onto the pathological images of the pathological slide to be diagnosed and the corresponding similar cases. After viewing the heatmap, doctors can intuitively see where the current pathological slide to be diagnosed has the strongest matching relationship with similar cases, such as areas of lesion feature aggregation, morphological abnormalities, or key tissue structures, thereby judging the reliability of the system's return of the case. Unlike traditional black-box classification models that directly output a diagnostic conclusion, this embodiment presents the most crucial pathological matching evidence in a visual manner, enabling doctors to independently verify the auxiliary diagnostic evidence based on the heatmap, reducing uncertainty about the system's output and improving clinical acceptance.

[0029] While providing interpretable information, the diagnostic reference generation module performs statistical analysis on a preset number of similar cases ranked at the top of the similar case list and generates a diagnostic reference summary. In this embodiment, the preset number is the top 10 similar cases. The diagnostic reference generation module first statistically analyzes the distribution of diagnostic tags in these 10 similar cases, then performs a weighted evaluation based on the similarity score of each case, and further performs diagnostic consistency analysis and confidence assessment, finally outputting a structured diagnostic reference summary. Taking a pathological slide to be diagnosed as an example, the system can generate statistical results such as "Among the 10 most similar cases returned, 8 diagnostic tags are for one subtype, and 2 diagnostic tags are for another subtype," and provide corresponding weighted support levels and confidence level hints. In this way, while viewing individual similar cases, doctors can also understand the overall concentration and diagnostic tendency of the current search results from a group statistical perspective. This process transforms discrete similar case information into quantitative diagnostic support information, further enhancing the system output from "individual case reference" to a dual auxiliary basis of "individual case + statistics," thereby increasing the practical value of the retrieval system in real clinical decision-making.

[0030] In this embodiment, the system forms a closed loop of "use-feedback-optimization" through a self-evolutionary optimization module. After each diagnostic task, new cases confirmed by doctors and their diagnostic tags can be added to the case retrieval database to enhance the database's coverage of actual clinical data. Simultaneously, the system records which one or more similar cases the doctor ultimately adopts as diagnostic references, analyzes the relationship between these adoptions and the current ranking strategy, and adjusts the retrieval parameters, ranking strategy, and case quality weights accordingly. By continuously accumulating feedback information, the system can gradually adapt to the usage habits and data distribution of different hospitals, improving the accuracy and stability of subsequent searches. The self-evolutionary optimization module can also include an evaluation submodule to perform periodic analysis of cross-domain retrieval accuracy, false positive distribution, and error types to determine whether the current retrieval parameters are reasonable, whether the case database additions are effective, and whether the system needs further adjustment. Since the optimization focuses on the case retrieval database, retrieval parameters, and ranking strategy, this embodiment avoids limiting the internal training process of the underlying unified representation model, thus maintaining clear system boundaries while still demonstrating the important invention point of the system's long-term iterative capability.

[0031] To further verify the technical effectiveness of this embodiment, a cross-domain pathological slide retrieval task was used for comparative testing. In the test, 1000 pathological slides of four different subtypes of the same cancer were used as the basic data for the case retrieval database in the source domain, and 300 pathological slides of four different subtypes of the same cancer from a medical institution were used as the retrieval data in the target domain. Target domain diagnostic labels were not used during training and retrieval; the diagnostic labels were only used for final performance evaluation. Retrieval was performed on all slides using both the "conch feature + average pooling" scheme and the scheme of this embodiment, and the retrieval accuracy for the four subtypes is shown in Table 1.

[0032] Table 1 Comparison of Cross-Domain Search Accuracy

[0033] As shown in Table 1, the retrieval accuracy of this embodiment for the four subtypes was 81.9%, 83.1%, 91.2%, and 79.8%, respectively, all significantly higher than the 48.6%, 65.2%, 78.3%, and 67.8% of the comparative scheme. These results indicate that, under conditions of both cross-institutional staining differences and equipment differences, simply using natural image pre-trained features plus average pooling is insufficient to guarantee the stable comparability of pathological slides from different sources within the same metric space. This embodiment, by invoking a unified cross-domain pathological image representation model and constructing a case retrieval database, an approximate nearest neighbor index, a two-stage retrieval system, heatmap interpretation, diagnostic reference summaries, and a feedback optimization closed loop, can significantly improve the accuracy of cross-domain pathological slide retrieval. Furthermore, looking at the results of each subtype, the improvement in subtypes 1 and 2 is particularly significant, indicating that this embodiment is more adaptable to scenarios where the source and target domains differ greatly and traditional methods are prone to feature drift. Subtypes 3 and 4 already have certain retrieval capabilities in comparison schemes, but this embodiment still achieves further improvements, indicating that this embodiment not only solves the cross-institutional comparability problem, but also transforms the simple retrieval process into a more complete clinical auxiliary diagnostic link through heatmap interpretation and diagnostic reference summaries.

[0034] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0035] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A retrieval-based AI-assisted diagnostic system for pathological slides, characterized in that, The system includes a data entry module, a unified feature encoding module, a case retrieval library construction module, a retrieval index construction module, a real-time retrieval module, an interpretability display module, a diagnostic reference generation module, and a self-evolutionary optimization module. The data entry module receives and structures pathological slide data and corresponding case information from different medical institutions. The unified feature encoding module invokes a pre-built cross-domain pathological image unified representation model to perform unified feature encoding on the entered pathological slides and the pathological slides to be diagnosed, obtaining slide-level feature vectors. The case retrieval library construction module associates and stores the slide-level feature vectors with corresponding case information and constructs a case retrieval library. The retrieval index construction module constructs an approximate nearest neighbor retrieval index based on the slide-level feature vectors. The real-time retrieval module performs candidate case retrieval and precise re-ranking based on query feature vectors to obtain a list of similar cases. The interpretability display module generates a retrieval heatmap of key similar regions. The diagnostic reference generation module generates a diagnostic reference summary. The self-evolutionary optimization module updates the case retrieval library and optimizes the retrieval strategy based on doctor feedback.

2. The AI-assisted diagnostic system for pathological slides based on retrieval according to claim 1, characterized in that, The unified feature encoding invocation module applies the same encoding rules to both the imported pathological slides and the pathological slides to be diagnosed, and outputs a fixed-dimensional slide-level feature vector, so that pathological slides from different sources are mapped to the same feature space.

3. The AI-assisted diagnostic system for pathological slides based on retrieval according to claim 1, characterized in that, The real-time retrieval module performs a two-stage retrieval. In the first stage, the approximate nearest neighbor retrieval index is used to retrieve a preset number of candidate cases from the case retrieval database. In the second stage, the similarity between the query feature vector and the corresponding slice-level feature vector of each candidate case is calculated and sorted according to similarity to obtain a list of similar cases.

4. The AI-assisted diagnostic system for pathological slides based on retrieval according to claim 3, characterized in that, After sorting by similarity, the real-time retrieval module performs a secondary re-sorting of candidate cases in conjunction with preset clinical rules, and outputs at least one of the following in the list of similar cases: diagnostic label, similarity score, key pathological image region, and clinical metadata.

5. The AI-assisted diagnostic system for pathological slides based on retrieval according to claim 1, characterized in that, The interpretability display module performs a similarity analysis on the regional correspondence between the pathological slide to be diagnosed and similar cases, identifies key similar regions for diagnosis, and overlays and renders the key similar regions for diagnosis as a heatmap on the pathological images of the pathological slide to be diagnosed and similar cases.

6. The AI-assisted diagnostic system for pathological slides based on retrieval according to claim 1, characterized in that, The diagnostic reference generation module performs statistical analysis on the distribution of diagnostic labels of a preset number of similar cases ranked at the top of the list of similar cases, and performs weighted evaluation, diagnostic consistency analysis and confidence evaluation in combination with the similarity scores of each similar case to generate a structured diagnostic reference summary.

7. The AI-assisted diagnostic system for pathological slides based on retrieval according to claim 1, characterized in that, When expanding the case retrieval database, the self-evolutionary optimization module adds new cases diagnosed by doctors and their diagnostic tags to the database. When optimizing the retrieval strategy, it records the reference behaviors of similar cases finally adopted by doctors and adjusts the retrieval parameters, ranking strategies, or case quality weights based on these reference behaviors.

8. The AI-assisted diagnostic system for pathological slides based on retrieval according to claim 1, characterized in that, The data entry module is connected to the case retrieval database construction module. The unified feature encoding invocation module is connected to both the case retrieval database construction module and the real-time retrieval module. The retrieval index construction module is connected to both the case retrieval database construction module and the real-time retrieval module. The interpretability display module and the diagnostic reference generation module are connected to the real-time retrieval module. The self-evolutionary optimization module is connected to the case retrieval database construction module, the retrieval index construction module, and the real-time retrieval module.

9. A retrieval-based AI-assisted diagnostic method for pathological slides, employing a retrieval-based AI-assisted diagnostic system for pathological slides as described in any one of claims 1-8, characterized in that, Includes the following steps: Step 1: Receive and structure pathological slide data and corresponding case information from different medical institutions; Step 2: Call the pre-built cross-domain pathological image unified representation model to perform unified feature encoding on the entered pathological slides, obtain slide-level feature vectors, and associate and store the slide-level feature vectors with the corresponding case information to build a case retrieval library; Step 3: Construct an approximate nearest neighbor retrieval index based on the slice-level feature vectors in the case retrieval database; Step 4: Use the same cross-domain pathological image unified representation model as in Step 2 to perform unified feature encoding on the pathological slides to be diagnosed, and obtain the query feature vector; Step 5: Based on the query feature vector, perform candidate case retrieval and precise reordering to obtain a list of similar cases; Step 6: Perform a similarity analysis on the regional correspondence between the pathological slides to be diagnosed and similar cases to generate a retrieval heatmap of key similar regions; Step 7: Generate a diagnostic reference summary based on the distribution of diagnostic tags, similarity scores, and case quality information in the list of similar cases; Step 8: Update the case search database and optimize the search strategy based on doctor feedback.