A clinically practical AI-assisted diagnostic system based on the sampling site of a patient case.
By using an AI-assisted diagnostic system based on the sampling sites of cases, and leveraging a self-attention mechanism for feature interaction and fusion reasoning, the system addresses the issue of the lack of consideration of the hierarchical structure of sampling sites in existing pathological AI analysis. This enables efficient and accurate simulation of the pathological diagnostic process and generation of diagnostic results.
Patent Information
- Application Number
- CN202510593105.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Existing AI-based pathology analysis methods fail to adequately consider the hierarchical structure of the sampling sites, resulting in high computational complexity, redundant reasoning, and strong dependence on annotations. They cannot accurately simulate the diagnostic process of pathologists and also have shortcomings in terms of efficiency and scalability.
The AI-assisted diagnostic system, which takes the sampling site of a case as the unit, obtains the case diagnosis information, determines the label information of the sampling site, performs image preprocessing and feature vector transformation, and uses a self-attention mechanism to perform feature interaction and fusion reasoning to generate AI-assisted diagnostic results.
It achieves efficient and accurate simulation of the pathologist's diagnostic process, reduces computational complexity and inference redundancy, reduces labeling dependence, and improves the ability and scalability of AI-assisted diagnosis.
Smart Images

Figure CN120473122B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of auxiliary diagnostic technology, and in particular to a clinically practical AI-assisted diagnostic system based on the sampling site of a medical case. Background Technology
[0002] Histological pathology diagnosis is a method of determining disease by observing the structure and cell morphology of patient tissue sections under a microscope, and it is the core of pathological diagnosis. It primarily relies on histological characteristics, such as cell morphology, arrangement, nuclear-cytoplasmic ratio, stroma changes, and vascular distribution, to classify and diagnose diseases. Histological pathology diagnosis is widely used in differentiating between benign and malignant tumors, classifying lesion types, and analyzing inflammatory and degenerative diseases.
[0003] Traditional histological diagnosis relies on staining techniques, such as hematoxylin and eosin (H&E) staining and special staining, supplemented by molecular marker detection techniques such as immunohistochemistry (IHC) and in situ hybridization (ISH), further improving diagnostic accuracy. With the rise of digital pathology and artificial intelligence technologies, histological pathology diagnosis is developing towards high-throughput, standardization, and intelligence, providing stronger support for clinical diagnosis and precision medicine. Computer vision, as an important branch of artificial intelligence, is dedicated to processing and understanding images by simulating the human visual system. The human visual recognition system perceives image details and, combined with experience, cognition, and contextual information, quickly identifies and classifies image content. Modern computer vision technology, through deep learning algorithms, especially convolutional neural networks (CNNs), mimics human recognition patterns, enabling it to learn and extract features from large amounts of labeled data, thereby achieving automatic recognition and analysis of image content. In pathological image analysis, computer vision technology is widely used in cell identification, tumor detection, and pathological feature extraction, improving the accuracy and efficiency of pathological image analysis. By combining computer vision with histological pathology diagnosis, automated diagnosis and real-time analysis of pathological images can be achieved, reducing the workload of pathologists while improving the speed and accuracy of diagnosis. The application of this technology is of significant practical importance, especially in high-throughput screening and remote consultations.
[0004] In the clinical histopathological diagnostic process, each case typically consists of one or more specimens, each containing one or more sampling sites. For example, a gastric specimen may include multiple sampling sites such as the antrum, angle, and lesser curvature of the body. Each sampling site will have one or more slides analyzed. Pathologists will comprehensively consider all slides from a given sampling site to form a comprehensive pathological diagnosis. Ultimately, based on the diagnostic results from each sampling site, the expert will produce a pathology report for the entire case.
[0005] Currently, computational pathology technology also attempts to mimic the recognition patterns of pathologists for image analysis and classification diagnosis. However, most existing mainstream technologies analyze single slides as units, ultimately merging the diagnostic results of all slides into a single pathology report. This method usually ignores the clinically critical level of the sampling site and fails to fully reflect the actual logic of pathological diagnosis. Unlike the analysis process of pathologists, this method directly merges the diagnostic results of multiple slides, failing to perform detailed classification and diagnosis according to the hierarchical structure of the sampling site. Therefore, this approach fails to strictly follow the clinical pathological diagnosis process and fails to fully utilize the important information level of the sampling site. In existing computational pathology technologies, pathological image analysis is usually performed by AI analysis on each slide as a unit, and the diagnostic results of multiple slides are finally superimposed to generate a pathology report. This approach has the following problems: (1) It does not conform to the logic of clinical diagnosis: Existing technologies mainly consider the spatial structure information and contextual relationships within a single slide, but do not fully consider the contextual relationships between slides. In clinical practice, when a pathologist diagnoses a sampling site, they usually need to comprehensively consider the relationships between multiple slides to arrive at the final diagnostic result. Therefore, the existing slide-based identification model fails to accurately simulate the pathologist's analysis process and ignores the hierarchical structure of the sampling site. (2) High computational complexity: When performing AI analysis on a slide-by-slide basis, a sampling site requires multiple calls to the AI model for reasoning, increasing computational complexity. However, in actual diagnosis, pathologists only need to put the slides together for comprehensive analysis to arrive at a diagnosis. Existing methods greatly reduce the potential of AI technology in improving diagnostic efficiency and may lead to inefficient collaboration between doctors and AI. (3) Redundant reasoning: In clinical pathology diagnosis, if a pathologist has found a malignant tumor on a slide, they usually make a diagnosis for the entire sampling site immediately without needing to confirm whether other slides have the same lesion. This indicates that the existing slide-level analysis has redundant reasoning, increases the invalid reasoning computation of the AI system, and wastes computational resources. (4) Strong labeling dependence: Current AI analysis methods require labeling of each slide, and even when using weakly supervised learning methods, a large amount of labeling work is required. This approach leads to pathology AI systems becoming overly reliant on labeled data, hindering their rapid expansion to different organs, lesion types, and large-scale datasets, thus limiting their universality and scalability in clinical practice. In summary, existing pathology AI analysis methods have significant shortcomings, failing to adequately simulate the workflow of pathologists and exhibiting numerous problems in terms of efficiency, inference redundancy, and label dependency. Summary of the Invention
[0006] This invention aims to at least partially solve one of the technical problems in the aforementioned technologies. Therefore, the objective of this invention is to propose a clinically practical AI-assisted diagnostic system based on the sampling site of a medical case, capable of accurately and efficiently simulating the diagnostic process of a pathologist and improving the capabilities of AI-assisted diagnosis.
[0007] To achieve the above objectives, embodiments of the present invention propose a clinically practical AI-assisted diagnostic system based on the sampling site of a medical case, comprising:
[0008] The acquisition module is used to acquire case diagnosis information;
[0009] The first determination module is used to determine the label information for each sampling site in each case based on the case diagnosis information;
[0010] The second determining module is used to determine the target glass slide corresponding to the target sampling location based on the label information;
[0011] The preprocessing module is used to preprocess the target glass slide to obtain the target glass slide in the form of feature vectors;
[0012] The fusion reasoning module is used to perform image fusion reasoning on the target glass slide, which has been converted into feature vector form, to obtain AI-assisted diagnostic results.
[0013] According to some embodiments of the present invention, the acquisition module acquires case diagnosis information from an information system, the case diagnosis information including the pathology number of the slide and the sampling site.
[0014] According to some embodiments of the present invention, the first determining module includes:
[0015] The classification module is used to classify each slide into its respective location based on the pathology number and the sampling site.
[0016] The first extraction module is used to extract the pathology report corresponding to the pathology number of each slide;
[0017] The splitting module is used to split a whole pathology report into pathological diagnoses based on the sampling site, and assign the diagnostic results to the corresponding sampling sites, based on the large language model.
[0018] The second extraction module is used to extract the labels of the required diseases from the diagnostic results of each sampling site using a large language model, and to determine the label information of each sampling site in each case.
[0019] According to some embodiments of the present invention, the fusion inference module includes:
[0020] The local region feature interaction module within the slide is used to perform contextual association of the features of the local region of the target slide, which has been converted into feature vector form, through a self-attention mechanism algorithm to obtain the first diagnostic result;
[0021] The slide-level feature interaction module is used to first shuffle the feature vectors of the target slide to ensure that the feature vectors placed together belong to different regions of the slide. Then, the self-attention mechanism algorithm is used again to connect the features of different regions of the entire slide in context to obtain the second diagnostic result.
[0022] The inter-slide feature interaction module is used to aggregate the feature vectors of the target slide into a whole feature vector using attention_pooling. The whole feature vector represents the feature vector of the entire slide. The feature vectors of all slides are put together and processed based on the self-attention mechanism algorithm to connect the features of different slides in the entire sampling site in context to obtain the third diagnostic result.
[0023] The AI-assisted diagnostic results are obtained based on the first, second, and third diagnostic results.
[0024] According to some embodiments of the present invention, the preprocessing module includes:
[0025] The generation module is used to obtain the pixel values of each pixel in the target slide and generate an associated region matrix for each pixel.
[0026] The third determining module is used to determine the segmentation index of each pixel based on the associated region matrix of each pixel.
[0027] The segmentation module is used to select pixels with a segmentation index greater than a preset segmentation index threshold as segmentation pixels, and to divide the target glass slide into blocks based on the segmentation pixels to obtain several sub-target glass slides.
[0028] The optimization module is used to extract the feature parameters of each sub-target slide and optimize them to obtain the target feature parameters;
[0029] The encoding module is used to encode the target feature parameters to obtain the feature vector of each sub-target slide, and then obtain the target slide in the form of feature vector.
[0030] According to some embodiments of the present invention, the generation module generates an associated region matrix for each pixel, including:
[0031] Generate the associated region matrix of the pixel in the i-th row and j-th column of the target glass slide;
[0032]
[0033] Among them, S i,jS is the pixel value of the pixel in the i-th row and j-th column; i-1,j-1 S is the pixel value of the pixel in the (i-1)th row and (j-1)th column; i-1,j S is the pixel value of the pixel in the (i-1)th row and jth column; i-1,j+1 S is the pixel value of the pixel in the (i-1)th row and (j+1)th column; i,j-1 S is the pixel value of the pixel in the i-th row and j-1-th column; i,j+1 S is the pixel value of the pixel in the i-th row and j+1-th column; i+1,j-1 S is the pixel value of the pixel in the (i+1)th row and (j-1)th column; i+1,j S is the pixel value of the pixel in the (i+1)th row and jth column; i+1,j+1 It is the pixel value of the pixel in the (i+1)th row and (j+1)th column.
[0034] According to some embodiments of the present invention, the third determining module determines the segmentation index of each pixel, including:
[0035] Determine the segmentation index of the pixel in the i-th row and j-th column of the target glass slide;
[0036]
[0037] Among them, A ij is the segmentation index of the pixel in the i-th row and j-th column of the target glass slide; D is the number of pixels included in the target glass slide.
[0038] According to some embodiments of the present invention, the optimization module includes:
[0039] The feature extractor is used to extract the feature parameters of each sub-target slide;
[0040]
[0041] Among them, F j (I i ) represents the feature extractor; P*Q represents the size parameters of the sub-target slide; I i (p,q) represents the sub-target slide I. i The pixel value at position (p, q); e is the natural constant; τ is the standard deviation controlling the Gaussian blur;
[0042] The feature optimizer is used to optimize the feature parameters to obtain the target feature parameters;
[0043]
[0044] Where R represents the target optimization feature parameters; N represents the number of feature parameters; λ represents the optimization coefficients; tanh represents the hyperbolic tangent function; ω j is the weight of the j-th feature extractor; M is the number of feature extractors.
[0045] According to some embodiments of the present invention, the local region feature interaction module within the slide uses a self-attention mechanism algorithm to contextually associate the features of the local region of the target slide, which have been converted into feature vector form, including:
[0046] The fourth determining module is used for:
[0047] Determine the grayscale value of each pixel in a local region of the target glass slide, and take the pixels with grayscale values greater than a preset threshold in the local region as local feature nodes;
[0048] Select any local feature node as the local feature node to be associated; calculate the distance between the local feature node to be associated and other local feature nodes, and take the minimum distance as the association distance; generate an associated region with the local feature node to be associated as the center and the association distance as the radius;
[0049] The fifth determination module is used to determine the associated region of each local feature node; and to determine the location information of the associated region of each local feature node.
[0050] The transformation module is used to form a vector composed of all high-frequency coefficients obtained by performing DCT transformation on the associated region of each local feature node in descending order, which serves as the local edge information vector of each local feature node.
[0051] The association module is used to perform contextual connections based on the location information of the associated region and the local edge information vector of each local feature node to obtain the association relationship;
[0052] The verification module is used to calculate the effective coefficient of the association relationship and compare it with the preset effective coefficient threshold; when the effective coefficient is determined to be greater than the preset effective coefficient threshold, the first diagnostic result is obtained.
[0053] According to some embodiments of the present invention, the verification module includes:
[0054] The sixth determining module is used to determine the associated dataset and several valid indicators corresponding to the association relationship;
[0055] The calculation module is used to calculate the sub-efficiency coefficients of the effective indicators in the associated dataset;
[0056]
[0057] Where K is the sub-efficiency coefficient of the effective index in the associated dataset; b a For each effective metric, F is the membership function of the effective metric on the associated dataset A; U is the universe of discourse of the effective metric; n is the number of all metrics on the associated dataset A; b x Let T(b) be the x-th metric in the associated dataset A;a b x ) represents the matching degree between the effective indicator and the x-th indicator in the associated dataset A;
[0058] Calculate the sub-efficiency coefficients of all effective indicators in the associated dataset and sum them. Use the ratio of the sum to the number of effective indicators as the effective coefficient of the association relationship.
[0059] This invention proposes a clinically practical AI-assisted diagnostic system based on the sampling site of a case, which can accurately and efficiently simulate the diagnostic process of a pathologist and improve the ability of AI-assisted diagnosis.
[0060] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0061] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0062] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0063] Figure 1 This is a block diagram of a clinically practical AI-assisted diagnostic system based on the sampling site of a case according to an embodiment of the present invention;
[0064] Figure 2 This is a block diagram of a first determining module according to an embodiment of the present invention;
[0065] Figure 3 This is a block diagram of a preprocessing module according to an embodiment of the present invention. Detailed Implementation
[0066] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0067] like Figure 1 As shown, this embodiment of the invention proposes a clinically practical AI-assisted diagnostic system based on the sampling site of a medical case, comprising:
[0068] The acquisition module is used to acquire case diagnosis information;
[0069] The first determination module is used to determine the label information for each sampling site in each case based on the case diagnosis information;
[0070] The second determining module is used to determine the target glass slide corresponding to the target sampling location based on the label information;
[0071] The preprocessing module is used to preprocess the target glass slide to obtain the target glass slide in the form of feature vectors;
[0072] The fusion reasoning module is used to perform image fusion reasoning on the target glass slide, which has been converted into feature vector form, to obtain AI-assisted diagnostic results.
[0073] The working principle of the above technical solution is as follows: Using natural language understanding, a pathology report of a case is broken down into pathological diagnoses for multiple sites. Disease fields are extracted from these diagnoses and converted into standard disease names, thus enabling the extraction of labels from the pathology report. Furthermore, by utilizing hierarchical multi-instance learning, pathological image data from a single sampling site are hierarchically fused from local-level features to slide-level features to site-level features, achieving feature fusion across multiple slides.
[0074] The beneficial effects of the above technical solution are as follows: Using the sampling site as the diagnostic unit, a single diagnostic result is derived by combining all slides from that site. The final pathology report lists and combines slides from multiple sampling sites, which better aligns with the diagnostic logic of pathologists. Since the analysis focuses on all slides from a single sampling site, only one inference is triggered, significantly reducing computational complexity and redundancy, and greatly improving analysis efficiency. By placing the analysis unit at the sampling site, it is not necessary for pathologists to label each slide; sufficiently accurate labels can be obtained from the pathology report issued by the doctor. This significantly reduces labeling time, allowing the solution to be quickly scaled to multiple diseases and large datasets. It can accurately and efficiently simulate the diagnostic process of pathologists, improving the capabilities of AI-assisted diagnosis.
[0075] According to some embodiments of the present invention, the acquisition module acquires case diagnosis information from an information system, the case diagnosis information including the pathology number of the slide and the sampling site.
[0076] like Figure 2 As shown, according to some embodiments of the present invention, the first determining module includes:
[0077] The classification module is used to classify each slide into its respective location based on the pathology number and the sampling site.
[0078] The first extraction module is used to extract the pathology report corresponding to the pathology number of each slide;
[0079] The splitting module is used to split a whole pathology report into pathological diagnoses based on the sampling site, and assign the diagnostic results to the corresponding sampling sites, based on the large language model.
[0080] The second extraction module is used to extract the labels of the required diseases from the diagnostic results of each sampling site using a large language model, and to determine the label information of each sampling site in each case.
[0081] The working principle of the above technical solution is as follows: Based on the pathology number and sampling site information of the slides, each slide is assigned to the corresponding case and its specific sampling site, forming a hierarchical structure based on the sampling site. The slide information list is read, and the pathology number and sampling site are parsed. The pathology number is associated with the specific case ID. Slides are grouped according to case ID and sampling site, constructing a hierarchical structure. A tree structure is determined with case ID as the root node and sampling site as the child nodes, with each node associated with a corresponding list of slides. The first extraction module traverses the slide list to obtain the pathology number of each slide. The pathology report database is queried based on the pathology number to obtain the corresponding report content. A mapping relationship between slide IDs and report content is constructed. The splitting module uses a Large Language Model (LLM) to split the complete pathology report into diagnostic segments based on the sampling site and assigns them to the corresponding sampling site nodes. For each pathology report, LLM is called for semantic analysis. Diagnostic content related to the sampling site in the report is identified, and key information (such as lesion type, description, etc.) is extracted. Based on the sampling site of the slide, diagnostic fragments are assigned to the corresponding sampling site nodes. A mapping relationship between sampling sites and diagnostic fragments is established. The second extraction module further extracts disease-related label information for each sampling site's diagnostic fragment using LLM, forming a structured label system. It iterates through the diagnostic fragments of each sampling site, calling LLM for label extraction. Based on the predefined disease label system, key information in the diagnostic fragments (such as "adenocarcinoma," "Grade I," etc.) is identified. The extracted label information is associated with the case ID and sampling site to form structured data. A mapping relationship between case ID, sampling site, and label information is established. The classification module generates a case-sampling site hierarchical structure. The first extraction module obtains the pathology report for each slide. The splitting module splits the pathology report into diagnostic fragments corresponding to the sampling sites. The second extraction module extracts disease label information from the diagnostic fragments. Finally, a structured mapping table of case ID, sampling site, and label information is output for use by subsequent modules.
[0082] The beneficial effects of the above technical solution are as follows: Labeling information for each sampling site in each case is determined based on the case diagnosis information. Such labels not only eliminate the need for manual annotation but also ensure accuracy, as they are extracted from authoritative pathology reports issued by doctors. The first determination module, through steps such as classification, extraction, splitting, and label extraction, realizes the transformation from slide information to structured label information, providing a high-quality data foundation for the AI-assisted diagnostic system. The design of this module fully considers the complexity and diversity of the medical field, ensuring the accuracy and interpretability of diagnostic information.
[0083] According to some embodiments of the present invention, the fusion inference module includes:
[0084] The local region feature interaction module within the slide is used to perform contextual association of the features of the local region of the target slide, which has been converted into feature vector form, through a self-attention mechanism algorithm to obtain the first diagnostic result;
[0085] The slide-level feature interaction module is used to first shuffle the feature vectors of the target slide to ensure that the feature vectors placed together belong to different regions of the slide. Then, the self-attention mechanism algorithm is used again to connect the features of different regions of the entire slide in context to obtain the second diagnostic result.
[0086] The inter-slide feature interaction module is used to aggregate the feature vectors of the target slide into a whole feature vector using attention_pooling. The whole feature vector represents the feature vector of the entire slide. The feature vectors of all slides are put together and processed based on the self-attention mechanism algorithm to connect the features of different slides in the entire sampling site in context to obtain the third diagnostic result.
[0087] The AI-assisted diagnostic results are obtained based on the first, second, and third diagnostic results.
[0088] The working principle of the above technical solution is as follows: The local region feature interaction module within a slide, targeting a single slide, captures the contextual relationships between its local region features through a self-attention mechanism to generate a first diagnostic result. The slide's feature vector is divided into multiple local region feature blocks, and the attention weight of each local region feature block with other feature blocks is calculated. The weighted summation generates an enhanced local feature representation. The enhanced local features are aggregated to form a preliminary slide-level diagnostic result, representing the feature correlation diagnosis of local regions within the slide. The slide-level feature interaction module, by shuffling the regional order of the slide feature vectors to eliminate prior spatial dependencies, utilizes a self-attention mechanism to capture the global contextual relationships between different regions within the slide, generating a second diagnostic result. The regional order of the slide feature vectors is randomly shuffled to ensure that the feature vectors belong to different regions of the slide. The attention weight of each region in the shuffled feature vectors with other regions is calculated. The weighted summation generates a globally enhanced feature representation. The globally enhanced features are aggregated to form a slide-level global diagnostic result, representing the global feature correlation diagnosis of different regions within the slide. The slide-to-slide feature interaction module aggregates the feature vectors of individual slides to generate a slide-level overall feature representation. Then, it uses a self-attention mechanism to capture the contextual relationships between different slides, generating a third diagnostic result. Attention pooling is applied to the feature vector of each slide to generate a slide-level overall feature vector. The overall feature vectors of all slides are used as input to calculate the attention weights between slides. Weighted summation based on these weights generates a global contextual feature representation between slides. The global features between slides are aggregated to form a diagnostic result at the sampling site level, representing the feature correlation diagnosis between different slides. Based on the first, second, and third diagnostic results, an AI-assisted diagnostic result is obtained, including: Result alignment: ensuring consistency of the three diagnostic results across diagnostic dimensions (e.g., lesion type, grade, etc.). Weighted fusion: weighted summation of the three diagnostic results based on preset weights or adaptive learning methods. Weights can be dynamically adjusted based on performance evaluation during model training. Conflict resolution: arbitrating conflicts in the diagnostic results (e.g., inconsistent judgments of the same lesion by different results). This can be resolved using voting mechanisms, confidence assessments, or expert rules.
[0089] The beneficial effects of the above technical solution are as follows: By integrating local features within the slide, global features at the slide level, and features between slides through self-attention and multi-scale feature interaction, a comprehensive AI-assisted diagnostic result is generated. The fusion inference module achieves comprehensive diagnostic analysis from local to global and from within the slide to between slides through multi-scale feature interaction and result fusion. It fully considers the complexity and diversity of pathological diagnosis, ensuring the accuracy and interpretability of the diagnostic results.
[0090] like Figure 3 As shown, according to some embodiments of the present invention, the preprocessing module includes:
[0091] The generation module is used to obtain the pixel values of each pixel in the target slide and generate an associated region matrix for each pixel.
[0092] The third determining module is used to determine the segmentation index of each pixel based on the associated region matrix of each pixel.
[0093] The segmentation module is used to select pixels with a segmentation index greater than a preset segmentation index threshold as segmentation pixels, and to divide the target glass slide into blocks based on the segmentation pixels to obtain several sub-target glass slides.
[0094] The optimization module is used to extract the feature parameters of each sub-target slide and optimize them to obtain the target feature parameters;
[0095] The encoding module is used to encode the target feature parameters to obtain the feature vector of each sub-target slide, and then obtain the target slide in the form of feature vector.
[0096] The working principle of the above technical solution is as follows: The association region matrix represents the relationship between a pixel and its neighboring pixels, which is used for subsequent segmentation index calculation. The third determination module is used to determine the segmentation index of each pixel based on the association region matrix of each pixel, which facilitates the subsequent segmentation of the target slide based on the segmentation index. Pixels with a segmentation index greater than a preset threshold are marked as segmented pixels, which are also important feature points. Based on the segmented pixels, the target slide is segmented into several sub-regions using a region growing algorithm or connected component analysis. Each sub-region generates a sub-target slide for subsequent feature extraction. The optimization module extracts the feature parameters of the sub-target slides based on deep learning methods (such as convolutional neural networks) and optimizes them to obtain target feature parameters, providing high-quality input for subsequent encoding. The encoding module uses word embedding or convolutional neural networks (CNN) to encode the target feature parameters into fixed-dimensional feature vectors. Pre-trained models (such as ResNet and EfficientNet) can be used for feature extraction and encoding. The encoded feature vectors are standardized (such as Z-score standardization) to generate feature vectors with fixed dimensions and units.
[0097] The beneficial effects of the above technical solution are as follows: by generating a pixel-related region matrix, calculating the segmentation index, dividing the slide into blocks, extracting and optimizing feature parameters, and encoding features, the target slide is transformed into a feature vector form, realizing a standardized transformation from pixel points to slide-level feature vectors, providing standardized input for subsequent diagnosis, and facilitating the accurate acquisition of the target slide in the form of feature vectors.
[0098] According to some embodiments of the present invention, the generation module generates an associated region matrix for each pixel, including:
[0099] Generate the associated region matrix of the pixel in the i-th row and j-th column of the target glass slide;
[0100]
[0101] Among them, S i,j S is the pixel value of the pixel in the i-th row and j-th column; i-1,j-1 S is the pixel value of the pixel in the (i-1)th row and (j-1)th column; i-1,j S is the pixel value of the pixel in the (i-1)th row and jth column; i-1,j+1 S is the pixel value of the pixel in the (i-1)th row and (j+1)th column; i,j-1 S is the pixel value of the pixel in the i-th row and j-1-th column; i,j+1 S is the pixel value of the pixel in the i-th row and j+1-th column; i+1,j-1 S is the pixel value of the pixel in the (i+1)th row and (j-1)th column; i+1,j S is the pixel value of the pixel in the (i+1)th row and jth column; i+1,j+1 It is the pixel value of the pixel in the (i+1)th row and (j+1)th column.
[0102] The working principle and beneficial effects of the above technical solution are as follows: An associated region matrix is generated for each pixel in the target slide. This associated region matrix contains information about the target pixel and its neighboring pixels, which can be used for subsequent image analysis tasks, such as determining the segmentation index of each pixel.
[0103] According to some embodiments of the present invention, the third determining module determines the segmentation index of each pixel, including:
[0104] Determine the segmentation index of the pixel in the i-th row and j-th column of the target glass slide;
[0105]
[0106] Among them, A ij is the segmentation index of the pixel in the i-th row and j-th column of the target glass slide; D is the number of pixels included in the target glass slide.
[0107] The working principle and beneficial effects of the above technical solution: Determinant | X ij The determinant is a scalar value that reflects the scaling factor of the linear transformation represented by the matrix. A segmentation index is determined for each pixel in the target slide. The segmentation index contains information about the relationship between the pixel and its neighboring pixels, and is normalized through determinant and logarithmic operations to accurately determine the structure and features of the image.
[0108] According to some embodiments of the present invention, the optimization module includes:
[0109] The feature extractor is used to extract the feature parameters of each sub-target slide;
[0110]
[0111] Among them, F j (I i ) represents the feature extractor; P*Q represents the size parameters of the sub-target slide; I i (p,q) represents the sub-target slide I. i The pixel value at position (p, q); e is the natural constant; τ is the standard deviation controlling the Gaussian blur;
[0112] The feature optimizer is used to optimize the feature parameters to obtain the target feature parameters;
[0113]
[0114] Where R represents the target optimization feature parameters; N represents the number of feature parameters; λ represents the optimization coefficients; tanh represents the hyperbolic tangent function; ω j is the weight of the j-th feature extractor; M is the number of feature extractors.
[0115] The working principle and beneficial effects of the above technical solution are as follows: The feature extractor is used to extract feature parameters from each sub-target slide. These feature parameters can capture the local or global characteristics of the sub-target slide. τ is the standard deviation controlling the Gaussian blur, used to adjust the locality or globality of feature extraction. It is a two-dimensional Gaussian function that calculates the distance between a pixel and the center point. Pixel values are weighted. Pixels closer to the center point have a higher weight, and pixels farther away have a lower weight. The weighted average is normalized by dividing by P×Q to ensure the feature parameter values are within a reasonable range. The choice of standard deviation τ significantly impacts the feature extraction results. A larger τ leads the feature extractor to focus more on global characteristics, while a smaller τ focuses more on local characteristics. The feature optimizer optimizes the feature parameters extracted by the feature extractor to obtain more robust and discriminative target feature parameters. λ is the optimization coefficient used to adjust the intensity of the optimization process; tanh is the hyperbolic tangent function used to introduce nonlinear transformations to enhance the discriminative power of the features; ω... j The weight of the j-th feature extractor is used to adjust the contribution of different feature extractors to the final optimized feature parameters; The feature parameters extracted by different feature extractors are fused to obtain a comprehensive feature representation for each sub-target slide. A nonlinear transformation is then applied to the comprehensive feature representation using the tanh function to introduce nonlinear characteristics and enhance the discriminative power of the features. The optimized feature parameters of all sub-target slides are averaged to obtain the final optimized target feature parameters. The weights of the feature extractor can be initialized randomly or based on prior knowledge. During training, the weights can be updated using optimization algorithms (such as gradient descent). A larger λ leads to a more aggressive optimization process, while a smaller λ results in a more conservative one. Through the collaborative work of the feature extractor and the feature optimizer, the optimization module can extract more robust and discriminative target feature parameters from the sub-target slides.
[0116] According to some embodiments of the present invention, the local region feature interaction module within the slide uses a self-attention mechanism algorithm to contextually associate the features of the local region of the target slide, which have been converted into feature vector form, including:
[0117] The fourth determining module is used for:
[0118] Determine the grayscale value of each pixel in a local region of the target glass slide, and take the pixels with grayscale values greater than a preset threshold in the local region as local feature nodes;
[0119] Select any local feature node as the local feature node to be associated; calculate the distance between the local feature node to be associated and other local feature nodes, and take the minimum distance as the association distance; generate an associated region with the local feature node to be associated as the center and the association distance as the radius;
[0120] The fifth determination module is used to determine the associated region of each local feature node; and to determine the location information of the associated region of each local feature node.
[0121] The transformation module is used to form a vector composed of all high-frequency coefficients obtained by performing DCT transformation on the associated region of each local feature node in descending order, which serves as the local edge information vector of each local feature node.
[0122] The association module is used to perform contextual connections based on the location information of the associated region and the local edge information vector of each local feature node to obtain the association relationship;
[0123] The verification module is used to calculate the effective coefficient of the association relationship and compare it with the preset effective coefficient threshold; when the effective coefficient is determined to be greater than the preset effective coefficient threshold, the first diagnostic result is obtained.
[0124] The working principle and beneficial effects of the above technical solution are as follows: The fourth determination module is responsible for determining the grayscale value of each pixel in the local region of the target slide, and filtering out local feature nodes based on the grayscale value, thereby determining the associated region as the influence region centered on the local feature node. The fifth determination module determines the positional information of the associated region of each local feature node, including relationships such as intersection, tangency, and overlap. A two-dimensional DCT transformation is performed on the associated region to obtain the coefficient matrix in the frequency domain. According to the frequency characteristics, the DCT coefficient matrix is divided into low-frequency, mid-frequency, and high-frequency regions. The coefficients of the high-frequency region are extracted as a high-frequency coefficient set. The high-frequency coefficient set is sorted in descending order to obtain an ordered high-frequency coefficient vector. The descendingly ordered high-frequency coefficient vector is used as the local edge information vector of the local feature node, which can effectively capture the edge and texture information of the image. Based on the positional information of the associated region of each local feature node, the spatial distance between each associated region is determined. The similarity between the local edge information vectors is calculated, and a feature similarity threshold is set. If the similarity of the local edge information vectors of two nodes is greater than the threshold, they are considered to be similar in features. The system employs weighted summation, logical AND / OR, and other methods to fuse two relationships, comprehensively considering spatial adjacency and feature similarity to construct the final association relationship. The effective coefficient of the association relationship is calculated and compared with a preset effective coefficient threshold; when the effective coefficient is determined to be greater than the preset threshold, a first diagnostic result is obtained. This improves the accuracy of the first diagnostic result.
[0125] In one embodiment, the method of using a self-attention mechanism algorithm to link the features of different regions of the entire slide in context in the slide-level feature interaction module, the method of using a self-attention mechanism algorithm to link the features of different slides in the entire sampling area in context in the inter-slide feature interaction module, and the method of using a self-attention mechanism algorithm to link the features of local regions of the target slide in context in the local region feature interaction module within the slide are consistent with the method of using a self-attention mechanism algorithm to link the features of local regions of the target slide in feature vector form in context in the local region feature interaction module within the slide.
[0126] According to some embodiments of the present invention, the verification module includes:
[0127] The sixth determining module is used to determine the associated dataset and several valid indicators corresponding to the association relationship;
[0128] The calculation module is used to calculate the sub-efficiency coefficients of the effective indicators in the associated dataset;
[0129]
[0130] Where K is the sub-efficiency coefficient of the effective index in the associated dataset; b a For each effective metric, F is the membership function of the effective metric on the associated dataset A; U is the universe of discourse of the effective metric; n is the number of all metrics on the associated dataset A; bx Let T(b) be the x-th metric in the associated dataset A; a b x ) represents the matching degree between the effective indicator and the x-th indicator in the associated dataset A;
[0131] Calculate the sub-efficiency coefficients of all effective indicators in the associated dataset and sum them. Use the ratio of the sum to the number of effective indicators as the effective coefficient of the association relationship.
[0132] The working principle of the above technical solution is as follows: Effective indicators include association strength, stability, and consistency. The calculation module is responsible for calculating the sub-effectiveness coefficients of the effective indicators in the associated dataset, and further calculating the effective coefficients of the association relationships. The membership function is used to describe the degree of membership of the effective indicators in the associated dataset. T(b a b x This is used to describe the similarity between effective indicators and other indicators in the associated dataset. The calculation module accurately calculates the sub-effectiveness coefficients of effective indicators in the associated dataset, calculates the sub-effectiveness coefficients of all effective indicators in the associated dataset, and sums them. The ratio of the sum to the number of effective indicators is used as the effective coefficient of the association relationship.
[0133] The beneficial effects of the above technical solution are as follows: The verification module calculates the sub-validity coefficients of the valid indicators in the associated dataset and further calculates the validity coefficients of the association relationship to verify the validity of the association relationship, thereby improving the accuracy of judging the size of the validity coefficients and the preset validity coefficient threshold, and making it easier to obtain accurate first diagnostic results.
[0134] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A clinically practical AI-assisted diagnostic system based on the sampling site of a medical case, characterized in that, include: The acquisition module is used to acquire case diagnosis information; The first determination module is used to determine the label information for each sampling site in each case based on the case diagnosis information; The second determining module is used to determine the target glass slide corresponding to the target sampling location based on the label information; The preprocessing module is used to preprocess the target glass slide to obtain the target glass slide in the form of feature vectors; The fusion reasoning module is used to perform image fusion reasoning on the target glass slide converted into feature vector form to obtain AI-assisted diagnostic results. The fusion reasoning module includes: The local region feature interaction module within the slide is used to perform contextual association of the features of the local region of the target slide, which has been converted into feature vector form, through a self-attention mechanism algorithm to obtain the first diagnostic result; The slide-level feature interaction module is used to first shuffle the feature vectors of the target slide to ensure that the feature vectors placed together belong to different regions of the slide. Then, the self-attention mechanism algorithm is used again to connect the features of different regions of the entire slide in context to obtain the second diagnostic result. The inter-slide feature interaction module is used to aggregate the feature vectors of the target slide into a whole feature vector using attention_pooling. The whole feature vector represents the feature vector of the entire slide. The feature vectors of all slides are put together and processed based on the self-attention mechanism algorithm to connect the features of different slides in the entire sampling site in context to obtain the third diagnostic result. AI-assisted diagnostic results are obtained based on the first, second, and third diagnostic results. The local region feature interaction module within the slide uses a self-attention mechanism algorithm to contextually associate the features of the local region of the target slide, which have been converted into feature vector form, including: The fourth determining module is used for: Determine the grayscale value of each pixel in a local region of the target glass slide, and take the pixels with grayscale values greater than a preset threshold in the local region as local feature nodes; Select any local feature node as the local feature node to be associated; calculate the distance between the local feature node to be associated and other local feature nodes, and take the minimum distance as the association distance; generate an associated region with the local feature node to be associated as the center and the association distance as the radius; The fifth determination module is used to determine the associated region of each local feature node; and to determine the location information of the associated region of each local feature node. The transformation module is used to form a vector composed of all high-frequency coefficients obtained by performing DCT transformation on the associated region of each local feature node in descending order, which serves as the local edge information vector of each local feature node. The association module is used to perform contextual connections based on the location information of the associated region and the local edge information vector of each local feature node to obtain the association relationship; The verification module is used to calculate the effective coefficient of the association relationship and compare it with the preset effective coefficient threshold; when the effective coefficient is determined to be greater than the preset effective coefficient threshold, the first diagnostic result is obtained.
2. The clinically practical AI-assisted diagnostic system based on the sampling site of a case as described in claim 1, characterized in that, The acquisition module obtains case diagnosis information from the information system, including the pathology number of the slide and the sampling site.
3. The clinically practical AI-assisted diagnostic system based on the sampling site of a case as described in claim 2, characterized in that, The first determining module includes: The classification module is used to classify each slide into its respective location based on the pathology number and the sampling site. The first extraction module is used to extract the pathology report corresponding to the pathology number of each slide; The splitting module is used to split a complete pathology report into pathological diagnoses based on the sampling site, and assign the diagnoses to the corresponding sampling sites, based on the large language model. The second extraction module is used to extract the labels of the required diseases from the diagnostic results of each sampling site using a large language model, and to determine the label information of each sampling site in each case.
4. The clinically practical AI-assisted diagnostic system based on the sampling site of a case as described in claim 1, characterized in that, The preprocessing module includes: The generation module is used to obtain the pixel values of each pixel in the target slide and generate an associated region matrix for each pixel. The third determining module is used to determine the segmentation index of each pixel based on the associated region matrix of each pixel. The segmentation module is used to select pixels with a segmentation index greater than a preset segmentation index threshold as segmentation pixels, and to divide the target glass slide into blocks based on the segmentation pixels to obtain several sub-target glass slides. The optimization module is used to extract the feature parameters of each sub-target slide and optimize them to obtain the target feature parameters; The encoding module is used to encode the target feature parameters to obtain the feature vector of each sub-target slide, and then obtain the target slide in the form of feature vector.
5. The clinically practical AI-assisted diagnostic system based on the sampling site of a case as described in claim 4, characterized in that, The generation module generates an associated region matrix for each pixel, including: Generate the associated region matrix of the pixel in the i-th row and j-th column of the target glass slide; ; in, Let be the pixel value of the pixel in the i-th row and j-th column; The pixel value of the pixel in the (i-1)th row and (j-1)th column; The pixel value of the pixel in the (i-1)th row and jth column; The pixel value of the pixel in the (i-1)th row and (j+1)th column; The pixel value of the pixel in the i-th row and j-1-th column; The pixel value of the pixel in the i-th row and j+1-th column; This is the pixel value of the pixel in the (i+1)th row and (j-1)th column. The pixel value of the pixel in the (i+1)th row and jth column; It is the pixel value of the pixel in the (i+1)th row and (j+1)th column.
6. The clinically practical AI-assisted diagnostic system based on the sampling site of a case as described in claim 5, characterized in that, The third determining module determines the segmentation index of each pixel, including: Determine the segmentation index of the pixel in the i-th row and j-th column of the target glass slide; ; in, The segmentation index of the pixel in the i-th row and j-th column of the target glass slide; The number of pixels included in the target glass slide; The determinant is a scalar value that reflects the scaling factor of the linear transformation represented by the matrix.
7. The clinically practical AI-assisted diagnostic system based on the sampling site of a case as described in claim 4, characterized in that, The optimization module includes: The feature extractor is used to extract the feature parameters of each sub-target slide; ; in, For feature extractors; These are the dimensional parameters of the sub-target glass slide; For the sub-target glass slide In position The pixel value; e is the natural constant; To control the standard deviation of Gaussian blur; The feature optimizer is used to optimize the feature parameters to obtain the target feature parameters; ; in, Optimize feature parameters for the target; The number of feature parameters; To optimize the coefficients; It is the hyperbolic tangent function; The weights of the j-th feature extractor; This represents the number of feature extractors.
8. The clinically practical AI-assisted diagnostic system based on the sampling site of a case as described in claim 1, characterized in that, The verification module includes: The sixth determining module is used to determine the associated dataset and several valid indicators corresponding to the association relationship; The calculation module is used to calculate the sub-efficiency coefficients of the effective indicators in the associated dataset; ; in, The sub-efficiency coefficient of the effective indicator in the associated dataset; As a valid indicator; For effective metrics in related datasets Membership function on; The domain of discourse in which the effective indicator resides; In the associated dataset The number of all indicators; Let x be the x-th metric in the associated dataset A; The degree of matching between the effective indicator and the x-th indicator in the associated dataset A; Calculate the sub-efficiency coefficients of all effective indicators in the associated dataset and sum them. Use the ratio of the sum to the number of effective indicators as the effective coefficient of the association relationship.
Citation Information
Patent Citations
Pathology auxiliary diagnosis device and method based on small sample learning method MAML
CN118248321A
Method and system for fusion-extracting whole slide pathology features based on multi-scale, system, electronic apparatus, and storage medium
JP2024027078A