Image feature matching method and device

By matching and fusing the full-slice potential features and classification label features, the problem of inconsistency between labels and text reports in image feature matching is solved, and efficient and accurate diagnostic result output is achieved.

CN120655945APending Publication Date: 2025-09-16HANGZHOU YICE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510746233.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the existing technology, the detailed classification labels and text report descriptions in the image feature matching method are the results generated from two branches respectively, and there is a lack of connection in the middle, which may cause the results to be mismatched with each other and fail to meet the accuracy and consistency requirements of clinical diagnosis.

Method used

By extracting the full-slice potential features and classification label features, performing feature matching and fusion, generating a text report, and fusing the classification probability in the classification module, the consistency of the label and text report is ensured.

Benefits of technology

It improves the accuracy of diagnosis and work efficiency, reduces the possibility of misjudgment and missed detection, and meets the accuracy and consistency requirements of clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655945A_ABST
    Figure CN120655945A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing technology, and discloses an image feature matching method and device, and the method comprises the steps: obtaining a full-slice digital image and a possible label option list; extracting full-slice features, full-slice potential features and classification label features from the full-slice digital image and the possible label option list; performing feature matching on the full-slice potential features and the classification label features, and decoding to generate a text report; carrying out feature fusion on the full-slice features and the classification tag features, and carrying out the output of a classification module; and according to the obtained classification module and the generated text report, obtaining an output label subdivision category and the text report. According to the image feature matching method designed by the invention, diagnosis can be automatically, quickly and accurately assisted in advance, the current situation of the current pathology department can be effectively improved, and the automatic diagnosis process of the pathology department is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and in particular to a method and device for image feature matching. Background Art

[0002] With breakthroughs in large-scale model technology, the application of artificial intelligence in histopathology-assisted diagnosis has evolved from the early, crude classification of benign and malignant tumors (e.g., differentiating between cancer and non-cancerous tissues) to a more refined, multi-dimensional approach. Current clinical needs have expanded to include subtyping, gene mutation prediction, and prognostic stratification.

[0003] At the same time, the volume of pathology diagnostic tasks is growing exponentially. According to statistics, a single tertiary hospital produces over 500,000 pathology slides annually, while my country faces a shortage of 100,000 pathologists. The average diagnosis time at primary hospitals is 3-5 days. Faced with this massive amount of data, AI-assisted systems have assumed a critical role. They use convolutional neural networks to automatically extract nuclear morphology parameters and combine multimodal data fusion techniques with pathology images, genomics, and electronic medical records, increasing diagnostic efficiency for common tumors by over 40%.

[0004] More notably, the intelligent report generation system based on the TRANSFORMER architecture automatically generates structured diagnostic opinions, significantly reducing report writing time while simultaneously increasing positive detection sensitivity to new heights. The application of the digital pathology platform has reduced remote consultation response time to under two hours. In the future, pathology diagnosis will establish a new working paradigm of "AI initial screening - physician review - intelligent reporting," effectively alleviating the global challenge of uneven distribution of medical resources.

[0005] Over the past two years, with the continuous advancement of technology and hardware, the field of histopathology has seen the emergence of single-channel full-slide feature extraction and downstream tasks. A representative example is PAIGEAI's PRISM model, which uses the Perceiver network to aggregate the feature embeddings of thousands of tiles extracted by Virchow to obtain full-slide features for downstream tasks. These downstream tasks consist of two branches: one for comparative learning to obtain detailed slice labels, and the other for decoding to obtain a textual report description of the slice. This approach best meets clinical needs and demonstrates the best experimental results. Logically, both branches occur after full-slide feature aggregation, providing the best representation of full-slide features. However, logically, the detailed category labels and textual report descriptions are generated independently from the two branches, with no connection other than a certain correlation between the features. This can lead to inconsistent results from the two branches, but this is unacceptable in real clinical applications.

[0006] Hospitals, medical examination centers, and other institutions face a daily demand for reviewing large numbers of tissue slides from various locations. However, manual review is inefficient, and pathologists face significant fatigue after prolonged review, making misdiagnosis and missed diagnosis more likely. Highly accurate tools are crucial for this task. During the review process, doctors must simultaneously label multiple slides from a single patient and compile each slide's label into a text report. This process requires ensuring both the accuracy of diagnostic labels and the consistency between the text report and the labels.

[0007] The PRISM model mentioned above aggregates features from the entire slice and then connects two task branches for detailed classification and report generation, respectively. In practice, mismatches between the report result "NORMAL COLONIC MUCOSA, NO PATHOLOGY DETECTED" and the label "POLYP" often occur. This is because the model only considers the generation of two types of results, but fails to consider the necessity of result consistency from a practical perspective. In actual use, doctors will first check the accuracy of the labels, then insert the labels into the text report, and finally integrate the reports of multiple slices into a single case report. At this time, consistency between the report results and the labels is necessary. Although large models can predict labels and text, the poor diagnostic consistency results in low practical auxiliary performance. Overall, reports need to describe labels realistically and accurately using richer text. Summary of the Invention

[0008] The present invention addresses the problem in the prior art that subdivision labels and text report descriptions are the results generated by two branches respectively, and there is no other connection between them except for a certain correlation between features. The results of the two branches may be different and do not match each other. A method and device for image feature matching are provided.

[0009] In order to solve the above technical problems, the present invention is solved by the following technical solutions:

[0010] A method for image feature matching, comprising:

[0011] Obtain a full-slide digital image and a list of possible labeling options;

[0012] Extracting full-slice features, full-slice potential features, and classification label features from a full-slice digital image and a list of possible label options;

[0013] The text report is generated by matching the full slice potential features with the classification label features and decoding them to generate a text report;

[0014] The output of the classification module is obtained by fusing the full slice features and the classification label features, and then outputting the classification module;

[0015] Output label segmentation categories and text reports, and obtain output label segmentation categories and text reports based on the obtained classification module and generated text reports.

[0016] As a preferred embodiment: output label subdivision categories and text reports, for the obtained classification module, obtain the classification probability through the classification probability, and then combine and generate the text output label subdivision categories and text reports.

[0017] Preferably, the calculation of the classification probability includes:

[0018] Acquisition of the first classification probability list: extracting full slice features and classification label features and obtaining the first classification probability list by comparing the classification branches;

[0019] Obtaining the second classification probability list, performing feature fusion on the full slice features and the classification label features to obtain the second classification probability list;

[0020] The classification probability is calculated by performing a weighted summation on the first classification probability list and the second classification probability list to obtain the classification probability.

[0021] As a preference:

[0022] The extraction of full-slice features includes: inputting a valid patch image of the full slice; extracting full-slice features based on the input valid patch image; outputting the extracted full-slice features, and outputting the features and potential features of the full slice;

[0023] The extraction of classification label features includes: the classification label list of all slices is [a1, a2, a3, a4, ...], where the classification label list dimension is k, and input into the language embedding and language decoding layer of BioGPT in the PRISM model to obtain the classification label features.

[0024] Preferably, obtaining the first classification probability list by comparing the classification branches includes:

[0025] Perform linear projection on the full slice features and classification label features obtained by the transformer structure respectively;

[0026] The matrix product of the linearly projected full slice features and the classification label features is calculated to obtain the first classification probability list.

[0027] Preferably, the feature fusion includes processing the input k*1024-dimensional embedding features through multiple layers of transposition and FC layers to output 1*1280-dimensional label embedding features;

[0028] The input 1*1280-dimensional full slice feature is added to the label embedding feature to obtain the updated 1*1280-dimensional feature;

[0029] After the operations of feature dimensionality reduction and dimensionality increase, the features are more deeply integrated and the final 1*1280-dimensional classification features are output.

[0030] In order to solve the above technical problems, the present invention also provides an image feature matching device, which is characterized by being implemented by the image feature matching method.

[0031] In order to solve the above technical problem, the present invention further provides a computer program product, characterized in that when running on a computer, the method described above is executed.

[0032] In order to solve the above technical problem, the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein the program is configured to implement the method described above when executed.

[0033] The present invention has significant technical effects due to the adoption of the above technical solutions:

[0034] The technical solution of the present invention extracts the encoding features of the full-slice potential features and the classification labels, fuses the full-slice potential features and the label features at the feature matching layer, outputs the full-slice potential features carrying the label features, and then decodes them to obtain a text report. On the other hand, the label features and the full-slice features are feature fused and input into the newly added classification module to obtain the classification probability of each subclass. Since PRISM itself obtains the classification probability of each subclass through comparative learning, the actual classification label of the full slice is finally derived by fusing the two classification probabilities. This process will facilitate the execution of tasks such as report writing, reduce labor costs, greatly improve work efficiency, reduce the possibility of misjudgment and missed detection, and ensure the accuracy of diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Flowchart of the present invention.

[0036] Figure 2 It is a structural schematic diagram of the feature matching of the present invention.

[0037] Figure 3 Schematic diagram of the structure of the feature fusion of the present invention.

[0038] Figure 4 This is a structural diagram of the classification module of the present invention. DETAILED DESCRIPTION

[0039] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0040] Example 1

[0041] A method for image feature matching sets the physical size of the patch input to the full-slide feature extraction module (a fixed module of the PRISM model) to 112 microns, thereby cropping patch images of different slides. Figure 1 In [1], all valid patch images of an arbitrary full-slice image are input, and the full-slice feature extraction module of the PRISM model directly outputs the full-slice features and latent features. The output full-slice features have a dimension of 1*1280, and the latent features have a dimension of 512*1280.

[0042] The other input is the set full-slice classification label list [a1, a2, a3, a4, ...], whose input dimension is K. It is input into the language embedding and language decoding layer of BioGPT in the PRISM model to obtain K*1024-dimensional label features. Based on the above input, it is necessary to establish the correlation between k-dimensional (generally 10 to 20)-dimensional classification labels and slice text descriptions to improve the consistency between the two, and propose a solution from the root from the perspective of model feature fusion and matching, so that the text and labels generated by the new model can meet the needs of clinical use. It helps in the execution of tasks such as report writing, reduces labor costs, greatly improves work efficiency, reduces the possibility of misjudgment and missed detection, and ensures the accuracy of diagnosis.

[0043] The feature is compared to achieve implicit feature matching. The specific implementation method is to linearly project the full slice feature obtained based on the transformer structure and the label encoding feature, and then perform a comparison (matrix product) calculation.

[0044] This branch inputs the K*1024-dimensional label feature into the linear projection layer (a single-layer FC network with a weight size of 1024*5120), and outputs the K*5120-dimensional label embedding feature. This feature is compared with the 1*5120-dimensional slice feature obtained by linear projection of the 1*1280-dimensional full-slice feature (the product of 1*5120 and K*5120 dimensions) and softmax probability calculation to obtain a 1*K-dimensional probability vector C1. The sum of the vectors is 1, and the value of each dimension represents the classification probability of the label at the corresponding position.

[0045] Feature matching: The purpose is to match and fuse the label features into the full-slice potential features, that is, the label embedding features of K*1024 dimensions and the full-slice potential features of 512*1280 dimensions are input into the feature matching layer at the same time. Figure 2 , after multiple rounds of matrix calculation output and label matching, the updated full-slice potential features of 512*1280 dimensions are obtained.

[0046] The entire feature matching branch obtains features from two paths and passes through the final fusion feature layer B to obtain a 512×1280 feature output. Each path involves feature dimension alignment, a cross-attention mechanism based on QKV feature interactions, and feature fusion using residual connections or matrix products. On path 1, the K×1024 label embedding features are first aligned. A linear projection layer is used to project the K×1024 label features to K×1280 dimensions (specifically, the FC layer of the nn.Linear(1024, 1280) structure), aligning their channel count with the full-slice features. This is then fed into the cross-attention layer A, which uses a multi-head attention layer to construct feature interactions. The full-slice features (512×1280) are used as the query, and the projected label features (K×1280) are used as the key / value. The attention weight matrix (512×K) is calculated by multiplying the transpose of the query and the key. Finally, the final 512×1280 fused features are obtained by multiplying the attention weight matrix with the value.

[0047] The 512×1280 dimension feature y1 output by the cross attention layer A is subjected to the residual connection layer operation with the 512×1280 feature y2 of the original input, specifically z1=y1+y2, and the output of this branch is directly input into the fusion feature layer B.

[0048] On another path 2, the label embedding features of K*1024 dimensions are aligned by the number of channels of the same feature dimension alignment layer, and then continuously enter two feature dimension reduction layers and one feature dimension increase layer, so that the label embedding features and the full-slice potential features achieve bilateral feature alignment, and the attention weight layer B is performed in a multi-head attention manner, followed by the fusion feature layer A, and the final 512×1280 dimensional feature output z2 is fused with the 512×1280 dimension z1 of the previous path 1 in the feature layer B.

[0049] Fusion classification is based on feature dimension alignment, residual connection, and linear projection mapping to carry out feature fusion.

[0050] The label embedding features are transformed into 1*1280 dimensional features that are completely aligned with the slice feature dimensions through continuous transposition and linear projection layers. Then, the residual connection of the features is performed, and linear projection mapping is applied to make the fused features more beneficial for subsequent classification operations. The classification label features are fused with the full slice features, and the classification probability of the label is calculated by the classification label features and the full slice features, so that the image and text information can be integrated to a new level. The K*1024 dimensional label embedding features and the 1*1280 dimensional full slice features are fused through a unique feature fusion layer (see Figure 3) are fused to output the updated full-slice features of 1*1280 dimension, and finally input into a specific classification module to output a 1*K dimensional probability vector, which also sums to 1.

[0051] Feature matching solves the problem of no label information guidance during text generation. Since the pathological slice data of different parts are likely to be composed of different labels, and if the text generation is only affected by the slide information, it is naturally very illogical, so it is necessary to fine-tune the feature state according to the content of the real-time label. The layer path 2 first receives the label feature embedding vector of K*1024 dimensions, and outputs 1280*512-dimensional features through 4 FC layers of different sizes. The 512*1280-dimensional latent features input on the other side are multiplied by the previously output 1280*512-dimensional features to obtain a 512*512-dimensional label and latent feature matching feature output. The feature continues to match the 512*1280-dimensional features. The output feature z2 and the 512*1280-dimensional output z1 obtained by the residual connection block on the other path 1 are jointly input into the fusion feature layer B to obtain the updated latent features. The fusion feature layer B specifically refers to z1+(the product of z2 and the transpose of z1)*z2. The specific structure of the feature matching layer is as follows Figure 2 shown.

[0052] Figure 1 The feature fusion layer in [1] fuses the BioGPT-encoded label features with the full-slice features, completing the fusion of multimodal information related to the current classification and feeding it into another classification module, providing a multimodal information foundation for subsequent classification modules. This layer primarily includes dimension alignment, residual connections, and feature depth fusion modules.

[0053] The feature fusion layer is as follows Figure 3 As shown in the figure, the input K*1024 dimensional embedding features are first processed through multiple layers of transposition and FC layers to output 1*1280 dimensional label embedding features to achieve the purpose of dimensional alignment. Then, the input 1*1280 dimensional full slice features are summed with the label embedding features, that is, the residual connection module, to obtain 1*1280 dimensional updated features. Finally, after the feature dimensionality reduction and dimensionality increase operations, the features are more deeply fused and the final 1*1280 dimensional classification features are output. The classification features are then input Figure 4 The classification module shown (which consists of 3 fully connected FC layers and a softmax layer) finally outputs a 1*K-dimensional probability vector C2, which also sums to 1.

[0054] Since both the comparative classification branch and the classification module output 1*K-dimensional probability vectors, the label classification value is obtained by taking the maximum value after weighted summation of the vectors during inference. Specifically, it is MAX(0.7*C1+0.3*C2). The probability values ​​0.7 and 0.3 can be adjusted based on actual conditions. Considering that the comparative classification is the original classification module in actual situations, it has the advantage of greater classification accuracy. In order to avoid disrupting its fundamental stability and ensure the accuracy of most classifications, the weight of C1 will be higher.

[0055] The complete steps for label classification and text report generation for whole-slide histopathology images are as follows: obtain a full-slide digital image and a list of possible label options; extract full-slide features and classification label features; obtain one of the classification probability lists by comparing the classification branches; generate a text report through feature matching and decoding; obtain another classification probability list by fusing the features of the full slide and the classification label; obtain the classification probability by weighted summation of the two classification probability lists; and finally output the label classification category and text report.

[0056] The following shows the labeling results of the text obtained by using the algorithm versions before and after the improvement. The text and labeling results before the improvement are output in the following table:

[0057] Table 1 Text and label prediction table before improvement

[0058]

[0059] It is obvious that there is a clear inconsistency between the text in the left column and the label in the right column. In the first example, when the label prediction is villous tubular adenoma, the text describes it as high-grade dysplasia. In contrast, in the table below, the improved text and label prediction are completely consistent.

[0060] Table 2 Improved text and label prediction table

[0061]

Claims

1. A method for image feature matching, comprising: Obtain a digital image of the whole slide and a list of possible labeling options; Extracting full-slice features, full-slice potential features, and classification label features from a full-slice digital image and a list of possible label options; The text report is generated by matching the full slice potential features with the classification label features and decoding them to generate a text report; The output of the classification module is obtained by fusing the full slice features and the classification label features, and then outputting the classification module; Output label segmentation categories and text reports, and obtain output label segmentation categories and text reports based on the obtained classification module and generated text reports.

2. The image feature matching method according to claim 1, wherein: Output label segmentation categories and text reports. For the obtained classification module, the classification probability is obtained through the classification probability, and then combined to generate text output label segmentation categories and text reports.

3. The image feature matching method according to claim 1, wherein: The calculation of classification probability includes: Acquisition of the first classification probability list: extracting full slice features and classification label features and obtaining the first classification probability list by comparing the classification branches; Obtaining the second classification probability list, performing feature fusion on the full slice features and the classification label features to obtain the second classification probability list; The classification probability is calculated by performing a weighted summation on the first classification probability list and the second classification probability list to obtain the classification probability.

4. The image feature matching method according to claim 1, wherein: The extraction of full-slice features includes: inputting a valid patch image of the full slice; extracting full-slice features based on the input valid patch image; outputting the extracted full-slice features, and outputting the features and potential features of the full slice; The extraction of classification label features includes: the classification label list of all slices is [a1, a2, a3, a4, ...], where the classification label list dimension is k, and input into the language embedding and language decoding layer of BioGPT in the PRISM model to obtain the classification label features.

5. The image feature matching method according to claim 1, wherein: The first classification probability list obtained by comparing the classification branches includes: Perform linear projection on the full slice features and classification label features obtained by the transformer structure respectively; The matrix product of the linearly projected full slice features and the classification label features is calculated to obtain the first classification probability list.

6. The image feature matching method according to claim 1, wherein: The feature fusion includes processing the input k*1024-dimensional embedding features through multiple layers of transposition and FC layers to output 1*1280-dimensional label embedding features; The input 1*1280-dimensional full slice feature is added to the label embedding feature to obtain the updated 1*1280-dimensional feature; After the operations of feature dimensionality reduction and dimensionality increase, the features are more deeply integrated and the final 1*1280-dimensional classification features are output.

7. A device for image feature matching, characterized in that: A device implemented by the image feature matching method described in any one of claims 1-6.

8. A computer program product, characterized in that When running on a computer, the method according to any one of claims 1 to 6 is executed.

9. A computer-readable storage medium having a computer program stored thereon, the program being configured to implement the method according to any one of claims 1 to 6 when executed.