Retrieval and matching method and system for similar lesions

By introducing a CT scan image encoder and auxiliary information encoder with VIT structure, combining CT scan image and patient information, the problem of only visual information in the existing method is solved, and more comprehensive lesion judgment and system optimization are achieved.

CN115080782BActive Publication Date: 2025-07-25SUN YAT SEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210664467.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-07-25
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

Most of the existing lesion detection and retrieval methods do not adopt the latest deep learning network architecture, and only consider the visual information of the CT scan results, while ignoring important information such as users' eating habits and living habits, resulting in inaccurate judgments.

Method used

Using a CT scan image encoder and auxiliary information encoder based on VIT structure, the patient's CT scan image and auxiliary information are fused through a sub-attention mechanism, the network is trained to output lesion judgment, and auxiliary diagnosis is performed in combination with similar cases, and the patient information management system is continuously optimized.

Benefits of technology

It achieves a more comprehensive lesion judgment, takes into account the patient's multi-faceted information, and improves the accuracy of diagnosis and systematic self-optimization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115080782B_ABST
    Figure CN115080782B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for retrieving and matching similar lesions. It includes: collecting various quantitative indicators of patients' physical examinations and CT image data of past patients' lesions, training a CT scan image encoder, training a CT scan image decoder, training an auxiliary information encoder, performing a CT scan on the patient to be diagnosed and inputting their information, obtaining the CT image of the lesion and inputting it into the CT scan image encoder, then inputting the auxiliary personal information of the patient to be diagnosed into the auxiliary information encoder, and finally combining relevant cases and the specific judgment of the CT scan image decoder on the lesion situation to assist doctors in diagnosing the patient. Finally, the diagnosis result is input into the patient information management system to continuously optimize the retrieval and matching system. The present invention can continuously optimize the patient information management system. Its lesion retrieval and matching system comprehensively considers various information that may affect the lesion attributes of patients, and introduces the VIT structure into the learning of lesion images, having greater potential to learn the lesion information of users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and information retrieval, and particularly to a method and system for retrieving and matching similar lesions. Background Art

[0002] In recent years, with the popularization of deep learning technology and the increasing maturity of medical imaging technology, applying related deep learning technologies to the medical field has become a major research topic. Currently, various medical diagnoses and relevant patient information have been digitized, which provides a basis for the large amount of data required in deep learning. Among them, a standardized and structured electronic medical record system has collected high-quality medical thoracic imaging data, which includes thoracic CT scan images of past patients in different positions. When doctors qualitatively diagnose the current patient's lesions, misdiagnosis often occurs due to subjective reasons. Reasonably using the previously collected CT image data and retrieving similar lesions from these data to assist doctors in diagnosis can largely avoid misdiagnosis.

[0003] With the huge impact of Transformer in the NLP field and the emergence of VIT in the past two years, the Transformer architecture has also started to stand out in the CV field. For supervised learning, the Transformer architecture has more powerful learning potential compared to traditional convolutional architectures. In addition, the attention mechanism in Transformer can effectively fuse various information. Compared with the information fusion methods of direct addition or channel splicing in convolutional networks, the Transformer architecture has more advantages in multi-information processing. In the medical field, determining a patient's lesions should not be limited to information related to a certain symptom of the user, but should be comprehensively judged in combination with aspects such as their living habits and eating habits.

[0004] One of the current existing technologies is the patent "A Lesion Matching Method and Device (CN110674885A)". This technology judges the size, shape, and relative position of each lesion, and then inputs the obtained data into the existing past medical record database for comparison to select lesion medical records with similar sizes and shapes to assist doctors in judgment. The disadvantage of this method is that since only the traditional geometric relationship between graphics is used and these information are not deeply mined, a lot of content useful for qualitatively diagnosing lesions may be ignored. At the same time, it is unreasonable to retrieve lesions only through one image without considering more personal information of the patient.

[0005] The second of the current prior arts is the lesion retrieval method based on a deep convolutional network in the paper "Deep lesion tracker: Monitoring lesions in 4D longitudinal imaging studies". This method first calibrates the lesion area of the patient, then uses the convolutional network to learn the CT images of the lesion areas in the database, and then uses the trained model to assist in the judgment of the lesion stars. Finally, the situation of the lesion area of the current patient is obtained to assist the doctor in making a judgment. The disadvantage of this method is that it does not use the retrieval function. Its judgment of the patient's lesion area is based on the situation of past patients learned by the network, and only the CT scan images of the patient's lesions are considered without taking into account more patient information.

[0006] The third of the current prior arts is the lesion detection method based on object detection in the paper "Medical Image Object Detection Based on Visual Attention Model". This method first uses the CT scan images of past users and the positions of the lesions as a data set to train the object detection network, and then inputs the CT scan images of the current patient into the object detection network to detect the position of the patient's lesion and the specific type of the lesion. The disadvantage of this solution is that the judgment of the lesion type is achieved by using the object detection network, and no special network is used to achieve it, which is likely to cause inaccurate judgment. At the same time, other personal information of the patient that may be useful for judging the lesion is not considered. Summary of the Invention

[0007] The object of the present invention is to overcome the deficiencies of the existing methods and propose a method and system for retrieving and matching similar lesions. The main problems solved by the present invention are that most of the existing lesion detection, retrieval and matching methods do not adopt the latest deep learning network architecture. At the same time, the existing lesion matching methods often only consider the visual information of the CT scan results and ignore many other important information such as the user's eating habits and living habits. Currently, the collation of past medical records is mostly based on specific cases rather than individuals, resulting in the lack of judgment information that may be beneficial to the qualitative determination of lesions. That is, how to establish a complete medical record file based on deep learning, starting from graphic recognition and the fusion of various types of information, and give a reasonable answer after comprehensively considering various patient information as much as possible when judging lesions.

[0008] To solve the above problems, the present invention proposes a method for retrieving and matching similar lesions, and the method includes:

[0009] Confirm the specific clinical symptoms of the patient and collect various quantitative indicators of the patient's physical examination.

[0010] Collect the CT images of the lesions of past patients and the lesion areas determined by doctors after judgment, mark the lesion areas, and record the specific type information of the lesions as category labels;

[0011] Train the CT scan image encoder. First, input the image obtained after target detection into the CT scan image encoder. After fusing the encoding of the output CT scan image and the representation result of the user auxiliary information, input it into the CT scan image decoder. After obtaining the similar lesion detection output and lesion judgment output, perform regression with the category label. When the accuracy drops until the network converges, complete the training of the CT scan image encoder;

[0012] Train the CT scan image decoder. Refer to the attention module in VIT (Vision Transformer), input the representation result of the user auxiliary information and the encoding of the CT scan image into the CT scan image decoder. Use the patient auxiliary information as K and V in the attention, and the encoding of the CT scan image as Q. Learn about the patient's illness situation by combining the information of the patient himself through the sub-attention mechanism, and directly output the judgment of the user's lesion, as a reference for the doctor to qualitatively determine the patient's condition. At the same time, input the lesion hidden code obtained by the CT scan image decoder into the similar lesion retrieval system to retrieve similar cases in the patient information file as a reference for the doctor to diagnose the patient's lesion;

[0013] Train the auxiliary information encoder. The auxiliary information encoder is based on the encoder structure of the transformer. First, convert the user's various physical indicators and living habit information into one-hot tensors, and then input the one-hot tensor information into the fully connected network (FC) for word embedding, similar to the processing method of natural language processing (NLP). Then, use the embedded tensor as the representation result of the user auxiliary information. Finally, input the representation result of the user auxiliary information and the encoding of the CT scan image into the CT image decoder based on the transformer, and train the auxiliary information encoder by gradient descent until it converges;

[0014] Input the information of the patient to be diagnosed into the lesion retrieval and matching system after completing the training of all sub-networks, and use it as an extended database for a supplementary data set;

[0015] Perform a CT scan on the patient to be diagnosed. Use several key imaging positions as the input to the lesion detection network FasterRCNN, and use the output lesion position as the basis for intercepting the image to intercept the original CT image of the lesion detection to obtain the CT image of the lesion;

[0016] Input the CT image of the lesion into the CT scan image encoder. At the same time, input the auxiliary personal information of the patient to be diagnosed into the auxiliary information encoder. Use the values of the two encoders as Q, K, and V of the sub-attention mechanism and input them into the CT scan image decoder respectively. The CT scan image decoder outputs the hidden code for lesion judgment and the specific judgment of the lesion respectively. The hidden code for lesion judgment is input into the similar lesion retrieval system to retrieve other relevant cases. Finally, assist the doctor in diagnosing the patient by combining the relevant cases and the specific judgment of the lesion situation by the CT scan image decoder;

[0017] Input the diagnosis result into the patient information management system as a supplement to the database. At the same time, continuously track the patient's situation and record the patient's subsequent treatment status as the label of the new dataset to continuously optimize the retrieval matching system.

[0018] Preferably, confirm the specific clinical symptoms of the patient and collect various quantitative indicators of the patient's physical examination, specifically:

[0019] Perform a CT scan on the user and collect the lesion area images of the patient from different imaging positions.

[0020] Preferably, collect the CT images of the lesions of past patients and the lesion areas determined after being judged by doctors, mark the lesion areas, and record the specific type information of the lesions as category labels, specifically:

[0021] Based on the collected object detection dataset, train the Faster R-CNN network until convergence. Stretch the lesion area images and adjust the contrast and brightness for data augmentation. After multiple trainings, use the object detection network with the best performance on the test set as the lesion area detection network of this method to test its accuracy in detecting the lesion position and ensure that its accuracy is close to or exceeds the detection indicators of normal doctors.

[0022] Preferably, train the auxiliary information encoder, specifically:

[0023] The auxiliary information encoder is based on the encoder structure of the Transformer. First, convert the patient's various physical indicators and living habit information into one-hot tensors, and then input the one-hot tensor information into the FC for word embedding. Then, use the embedded tensor as the representation result of the user's auxiliary information;

[0024] Use the representation result of the user auxiliary information as the K and V of the Transformer-based CT image decoder, and use the encoding of the CT scan image as Q. Input them into the Transformer decoder, and through the decoder and the subsequent similar lesion retrieval system, obtain the judgment of the nature of the lesion by the entire network, and retrieve similar lesions as the output of the network structure. Do regression on the output and the previous cases of different users collected, and train the auxiliary information encoder by gradient descent until it converges.

[0025] Correspondingly, the present invention also provides a system for retrieving and matching similar lesions, including:

[0026] A data collection unit for confirming the specific clinical symptoms of the patient and collecting various quantitative indicators of the patient's physical examination; collecting the CT images of the lesions of past patients and the lesion areas determined after the doctor's judgment, marking the lesion areas, and recording the specific type information of the lesions as category labels;

[0027] A CT scan image encoder training unit for training the CT scan image encoder. First, input the image obtained after target detection into the CT scan image encoder, fuse the encoding of the CT scan image output by it and the representation result of the user auxiliary information, and then input it into the CT scan image decoder. After obtaining the similar lesion detection output and the lesion judgment output, do regression with the category label. When the accuracy drops until the network converges, complete the training of the CT scan image encoder;

[0028] A CT scan image decoder training unit for training the CT scan image decoder. Referring to the attention module in VIT, input the representation result of the user auxiliary information and the encoding of the CT scan image into the CT scan image decoder. Use the patient auxiliary information as K and V in the attention, and the encoding of the CT scan image as Q. Learn about the patient's illness through the sub-attention mechanism combined with the patient's own information, and directly output the judgment of the user's lesion as a reference for the doctor to qualitatively diagnose the patient's condition. At the same time, input the lesion hidden code obtained by the CT scan image decoder into the similar lesion retrieval system to retrieve similar cases in the patient information file as a reference for the doctor to diagnose the patient's lesion;

[0029] An auxiliary information encoder training unit for training an auxiliary information encoder. The auxiliary information encoder is based on the encoder structure of the Transformer. First, various physical indicators and lifestyle information of the user are converted into one-hot tensors. Then, the one-hot tensor information is input into the FC for word embedding. Next, the embedded tensor is used as the representation result of the user's auxiliary information. Finally, the representation result of the user's auxiliary information and the encoding of the CT scan image are input into the CT image decoder based on the Transformer, and the auxiliary information encoder is trained by gradient descent until it converges;

[0030] A unit for inputting information of a patient to be tested, which is used to input the information of the patient to be diagnosed into the lesion retrieval and matching system after all sub-networks are trained, and use it as an extended database for a supplementary data set;

[0031] A lesion CT image generation unit for performing a CT scan on the patient to be diagnosed, taking several key imaging positions as the input of the lesion detection network FasterRCNN, and using the output lesion position as the basis for intercepting the image to intercept the original CT image of the lesion detection, so as to obtain the CT image of the lesion;

[0032] An auxiliary diagnosis unit for inputting the CT image of the lesion into the CT scan image encoder, and at the same time inputting the auxiliary personal information of the patient to be diagnosed into the auxiliary information encoder. The values of the two encoders are used as Q, K, and V of the sub-attention mechanism and input into the CT scan image decoder respectively. The CT scan image decoder outputs the hidden code for lesion judgment and the specific judgment of the lesion respectively. The hidden code for lesion judgment is input into the similar lesion retrieval system to retrieve other relevant cases. Finally, the relevant cases and the specific judgment of the lesion situation by the CT scan image decoder are combined to assist the doctor in diagnosing the patient;

[0033] A retrieval and matching system optimization unit for inputting the diagnosis result into the patient information management system as a supplement to the database, and at the same time continuously tracking the situation of the patient, recording the subsequent treatment status of the patient as the label of a new data set to continuously optimize the retrieval and matching system.

[0034] Implementing the present invention has the following beneficial effects:

[0035] The present invention designs a patient information management system that can be continuously optimized. Taking specific patients as units, it records various information including the diseases suffered by the patients, and can continuously generate new data to train the sub-networks of the system itself as the system is applied, so as to optimize the parameters of the entire network. The lesion retrieval and matching system not only considers the current disease conditions of the users, but also considers various information that may affect the lesion attributes of the users, such as living habits and eating habits, taking into account more influencing information factors than other vision-based retrieval and matching systems. The present invention introduces the recently emerged VIT structure into the learning of lesion images, and has greater potential to learn the lesion information of users than other convolution-based deep learning methods. Brief Description of the Drawings

[0036] Figure 1 is a flowchart of the method for retrieving and matching similar lesions in an embodiment of the present invention;

[0037] Figure 2 is a structural diagram of the system for retrieving and matching similar lesions in an embodiment of the present invention. Detailed Embodiments

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] Figure 1 is a flowchart of the method for retrieving and matching similar lesions in an embodiment of the present invention, as Figure 1 shown, the method includes:

[0040] S1. Confirm the specific clinical symptoms of the patient and collect various quantitative indicators of the patient's physical examination;

[0041] S2. Collect the past patients' lesion CT images and the lesion areas determined after being judged by doctors, mark the lesion areas, and record the specific type information of the lesions as category labels;

[0042] S3. Train the CT scan image encoder. First, input the image obtained after target detection into the CT scan image encoder, fuse the encoded CT scan image output and the representation result of the user auxiliary information, and then input it into the CT scan image decoder. After obtaining the similar lesion detection output and the lesion judgment output, perform regression with the category label. When the accuracy drops until the network converges, the training of the CT scan image encoder is completed;

[0043] S4. Train the CT scan image decoder. Referring to the attention module in VIT, input the representation result of the user auxiliary information and the encoding of the CT scan image into the CT scan image decoder. Use the patient auxiliary information as K and V in the attention, and the encoding of the CT scan image as Q. Learn about the patient's disease condition by combining the patient's own information through the sub-attention mechanism, and directly output the judgment of the user's lesion as a reference for the doctor to qualitatively diagnose the patient's condition. At the same time, input the lesion hidden code obtained by the CT scan image decoder into the similar lesion retrieval system to retrieve similar cases in the patient information file as a reference for the doctor to confirm the patient's lesion;

[0044] S5. Train the auxiliary information encoder. The auxiliary information encoder is based on the encoder structure of the transformer. First, convert the user's various physical indicators and living habit information into one-hot tensors, then input the one-hot tensor information into the FC for word embedding. Next, use the embedded tensor as the representation result of the user auxiliary information. Finally, input the representation result of the user auxiliary information and the encoding of the CT scan image into the CT image decoder based on the transformer, and train the auxiliary information encoder by gradient descent until it converges;

[0045] S6. Input the information of the patient to be diagnosed into the lesion retrieval and matching system after all sub-networks are trained, and use it as an extended database for a supplementary data set;

[0046] S7. Conduct a CT scan on the patient to be diagnosed. Use several key imaging positions as the input to the lesion detection network FasterRCNN, and use the output lesion position as the basis for intercepting the image to intercept the original CT image of the lesion detection to obtain the CT image of the lesion;

[0047] S8. Input the CT image of the lesion into the CT scan image encoder. At the same time, input the auxiliary personal information of the patient to be diagnosed into the auxiliary information encoder. Use the values of the two encoders as Q, K, and V of the sub-attention mechanism and input them into the CT scan image decoder respectively. The CT scan image decoder outputs the hidden code for lesion judgment and the specific judgment of the lesion respectively. The hidden code for lesion judgment is input into the similar lesion retrieval system to retrieve other relevant cases. Finally, assist the doctor in diagnosing the patient by combining the relevant cases and the specific judgment of the lesion situation by the CT scan image decoder;

[0048] S9. Input the diagnosis results into the patient information management system as a supplement to the database, and continuously track the patient's condition, record the patient's subsequent treatment status as the label of the new dataset, so as to continuously optimize the retrieval and matching system.

[0049] Step S1 is as follows:

[0050] S1-1. Perform a CT scan on the user and collect images of the lesion area of the patient from different imaging positions.

[0051] Step S2 is as follows:

[0052] S2-1. Based on the collected target detection dataset, train the Faster R-CNN network until convergence. Stretch and adjust the contrast and brightness of the lesion area images for data augmentation. After multiple trainings, use the target detection network with the best performance on the test set as the lesion area detection network of this method to test its detection accuracy of the lesion location and ensure that its accuracy is close to or exceeds the detection indicators of normal doctors.

[0053] Step S5 is as follows:

[0054] S5-1. The auxiliary information encoder is based on the encoder structure of the Transformer. First, convert the user's various physical indicators and lifestyle information into one-hot tensors, and then input the one-hot tensor information into the FC for word embedding. Then, use the embedded tensor as the representation result of the user's auxiliary information.

[0055] S5-2. Use the representation result of the user's auxiliary information as the K and V of the Transformer-based CT image decoder, and use the encoding of the CT scan image as Q. Input them into the Transformer decoder, and through the decoder and the subsequent similar lesion retrieval system, obtain the judgment of the nature of the lesion by the entire network, and retrieve similar lesions as the output of the network structure. Perform regression on the output and the previous cases of different users collected, and train the auxiliary information encoder by gradient descent until it converges.

[0056] Correspondingly, the present invention also provides a retrieval and matching system for similar lesions, as Figure 2 shown, including:

[0057] The data collection unit 1 is used to confirm the specific clinical symptoms of the patient and collect various quantitative indicators of the patient's physical examination; collect the CT images of the lesions of past patients and the lesion areas determined by doctors' judgments, mark the lesion areas, and record the specific type information of the lesions as category labels.

[0058] Specifically, a CT scan is performed on the user, and images of the lesion area of the patient are collected from different imaging positions; based on the collected target detection data set, the Faster RCNN network is trained until convergence. The images of the lesion area are stretched and the contrast and brightness are adjusted for data augmentation. After multiple trainings, the target detection network with the best performance on the test set is used as the lesion area detection network of this method to test its accuracy in detecting the lesion position and ensure that its accuracy is close to or exceeds the detection index of a normal doctor.

[0059] The CT scan image encoder training unit 2 is used to train the CT scan image encoder. First, the image obtained after target detection is input into the CT scan image encoder. After fusing the encoded CT scan image output by it and the representation result of the user auxiliary information, it is input into the CT scan image decoder. After obtaining the similar lesion detection output and the lesion judgment output, regression is performed with the category label. When the accuracy drops until the network converges, the training of the CT scan image encoder is completed;

[0060] The CT scan image decoder training unit 3 is used to train the CT scan image decoder. Referring to the attention module in VIT, the representation result of the user auxiliary information and the encoding of the CT scan image are input into the CT scan image decoder. The patient auxiliary information is used as K and V in the attention, and the encoding of the CT scan image is used as Q. The information of the patient himself is learned through the sub-attention mechanism to learn about the patient's disease condition, and the judgment of the user's lesion is directly output as a reference for the doctor to qualitatively diagnose the patient's condition. At the same time, the lesion hidden code obtained by the CT scan image decoder is input into the similar lesion retrieval system to retrieve similar cases in the patient information file as a reference for the doctor to diagnose the patient's lesion;

[0061] The auxiliary information encoder training unit 4 is used to train the auxiliary information encoder. The auxiliary information encoder is based on the encoder structure of the transformer. First, various physical indicators and lifestyle information of the user are converted into one-hot tensors, and then the one-hot tensor information is input into the FC for word embedding. Then, the embedded tensor is used as the representation result of the user auxiliary information. Finally, the representation result of the user auxiliary information and the encoding of the CT scan image are input into the CT image decoder based on the transformer, and the auxiliary information encoder is trained by gradient descent until it converges;

[0062] Specifically, first, various physical indicators and lifestyle information of the user are converted into one-hot tensors. Then, the one-hot tensor information is input into the FC for word embedding. Next, the embedded tensor is used as the representation result of the user auxiliary information. The representation result of the user auxiliary information is used as the K and V of the transformer-based CT image decoder, and the encoding of the CT scan image is used as Q, which is input into the transformer decoder. Through the decoder and the subsequent similar lesion retrieval system, the entire network makes a judgment on the nature of the lesion, and retrieves similar lesions as the output of the network structure. The output is regressed with the previous cases of different users collected, and the auxiliary information encoder is trained by gradient descent until it converges.

[0063] The information input unit 5 of the patient to be tested is used to input the information of the patient to be diagnosed into the lesion retrieval and matching system after all sub-networks are trained, and use it as an extended database for a supplementary data set.

[0064] The lesion CT image generation unit 6 is used to perform a CT scan on the patient to be diagnosed, take several key imaging positions as the input of the lesion detection network FasterRCNN, and use the output lesion position as the basis for intercepting the image to intercept the original CT image of the lesion detection, so as to obtain the CT image of the lesion.

[0065] The auxiliary diagnosis unit 7 is used to input the CT image of the lesion into the CT scan image encoder, and at the same time input the auxiliary personal information of the patient to be diagnosed into the auxiliary information encoder. The values of the two encoders are used as Q, K, and V of the sub-attention mechanism and input into the CT scan image decoder respectively. The CT scan image decoder outputs the hidden code for lesion judgment and the specific judgment of the lesion respectively. The hidden code for lesion judgment is input into the similar lesion retrieval system to retrieve relevant other cases. Finally, combined with the relevant cases and the specific judgment of the CT scan image decoder on the lesion situation, it assists the doctor in diagnosing the patient.

[0066] The retrieval and matching system optimization unit 8 is used to input the diagnosis result into the patient information management system as a supplement to the database, and continuously track the situation of the patient, record the subsequent treatment status of the patient as the label of the new data set, so as to continuously optimize the retrieval and matching system.

[0067] Therefore, through a patient information management system that can be continuously optimized, the present invention records various information including the diseases suffered by patients in units of specific patients, and can continuously generate new data to train the sub-network of the system itself as the system is applied, so as to optimize the parameters of the entire network. The lesion retrieval and matching system of the present invention not only considers the current disease conditions of users, but also considers various information that may affect the lesion attributes of users, such as living habits and eating habits, considering more influencing information factors than other vision-based retrieval and matching systems. The present invention introduces the recently emerged VIT structure into the learning of lesion images, and has greater potential to learn the lesion information of users than other convolutional-based deep learning methods.

[0068] The above has introduced in detail the method and system for retrieving and matching similar lesions provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for retrieving and matching similar lesions, characterized in that The method includes: Confirm the specific clinical symptoms of the patient and collect various quantitative indicators of the patient's physical examination; Collect the CT images of the lesions of past patients and the lesion areas determined by doctors' judgments, label the lesion areas, and record the specific type information of the lesions as category labels; Convert the patient's various physical indicators and lifestyle information into one-hot tensors, then input the one-hot tensor information into the fully connected network FC for word embedding, and then use the embedded tensor as the representation result of the patient's auxiliary information; Train the CT scan image encoder. First, input the image obtained after object detection into the CT scan image encoder, fuse the encoded CT scan image output and the representation result of the patient's auxiliary information, and then input it into the CT scan image decoder. After obtaining the similar lesion detection output and lesion judgment output, perform regression with the category label. When the accuracy drops until the network converges, complete the training of the CT scan image encoder; Train the CT scan image decoder. Referring to the attention module in VIT, input the representation result of the patient's auxiliary information and the encoded CT scan image into the CT scan image decoder. Use the patient's auxiliary information as K and V in the attention, and the encoded CT scan image as Q. Learn the patient's disease condition by combining the patient's own information through the sub-attention mechanism, and directly output the judgment of the patient's lesion as a reference for the doctor to qualitatively diagnose the patient's condition. At the same time, input the lesion hidden code obtained by the CT scan image decoder into the similar lesion retrieval system to retrieve similar cases in the patient information file as a reference for the doctor to confirm the patient's lesion; Train the auxiliary information encoder. The auxiliary information encoder is based on the encoder structure of the transformer. Use the representation result of the patient's auxiliary information as K and V of the CT image decoder based on the transformer, and the encoded CT scan image as Q, and input it into the transformer decoder. Then, through the decoder and the subsequent similar lesion retrieval system, obtain the judgment of the nature of the lesion by the entire network, and retrieve similar lesions as the output of the network structure. Perform regression with the collected cases of different patients before, and train the auxiliary information encoder by gradient descent until it converges; Input the information of the patient to be diagnosed into the lesion retrieval and matching system after completing the training of all sub-networks, and use it as an extended database for a supplementary data set; Perform a CT scan on the patient to be diagnosed. Use several key imaging positions as the input of the lesion detection network FasterRCNN, and use the output lesion position as the basis for intercepting the image to intercept the original CT image of the lesion detection to obtain the CT image of the lesion; Input the CT image of the lesion into the CT scan image encoder. At the same time, input the auxiliary personal information of the patient to be diagnosed into the auxiliary information encoder. Use the values of the two encoders as Q, K, and V of the sub-attention mechanism and input them into the CT scan image decoder respectively. The CT scan image decoder outputs the hidden code for lesion judgment and the specific judgment of the lesion respectively. The hidden code for lesion judgment is input into the similar lesion retrieval system to retrieve other relevant cases. Finally, combine the relevant cases and the specific judgment of the lesion situation by the CT scan image decoder to assist doctors in diagnosing the patient; Input the diagnosis result into the patient information management system as a supplement to the database. At the same time, continuously track the patient's situation and record the patient's subsequent treatment status as the label of the new dataset to continuously optimize the retrieval matching system.

2. The method for retrieving and matching similar lesions according to claim 1, wherein Confirm the specific clinical symptoms of the patient and collect various quantitative indicators of the patient's physical examination. Specifically: Perform a CT scan on the patient and collect the lesion area images of the patient from different positions.

3. The method for retrieving and matching similar lesions according to claim 1, characterized in that, Collect the CT images of the lesions of past patients and the lesion areas determined after being judged by doctors, mark the lesion areas, and record the specific type information of the lesions as category labels. Specifically: Based on the collected object detection dataset, train the Faster RCNN network until convergence. Stretch and adjust the contrast and brightness of the lesion area images for data augmentation. After multiple trainings, use the object detection network with the best effect on the test set as the lesion area detection network of this method to test its accuracy in detecting the lesion position and ensure that its accuracy is close to or exceeds the detection indicators of normal doctors.

4. A retrieval and matching system for similar lesions, characterized in that, The system includes: A data collection unit, which is used to confirm the specific clinical symptoms of the patient and collect various quantitative indicators of the patient's physical examination; collect the CT images of the lesions of past patients and the lesion areas determined after being judged by doctors, mark the lesion areas, and record the specific type information of the lesions as category labels; convert the patient's various physical indicators and lifestyle information into one-hot tensors, then input the one-hot tensor information into the fully connected network FC for word embedding, and then use the embedded tensor as the representation result of the patient's auxiliary information; A CT scan image encoder training unit, which is used to train the CT scan image encoder. First, input the image obtained after object detection into the CT scan image encoder, fuse the encoded CT scan image output and the representation result of the patient's auxiliary information, and then input them into the CT scan image decoder. After obtaining the similar lesion detection output and the lesion judgment output, perform regression with the category label. When the accuracy drops to network convergence, complete the training of the CT scan image encoder; A CT scan image decoder training unit for training a CT scan image decoder. Referring to the attention module in ViT, the representation result of the patient auxiliary information and the encoding of the CT scan image are input into the CT scan image decoder. The patient auxiliary information is used as K and V in the attention, and the encoding of the CT scan image is used as Q. Through the sub-attention mechanism, the patient's disease condition is learned by combining the patient's own information, and a judgment on the patient's lesion is directly output as a reference for the doctor to qualitatively diagnose the patient's condition. At the same time, the lesion latent code obtained by the CT scan image decoder is input into the similar lesion retrieval system to retrieve similar cases in the patient information file as a reference for the doctor to diagnose the patient's lesion; An auxiliary information encoder training unit for training an auxiliary information encoder. The auxiliary information encoder is based on the encoder structure of the Transformer. The representation result of the patient auxiliary information is used as K and V of the Transformer-based CT image decoder, and the encoding of the CT scan image is used as Q and input into the Transformer decoder. Through the decoder and the subsequent similar lesion retrieval system, a judgment on the nature of the lesion is obtained for the entire network, and similar lesions are retrieved as the output of the network structure. The output is regressed with the previous cases of different patients collected, and the auxiliary information encoder is trained by the gradient descent method until it converges; A unit for inputting information of a patient to be tested, which is used to input the information of the patient to be diagnosed into the lesion retrieval and matching system after all sub-networks are trained, and use it as an extended database for a supplementary data set; A lesion CT image generation unit for performing a CT scan on the patient to be diagnosed, taking several key imaging positions as the input of the lesion detection network Faster RCNN, and using the output lesion position as the basis for intercepting the image to intercept the original CT image of the lesion detection to obtain the CT image of the lesion; An auxiliary diagnosis unit for inputting the CT image of the lesion into the CT scan image encoder, and at the same time inputting the auxiliary personal information of the patient to be diagnosed into the auxiliary information encoder. The values of the two encoders are used as Q, K, and V of the sub-attention mechanism and input into the CT scan image decoder respectively. The CT scan image decoder outputs the latent code for lesion judgment and the specific judgment of the lesion respectively. The latent code for lesion judgment is input into the similar lesion retrieval system to retrieve other relevant cases. Finally, the relevant cases and the specific judgment of the lesion situation by the CT scan image decoder are combined to assist the doctor in diagnosing the patient; A unit for optimizing the retrieval and matching system, which is used to input the diagnosis result into the patient information management system as a supplement to the database, and continuously track the patient's situation, record the patient's subsequent treatment status as the label of a new data set to continuously optimize the retrieval and matching system.

5. The retrieval and matching system for similar lesions according to claim 4, characterized in that, The data collection unit needs to perform a CT scan on the patient and collect images of the lesion area of the patient from different imaging positions; based on the collected target detection data set, train the Faster R-CNN network until convergence, stretch the lesion area images and adjust the contrast and brightness for data augmentation. After multiple trainings, use the target detection network with the best performance on the test set as the lesion area detection network of this method to test its accuracy in detecting the lesion position and ensure that its accuracy is close to or exceeds the detection indicators of normal doctors.

Citation Information

Patent Citations

  • Focus matching method and device

    CN110674885A

  • Auxiliary cancer diagnosis method based on digital pathological images

    CN105975793A

  • Illness state visual prediction system and method, computer equipment and storage medium

    CN111785376A