Multi-modal medical image anomaly detection method, medium and equipment
By receiving medical images in the neural network model and generating text description data sets, combining attention mechanism and multi-layer feature adapter, the accuracy problem when the number of samples is small in medical image abnormality detection is solved, and efficient abnormality detection and segmentation is achieved.
Patent Information
- Application Number
- CN202510209647.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When detecting medical image abnormalities, the text description and abnormal area positioning are not accurate enough when the number of samples is small.
By receiving medical images, inputting the trained neural network model, outputting abnormal detection results, including text description information and abnormal feature maps. When training, the neural network model generates text description data sets based on templates and large language models, extracts key features of the image through attention mechanism, and uses multi-layer feature adapters for dynamic learning, and calculates cosine similarity to obtain text description information and anomaly feature maps.
It improves the accuracy and efficiency of medical image abnormality detection, can achieve efficient abnormality detection and segmentation with few sample data, and reduces the cost of collecting and labeling of medical data.
Smart Images

Figure CN120047749A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a multi-modal medical image anomaly detection method, medium, and device. Background Art
[0002] Medical anomaly detection is a crucial task in the field of computer vision. It focuses on identifying deviations in medical image data that are different from most medical images and locating the positions of these lesions, which can provide auxiliary insights for medical staff in medical decision-making and improve the efficiency of medical personnel. However, in practical applications, it is often difficult to collect medical data, and a large number of professionals are required for annotation. This makes existing methods often require thousands or more data for model training, increasing the development cost and resource consumption. Therefore, how to use few-shot data for model training while achieving good performance can also avoid the collection and annotation of medical data and cost consumption. In addition, in different medical fields, the huge differences in medical images make it difficult to have a general generalization model that performs well on all data, and most of them need to be trained or adjusted for specific data.
[0003] Currently, large-scale pre-trained vision-language models have demonstrated their powerful generalization ability by mapping natural images and raw text to a unified space and have performed excellently in different downstream tasks. Nevertheless, existing methods rely on fixed text descriptions of images. The design of manual prompts not only requires a large amount of domain knowledge and time cost but also faces the problem of semantic ambiguity. In addition, based on vision-language models, global anomalies can be obtained by aligning image-level visual features with anomaly prompts, but they still have deficiencies in accurate anomaly localization. Therefore, how to use vision-language models to obtain local anomalies has become a crucial issue. Summary of the Invention
[0004] In view of the above problems, the present invention provides a technical solution for multi-modal medical image anomaly detection to solve the problem that existing neural network models have inaccurate text descriptions and anomaly region localization when detecting anomalies in medical images with a small number of samples.
[0005] To achieve the above object, in a first aspect, the present application provides a multi-modal medical image anomaly detection method, including the following steps:
[0006] Receive a medical image;
[0007] Input the medical image into a trained neural network model, and output an anomaly detection result corresponding to the medical image. The anomaly detection result includes text description information on whether the image is abnormal and an anomaly feature map;
[0008] When the neural network model is trained, it includes the following steps:
[0009] S1: Generate text description datasets of different categories based on templates and large language models. The text description datasets include a normal text description set, an abnormal text description set, a normal text generation set, and an abnormal text generation set;
[0010] S2: Receive medical sample images, calculate the cosine similarities between the medical sample images and the normal text description set, the abnormal text description set, the normal text generation set, and the abnormal text generation set respectively, and calculate correlation scores based on the cosine similarities;
[0011] If the calculated correlation score is greater than 0, retain the corresponding text description; otherwise, delete the corresponding text description until all the data in the text description datasets are traversed, and obtain a filtered and accurate text description dataset that is most relevant to the currently input medical sample image. The accurate text description dataset includes an accurate normal text description set and an accurate abnormal text description set;
[0012] S3: Extract the key features of the medical sample image through an attention mechanism, input the key features into a multi-layer feature adapter for dynamic learning to obtain deep features, and calculate the cosine similarity between the deep features and the data in the accurate text description dataset to obtain the text description information and abnormal feature map corresponding to the current medical sample image.
[0013] Optionally, the normal text description set and the abnormal text description set are generated based on templates, and the template is a text description information describing a certain disease type and health status. The normal text generation set and the abnormal text generation set are generated based on the feedback results of the large language model for the question information, and the question information includes a certain disease type and health status;
[0014] Mark the normal text description set, the abnormal text description set, the normal text generation set, and the abnormal text generation set as Tn, Ta, Gn, and Ga respectively. Then the calculation of the cosine similarity includes:
[0015] d a (x) = <f(x), g(t a )>, t a ∈{T a , G a};
[0016] d n (x) = <f(x), g(t n )>, t n ∈{T n , G n};
[0017] Among them, d a (x) represents the cosine similarity between the input medical sample image x and the abnormal text dataset, and d n (x) represents the cosine similarity between the input medical sample image x and the normal text dataset.
[0018] f(x) ∈ R d , where f(x) refers to the visual embedding features obtained by the pre-trained visual encoder in the neural network model, d is the dimension of the latent space, and R d represents a real vector space of dimension, and g(t a ) refers to the abnormal text embedding features obtained by the pre-trained text encoder, and g(t n ) refers to the normal text embedding features obtained by the pre-trained text encoder. < > represents the cosine similarity function, and the cosine similarity function is expressed as follows:
[0019]
[0020] Among them, f(x)·g(t) represents the inner product of f(x) and g(t), ||f(x)|| represents the norm of f(x), and ||g(t)|| is the norm of g(t).
[0021] Optionally, the correlation score is calculated based on the cosine similarity through the following formula:
[0022]
[0023] S(t) represents the correlation score, and its value range is [0,1). Among them, L is expressed as follows:
[0024] L = ||D(d x (t), d n (x)) - D(d x (t), d a (x))||;
[0025] Among them, d x (t) refers to the cosine similarity between the text description t and the medical sample image x, k is a hyperparameter that controls the slope of the function, and the function D() is used to calculate d x (t) and d n (t), or d x (t) and d a (x) distances, which are expressed as follows:
[0026] D(d x (t), d n (x)) = max(0, max(min{d n (x)}-d x(t), d x (t)) - max{d n (x)}));
[0027] D(d x (t), da(x)) = max(0, max(min{d a (x} - d x (t), d x (t) - max{d a (x}))。
[0028] Optionally, the key features of the medical sample image are extracted through an attention mechanism and calculated according to the following formula:
[0029] Attn(Q, K, V) = softmax(Q · K · scale) · V;
[0030] where Q, K, and V are vector representations obtained through specific transformations or extracted from the input data, used to calculate the attention weights and obtain the weighted output, softmax represents the activation function, used to convert the input values into a probability distribution, and scale represents the scaling factor.
[0031] Optionally, the key features are input into a multi-layer feature adapter for dynamic learning to obtain deep features including:
[0032]
[0033] Z l-1 = [t′ cls ; t′ 1 ; t′ 2 , …; t′ T ;
[0034]
[0035] Z l = Proj. l (Attn(V l , V l , V l )) + z l-1 。
[0036] where represents the output of the L - 1 layer of the visual encoder, Z l-1 represents the local feature output of the L - 1 layer, QKV_Proj. l and Proj. l represent the QKV projection and the output projection respectively, and the output of the last layer is the original output Z o and the local output Z, and the category feature Z o[0] ∈ R d For image - level anomaly detection, the local feature map Z[1:] ∈ R Txd For pixel - level anomaly detection;
[0037] For the input medical sample image x ∈ R hxwx3 , the medical sample image x is converted into the feature space Z l ∈ R Gxd , l ∈ {1, 2, 3, 4}, where h and w represent the size of the image, G represents the number of grids, and d represents the feature dimension;
[0038] The multi - layer feature adapter includes four adapters, denoted as A l (·), l ∈ {1, 2, 3, 4}, at each stage S l , l ∈ {1, 2, 3, 4}, the feature adapter is integrated into the feature Z through two linear transformation layers l as follows:
[0039]
[0040] where Relu represents the activation function and Linear represents the linear transformation layer;
[0041] After obtaining the features passing through the multi - layer feature adapter, residual connections are used to retain the pre - trained original features. Then the adapted feature of the L - th layer can be expressed as:
[0042] Z′ l = αA l (Z l ) T +(1 - α)Z l , l ∈ {1, 2, 3, 4} (11)
[0043] where Z′ l serves as the input for the next stage S l+1 , and the constant value α serves as the residual ratio to adjust the degree of retaining the original features;
[0044] The feature adapter is a two - branch feature adapter, generating two parallel features F cls,l and F seg,l at each layer, where F cls,l represents the feature for the classification task generated by the two - branch feature adapter at the L - th layer, used to indicate whether the input medical sample image x is abnormal, and F seg,l represents the feature for the segmentation task generated by the two - branch feature adapter at the layer, used to segment the abnormal region features;
[0045] Calculate the cosine similarity between the depth features and the data in the accurate text description dataset to obtain the text description information and abnormal feature map corresponding to the current medical sample image, including:
[0046]
[0047] Where F ∈ {F cls,l , F seg,l}, τ is the temperature hyperparameter used to adjust the output distribution of the softmax function, F T represents the text description feature data in the accurate text description dataset, F n represents the normal text description feature data, F a represents the abnormal text description feature data, exp<> represents the exponential function;
[0048]
[0049] Where, C 1 represents the classification index of the abnormal type, S 1 represents the segmentation index of the abnormal feature map, BI() represents the image processing function used to reshape the abnormal map into and restore it to the resolution of the original input medical sample image using bilinear interpolation, G represents the number of grids;
[0050]
[0051]
[0052] Where, m represents the minimum distance between the features of each level recorded in the multi-level feature memory bank G, the function Dist() represents the cosine distance, and the final predicted classification and segmentation results are:
[0053] C pre = βC 1 + (1 - β)C 2 ;
[0054] S pre = βS 1 + (1 - β)S 2 ;
[0055] Where, C pre is the finally determined image abnormal type, S pre is the finally determined abnormal feature map, and β represents the weight value.
[0056] Optionally, the loss function of the neural network model is obtained in the following manner:
[0057] L 1 = θ 1Dice(softmax(F seg,l ,F T ),S gt +θ 2 Focal(softmax(F seg,l ,F T ),S gt )+θ 3 BCE(max(softmax(F cls,l ,F T )),C gt );
[0058] Among them, the Dice loss is used to measure the similarity between two samples and is defined as:
[0059]
[0060] The larger the Dice coefficient, the higher the similarity;
[0061] Focal(softmax(F seg,l ,F T ),S gt )=-αS gt (1 - softmax(F seg,l ,F T )) γ log(softmax(F seg,l ,F T )) - α(1 - S gt )S gt (1 - softmax(F seg,l ,F T )) γ log(softmax(F seg,l ,F T ));
[0062] Among them, α is the class balance factor, which is used to reduce the impact of class imbalance; γ is the focal factor, which is used to adjust the weights of easy and difficult samples; S gt represents the true label;
[0063] BCE(max(softmax(F cls,l ,F T )),C gt )= - [C gt log(max(softmax(F cls,l ,F T )))+(1 - C gt )log(1 - max(softmax(F cls,l ,F T )))];
[0064] The BCE loss is used to measure the difference between the predicted probability distribution and the true label, where C gt is the true label, taking values of 0 or 1, and max(softmax(F cls,l , F T )) is the probability predicted by the model. The smaller the loss, the better the model performance;
[0065] The total loss of the model is expressed as follows:
[0066] where L1 represents the total loss of the hierarchy, and θ 1 , θ 2 , θ 3 represent hyperparameters, which are optimized for the neural network model by being set to different ratios.
[0067] Optionally, when the neural network model is being trained, it further includes the following steps:
[0068] The accuracy of the neural network model is judged by AUC, and the calculation formula is as follows:
[0069]
[0070] where TPR represents the true positive rate, FPR represents the false positive rate, TP represents the number of pixels correctly detected as abnormal in the input medical sample image, FN represents the number of abnormal pixels misdetected as normal in the input medical sample image, FP represents the number of normal pixels misdetected as abnormal in the medical sample image, and TN represents the number of pixels correctly detected as normal in the medical sample image;
[0071]
[0072] In a second aspect, the present application provides a multi-modal medical image anomaly detection system, including:
[0073] An image receiving module for receiving medical images;
[0074] An anomaly detection module for inputting the medical image into the trained neural network model and outputting the anomaly detection result corresponding to the medical image, where the anomaly detection result includes text description information on whether the image is abnormal and an anomaly feature map;
[0075] When the neural network model is being trained, it includes:
[0076] A text data generation module for generating text description datasets of different categories based on templates and large language models, where the text description datasets include a normal text description set, an abnormal text description set, a normal text generation set, and an abnormal text generation set;
[0077] A correlation analysis module for receiving medical sample images, calculating the cosine similarities between the medical sample images and the normal text description set, the abnormal text description set, the normal text generation set, and the abnormal text generation set respectively, and calculating correlation scores based on the cosine similarities;
[0078] If the calculated correlation score is greater than 0, the corresponding text description is retained; otherwise, the corresponding text description is deleted until all the data in the text description datasets are traversed, obtaining a filtered and accurate text description dataset most relevant to the currently input medical sample image, where the accurate text description dataset includes an accurate normal text description set and an accurate abnormal text description set;
[0079] An attention mechanism module for extracting key features of the medical sample image through the attention mechanism;
[0080] A multi-layer feature adapter for dynamically learning the key features to obtain deep features, and calculating the cosine similarity between the deep features and the data in the accurate text description dataset to obtain text description information and abnormal feature maps corresponding to the current medical sample image.
[0081] In a third aspect, the present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, implement the method described in the first aspect.
[0082] In a fourth aspect, the present invention also provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and the one or more computer program instructions are executed by the processor to implement the method described in the first aspect.
[0083] Different from the prior art, the above solution provides a multi-modal medical image anomaly detection method, medium and device. The method includes: receiving a medical image, inputting the medical image into a trained neural network model, and outputting an anomaly detection result corresponding to the medical image. When training the neural network model, a text description data set is generated based on a large language model and templates. By learning the latent distances and relationships between texts, the interval boundaries are implicitly optimized to distinguish between the normal text description prompt interval and the abnormal prompt text interval, and a more accurate text description data set is screened out, thereby alleviating the semantic ambiguity problem. In addition, the present application also introduces an attention mechanism. Without changing the original architecture of the pre-trained model, local features of the image visual features are obtained, and a multi-layer adapter is used to fine-tune the vision-language model to meet the needs of medical image anomaly detection. Anomaly detection and segmentation are achieved by obtaining the global anomaly score and the local anomaly map based on the calculated visual features and text features.
[0084] The above description of the invention content is only an overview of the technical solution of the present invention. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of the present invention, and thus can be implemented according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features and advantages of the present invention more easily understood, the following is described in conjunction with the specific embodiments and drawings of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] The drawings are only used to illustrate the principles, implementation methods, applications, features and effects of the specific embodiments of the present invention and other related contents, and should not be considered as a limitation of the present invention.
[0086] In the accompanying drawings of the specification:
[0087] Figure 1 is a flowchart of the multi-modal medical image anomaly detection method according to the first exemplary embodiment of the present invention;
[0088] Figure 2 is a flowchart of the multi-modal medical image anomaly detection method according to the second exemplary embodiment of the present invention;
[0089] Figure 3 is a schematic diagram of the modules of the multi-modal medical image anomaly detection system according to a specific embodiment of the present invention;
[0090] Figure 4 is a schematic diagram of the modules of the neural network model according to a specific embodiment of the present invention;
[0091] Figure 5 is a schematic diagram of an electronic device according to a specific embodiment of the present invention;
[0092] Figure 6The schematic diagram of the neural network model training involved in a specific embodiment of the present invention;
[0093] Figure 7 The schematic flowchart of the neural network model testing involved in a specific embodiment of the present invention;
[0094] Figure 8 The schematic diagram of the principle of the attention mechanism involved in a specific embodiment of the present invention;
[0095] Figure 9 The schematic diagram of the principle of the multi - layer feature adapter involved in a specific embodiment of the present invention;
[0096] Figure 10 The detailed flowchart of the text boundary optimization involved in a specific embodiment of the present invention.
[0097] The descriptions of the reference numerals involved in the above - mentioned drawings are as follows:
[0098] 10. Electronic device;
[0099] 101. Processor;
[0100] 102. Storage medium.
[0101] 30. Multi - modal medical image anomaly detection system;
[0102] 301. Image receiving module;
[0103] 302. Anomaly detection module;
[0104] 40. Neural network model;
[0105] 401. Text data generation module;
[0106] 402. Correlation analysis module;
[0107] 403. Attention mechanism module;
[0108] 404. Multi - layer feature adapter. Specific embodiments
[0109] To illustrate in detail the possible application scenarios, technical principles, specific implementable solutions, achievable purposes and effects of the present invention, the following is a detailed description with reference to the specific examples listed and in conjunction with the drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of the present invention, so they are only examples and cannot be used to limit the protection scope of the present invention.
[0110] References to "embodiments" in this document mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The term "embodiment" appearing in various positions in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in the present invention, as long as there is no technical contradiction or conflict, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.
[0111] Unless otherwise defined, the meanings of the technical terms used in this document are the same as those commonly understood by those skilled in the technical field to which the present invention belongs; the use of the relevant terms in this document is only for describing specific embodiments and is not intended to limit the present invention.
[0112] In the description of the present invention, the term "and / or" is an expression used to describe the logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: the existence of A, the existence of B, and the simultaneous existence of A and B. In addition, the character " / " in this document generally represents an "or" logical relationship between the associated objects before and after.
[0113] In the present invention, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantitative, primary-secondary, or sequential relationship between these entities or operations.
[0114] Without further limitation, in the present invention, the open-ended expressions such as "including", "comprising", "having", or other similar expressions used in the statements are intended to cover non-exclusive inclusion. These expressions do not exclude the possibility that there may be additional elements in the process, method, or product including the said elements, so that the process, method, or product including a series of elements may include not only those defined elements, but also other elements not explicitly listed, or elements inherent to such process, method, or product.
[0115] In the present invention, expressions such as "greater than", "less than", "exceeding", etc. are understood not to include the number itself; expressions such as "above", "below", "within", etc. are understood to include the number itself. In addition, in the description of the embodiments of the present invention, the meaning of "multiple" is two or more (including two), and similar expressions related to "many", such as "multiple groups", "multiple times", etc., are understood in the same way, unless otherwise specifically defined.
[0116] In a first aspect, as Figure 1 shown, the present application provides a multi-modal medical image anomaly detection method, including the following steps:
[0117] Step S201: Receive a medical image;
[0118] Step S202: Input the medical image into a trained neural network model, and output an anomaly detection result corresponding to the medical image. The anomaly detection result includes text description information on whether the image is abnormal and an anomaly feature map.
[0119] In short, the medical image is the medical image to be analyzed and processed. The abnormal text description information includes two parts. One is to analyze whether the current medical image is abnormal. The other is when the medical image is an abnormal image, it will explain the abnormal type, abnormal area, etc.
[0120] To improve the accuracy of the anomaly detection result recognized by the neural network model, as Figure 2 shown, in this application, when the neural network model is trained, it includes the following steps:
[0121] S1: Generate text description datasets of different categories based on templates and large language models. The text description datasets include a normal text description set, an abnormal text description set, a normal text generation set, and an abnormal text generation set.
[0122] In step S1, the normal text description set and the abnormal text description set are generated based on templates. The template is text description information describing a certain disease type and health status. The medical image category refers to different disease types, body parts, or pathological states involved in the medical image, etc. For example, it may include brain disease categories (such as brain tumors, cerebral hemorrhages, etc.), chest disease categories (such as lung cancer, pneumonia, etc.), bone disease categories (such as fractures, osteoporosis, etc.), etc. In this application, when generating the text description set based on the template, the template of the text description can be "a healthy picture of [category]" or "a diseased picture of [category]". The [category] in this range is the specific medical image category. By clarifying the category, text information related to this category can be constructed targeted, so as to better associate and analyze with the medical images of the corresponding category in the follow-up, so as to realize the anomaly detection and segmentation of different types of medical images.
[0123] The normal text generation set and the abnormal text generation set are generated based on the feedback results of the large language model for the question information, where the question information includes a certain disease type and health status. Specifically, the normal text generation set and the abnormal text generation set can be obtained by asking questions to the large language model. For example, questions such as "[Category] What abnormalities may occur?" and "Describe the situation of normal / abnormal images regarding [Category]" can be asked to the large language model, and the large language model generates corresponding text descriptions based on the knowledge and language patterns it has learned. The normal text generation set Gn can be constructed based on the answers in the normal state, and the abnormal text generation set Ga can be constructed based on the answers in the abnormal state. Utilizing the powerful language generation ability of the large language model, more diverse and rich text descriptions can be obtained, making up for the limitations that may exist in generating only by templates, further expanding the data source of text descriptions for analyzing medical images, and helping to improve the adaptability and accuracy of abnormal analysis of medical images in different application scenarios.
[0124] S2: Receive a medical sample image, calculate the cosine similarity between the medical sample image and the normal text description set, abnormal text description set, normal text generation set, and abnormal text generation set respectively, calculate the correlation score based on the cosine similarity. If the calculated correlation score is greater than 0, retain the corresponding text description; otherwise, delete the corresponding text description until all the data in the text description datasets are traversed, obtaining a filtered and precise text description dataset that is most relevant to the currently input medical sample image. The precise text description dataset includes a precise normal text description set and a precise abnormal text description set.
[0125] In step S2, mark the normal text description set, abnormal text description set, normal text generation set, and abnormal text generation set as Tn, Ta, Gn, and Ga respectively. Then, calculating the cosine similarity includes:
[0126] d a (x) = <f(x), g(t a )>, t a ∈ {T a , G a};
[0127] d n (x) = <f(x), g(t n )>, t n ∈ {T n , G n};
[0128] Among them, d a (x) represents the cosine similarity between the input medical sample image x and the abnormal text dataset, d nCosine similarity between the input medical sample image x and the normal text dataset is denoted as (x).
[0129] f(x) ∈ R d , where f(x) refers to the visual embedding features obtained by the pre-trained visual encoder in the neural network model, d is the dimension of the latent space, and R d represents a real vector space of dimension, and g(t a ) refers to the abnormal text embedding features obtained by the pre-trained text encoder, and g(t n ) refers to the normal text embedding features obtained by the pre-trained text encoder. < > represents the cosine similarity function, and the cosine similarity function is expressed as follows:
[0130]
[0131] where f(x)·g(t) represents the inner product of f(x) and g(t), ||f(x)|| represents the norm of f(x), and the calculation formula is: ||g(t)|| is the norm of g(t), and the calculation formula is similar to that of the norm of f(x).
[0132] The pre-trained visual encoder is a neural network model pre-trained with a large amount of image data. Its main function is to convert the input visual image data into a feature representation with semantic information. For example, it can receive a medical image as input and output a vector of a fixed dimension to represent the features of this image.
[0133] The pre-trained text encoder is a neural network model pre-trained for text data. It can convert the input text into a feature vector containing semantic information. For example, it can receive a medical image as input and output a vector to indicate whether this medical image is abnormal, and better identify the abnormal regions in the image through the disease description in the text.
[0134] The visual embedding features are the output results of the pre-trained visual encoder processing the input medical image x. These features are high-dimensional vectors, and their purpose is to represent the information of the original medical image in a form that is more conducive to subsequent processing and analysis.
[0135] By calculating the cosine similarity, the correlation between medical sample images and text descriptions can be measured. Ideally, the distances between normal text descriptions and medical images, and between abnormal text descriptions and medical images, should be in two non-overlapping interval ranges. However, due to the uncertainty of text description prompts generated by large language models, some abnormal prompt features are closer to normal prompts, resulting in an overlap between the normal text description prompt interval and the abnormal text description prompt interval. These prompts will have a negative impact on the performance of the neural network model.
[0136] Based on this, the present application quantifies the impact of abnormalities in normal prompts through the logits function. Specifically, the correlation score is calculated based on the cosine similarity through the following formula:
[0137]
[0138] S(t) represents the correlation score, with a value range of [0,1). Among them, L is as follows:
[0139] L = |D(d x (t), d n (x)) - D(d x (t), d a (x))||;
[0140] Among them, d x (t) refers to the cosine similarity between the text description t and the medical sample image x. k is a hyperparameter that controls the slope of the function (usually the value of k is set to 1). The function D() is used to calculate the distance between d x (t) and d n (x), or between d x (t) and d a (x), which is expressed as follows:
[0141] D(d x (t), d n (x)) = max(0, max(min{d n (x)}-d x (t), d x (t) - max{d m (x}));
[0142] D(d x (t), d a (x)) = max(0, max(min{d a (x)}-d x (t), d x (t)) - max{d a (x)}));
[0143] Generally, if the distance between the text description prompt and the input medical image spans the positive and negative anomaly intervals, the score stabilizes at 0, indicating a poor correlation between the text description prompt and the input medical image. Therefore, in practical applications, this application filters out the intervals where the correlation score is less than or equal to 0, and only retains the text description prompts in the non-overlapping intervals, that is, T = {t′ a , t′ n}, where t′ a represents the filtered abnormal description text (i.e., the precise abnormal text description set), and t′ n represents the filtered normal description text (i.e., the precise normal text description set). Based on this screening process, a unique set of text features most suitable for the input image can be adaptively selected for each medical image, thereby improving the effectiveness of the model.
[0144] S3: Extract the key features of the medical sample image through the attention mechanism, input the key features into a multi-layer feature adapter for dynamic learning to obtain deep features, and calculate the cosine similarity between the deep features and the data in the precise text description dataset to obtain the text description information and abnormal feature map corresponding to the current medical sample image.
[0145] Some vision-language models (such as CLIP) have poor adaptability in guiding image localization tasks because the global features of Q-K self-attention ignore the importance of local features, which is a fatal defect in pixel-level detection and segmentation. This application introduces a V-V local attention mechanism to enhance the influence of local features without affecting the network structure of the original model. The calculation formula of Q-K self-attention is as follows:
[0146] Attn(Q, K, V) = softmax(Q·K·scale)·V;
[0147] Among them, Q, K, and V are vector representations obtained through specific transformations or extracted from the input data, used to calculate the attention weights and obtain the weighted output. Softmax represents the activation function, used to convert the input numerical values into a probability distribution, and scale represents the scaling factor.
[0148] In some embodiments, inputting the key features into a multi-layer feature adapter for dynamic learning to obtain deep features includes:
[0149]
[0150] Z l-1 = [t′ cls ; t′ 1 ; t′ 2 , …; t′T ;
[0151]
[0152] Z l = Proj. l (Attn(V l , V l , V l )) + Z l-1 ;
[0153] Among them, among them represents the output of the L-1 layer of the visual encoder, and Z l-1 represents the local feature output of the L-1 layer, and QKV_Proj. l and Proj. l represent QKV projection and output projection respectively. The output of the last layer is the original output Z o and the local output Z, and the class feature Z o [0] ∈ R d is used for image-level anomaly detection, and the local feature map Z[1:] ∈ R Txd is used for pixel-level anomaly detection;
[0154] For the input medical sample image x ∈ R hxwx3 , the medical sample image x is converted into the feature space Z l ∈ R Gxd , l ∈ {1, 2, 3, 4}, where h and w represent the size of the image, G represents the number of grids, and d represents the feature dimension;
[0155] The multi-layer feature adapter includes four adapters, denoted as A l (·), l ∈ {1, 2, 3, 4}. At each stage S l , l ∈ {1, 2, 3, 4}, the feature adapter is integrated into the feature Z l through two linear transformation layers, which is expressed as follows:
[0156]
[0157] Among them, ReLu represents the activation function, and Linear represents the linear transformation layer;
[0158] After obtaining the features passing through the multi-layer feature adapter, the residual connection is used to retain the pre-trained original features. Then the adapted feature of the L-th layer can be expressed as:
[0159] Z’ l = αA l (Z l ) T + (1 - α)Z l, l ∈ {1, 2, 3, 4} (11)
[0160] Among them, Z′ l As the input for the next stage S l+1 , the constant value α is used as the residual ratio to adjust the degree of retaining the original features;
[0161] The feature adapter is a two-branch feature adapter, generating two parallel features F cls,l and F seg,l at each layer. Among them, F cls,l represents the feature for the classification task generated by the two-branch feature adapter at the L-th layer, used to indicate whether the input medical sample image x is abnormal, and F seg,l represents the feature for the segmentation task generated by the two-branch feature adapter at the -th layer, used to segment the abnormal region features;
[0162] Performing cosine similarity calculation on the depth feature and the data in the accurate text description dataset, the text description information and abnormal feature map corresponding to the current medical sample image include:
[0163]
[0164] Among them, F ∈ {F cls,l , F seg,l}, τ is the temperature hyperparameter, used to adjust the output distribution of the softmax function, F T represents the text description feature data in the accurate text description dataset, F n represents the normal text description feature data, F a represents the abnormal text description feature data, and exp<> represents the exponential function;
[0165]
[0166] Among them, C 1 represents the classification index of the abnormal type, S 1 represents the segmentation index of the abnormal feature map, and BI() represents an image processing function, used to reshape the abnormal map into and restore it to the resolution of the original input medical sample image using bilinear interpolation. G represents the number of grids;
[0167]
[0168] Among them, m represents the minimum distance between the features of each level recorded in the multi-level feature memory bank G, and the function Dist() represents the cosine distance. The final predicted classification and segmentation results are:
[0169] C pre = βC1 +(1 - β)C 2 ;
[0170] S pre = βS 1 +(1 - β)S 2 ;
[0171] Wherein, C pre is the finally determined image anomaly type, S pre is the finally determined anomaly feature map, and β represents the weight value.
[0172] The above solution uses all multi-level visual features of several normal medical images to construct a multi-level feature memory bank G to facilitate feature comparison with normal features. By comparing the minimum distance between the test features and the feature of each level memory bank through the nearest neighbor search process, the characteristics of anomaly detection and segmentation are improved.
[0173] In some embodiments, the loss function of the neural network model is obtained according to the following manner:
[0174] L 1 = θ 1 Dice(softmax(F seg,l , F T ), S gt ) + θ 2 Focal(softmax(F seg,l , F T ), S gt ) + θ 3 BCE(max(softmax(F cls,l , F T ))), C gt );
[0175] Wherein, the Dice loss is used to measure the similarity between two samples and is defined as:
[0176]
[0177] The larger the Dice coefficient, the higher the similarity;
[0178] Focal(softmax(F seg,l , F T ), S gt ) = -αS gt (1 - softmax(F seg,l , F T )) γ log(softmax(F seg,l , F T )) - α(1 - S gt )Sgt (1 - softmax(F seg,l , F T )) γ log(softmax(F seg,l , F T ));
[0179] Among them, α is the class balance factor, which is used to reduce the impact of class imbalance; γ is the focal factor, which is used to adjust the weights of easy and hard samples; S gt represents the true label;
[0180] BCE(max(softmax(F cls,l , F T ), C gt ) = -[C gt log(max(softmax(F cls,l , F T ))) + (1 - C gt )log(1 - max(softmax(F cls,l , F T )))];
[0181] The BCE loss is used to measure the difference between the predicted probability distribution and the true label. Among them, C gt is the true label, taking values of 0 or 1, and max(softmax(F cls,l , F T )) is the probability predicted by the model. The smaller the loss, the better the model performance;
[0182] The total loss of the model is expressed as follows:
[0183] Among them, L1 represents the total hierarchical loss, and θ 1 , θ 2 , θ 3 represent hyperparameters, which are optimized for the neural network model by being set to different ratios.
[0184] After the above training, a trained neural network model can be obtained. When the neural network model of this application is applied to medical image anomaly analysis, it has the following advantages:
[0185] (1) It can effectively utilize text description information, learn the relationship with medical images, and improve the recognition accuracy;
[0186] (2) It can obtain high efficiency in accuracy by only inputting a small number of samples (such as 2, 4, 8), making up for the defect that the accuracy of the model trained due to less medical image data is not high, and providing help for medical staff.
[0187] (3) Using lightweight multi-layer feature adapters consumes fewer resources and has a fast inference speed, which can be widely applied in actual processes.
[0188] As Figures 6 - 10 shown, this application provides a method for iteratively training a neural network based on a small number of medical abnormal image samples, specifically including the following steps:
[0189] Take Figure 6 the 12-layer image encoder of the multi-modal pre-trained model shown as the basic feature extraction module. Based on this, without changing the original neural network architecture, introduce the V-V attention mechanism. The features obtained after every 3 layers of the V-V attention mechanism for the input medical sample image are input to the corresponding level of the adapter for dynamic learning. As the number of layers increases, more and more deep feature details are retained. Then, calculate the cosine similarity between the features output by the multi-layer adapters (i.e., deep features) and the accurate text description feature set.
[0190] At the same time, use the loss function described above as the loss for image classification and segmentation, and use the SGD optimization algorithm to perform backpropagation correction on the neural network. Iteratively repeat the above steps. When the loss no longer drops significantly or reaches the maximum number of training times, stop training to obtain the trained neural network model.
[0191] For the trained neural network model, when receiving any medical image x (x ∈ X), the neural network model can predict an anomaly score and a pixel-level anomaly feature map. Then, according to the minimum distance between the medical image x and the features of each level of the memory bank, obtain the anomaly score and pixel-level anomaly feature map based on normalcy. Combine this anomaly score and pixel-level anomaly feature map with these two parameters obtained by the model to get the final anomaly score and anomaly feature map. By judging whether the final anomaly score is greater than a certain threshold, determine whether the current medical image x exists, and by comparing the input medical image with the real image pixel by pixel, judge the abnormal area and segment and output it.
[0192] In some embodiments, when the neural network model is being trained, it further includes the following steps:
[0193] Judge the accuracy of the neural network model through AUC, and the calculation formula is as follows:
[0194]
[0195] Among them, TPR represents the true positive rate, FPR represents the false positive rate, TP represents the number of pixels correctly detected as abnormal in the input medical sample image, FN represents the number of abnormal pixels misdetected as normal in the input medical sample image, FP represents the number of normal pixels misdetected as abnormal in the medical sample image, and TN represents the number of pixels correctly detected as normal in the medical sample image;
[0196]
[0197] In a second aspect, as Figure 3 shown, the present application provides a multi-modal medical image abnormality detection system 30, including:
[0198] An image receiving module 301 for receiving medical images;
[0199] An abnormality detection module 302 for inputting the medical image into a trained neural network model and outputting an abnormality detection result corresponding to the medical image, where the abnormality detection result includes text description information on whether the image is abnormal and an abnormality feature map;
[0200] As Figure 4 shown, when the neural network model 40 is being trained, it includes:
[0201] A text data generation module 401 for generating text description data sets of different categories based on templates and large language models, where the text description data sets include a normal text description set, an abnormal text description set, a normal text generation set, and an abnormal text generation set;
[0202] A correlation analysis module 402 for receiving a medical sample image, calculating the cosine similarity between the medical sample image and the normal text description set, the abnormal text description set, the normal text generation set, and the abnormal text generation set respectively, and calculating a correlation score based on the cosine similarity;
[0203] If the calculated correlation score is greater than 0, the corresponding text description is retained, otherwise the corresponding text description is deleted until all the data in the text description data sets are traversed, obtaining a filtered precise text description data set most relevant to the currently input medical sample image, where the precise text description data set includes a precise normal text description set and a precise abnormal text description set;
[0204] An attention mechanism module 403 for extracting key features of the medical sample image through the attention mechanism;
[0205] The multi-layer feature adapter 404 is used to dynamically learn the key features to obtain depth features, and calculate the cosine similarity between the depth features and the data in the accurate text description dataset to obtain the text description information and abnormal feature maps corresponding to the current medical sample image.
[0206] In a second aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the multi-modal medical image anomaly detection method as described in the first aspect of the present invention.
[0207] Among them, the computer-readable storage medium may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories.
[0208] The non-volatile memory may be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD ROM, Compact Disc Read Only Memory); the magnetic surface memory may be a disk memory or a tape memory.
[0209] The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM). The computer-readable storage medium described in the embodiments of the present invention is intended to include these and any other suitable types of memory.
[0210] As Figure 5 shown, in a third aspect, the present invention provides an electronic device 10, including a processor 101 and a storage medium 102, where a computer program is stored on the storage medium, and when the computer program is executed by the processor, the multi-modal medical image anomaly detection method described in the first aspect of the present invention is implemented.
[0211] In some embodiments, the processor may be implemented by software, hardware, firmware, or a combination thereof, and may use at least one of a circuit, a single or multiple application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, and microprocessors, so that the processor can execute some steps, all steps, or any combination of the steps in the multi-modal medical image anomaly detection method described in the various embodiments of the present application.
[0212] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of the present invention, the patent protection scope of the present invention cannot be limited thereby. Any technical solution obtained by replacing or modifying an equivalent structure or equivalent process based on the essential concept of the present invention and using the content recorded in the text and drawings of the specification of the present invention, as well as directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, are all included in the patent protection scope of the present invention.
Claims
1. A multimodal medical image anomaly detection method, characterized in that: The following steps are involved: receiving medical images; Inputting the medical image into the trained neural network model, and outputting an abnormality detection result corresponding to the medical image, wherein the abnormality detection result includes text description information of whether the image is abnormal and an abnormality feature map; When the neural network model is trained, the following steps are included: S1: Generate text description datasets of different categories based on a template and a large language model, wherein the text description datasets include a normal text description set, an abnormal text description set, a normal text generation set, and an abnormal text generation set; S2: receiving a medical sample image, calculating the cosine similarity between the medical sample image and a normal text description set, an abnormal text description set, a normal text generation set, and an abnormal text generation set, respectively, and calculating a correlation score based on the cosine similarity. If the calculated correlation score is greater than 0, the corresponding text description is retained, otherwise the corresponding text description is deleted until the data traversal in all text description data sets is completed, and a filtered accurate text description data set that is most relevant to the currently input medical sample image is obtained, wherein the accurate text description data set includes an accurate normal text description set and an accurate abnormal text description set; S3: extracting the key features of the medical sample image through the attention mechanism, inputting the key features into the multi-layer feature adapter for dynamic learning to obtain deep features, and performing cosine similarity calculation between the deep features and the data in the precise text description dataset to obtain the text description information and abnormal feature map corresponding to the current medical sample image.
2. The multimodal medical image anomaly detection method according to claim 1, characterized in that: The normal text description set and the abnormal text description set are generated based on a template, wherein the template is text description information describing a certain disease type and health status, and the normal text generation set and the abnormal text generation set are generated based on feedback results of a large language model for question information, wherein the question information includes a certain disease type and health status; The normal text description set, abnormal text description set, normal text generation set and abnormal text generation set are marked as Tn, Ta, Gn and Ga respectively, and the calculation of cosine similarity includes: d a (x)=<f(x),g(t a )>,t a ∈{T a ,G a }; d n (x)=<f(x),g(t n )>,t n ∈{T n ,G n }; Among them, d a (x) represents the cosine similarity between the input medical sample image x and the abnormal text dataset, d n (x) represents the cosine similarity between the input medical sample image x and the normal text dataset, f(x)∈R d , f(x) refers to the visual embedding features obtained by the pre-trained visual encoder in the neural network model, d is the dimension of the latent space, and R d represents a one-dimensional real vector space, g(t a ) refers to the abnormal text embedding feature obtained by the pre-trained text encoder, g(t n ) refers to the normal text embedding feature obtained by the pre-trained text encoder, and <> represents the cosine similarity function, which is expressed as follows: Where f(x)·g(t) represents the inner product of f(x) and g(t), ||f(x)|| represents the norm of f(x), and ||g(t)|| is the norm of g(t).
3. The multimodal medical image anomaly detection method according to claim 2, characterized in that: The relevance score based on the cosine similarity is calculated using the following formula: S(t) represents the correlation score, which ranges from [0,1), where L is as follows: L=||D(d x (t),d n (x))-D(d x (t),d a (x)||; Among them, d x (t) refers to the cosine similarity between the text description t and the medical sample image x, k is a hyperparameter that controls the slope of the function, and the function D() is used to calculate d x (t) and d n (x), or d x (t) and d a The distance of (x) is expressed as follows: D(d x (t),d n (x))=max(0,max(min{d n (x)}-d x (t),d x (t)-max{d n (x)})); D(d x (t),d a (x))=max(0,max(min{d a (x)}-d x (t),d x (t)-max{d a (x)}))。 4. The multimodal medical image anomaly detection method according to claim 1, wherein: The key features of the medical sample image are extracted by the attention mechanism and calculated according to the following formula: Attn(Q,K,V)=softmax(Q·K·scale)·V; Among them, Q, K, V are vector representations that have been specifically transformed or extracted from the input data and are used to calculate the attention weights and obtain the weighted outputs. Softmax represents the activation function, which is used to convert the input values into probability distributions. Scale represents the scaling factor.
5. The multimodal medical image anomaly detection method according to claim 4, characterized in that: The key features are input into the multi-layer feature adapter for dynamic learning, and the deep features obtained include: Z l- l=[t′ cls ;t′1;t′2…;t′ T ]; From l =Project. l (Attn(V l ,In l ,In l ))+z l-1 ; Among them, represents the output of the L-1 layer of the visual encoder, Z l-1 Represents the local feature output of L-1 layer, QKV_Proj. l and Proj. l Represents QKV projection and output projection respectively, and the last layer outputs the original output Z o and local output Z, category feature Z o [0]∈R d For image-level anomaly detection, the local feature map Z[1:]∈R Txd Used for pixel-level anomaly detection; For the input medical sample image x∈R kxwx3 , convert the medical sample image x into feature space Z l ∈R Gxd , l∈{1, 2, 3, 4}, where h and w represent the size of the image, G represents the number of grids, and d represents the feature dimension; The multi-layer feature adapter includes four adapters, denoted as A l (·), l∈{1, 2, 3, 4}, in each stage S l , l∈{1, 2, 3, 4}, the feature adapter is integrated into the feature Z through two linear transformation layers l In, it is expressed as follows: Among them, ReLu represents the activation function, and Linear represents the linear transformation layer; After obtaining the features of the multi-layer feature adapter, the residual connection is used to retain the pre-trained original features, and the adapted features of the Lth layer can be expressed as: Z’ l =αA l (Z l ) T +(1-α)Z l ,l∈{1,2,3,4} (11) Among them, Z′ l As the next stage S l+1 The constant value α is used as the residual ratio to adjust the degree of retaining the original features; The feature adapter is a dual-branch feature adapter that generates two parallel features F at each layer. cls,l and F seg,l , where F cls,l represents the features for classification tasks generated by the dual-branch feature adapter at layer L, which is used to indicate whether the input medical sample image x is abnormal. seg,l It represents the features for segmentation tasks generated by the dual-branch feature adapter at the first layer, which is used to segment out the abnormal area features; The cosine similarity calculation is performed between the deep feature and the data in the precise text description data set to obtain the text description information and the abnormal feature map corresponding to the current medical sample image, including: where F∈{F cls,l ,F seg,l },τ is the temperature hyperparameter used to adjust the output distribution of the softmax function, F T represents the text description feature data in the precise text description dataset, F n Represents normal text description feature data, F a Indicates abnormal text description feature data, exp<> indicates exponential function; Among them, C1 represents the classification index of the abnormal type, S1 represents the segmentation index of the abnormal feature map, and BI() represents the image processing function, which is used to reshape the abnormal map into And use bilinear interpolation to restore it to the resolution of the original input medical sample image, G represents the number of grids; Among them, m represents the minimum distance between the features of each level recorded in the multi-level feature memory G, and the function Dist() represents the cosine distance. The final predicted classification and segmentation results are: C pre =βC1+(1-β)C2; S pre =βS1+(1-β)S2; Among them, C pre is the final determined image abnormality type, S pre is the final abnormal feature map, and β represents the weight value.
6. The multimodal medical image anomaly detection method according to claim 5, characterized in that: The loss function of the neural network model is obtained as follows: L1=θ1Dice(softmax(F seg,l ,F T ),S gt )+θ2Focal(softmax(F seg,l ,F T ),S gt )+θ3BCE(max(softmax(F cls,l ,F T )),C gt ); Among them, Dice loss is used to measure the similarity between two samples and is defined as: The larger the Dice coefficient, the higher the similarity; Focal(softmax(F seg,l ,F T ),S gt )=-αS gt (1-softmax(F seg,l ,F T )) γ log(softmax(F seg,l ,F T ))-α(1-S gt )S gt (1-softmax(F seg,l ,F T )) γ log(softmax(F seg,l ,F T )); Among them, α is the category balance factor, which is used to reduce the impact of category imbalance; γ is the focus factor, which is used to adjust the weight of difficult and easy samples; S gt represents the true label; BCE(max(softmax(F cls,l ,F T )),C gt )=-[C gt log(max(softmax(F cls,l ,F T )))+(1-C gt )log(1-max(softmax(F cls,l ,F T ))]; BCE loss is used to measure the difference between the predicted probability distribution and the true label, where C gt is the true label, which takes the value of 0 or 1, max(softmax(F cls,l , F T )) is the probability predicted by the model. The smaller the loss, the better the model performance. The total model loss is expressed as follows: Among them, L1 represents the total loss of the layer, θ1, θ2, θ3 represent hyperparameters, and the neural network model is optimized by being set to different proportions.
7. The multimodal medical image anomaly detection method according to claim 6, characterized in that: When the neural network model is trained, the following steps are also included: The accuracy of the neural network model is judged by AUC, and the calculation formula is as follows: Wherein, TPR represents the true positive rate, FPR represents the false positive rate, TP represents the number of pixels correctly detected as abnormal in the input medical sample image, FN represents the number of abnormal pixels incorrectly detected as normal in the input medical sample image, FP represents the number of normal pixels incorrectly detected as abnormal in the medical sample image, and TN represents the number of pixels correctly detected as normal in the medical sample image; 8. A multimodal medical image anomaly detection system, characterized in that: include: An image receiving module, used for receiving medical images; An anomaly detection module, used to input the medical image into the trained neural network model, and output an anomaly detection result corresponding to the medical image, wherein the anomaly detection result includes text description information of whether the image is abnormal and an abnormal feature map; When training, the neural network model includes: A text data generation module, used to generate text description data sets of different categories based on a template and a large language model, wherein the text description data sets include a normal text description set, an abnormal text description set, a normal text generation set, and an abnormal text generation set; A correlation analysis module, configured to receive a medical sample image, calculate cosine similarities between the medical sample image and a normal text description set, an abnormal text description set, a normal text generation set, and an abnormal text generation set, and calculate a correlation score based on the cosine similarities; If the calculated correlation score is greater than 0, the corresponding text description is retained, otherwise the corresponding text description is deleted until the data traversal in all text description data sets is completed, and the filtered accurate text description data set most relevant to the currently input medical sample image is obtained, and the accurate text description data set includes an accurate normal text description set and an accurate abnormal text description set; An attention mechanism module, used for extracting key features of the medical sample image through an attention mechanism; A multi-layer feature adapter is used to dynamically learn the key features to obtain deep features, and to calculate the cosine similarity between the deep features and the data in the precise text description data set to obtain the text description information and abnormal feature map corresponding to the current medical sample image.
9. A computer-readable storage medium storing computer program instructions, characterized in that: The computer program instructions implement the method of any one of claims 1 to 7 when executed by a processor.
10. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal data exception identification method and device, electronic equipment and storage medium
CN118379755A
Zero sample anomaly detection method and device based on image-text pre-training model
CN118864876A
Model training method and device, industrial image defect detection method and device and storage medium
CN119068305A
Zero sample image anomaly detection method and device
CN119130931A
Cited By
Medical anomaly detection method based on graph language feature guidance
CN120413046A
Industrial anomaly detection method based on multi-agent prompt learning
CN120852894A
An industrial anomaly detection method based on multi-agent hint learning
CN120852894B
Abnormality detection method and device applied to inspection robot, equipment and medium
CN121708535A