Image abnormal state classification method and device, computer equipment and storage medium
Through comparative learning and feature fusion technology, the problem of insufficient diagnostic accuracy caused by isolated processing of mammary X-ray images and text information is solved, and the accurate classification of early diagnosis of breast cancer is achieved, and the accuracy of classification of image abnormal states is improved.
Patent Information
- Application Number
- CN202510032129.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, mammary X-ray images and clinical text information are processed as isolated information sources, resulting in insufficient accuracy of early diagnosis of breast cancer, and the labeling of medical image data is complex and expensive, so it is impossible to effectively classify and identify breast abnormalities.
By obtaining exception image information and text information, mask processing and comparison learning, using image comparison encoder and text encoder for parameter sharing and supervision adjustment, and combining pre-trained exception classifiers for feature extraction and fusion, to achieve accurate classification of abnormal states.
It improves the accuracy of early diagnosis of breast cancer, realizes effective classification of abnormal image information and text information, and improves the accuracy of classification of abnormal image status.
Smart Images

Figure CN119942206A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, specifically to the field of digital medicine, and in particular to a method, device, computer equipment and storage medium for classifying abnormal image states. Background Art
[0002] In the field of digital medicine, early diagnosis of breast cancer has always been an important topic in clinical practice, and its accuracy is directly related to the treatment effect and quality of life of patients. Currently, mammography, as the gold standard for breast cancer screening, can capture subtle structural changes in breast tissue, but relying solely on image information for diagnosis still faces many limitations. The professional knowledge and experience of clinicians play an indispensable role in the diagnostic process. However, how to effectively integrate this valuable knowledge with image information has become a key challenge to improving diagnostic accuracy.
[0003] Traditional methods often treat breast X-ray images and clinical text information as two isolated information sources, ignoring the potential inherent connection between them, resulting in a lack of comprehensiveness and accuracy in the diagnosis results. In addition, the annotation process of medical image data is complex and costly, making it impossible to accurately and effectively classify and identify breast abnormalities when the annotated data is limited, which brings certain difficulties to breast diagnosis. Summary of the invention
[0004] The purpose of the embodiments of the present application is to propose a method, device, computer equipment and storage medium for classifying abnormal image states, so as to solve the problem that the abnormal states described by abnormal image information and abnormal text information cannot be effectively and accurately classified.
[0005] In order to solve the above technical problems, the embodiment of the present application provides a method for classifying abnormal image states, which adopts the following technical solutions:
[0006] Obtain abnormal image information and abnormal text information;
[0007] Performing mask processing on the abnormal image information to obtain an input image pair, and performing contrast learning based on the input image pair to obtain an image contrast encoder;
[0008] The image contrast encoder and the preset text encoder are parameter-shared to obtain a shared text encoder, and the abnormal image information and the abnormal text information are input into the image contrast encoder and the shared text encoder for comparative learning to obtain an optimized image encoder and an optimized text encoder;
[0009] Acquire abnormal annotated data, and perform supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotated data to obtain an adjusted image encoder and an adjusted text encoder;
[0010] Acquire target abnormal image information and target abnormal text information, input the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction, and obtain effective image features and effective text features;
[0011] The effective image features and the effective text features are fused to obtain effective fused features, and the information of the object to be tested is obtained. The information of the object to be tested and the effective fused features are input into a pre-trained abnormality classifier for classification to obtain an image abnormality state classification result.
[0012] Furthermore, the step of obtaining abnormal image information and abnormal text information specifically includes:
[0013] Acquire the abnormal object number information, and extract the abnormal analysis report data from the database according to the abnormal object number information;
[0014] Performing image extraction and text extraction on the abnormal analysis report data to obtain initial abnormal image information and initial abnormal text information;
[0015] The initial abnormal image information and the initial abnormal text information are preprocessed to obtain the abnormal image information and the abnormal text information.
[0016] Furthermore, the step of performing mask processing on the abnormal image information to obtain an input image pair, and performing contrast learning based on the input image pair to obtain an image contrast encoder specifically includes:
[0017] Randomly masking the abnormal image information to generate an input image pair, wherein the input image pair includes a first mask image and a second mask image;
[0018] Inputting the first mask image and the second mask image into a first image encoder and a second image encoder that share a weight respectively for feature extraction to obtain a first image feature representation and a second image feature representation;
[0019] Performing low-dimensional projection on the first image feature representation and the second image feature representation through a first projection block and a second projection block that share weights, respectively, to obtain a first low-dimensional feature representation and a second low-dimensional feature representation;
[0020] Performing loss calculation according to the first low-dimensional feature representation and the second low-dimensional feature representation to obtain an image loss result;
[0021] The first image encoder and the second image encoder are optimized according to the image loss result to obtain the image contrast encoder.
[0022] Furthermore, the step of sharing parameters of the image contrast encoder and the preset text encoder to obtain a shared text encoder, and inputting the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for comparative learning to obtain an optimized image encoder and an optimized text encoder specifically includes:
[0023] Inputting the abnormal image information into the image contrast encoder to obtain abnormal image coding features;
[0024] Obtain a preset text encoder, and perform weight sharing between the image comparison encoder and the preset text encoder to obtain the shared text encoder;
[0025] Inputting the abnormal text information into the shared text encoder to obtain abnormal text encoding features;
[0026] Performing low-dimensional projection on the abnormal text encoding features to obtain abnormal text low-dimensional features;
[0027] Globally pooling the abnormal image encoding features and the abnormal text low-dimensional features to obtain a first pooling feature and a second pooling feature;
[0028] The image contrast encoder and the shared text encoder are optimized according to the first pooling feature and the second pooling feature to obtain the optimized image encoder and the optimized text encoder.
[0029] Further, the step of obtaining the abnormal annotated data, and performing supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotated data to obtain the steps of adjusting the image encoder and adjusting the text encoder specifically includes:
[0030] Obtaining a data extraction identifier, and extracting the abnormal annotated data from a database according to the data extraction identifier;
[0031] Parsing the abnormal annotation data to obtain abnormal image annotation data and abnormal text annotation data;
[0032] The abnormal image annotated data and the abnormal text annotated data are used as supervisory signals, and the optimized image encoder and the optimized text encoder are fine-tuned with the supervisory signals to obtain the adjusted image encoder and the adjusted text encoder.
[0033] Further, the step of obtaining target abnormal image information and target abnormal text information, inputting the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction, and obtaining effective image features and effective text features specifically includes:
[0034] Obtaining a target object number, and extracting the target abnormal image information and the target abnormal text information from a database according to the target object number;
[0035] Inputting the target abnormal image information into the adjustment image encoder to extract image features to obtain the effective image features;
[0036] The target abnormal text information is input into the adjustment text encoder to extract text features to obtain the effective text features.
[0037] Furthermore, before the step of fusing the effective image features and the effective text features to obtain effective fused features, obtaining information of the object to be tested, inputting the information of the object to be tested and the effective fused features into a pre-trained abnormality classifier for classification, and obtaining the image abnormality state classification result, the following steps are also included:
[0038] Obtaining sample test object information and sample fusion features;
[0039] quantifying the influence of the sample to-be-tested object information on the sample fusion feature according to a pre-built object feature association model, and obtaining a personalized fusion feature weight;
[0040] Calculating a weighted feature vector according to the personalized fusion feature weight and the sample fusion feature;
[0041] Constructing a basic classifier based on an ensemble learning algorithm, and adjusting the weight of the basic classifier according to the weighted feature vector to obtain a personalized ensemble classifier;
[0042] Acquire sample abnormal object data, train the personalized integrated classifier according to the sample abnormal object data, and optimize the personalized integrated classifier based on cross-validation to obtain the abnormal classifier.
[0043] In order to solve the above technical problems, the embodiment of the present application further provides an image abnormal state classification device, which adopts the following technical solution:
[0044] An information acquisition module, used to acquire abnormal image information and abnormal text information;
[0045] A first contrast learning module is used to perform mask processing on the abnormal image information to obtain an input image pair, and perform contrast learning based on the input image pair to obtain an image contrast encoder;
[0046] A second contrast learning module is used to share parameters of the image contrast encoder and the preset text encoder to obtain a shared text encoder, and input the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder;
[0047] A supervision adjustment module, used for acquiring abnormal annotation data, and performing supervision adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotation data to obtain an adjusted image encoder and an adjusted text encoder;
[0048] A feature extraction module, used for acquiring target abnormal image information and target abnormal text information, inputting the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction, and obtaining effective image features and effective text features;
[0049] The abnormal classification module is used to fuse the effective image features and the effective text features to obtain effective fusion features, and obtain the information of the object to be tested, and input the information of the object to be tested and the effective fusion features into a pre-trained abnormal classifier for classification to obtain the image abnormal state classification result.
[0050] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0051] A computer device comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the image abnormal state classification method as described in any one of the above items are implemented.
[0052] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0053] A computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the image abnormal state classification method as described in any one of the above items are implemented.
[0054] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: the embodiment obtains abnormal image information and abnormal text information; performs mask processing on the abnormal image information to obtain an input image pair, and performs contrast learning based on the input image pair to obtain an image contrast encoder; shares parameters between the image contrast encoder and the preset text encoder to obtain a shared text encoder, and inputs the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder; obtains abnormal annotation data, and performs supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotation data to obtain an adjusted image encoder and an adjusted text encoder; obtains target abnormal image information and target abnormal text information, and inputs the target abnormal image information and the target abnormal text information into the adjusted image encoder and the adjusted text encoder for feature extraction to obtain effective image features and effective text features; fuses the effective image features and the effective text features to obtain effective fusion features, and obtains information of the object to be measured, and inputs the information of the object to be measured and the effective fusion features into a pre-trained abnormal classifier for classification to obtain an image abnormal state classification result. This effectively achieves accurate classification of abnormal states described by abnormal image information and abnormal text information, and improves the accuracy of image abnormal state classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the scheme in the present application, a brief introduction is given below to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0057] Figure 2 A flowchart of an embodiment of a method for classifying abnormal image states according to the present application;
[0058] Figure 3 is a structural schematic diagram of an embodiment of an image abnormal state classification device according to the present application;
[0059] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0061] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor does it refer to non-related or alternative embodiments that are mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0062] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0063] like Figure 1 As shown, the system architecture 100 may include a terminal device 101, a network 102 and a server 103. The terminal device 101 may be a laptop 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links or optical fiber cables.
[0064] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0065] The terminal device 101 can be any electronic device with a display screen and supporting web browsing. In addition to a laptop computer 1011, a tablet computer 1012 or a mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV), a laptop computer, a desktop computer, etc.
[0066] The server 103 may be a server that provides various services, such as a background server that provides support for a web page displayed on the terminal device 101 .
[0067] It should be noted that the image abnormal state classification method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the image abnormal state classification device is generally set in the server / terminal device.
[0068] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0069] Continue to refer Figure 2 , shows a flow chart of an embodiment of a method for classifying abnormal image states according to the present application. The method for classifying abnormal image states comprises the following steps:
[0070] Step S10, obtaining abnormal image information and abnormal text information;
[0071] In this embodiment, the abnormal image information refers to a mammographic image, which includes image data of multiple patients from hospitals in different geographical locations. The mammographic image includes image data of two categories in a breast imaging report and data system. Specifically, the abnormal image information may include mammographic images of 4557 patients, of which 1179 are BI-RADS3 category and 3378 are BI-RADS4 category. The abnormal text information is the abnormal confirmation result corresponding to the abnormal image information. The abnormal confirmation result is in the form of text, which may be a specific abnormal confirmation report paragraph, or an abnormal confirmation result field. For example, the BI-RADS of the patient is category 3.
[0072] Step S20, performing mask processing on the abnormal image information to obtain an input image pair, and performing contrast learning based on the input image pair to obtain an image contrast encoder;
[0073] In this embodiment, mask processing is an image processing technique used to highlight specific areas in an image or hide other areas. The mask can be generated using an automatic segmentation algorithm. The mask itself is a binary image of the same size as the original image (usually composed of 0 and 255 or 0 and 1, depending on how the image is represented), where white (or high-value) areas represent areas of interest, while black (or low-value) areas represent areas of no interest. The automatic segmentation algorithm uses computer vision and machine learning algorithms to automatically segment areas of interest from an image. The algorithm can perform segmentation based on features such as grayscale value, texture, and shape of the image. In this embodiment, the above-mentioned automatic segmentation algorithm can use threshold segmentation. In this embodiment, contrastive learning is unimodal contrastive learning, which is a mutual learning process based on image loss between feature representations of input image pairs. The goal is to make the feature representations of similar samples close to each other in a low-dimensional space, while the feature representations of dissimilar samples are far away from each other, thereby obtaining an image contrast encoder that can effectively extract image features.
[0074] Step S30, sharing parameters of the image contrast encoder and the preset text encoder to obtain a shared text encoder, and inputting the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder;
[0075] In this embodiment, the image contrast encoder can be a deep learning model, specifically a convolutional neural network (CNN) model, which can convert the input abnormal image information into a high-dimensional feature vector. The text encoder can adopt a Transformer model, specifically a BERT model, which can learn the language representation in the input abnormal text information to obtain the corresponding language feature vector. Parameter sharing can be achieved by introducing the same parameters between the feature extraction layer, intermediate representation layer, processing layer, etc. of the image encoder and the text encoder, and can be specifically achieved by designing the above layers of the two encoders to be the same architecture and weights. In this embodiment, contrast learning is multimodal contrast learning, which is based on the features corresponding to the abnormal image information and the abnormal text information to optimize the image contrast encoder and the shared text encoder, and the optimization can be achieved by an optimizer.
[0076] Step S40, obtaining abnormal annotated data, and performing supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotated data to obtain an adjusted image encoder and an adjusted text encoder;
[0077] In this embodiment, the abnormal annotated data includes abnormal image annotated data and abnormal text annotated data, wherein the abnormal image annotated data is annotated abnormal image information, and the abnormal text annotated data is annotated abnormal text information. The supervised adjustment can be achieved by constructing a supervised learning framework, which aims to use the annotated data to guide the learning process of the model, thereby optimizing the model parameters to improve its performance. The construction of the supervised learning framework includes the steps of defining a loss function, designing a fusion mechanism, and constructing an optimizer.
[0078] Step S50, obtaining target abnormal image information and target abnormal text information, inputting the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction, and obtaining effective image features and effective text features;
[0079] In this embodiment, the target abnormal image information and the target abnormal text information refer to the breast X-ray image and abnormal confirmation result report corresponding to the target patient to be identified and classified, by inputting the breast X-ray image of the target patient into the adjustment image encoder for feature extraction to obtain effective image features, and inputting the abnormal confirmation result report of the target patient into the adjustment text encoder for feature extraction to obtain effective text features.
[0080] Step S60, fusing the effective image features and the effective text features to obtain effective fused features, and obtaining information of the object to be tested, inputting the information of the object to be tested and the effective fused features into a pre-trained abnormality classifier for classification, and obtaining an image abnormality state classification result.
[0081] In this embodiment, the effective fusion feature is the feature information obtained by fusing the effective image feature and the effective text feature, and the object information to be tested is the individual information corresponding to the target patient, which is used to effectively classify the specific conditions of different patients. The object information to be tested includes age, breast density, family history, etc., which can be obtained from the medical records of the target patient collected from the medical data system. The pre-trained abnormal classifier is a classifier for classifying abnormal image states, and the classifier can be an integrated classifier, wherein the integrated classifier, also known as integrated learning or classifier integration, refers to a method of combining the prediction results of multiple base classifiers to improve the overall classification performance. The abnormal image state classification results correspond to BI-RADS3 and 4 categories, wherein BI-RADS 3 category refers to the probability of breast malignancy ≤2%, and BI-RADS 4 category refers to the probability of breast malignancy between 2% and 95%. By obtaining the abnormal image state classification results, a reliable reference basis is provided for subsequent abnormal diagnosis.
[0082] In this embodiment, the above method can be applied to a medical service system, in which the image abnormality status is classified according to the target abnormal image information and the target abnormal text information, so as to effectively obtain the corresponding image abnormality status classification result. Specifically, in this embodiment, the medical service system can be one or more of a medical insurance system and a disease insurance system, the abnormal image information and the abnormal text information are medical-related image data and text data containing the disease condition, the abnormal image information, the abnormal text information, the abnormal annotation data, the target abnormal image information, the target abnormal text information, and the information of the object to be measured are all stored in the medical insurance system and the disease insurance system and obtained from the database of the above system, and the effective image abnormality status classification result is generated by the above system through the processing of the method of this embodiment and stored in the database of the above system.
[0083] This embodiment obtains abnormal image information and abnormal text information; performs mask processing on the abnormal image information to obtain an input image pair, and performs contrast learning based on the input image pair to obtain an image contrast encoder; shares parameters between the image contrast encoder and a preset text encoder to obtain a shared text encoder, and inputs the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder; obtains abnormal annotation data, and performs supervised adjustment on the optimized image encoder and the optimized text encoder based on the abnormal annotation data to obtain an adjusted image encoder and an adjusted text encoder; obtains target abnormal image information and target abnormal text information, and inputs the target abnormal image information and the target abnormal text information into the adjusted image encoder and the adjusted text encoder for feature extraction to obtain effective image features and effective text features; fuses the effective image features and the effective text features to obtain effective fused features, and obtains information about an object to be tested, and inputs the information about the object to be tested and the effective fused features into a pre-trained abnormal classifier for classification to obtain an image abnormal state classification result. This effectively achieves accurate classification of abnormal states described by abnormal image information and abnormal text information, and improves the accuracy of image abnormal state classification.
[0084] In some optional implementations of this embodiment, the obtaining of abnormal image information and abnormal text information includes the following steps:
[0085] Acquire the abnormal object number information, and extract the abnormal analysis report data from the database according to the abnormal object number information;
[0086] In this embodiment, the abnormal object number information is identification information corresponding to the abnormal analysis report data, and the abnormal object number information is used as a query condition to perform traversal matching in the database to extract the corresponding abnormal analysis report data.
[0087] Performing image extraction and text extraction on the abnormal analysis report data to obtain initial abnormal image information and initial abnormal text information;
[0088] In this embodiment, the abnormal analysis report data is a medical report that records abnormal images and corresponding abnormal texts. The initial abnormal image information can be extracted from the abnormal analysis report data based on image processing software, and the initial abnormal text information can be extracted from the abnormal analysis report data based on OCR technology.
[0089] The initial abnormal image information and the initial abnormal text information are preprocessed to obtain the abnormal image information and the abnormal text information.
[0090] In this embodiment, the preprocessing of the initial abnormal image information includes grayscale conversion, noise removal, image enhancement, morphological processing, edge detection, etc. The preprocessing of the initial abnormal text information includes text cleaning, word segmentation, part-of-speech tagging, redundant information removal, text normalization, etc. By performing the above preprocessing on the initial abnormal image information and the initial abnormal text information, the abnormal image information and the abnormal text information are effectively obtained.
[0091] This embodiment obtains the abnormal object number information, and extracts the abnormal analysis report data from the database according to the abnormal object number information; performs image extraction and text extraction on the abnormal analysis report data to obtain initial abnormal image information and initial abnormal text information; and pre-processes the initial abnormal image information and the initial abnormal text information to obtain the abnormal image information and the abnormal text information. Thus, standard and effective abnormal image information and abnormal text information are effectively obtained to provide reliable data support for subsequent processing steps.
[0092] In some optional implementations of this embodiment, the step of performing mask processing on the abnormal image information to obtain an input image pair, and performing contrast learning based on the input image pair to obtain an image contrast encoder includes the following steps:
[0093] Randomly masking the abnormal image information to generate an input image pair, wherein the input image pair includes a first mask image and a second mask image;
[0094] In this embodiment, the first mask image is an image obtained by applying a first randomly generated mask to the original abnormal image. The second mask image is an image obtained by using a second random mask different from the first mask to block another part of the original image. By using the image processing library to create a blank mask (usually a full zero matrix) of the same size as the image, and drawing a randomly shaped area on the mask, setting the pixel value of the area to 1 (or other non-zero value, indicating the mask area), the above mask generation process is run multiple times to generate multiple different masks. A random seed can be set to ensure repeatability, and the size and position of the mask can be randomly selected on the image, but it is necessary to ensure that the mask covers a portion of the abnormal area of the image.
[0095] Inputting the first mask image and the second mask image into a first image encoder and a second image encoder that share a weight respectively for feature extraction to obtain a first image feature representation and a second image feature representation;
[0096] In this embodiment, the shared weight means that the first image encoder and the second image encoder have the same network structure and weight, so that the first image encoder and the second image encoder process the input image in exactly the same way, but with different inputs. The first image encoder and the second image encoder may use a convolutional neural network (CNN), and obtain a first image feature representation by inputting the first mask image into the first image encoder for feature extraction, and obtain a second image feature representation by inputting the second mask image into the second image encoder for feature extraction.
[0097] Performing low-dimensional projection on the first image feature representation and the second image feature representation through a first projection block and a second projection block that share weights, respectively, to obtain a first low-dimensional feature representation and a second low-dimensional feature representation;
[0098] In this embodiment, the shared weight means that the first projection block and the second projection block have the same network structure and weight, and the first low-dimensional feature representation is obtained by inputting the first image feature representation into the first projection block for low-dimensional projection, so as to map the first image feature representation into a preset low-dimensional vector space. Similarly, the second low-dimensional feature representation is obtained by inputting the second image feature representation into the second projection block for low-dimensional projection, so as to map the second image feature representation into a preset low-dimensional vector space.
[0099] Performing loss calculation according to the first low-dimensional feature representation and the second low-dimensional feature representation to obtain an image loss result;
[0100] In this embodiment, the loss calculation is implemented based on the contrast loss function optimization model. Specifically, the contrast loss function optimization model is L2 loss. L2 loss (also known as mean square error MSE) is a commonly used loss function used to measure the Euclidean distance between two vectors. The corresponding loss value is obtained by calculating the square root of the sum of the squares of the difference between the first low-dimensional feature representation and the second low-dimensional feature representation of the two vectors, and the loss value is the image loss result.
[0101] The first image encoder and the second image encoder are optimized according to the image loss result to obtain the image contrast encoder.
[0102] In this embodiment, according to the calculated image loss result, the back propagation algorithm and gradient descent (or its variant) are used to update the weights of the first image encoder and the second image encoder. The weight update process can be implemented by an optimizer (such as SGD, Adam, etc.), and the optimizer adjusts the parameters of the model according to the gradient of the loss function to minimize the loss. When the weight update of the above-mentioned image encoder is completed, the optimization of the first image encoder and the second image encoder is completed, and the image contrast encoder is obtained.
[0103] This embodiment generates an input image pair by randomly masking the abnormal image information, wherein the input image pair includes a first mask image and a second mask image; the first mask image and the second mask image are respectively input into a first image encoder and a second image encoder that share weights for feature extraction to obtain a first image feature representation and a second image feature representation; the first image feature representation and the second image feature representation are respectively low-dimensionally projected through a first projection block and a second projection block that share weights to obtain a first low-dimensional feature representation and a second low-dimensional feature representation; loss calculation is performed based on the first low-dimensional feature representation and the second low-dimensional feature representation to obtain an image loss result; the image encoder is optimized based on the image loss result, thereby effectively obtaining an image contrast encoder that has been optimized through image unimodal contrast learning, so as to facilitate subsequent multimodal contrast learning processing.
[0104] In some optional implementations of this embodiment, the step of sharing parameters of the image contrast encoder and the preset text encoder to obtain a shared text encoder, and inputting the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for comparative learning to obtain an optimized image encoder and an optimized text encoder includes the following steps:
[0105] Inputting the abnormal image information into the image contrast encoder to obtain abnormal image coding features;
[0106] In this embodiment, the abnormal image information is input into the image contrast encoder for feature extraction to obtain abnormal image coding features, which include morphological features, texture features, density distribution features, etc. Among them, the morphological features include: mass features, calcification features, structural distortion features, asymmetry features, etc., the texture features include: fibrous tissue arrangement features, density uniformity features, fine structure texture features, etc., and the density distribution features include breast tissue uniformity features, high-density area distribution features, etc.
[0107] Obtain a preset text encoder, and perform weight sharing between the image comparison encoder and the preset text encoder to obtain the shared text encoder;
[0108] In this embodiment, the preset text encoder can adopt the BERT model of Transformer, and the BERT model is trained by preset sample text data so that it can extract useful text features to obtain the preset text encoder. The shared text encoder is obtained by performing weight sharing processing on the feature extraction layer, intermediate representation layer, processing layer, etc. of the image contrast encoder and the preset text encoder.
[0109] Inputting the abnormal text information into the shared text encoder to obtain abnormal text encoding features;
[0110] In this embodiment, the abnormal text information is input into a shared text encoder for feature extraction to obtain abnormal text encoding features, which include keyword features (such as "cancer", "tumor", "hyperplasia", etc.), semantic features, emotional features (such as "suspected", "possible", "confirmed" and other words), characteristic disease-related features (such as breast cancer, breast hyperplasia, etc.), etc.
[0111] Performing low-dimensional projection on the abnormal text encoding features to obtain abnormal text low-dimensional features;
[0112] In this embodiment, the abnormal text encoding features can be low-dimensionally projected by principal component analysis, and the principal component analysis includes the following steps: calculating the covariance matrix between the abnormal text encoding features, wherein each element of the covariance matrix represents the covariance between two features, that is, the degree of their common change; performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors; according to the size of the eigenvalues, selecting the eigenvectors corresponding to the first k largest eigenvalues as the principal components; projecting the abnormal text encoding features into a low-dimensional space composed of the selected principal components, and obtaining a low-dimensional feature matrix by multiplying the original data matrix with the principal component matrix. The low-dimensional feature matrix is the low-dimensional feature of the abnormal text.
[0113] Globally pooling the abnormal image encoding features and the abnormal text low-dimensional features to obtain a first pooling feature and a second pooling feature;
[0114] In this embodiment, global pooling is a technique for converting a feature map into a fixed-length feature vector, and the abnormal image coding feature may be a feature map extracted by CNN. Global pooling may be used to convert the abnormal image coding feature into a first pooling feature of a fixed length, and global pooling may be used to convert the abnormal text low-dimensional feature into a second pooling feature.
[0115] The image contrast encoder and the shared text encoder are optimized according to the first pooling feature and the second pooling feature to obtain the optimized image encoder and the optimized text encoder.
[0116] In this embodiment, the contrast loss between the first pooled features and the second pooled features is calculated by combining a multimodal loss function that combines different loss terms from images and texts. The contrast loss is then used to encourage similar image-text pairs to be closer in the feature space and dissimilar image-text pairs to be farther away. This is how the image contrast encoder and the shared text encoder are optimized and adjusted. This optimization and adjustment step can be achieved by updating the weights of the encoder using a back-propagation algorithm and a preset optimizer (such as SGD, Adam, etc.).
[0117] This embodiment obtains abnormal image coding features by inputting the abnormal image information into the image contrast encoder; obtains a preset text encoder, and weight-shares the image contrast encoder and the preset text encoder to obtain the shared text encoder; inputs the abnormal text information into the shared text encoder to obtain abnormal text coding features; performs low-dimensional projection on the abnormal text coding features to obtain abnormal text low-dimensional features; performs global pooling on the abnormal image coding features and the abnormal text coding features to obtain first pooling features and second pooling features; optimizes the image contrast encoder and the shared text encoder according to the first pooling features and the second pooling features to obtain the optimized image encoder and the optimized text encoder. In this way, the optimized image encoder and the optimized text encoder after multimodal learning optimization based on the abnormal image information and the abnormal text information are effectively obtained to facilitate subsequent supervised adjustment processing.
[0118] In some optional implementations of this embodiment, the step of obtaining the abnormal annotated data, and performing supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotated data to obtain the adjusted image encoder and the adjusted text encoder comprises the following steps:
[0119] Obtaining a data extraction identifier, and extracting the abnormal annotated data from a database according to the data extraction identifier;
[0120] In this embodiment, the data extraction identifier is the identification information corresponding to the abnormal annotated data, and the corresponding abnormal annotated data is extracted by traversing and matching in the database using the data extraction identifier as a query condition. The abnormal annotated data is annotated abnormal image information and abnormal text information, and the annotation can be achieved by manually annotating the abnormal image information and abnormal text information in advance.
[0121] Parsing the abnormal annotation data to obtain abnormal image annotation data and abnormal text annotation data;
[0122] In this embodiment, the abnormal image annotation data carries corresponding annotation tags, such as disease type, stage, severity, etc. The abnormal text annotation data also carries corresponding standard tags, such as classification, keyword extraction, sentiment analysis or other forms of annotation.
[0123] The abnormal image annotated data and the abnormal text annotated data are used as supervisory signals, and the optimized image encoder and the optimized text encoder are fine-tuned with the supervisory signals to obtain the adjusted image encoder and the adjusted text encoder.
[0124] In this embodiment, the supervisory signal exists in a form that can be directly used for training the model, such as a label vector of One-Hot Encoding, a real number label or an embedded vector. Fine-tuning is a process of making small-scale adjustments based on a pre-trained model to adapt to new tasks or new data. By taking abnormal image annotation data (such as disease labels) as the target output, the parameters of the image encoder are adjusted through the back propagation algorithm to minimize the difference between the predicted label and the actual label, so as to achieve supervised fine-tuning of the optimized image encoder and obtain an adjusted image encoder. Similarly, by taking abnormal text annotation data (such as classification labels, keywords, etc.) as the target output, the parameters of the text encoder are adjusted to optimize performance, so as to achieve supervised fine-tuning of the optimized text encoder and obtain an adjusted text encoder.
[0125] This embodiment obtains a data extraction identifier, extracts the abnormal annotated data from the database according to the data extraction identifier; parses the abnormal annotated data to obtain abnormal image annotated data and abnormal text annotated data; uses the abnormal image annotated data and the abnormal text annotated data as supervisory signals, and uses the supervisory signals to fine-tune the optimized image encoder and the optimized text encoder to obtain the adjusted image encoder and the adjusted text encoder. This effectively implements supervised fine-tuning of the optimized image encoder and the optimized text encoder according to the abnormal annotated data, so as to further optimize the adjusted image encoder and the adjusted text encoder, and facilitates the subsequent feature extraction processing of the target abnormal image information and the target abnormal text information.
[0126] In some optional implementations of this embodiment, the acquiring of target abnormal image information and target abnormal text information, inputting the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction, and obtaining valid image features and valid text features includes the following steps:
[0127] Obtaining a target object number, and extracting the target abnormal image information and the target abnormal text information from a database according to the target object number;
[0128] In this embodiment, the target object number is identification information corresponding to the target abnormal image information and the target abnormal text information, and the target abnormal image information and the target abnormal text information are extracted by traversing and matching in the database using the target object number as a query condition. The target object number can be a system number or number ID corresponding to the target patient in the medical database.
[0129] Inputting the target abnormal image information into the adjustment image encoder to extract image features to obtain the effective image features;
[0130] In this embodiment, the target abnormal image information is a breast X-ray image actually collected from the target patient. The target abnormal image information is input into the adjustment image encoder to extract image features, so as to obtain corresponding effective image features.
[0131] The target abnormal text information is input into the adjustment text encoder to extract text features to obtain the effective text features.
[0132] In this embodiment, the target abnormal text information is a text of abnormal confirmation result that truly describes the target patient. The target abnormal text information is input into an adjustment text encoder to perform text feature extraction to obtain corresponding valid text features.
[0133] This embodiment obtains the target object number, extracts the target abnormal image information and the target abnormal text information from the database according to the target object number; inputs the target abnormal image information into the adjustment image encoder for image feature extraction to obtain the effective image feature; inputs the target abnormal text information into the adjustment text encoder for text feature extraction to obtain the effective text feature. Thus, effective image feature extraction and text feature extraction are achieved according to the adjusted image encoder and the adjusted text encoder after supervised fine-tuning, so as to facilitate subsequent feature fusion and image abnormal state classification.
[0134] In some optional implementations of this embodiment, before the effective image features and the effective text features are fused to obtain effective fused features, and the information of the object to be tested is obtained, and the information of the object to be tested and the effective fused features are input into a pre-trained abnormality classifier for classification, and before obtaining the image abnormal state classification result, the following steps are also included:
[0135] Obtaining sample test object information and sample fusion features;
[0136] In this embodiment, the sample test object information includes historical patient age, patient breast density, patient family medical history, etc. The sample fusion features include sample abnormal image features and sample abnormal text features, which are feature information extracted from historical abnormal image information and historical abnormal text information.
[0137] quantifying the influence of the sample to-be-tested object information on the sample fusion feature according to a pre-built object feature association model, and obtaining a personalized fusion feature weight;
[0138] In this embodiment, the sample object information and sample fusion features are passed as input data to the object feature association model, and the object feature association model is used to calculate the correlation or relevance between the sample object information and each sample fusion feature. The object feature association model can use a deep learning neural network model, and the model is trained by a labeled data set containing historical object information and corresponding fusion features, wherein the input layer of the deep learning neural network model is used to receive the sample object information and sample fusion features, the hidden layer is used to learn the complex relationship in the input data and extract features, and the output layer is used to calculate the degree of influence of each feature on the object information to be tested, and the degree of influence is the above-mentioned degree of association. Then, according to the calculation result of the degree of association, a weight value is assigned to each sample fusion feature to obtain a personalized fusion feature weight. The size of the personalized fusion feature weight reflects the relative importance of the feature in a specific patient object.
[0139] Calculating a weighted feature vector according to the personalized fusion feature weight and the sample fusion feature;
[0140] In this embodiment, a weighted feature vector reflecting a comprehensive indicator of the patient's status is obtained by multiplying the value of each fused feature by its corresponding weight.
[0141] Constructing a basic classifier based on an ensemble learning algorithm, and adjusting the weight of the basic classifier according to the weighted feature vector to obtain a personalized ensemble classifier;
[0142] In this embodiment, an integrated learning algorithm is used to construct multiple basic classifiers, and the weights of the basic classifiers are adaptively adjusted according to the weighted feature vector to form a personalized integrated classification model. The steps of constructing basic classifiers and adjusting weights include: using a random forest integrated learning algorithm to train multiple basic classifiers using a training data set. These classifiers can be different types of machine learning models, or they can be the same type of models but trained with different parameters or feature subsets. According to the patient's specific information, a personalized image feature vector is extracted and calculated. The performance of each basic classifier, such as accuracy, recall rate, F1 score, etc., is evaluated on the validation data set. According to the performance of the basic classifier on the validation data set, its weight in the final decision is dynamically adjusted. The weight adjustment strategy can be weighted according to the accuracy or F1 score of the classifier.
[0143] Acquire sample abnormal object data, train the personalized integrated classifier according to the sample abnormal object data, and optimize the personalized integrated classifier based on cross-validation to obtain the abnormal classifier.
[0144] In this embodiment, the sample abnormal object data is the abnormal image information and abnormal text information collected from the patient object in history. In the training process of the personalized integrated classifier, a regularization term can be introduced to control the complexity of the model and prevent overfitting. The regularization term can be L1 regularization, L2 regularization or other forms of regularization. In this embodiment, L1 regularization is used. According to the performance of the model, the strength of the regularization term is adjusted to balance the complexity and generalization ability of the model. Cross-validation uses k-fold cross-validation or other cross-validation methods to divide the training data set into multiple subsets, and train and cross-validate them respectively to evaluate the classification performance of the model. The results of cross-validation include accuracy, recall rate, F1 score, etc. According to the results of cross-validation, the parameter range to be optimized is determined, including the number of basic classifiers, the depth of each classifier, the learning rate and other hyperparameters. A genetic algorithm is selected to search for the optimal parameter combination in the parameter space. The final classifier model is trained using the optimal parameter combination, and its performance is evaluated on the test data set. Thus, an abnormal classifier that can classify the abnormal state corresponding to the effective fusion feature is obtained.
[0145] This embodiment obtains sample test object information and sample fusion features; quantifies the influence of the sample test object information on the sample fusion features according to a pre-constructed object feature association model to obtain personalized fusion feature weights; calculates weighted feature vectors according to the personalized fusion feature weights and the sample fusion features; constructs a basic classifier based on an ensemble learning algorithm, and adjusts the weight of the basic classifier according to the weighted feature vector to obtain a personalized integrated classifier; obtains sample abnormal object data, trains the personalized integrated classifier according to the sample abnormal object data, and optimizes the personalized integrated classifier based on cross-validation, thereby effectively obtaining an abnormal classifier that can perform personalized image abnormal state classification according to effective fusion features and test object information.
[0146] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0147] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0148] Further references Figure 3 , as a response to the above Figure 1 The present application provides an embodiment of an image abnormal state classification device, and the device embodiment is similar to Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0149] like Figure 3As shown, the image abnormal state classification device 700 described in this embodiment includes: an information acquisition module 701, a first contrast learning module 702, a second contrast learning module 703, a supervision adjustment module 704, a feature extraction module 705, and an abnormal classification module 706. Among them:
[0150] Information acquisition module 701, used to acquire abnormal image information and abnormal text information;
[0151] A first contrast learning module 702 is used to perform mask processing on the abnormal image information to obtain an input image pair, and perform contrast learning based on the input image pair to obtain an image contrast encoder;
[0152] The second contrast learning module 703 is used to share parameters of the image contrast encoder and the preset text encoder to obtain a shared text encoder, and input the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder;
[0153] A supervisory adjustment module 704 is used to obtain abnormal annotation data, and perform supervisory adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotation data to obtain an adjusted image encoder and an adjusted text encoder;
[0154] A feature extraction module 705 is used to obtain target abnormal image information and target abnormal text information, and input the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction to obtain valid image features and valid text features;
[0155] The abnormal classification module 706 is used to fuse the effective image features and the effective text features to obtain effective fusion features, and obtain the information of the object to be tested, input the information of the object to be tested and the effective fusion features into a pre-trained abnormal classifier for classification, and obtain the image abnormal state classification result.
[0156] By adopting the above-mentioned image abnormal state classification device, this embodiment can obtain abnormal image information and abnormal text information; perform mask processing on the abnormal image information to obtain an input image pair, and perform contrast learning based on the input image pair to obtain an image contrast encoder; share parameters between the image contrast encoder and the preset text encoder to obtain a shared text encoder, and input the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder; obtain abnormal annotation data, and perform supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotation data to obtain an adjusted image encoder and an adjusted text encoder; obtain target abnormal image information and target abnormal text information, and input the target abnormal image information and the target abnormal text information into the adjusted image encoder and the adjusted text encoder for feature extraction to obtain effective image features and effective text features; fuse the effective image features and the effective text features to obtain effective fusion features, and obtain information of the object to be tested, and input the information of the object to be tested and the effective fusion features into a pre-trained abnormal classifier for classification to obtain an image abnormal state classification result. This effectively achieves accurate classification of abnormal states described by abnormal image information and abnormal text information, and improves the accuracy of image abnormal state classification.
[0157] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0158] The computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 8 with components 81-83, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.
[0159] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with a user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device.
[0160] The memory 81 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 81 can be an internal storage unit of the computer device 8, such as a hard disk or memory of the computer device 8. In other embodiments, the memory 81 can also be an external storage device of the computer device 8, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. equipped on the computer device 8. Of course, the memory 81 can also include both the internal storage unit of the computer device 8 and its external storage device. In this embodiment, the memory 81 is generally used to store the operating system and various application software installed on the computer device 8, such as computer-readable instructions of the abnormal image state classification method, etc. In addition, the memory 81 can also be used to temporarily store various types of data that have been output or are to be output.
[0161] The processor 82 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 82 is generally used to control the overall operation of the computer device 8. In this embodiment, the processor 82 is used to run computer-readable instructions or process data stored in the memory 81, such as computer-readable instructions for running the image abnormal state classification method.
[0162] The network interface 83 may include a wireless network interface or a wired network interface. The network interface 83 is generally used to establish a communication connection between the computer device 8 and other electronic devices.
[0163] By adopting the above-mentioned computer device, this embodiment can obtain abnormal image information and abnormal text information; perform mask processing on the abnormal image information to obtain an input image pair, and perform contrast learning based on the input image pair to obtain an image contrast encoder; share parameters between the image contrast encoder and the preset text encoder to obtain a shared text encoder, and input the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder; obtain abnormal annotation data, and perform supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotation data to obtain an adjusted image encoder and an adjusted text encoder; obtain target abnormal image information and target abnormal text information, and input the target abnormal image information and the target abnormal text information into the adjusted image encoder and the adjusted text encoder for feature extraction to obtain effective image features and effective text features; fuse the effective image features and the effective text features to obtain effective fused features, and obtain information about the object to be tested, and input the information about the object to be tested and the effective fused features into a pre-trained abnormal classifier for classification to obtain an image abnormal state classification result. This effectively achieves accurate classification of abnormal states described by abnormal image information and abnormal text information, and improves the accuracy of image abnormal state classification.
[0164] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned image abnormality state classification method.
[0165] By adopting the above-mentioned computer-readable storage medium, this embodiment can obtain abnormal image information and abnormal text information; perform mask processing on the abnormal image information to obtain an input image pair, and perform contrast learning based on the input image pair to obtain an image contrast encoder; share parameters between the image contrast encoder and the preset text encoder to obtain a shared text encoder, and input the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder; obtain abnormal annotation data, and perform supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotation data to obtain an adjusted image encoder and an adjusted text encoder; obtain target abnormal image information and target abnormal text information, and input the target abnormal image information and the target abnormal text information into the adjusted image encoder and the adjusted text encoder for feature extraction to obtain effective image features and effective text features; fuse the effective image features and the effective text features to obtain effective fusion features, and obtain information about the object to be tested, and input the information about the object to be tested and the effective fusion features into a pre-trained abnormal classifier for classification to obtain an image abnormal state classification result. This effectively achieves accurate classification of abnormal states described by abnormal image information and abnormal text information, and improves the accuracy of image abnormal state classification.
[0166] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0167] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.
[0168] The non-Company software tools or components appearing in the embodiments of this application are merely examples and do not represent actual use.
Claims
1. A method for classifying abnormal image states, characterized in that: The steps include: Obtain abnormal image information and abnormal text information; Performing mask processing on the abnormal image information to obtain an input image pair, and performing contrast learning based on the input image pair to obtain an image contrast encoder; The image contrast encoder and the preset text encoder are parameter-shared to obtain a shared text encoder, and the abnormal image information and the abnormal text information are input into the image contrast encoder and the shared text encoder for comparative learning to obtain an optimized image encoder and an optimized text encoder; Acquire abnormal annotated data, and perform supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotated data to obtain an adjusted image encoder and an adjusted text encoder; Acquire target abnormal image information and target abnormal text information, input the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction, and obtain effective image features and effective text features; The effective image features and the effective text features are fused to obtain effective fused features, and the information of the object to be tested is obtained. The information of the object to be tested and the effective fused features are input into a pre-trained abnormality classifier for classification to obtain an image abnormality state classification result.
2. The image abnormality classification method according to claim 1, characterized in that: The step of obtaining abnormal image information and abnormal text information specifically includes: Acquire the abnormal object number information, and extract the abnormal analysis report data from the database according to the abnormal object number information; Performing image extraction and text extraction on the abnormal analysis report data to obtain initial abnormal image information and initial abnormal text information; The initial abnormal image information and the initial abnormal text information are preprocessed to obtain the abnormal image information and the abnormal text information.
3. The image abnormality classification method according to claim 1, characterized in that: The step of performing mask processing on the abnormal image information to obtain an input image pair, and performing contrast learning based on the input image pair to obtain an image contrast encoder specifically includes: Randomly masking the abnormal image information to generate an input image pair, wherein the input image pair includes a first mask image and a second mask image; Inputting the first mask image and the second mask image into a first image encoder and a second image encoder that share a weight respectively for feature extraction to obtain a first image feature representation and a second image feature representation; Performing low-dimensional projection on the first image feature representation and the second image feature representation through a first projection block and a second projection block that share weights, respectively, to obtain a first low-dimensional feature representation and a second low-dimensional feature representation; Performing loss calculation according to the first low-dimensional feature representation and the second low-dimensional feature representation to obtain an image loss result; The first image encoder and the second image encoder are optimized according to the image loss result to obtain the image contrast encoder.
4. The image abnormality classification method according to claim 1, characterized in that: The step of sharing parameters of the image contrast encoder and the preset text encoder to obtain a shared text encoder, and inputting the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for comparative learning to obtain an optimized image encoder and an optimized text encoder specifically includes: Inputting the abnormal image information into the image contrast encoder to obtain abnormal image coding features; Obtain a preset text encoder, and perform weight sharing between the image comparison encoder and the preset text encoder to obtain the shared text encoder; Inputting the abnormal text information into the shared text encoder to obtain abnormal text encoding features; Performing low-dimensional projection on the abnormal text encoding features to obtain abnormal text low-dimensional features; Globally pooling the abnormal image encoding features and the abnormal text low-dimensional features to obtain a first pooling feature and a second pooling feature; The image contrast encoder and the shared text encoder are optimized according to the first pooling feature and the second pooling feature to obtain the optimized image encoder and the optimized text encoder.
5. The image abnormality classification method according to claim 1, characterized in that: The step of obtaining abnormal annotated data, and performing supervised adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotated data to obtain an adjusted image encoder and an adjusted text encoder specifically includes: Obtaining a data extraction identifier, and extracting the abnormal annotated data from a database according to the data extraction identifier; Parsing the abnormal annotation data to obtain abnormal image annotation data and abnormal text annotation data; The abnormal image annotated data and the abnormal text annotated data are used as supervisory signals, and the optimized image encoder and the optimized text encoder are fine-tuned with the supervisory signals to obtain the adjusted image encoder and the adjusted text encoder.
6. The image abnormality classification method according to claim 1, characterized in that: The step of obtaining target abnormal image information and target abnormal text information, inputting the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction, and obtaining effective image features and effective text features specifically includes: Obtaining a target object number, and extracting the target abnormal image information and the target abnormal text information from a database according to the target object number; Inputting the target abnormal image information into the adjustment image encoder to extract image features to obtain the effective image features; The target abnormal text information is input into the adjustment text encoder to extract text features to obtain the effective text features.
7. The image abnormality classification method according to claim 1, characterized in that: Before the step of fusing the effective image features and the effective text features to obtain effective fusion features, obtaining information of the object to be tested, inputting the information of the object to be tested and the effective fusion features into a pre-trained abnormality classifier for classification, and obtaining the image abnormal state classification result, the following steps are also included: Obtaining sample test object information and sample fusion features; quantifying the influence of the sample to-be-tested object information on the sample fusion feature according to a pre-built object feature association model, and obtaining a personalized fusion feature weight; Calculating a weighted feature vector according to the personalized fusion feature weight and the sample fusion feature; Constructing a basic classifier based on an ensemble learning algorithm, and adjusting the weight of the basic classifier according to the weighted feature vector to obtain a personalized ensemble classifier; Acquire sample abnormal object data, train the personalized integrated classifier according to the sample abnormal object data, and optimize the personalized integrated classifier based on cross-validation to obtain the abnormal classifier.
8. An image abnormal state classification device, characterized in that: include: An information acquisition module, used to acquire abnormal image information and abnormal text information; A first contrast learning module is used to perform mask processing on the abnormal image information to obtain an input image pair, and perform contrast learning based on the input image pair to obtain an image contrast encoder; A second contrast learning module is used to share parameters of the image contrast encoder and the preset text encoder to obtain a shared text encoder, and input the abnormal image information and the abnormal text information into the image contrast encoder and the shared text encoder for contrast learning to obtain an optimized image encoder and an optimized text encoder; A supervision adjustment module, used for acquiring abnormal annotation data, and performing supervision adjustment on the optimized image encoder and the optimized text encoder according to the abnormal annotation data to obtain an adjusted image encoder and an adjusted text encoder; A feature extraction module, used for acquiring target abnormal image information and target abnormal text information, inputting the target abnormal image information and the target abnormal text information into the adjustment image encoder and the adjustment text encoder for feature extraction, and obtaining effective image features and effective text features; The abnormal classification module is used to fuse the effective image features and the effective text features to obtain effective fusion features, and obtain the information of the object to be tested, and input the information of the object to be tested and the effective fusion features into a pre-trained abnormal classifier for classification to obtain the image abnormal state classification result.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the abnormal image state classification method according to any one of claims 1 to 7 when executing the computer-readable instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the image abnormal state classification method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Automatic driving obstacle intention prediction and avoidance method and system
CN120635865A
Network security test coverage assessment method and system based on deep learning
CN121037125A