An image language fusion prediction learning method for cervical cytology unbalanced data
By segmenting whole cervical cytology slides into patch images and filtering out suspicious positive cell images, and extracting and fusing features related to clinical text, the problem of information loss or confusion in multimodal data prediction is solved, thereby improving the accuracy of cervical lesion category prediction.
Patent Information
- Application Number
- CN202411973266.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing methods for predicting cervical lesion categories fail to fully utilize the complementarity and interrelationships between multimodal data, resulting in information loss or confusion and low accuracy.
The whole cervical cytology slide image was segmented into multiple patch images, suspicious positive cell images were screened out, and features related to clinical text were extracted from the features of these images. The image and text features were then fused to predict the cervical lesion category.
By fully utilizing the complementarity and interrelationship between images and text, and avoiding information loss or confusion, the accuracy of cervical lesion category prediction is improved.
Smart Images

Figure CN119784734B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to an image language fusion prediction learning method for cervical cytology unbalanced data. BACKGROUND
[0002] Thanks to the development of fully digital automatic scanners, the cervical cytology-based whole slide image (WSI) computing cervical pathology classification algorithm has made great progress, greatly reducing the burden of pathologists. However, these methods only use the whole slide image, ignoring the different levels and scales of pathological information provided by computer tomography (CT), magnetic resonance imaging (MRI), fluorescence imaging, clinical case information, diagnosis text, etc., hindering the development of more accurate pathological diagnosis and more accurate disease classification methods. In order to utilize information from different modalities, it is necessary to jointly model multiple modalities, reveal the potential associations and implicit knowledge between modalities, and extract more comprehensive and discriminative features based on the complementarity between different modalities. In traditional multi-modal data joint modeling, one of the commonly used methods is feature fusion. This method extracts feature representations of each modality data, and then combines or fuses these features to obtain a comprehensive representation. However, due to insufficient model capacity and simple fusion or splicing between modalities, it may not be able to fully utilize their complementarity and mutual relationship, resulting in information loss or confusion, poor modeling effect, and thus low accuracy of cervical lesion class prediction. SUMMARY
[0003] Therefore, it is necessary to provide an image language fusion prediction learning method for cervical cytology unbalanced data, which can improve the accuracy of cervical lesion class prediction.
[0004] In a first aspect, the present application provides an image language fusion prediction learning method for cervical cytology unbalanced data, which comprises:
[0005] Obtaining a cervical cytology whole slide image, and cutting the cervical cytology whole slide image into a plurality of patch images;
[0006] Detecting the plurality of patch images to screen a plurality of suspicious positive cell images from the plurality of patch images;
[0007] Obtaining clinical text corresponding to the cervical cytology whole slide image;
[0008] From the first image features of the plurality of suspicious positive cell images, extracting a plurality of second image features related to the clinical text;
[0009] The second image features and text features of the clinical text are fused to obtain image language fusion features, and a cervical lesion category to which the cervical cytology whole slide image belongs is predicted based on the image language fusion features.
[0010] In a second aspect, the present application provides an image language fusion prediction learning device for cervical cytology unbalanced data, the device comprising:
[0011] An acquisition module is configured to acquire a cervical cytology whole slide image and cut the cervical cytology whole slide image into a plurality of patch images;
[0012] A detection module is configured to detect the plurality of patch images to screen a plurality of suspicious positive cell images from the plurality of patch images;
[0013] The acquisition module is further configured to acquire clinical text corresponding to the cervical cytology whole slide image;
[0014] An extraction module is configured to extract a plurality of second image features related to the clinical text from first image features of the plurality of suspicious positive cell images;
[0015] A prediction module is configured to fuse the plurality of second image features and text features of the clinical text to obtain image language fusion features, and predict a cervical lesion category to which the cervical cytology whole slide image belongs based on the image language fusion features.
[0016] In a third aspect, the present application provides a computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing steps in each method embodiment of the present application when executing the computer program.
[0017] In a fourth aspect, the present application provides a computer readable storage medium storing a computer program, and the computer program implementing steps in each method embodiment of the present application when executed by a processor.
[0018] In a fifth aspect, the present application provides a computer program product comprising a computer program, and the computer program implementing steps in each method embodiment of the present application when executed by a processor.
[0019] The aforementioned image-language fusion prediction learning method for imbalanced cervical cytology data involves: acquiring a whole cervical cytology slide image and segmenting it into multiple patch images; detecting these patch images to identify multiple suspicious positive cell images; obtaining the clinical text corresponding to the whole cervical cytology slide image; extracting multiple second image features related to the clinical text from the first image features of the multiple suspicious positive cell images; fusing the multiple second image features and the text features of the clinical text to obtain image-language fusion features; and predicting the cervical lesion category of the whole cervical cytology slide image based on these image-language fusion features. Compared to traditional methods that rely solely on whole cervical cytology slides for cervical lesion category prediction, or simply fuse the image features of whole cervical cytology slides with the textual features of corresponding text, this application first segments the whole cervical cytology slide into multiple patch images, then filters out multiple suspicious positive cell images from these patch images, and further extracts multiple second image features closely related to clinical text from the first image features of these multiple suspicious positive cell images. These extracted second image features are then fused with the textual features of the clinical text, and cervical lesion category prediction is performed based on the fused features. This approach fully utilizes the complementarity and interrelationship between images and text, avoids information loss or confusion, and improves the accuracy of cervical lesion category prediction. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating an image-language fusion prediction learning method for cervical cytology imbalance data in one embodiment.
[0021] Figure 2 This is a schematic diagram illustrating the training process of a cervical lesion prediction model in one embodiment;
[0022] Figure 3 This is a schematic diagram of the network architecture of an image-text alignment network in one embodiment;
[0023] Figure 4 This is a block diagram of an image-language fusion predictive learning device for cervical cytology imbalance data in one embodiment.
[0024] Figure 5 This is an internal structural diagram of a computer device in one embodiment;
[0025] Figure 6 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation
[0026] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not used to limit the present application.
[0027] In one embodiment, as shown in Figure 1 An image language fusion prediction learning method for cervical cytology unbalanced data is provided, which can be applied to a computer device, which can be a terminal or a server, that is, the method can be executed by the terminal or the server alone, or can be realized through interaction between the terminal and the server. This embodiment takes the method applied to the computer device as an example for description, which includes the following steps:
[0028] Step 102, acquiring a cervical cytology whole slide image, and cutting the cervical cytology whole slide image into a plurality of patch images.
[0029] It can be understood that in a medical scenario, the pixel size of the cervical cytology whole slide image is extremely large, and before performing cervical lesion category prediction, the computer device can first cut the cervical cytology whole slide image into a plurality of patch images. For example, the cervical cytology whole slide image can be cut into a plurality of 1024x1024 size patch images.
[0030] Step 104, detecting the plurality of patch images to screen a plurality of suspicious positive cell images from the plurality of patch images.
[0031] In one embodiment, the computer device can input the plurality of patch images into a trained cell detection model to screen a plurality of suspicious positive cell images from the plurality of patch images through the cell detection model. It can be understood that the number of suspicious positive cell images is less than the number of patch images.
[0032] Step 106, acquiring clinical text corresponding to the cervical cytology whole slide image.
[0033] The clinical text corresponding to the cervical cytology whole slide image is a text used in medical clinics to describe the object to which the cervical cytology whole slide image belongs.
[0034] In one embodiment, the clinical text corresponding to the cervical cytology whole slide image can be at least one of a case, a medical history or an age of the object to which the cervical cytology whole slide image belongs.
[0035] Step 108, extracting a plurality of second image features related to the clinical text from first image features of the plurality of suspicious positive cell images.
[0036] In an embodiment, the computer device can input the first image features of the plurality of suspicious positive cell images into the trained image-text alignment network to extract a plurality of second image features closely related to the clinical text from the first image features of the plurality of suspicious positive cell images through the image-text alignment network. It can be understood that the number of second image features is less than that of first image features.
[0037] In step 110, the second image features and the text features of the clinical text are fused to obtain image-language fusion features, and the cervical lesion category to which the cervical cytology whole slide image belongs is predicted based on the image-language fusion features.
[0038] In an embodiment, the computer device can splice the second image features and the text features of the clinical text to obtain image-language fusion features, and predict the cervical lesion category to which the cervical cytology whole slide image belongs based on the image-language fusion features.
[0039] In the above image-language fusion prediction learning method for cervical cytology unbalanced data, the cervical cytology whole slide image is obtained, and the cervical cytology whole slide image is cut into a plurality of patch images; the plurality of patch images are detected to screen a plurality of suspicious positive cell images from the plurality of patch images; the clinical text corresponding to the cervical cytology whole slide image is obtained; a plurality of second image features related to the clinical text are extracted from first image features of the plurality of suspicious positive cell images; the second image features and the text features of the clinical text are fused to obtain image-language fusion features, and the cervical lesion category to which the cervical cytology whole slide image belongs is predicted based on the image-language fusion features. Compared with the traditional cervical lesion category prediction based only on the cervical cytology whole slide image, or the simple fusion of the image features of the cervical cytology whole slide image and the text features of the corresponding text, the present application cuts the cervical cytology whole slide image into a plurality of patch images, and screens a plurality of suspicious positive cell images from the plurality of patch images, and then extracts a plurality of second image features closely related to the clinical text from the first image features of the plurality of suspicious positive cell images, and fuses the extracted second image features and the text features of the clinical text, so as to predict the cervical lesion category based on the fused features. This can make full use of the complementarity and mutual relationship between image and text, avoid information loss or confusion, and improve the accuracy of cervical lesion category prediction.
[0040] In an embodiment, the cervical lesion category is predicted by a trained cervical lesion prediction model; the cervical lesion prediction model comprises a trained visual network, a trained image-text alignment network, and a trained prediction network; the first image features are extracted by the visual network in the cervical lesion prediction model; the second image features are extracted by the image-text alignment network in the cervical lesion prediction model; and the cervical lesion category is predicted by the prediction network in the cervical lesion prediction model.
[0041] In the above embodiment, the cervical lesion category to which the cervical cytology whole slide image belongs is predicted by a trained cervical lesion prediction model, and the cervical lesion prediction model is trained by a positive and negative sample balancing method, which can further improve the accuracy of cervical lesion category prediction.
[0042] In an embodiment, the image-text alignment network comprises an image processing unit and a text processing unit; the image processing unit and the text processing unit share a feedforward layer parameter; the method further comprises an image-text alignment network training step; the image-text alignment network training step comprises: obtaining a plurality of first image-text sample pairs; each first image-text sample pair comprises a plurality of first sample suspicious positive cell images and corresponding first sample clinical text; for each first image-text sample pair, inputting the first sample image features of the plurality of first sample suspicious positive cell images and the first sample clinical text in the first image-text sample pair into the image-text alignment network to be trained, so as to extract, by the image processing unit in the image-text alignment network to be trained, a plurality of image features related to the first sample clinical text from the first sample image features of the plurality of first sample suspicious positive cell images, obtain second sample image features corresponding to the first image-text sample pair, extract features of the first sample clinical text by the text processing unit in the image-text alignment network to be trained, and obtain sample text features corresponding to the first image-text sample pair; determining a first loss value according to the second sample image features and the sample text features corresponding to each of the plurality of first image-text sample pairs; training the image-text alignment network to be trained according to the first loss value to obtain a trained image-text alignment network.
[0043] In the above embodiment, during the training of the image-text alignment network, the similarity of the same sample pair is increased, and the similarity of the non-paired sample is reduced, so that the image features and the text features are aligned, a joint image and text representation space is constructed, and thus the second image features closely related to the clinical text can be accurately extracted from the first image features of the suspicious positive cell images, thereby further improving the accuracy of cervical lesion category prediction.
[0044] In an embodiment, the image-text alignment network to be trained is trained according to the first loss value, and a trained image-text alignment network is obtained, including: obtaining category token features; for each first image-text sample pair, the category token features are interacted with first sample image features of a first sample suspicious positive cell image in the first image-text sample pair and sample text features corresponding to first sample clinical text to obtain cervical lesion prediction task features, the cervical lesion prediction task features are input into a fully connected classifier to predict first cervical lesion category probabilities corresponding to the first image-text sample pair; second loss values are determined according to the first cervical lesion category probabilities corresponding to the plurality of first image-text sample pairs respectively; the image-text alignment network to be trained is trained according to the first loss value and the second loss value, and the trained image-text alignment network is obtained.
[0045] In an embodiment, the computer device can determine a target loss value according to the first loss value and the second loss value, and train the image-text alignment network to be trained based on the target loss value to obtain the trained image-text alignment network. The target loss value can be calculated by the following formula:
[0046] ;
[0047] ;
[0048] ;
[0049] wherein i represents the i th first image-text sample pair in the n first image-text sample pairs, represents the second sample image features corresponding to the i th first image-text sample pair, represents the sample text features corresponding to the i th first image-text sample pair, represents the n first image-text sample pairs other than the i th first image-text sample pair, and t represents a preset temperature coefficient, represents the first loss value. Herein, represents the first cervical lesion category probability corresponding to the i th first image-text sample pair in the n first image-text sample pairs, represents the second loss value, represents the target loss value.
[0050] In the above embodiments, the target loss value is determined according to the first loss value and the second loss value, and the image-text alignment network to be trained is trained based on the target loss value, which can further improve the performance of the trained image-text alignment network, and can more accurately extract second image features closely related to clinical text from first image features of suspicious positive cell images, thereby further improving the accuracy of cervical lesion category prediction.
[0051] In an embodiment, the trained image-text alignment network is a preliminary trained image-text alignment network; the method further comprises a cervical lesion prediction model training step; the cervical lesion prediction model training step comprises: obtaining a plurality of second image-text sample pairs; each second image-text sample pair comprises a plurality of second sample suspicious positive cell images and corresponding second sample clinical texts; the second sample suspicious positive cell image is obtained by covering a mask on a central region of the first sample suspicious positive cell image in the first image-text sample pair; the second sample clinical text is the first sample clinical text in the corresponding first image-text sample pair; for each second image-text sample pair, the plurality of second sample suspicious positive cell images and the second sample clinical text in the second image-text sample pair are input into the cervical lesion prediction model to be trained, so as to extract third sample image features of the plurality of second sample suspicious positive cell images through a visual network to be trained in the cervical lesion prediction model to be trained, extract a plurality of image features related to the second sample clinical text from the third sample image features of the plurality of second sample suspicious positive cell images through a preliminary trained image-text alignment network to be trained in the cervical lesion prediction model to be trained, obtain fourth sample image features corresponding to the second image-text sample pair, and predict a second cervical lesion category probability corresponding to the second image-text sample pair based on the fourth sample image features and sample text features of the second sample clinical text through a prediction network in the cervical lesion prediction model to be trained; a third loss value is determined according to the cervical lesion category probabilities corresponding to the plurality of second image-text sample pairs respectively; for each first image-text sample pair, the plurality of first sample suspicious positive cell images and the first sample clinical text in the first image-text sample pair are input into the cervical lesion prediction model to be trained, so as to extract fifth sample image features of the plurality of first sample suspicious positive cell images through the visual network to be trained in the cervical lesion prediction model to be trained, extract a plurality of image features related to the first sample clinical text from the fifth sample image features of the plurality of first sample suspicious positive cell images through the preliminary trained image-text alignment network to be trained in the cervical lesion prediction model to be trained, obtain sixth sample image features corresponding to the first image-text sample pair, and predict a third cervical lesion category probability corresponding to the first image-text sample pair based on the sixth sample image features and sample text features of the first sample clinical text through the prediction network in the cervical lesion prediction model to be trained; a fourth loss value is determined according to the third cervical lesion category probabilities corresponding to the plurality of first image-text sample pairs respectively; the cervical lesion prediction model to be trained is trained according to the third loss value and the fourth loss value, and a trained cervical lesion prediction model is obtained.
[0052] In an embodiment, the third loss value can be calculated by the following formula:
[0053] ;
[0054] wherein, here, denotes the second cervical lesion category probability corresponding to the i-th second image-text sample pair in the n second image-text sample pairs, denotes the third loss value.
[0055] In the above embodiment, in the medical scene, the data tends to be long-tail distribution according to the category, that is, the category is unbalanced. By covering the center area of the first sample suspicious positive cell image with a mask and adding the same number of negative samples to balance the training process, the problem that the model pays too much attention to the majority class and causes poor performance of the minority class is avoided, thereby further improving the accuracy of cervical lesion category prediction.
[0056] In one embodiment, a plurality of patch images are detected to screen a plurality of suspicious positive cell images from the plurality of patch images, including: inputting the plurality of patch images into a trained cell detection model to determine the detection frame of the positive cell in the plurality of patch images and its confidence by the cell detection model; selecting the region corresponding to the detection frame of the positive cell with a preset number of confidence as the plurality of suspicious positive cell images from the detection frame of the positive cell in the plurality of patch images.
[0057] In the above embodiment, by the cell detection model, a plurality of suspicious positive cell images with a preset number of confidence are detected from a plurality of patch images as training data for a subsequent cervical lesion prediction model, which can further improve the efficiency and accuracy of the model training.
[0058] In one embodiment, as Figure 2As shown, the cervical lesion prediction model comprises a trained vision network (VisionTransformer), a trained image-text alignment network (Q-Former), and a prediction network (large language model). The computer device can obtain a cervical cytology whole slide image, and cut the cervical cytology whole slide image into a plurality of patch images; input the plurality of patch images into the trained cell detection model to determine the detection frame and confidence of positive cells in the plurality of patch images through the cell detection model; from the plurality of detected cells, select the top k confidence as a plurality of suspicious positive cell images (i.e. Topk). The computer device can input the plurality of suspicious positive cell images into the vision network to obtain the corresponding image features, input the corresponding image features of the plurality of suspicious positive cell images and a pre-initialized set of image feature queries (Queries) into the image-text alignment network to extract a part of image features related to the clinical text from the corresponding image features of the plurality of suspicious positive cell images, splice the extracted image features, and input the spliced image features and the clinical text corresponding to the cervical cytology whole slide image into the prediction network to predict the cervical lesion category of the cervical cytology whole slide image (i.e. output the cervical lesion category in the form of text). It can be understood that in the training stage of the cervical lesion prediction model, the clinical text corresponding to the cervical cytology whole slide image used for training needs to be input into the image-text alignment network for training, so that the image-text alignment network has the ability to extract a part of image features related to the clinical text. Therefore, in the application stage of the cervical lesion prediction model, the clinical text corresponding to the cervical cytology whole slide image does not need to be input into the image-text alignment network.
[0059] In one embodiment, as shown in Figure 3 The image-text alignment network (Q-Former) comprises two branches, one for processing image processing units (i.e. image Transformer) for processing learnable image feature queries query and one for processing text processing units (i.e. text Transformer) for processing clinical text, and both are stacked by multiple blocks, wherein the constituent blocks of the image processing unit comprise self-attention layers, cross-attention layers and feedforward layers, and each cross-attention layer interacts with the image features output by the Vision Transformer. The constituent blocks of the text processing unit are less cross-attention layers, and the image processing unit and the text processing unit share parameters.
[0060] It should be understood that although each step in the flowchart of each of the above embodiments is shown in sequence, these steps are not necessarily executed in sequence. Unless explicitly stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least some of the steps in each of the above embodiments can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these sub-steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or sub-steps or stages of other steps.
[0061] In one embodiment, as shown in Figure 4 An image language fusion prediction learning device 400 for cervical cytology unbalanced data is provided, which specifically comprises:
[0062] An acquisition module 402 is configured to acquire a cervical cytology whole slide image, and cut the cervical cytology whole slide image into a plurality of patch images;
[0063] A detection module 404 is configured to detect the plurality of patch images to screen a plurality of suspicious positive cell images from the plurality of patch images;
[0064] The acquisition module 402 is further configured to acquire clinical text corresponding to the cervical cytology whole slide image;
[0065] An extraction module 406 is configured to extract a plurality of second image features related to the clinical text from first image features of the plurality of suspicious positive cell images;
[0066] A prediction module 408 is configured to fuse the plurality of second image features and text features of the clinical text to obtain image language fusion features, and predict a cervical lesion category to which the cervical cytology whole slide image belongs based on the image language fusion features.
[0067] In one embodiment, the cervical lesion category is predicted by a trained cervical lesion prediction model; the cervical lesion prediction model comprises a trained visual network, a trained image text alignment network, and a trained prediction network; the first image features are extracted by the visual network in the cervical lesion prediction model; the second image features are extracted by the image text alignment network in the cervical lesion prediction model; and the cervical lesion category is predicted by the prediction network in the cervical lesion prediction model.
[0068] In one embodiment, the image text alignment network comprises an image processing unit and a text processing unit; the image processing unit and the text processing unit share feedforward layer parameters; and the device further comprises:
[0069] The training module is configured to obtain a plurality of first image-text sample pairs, each of which comprises a plurality of first sample suspicious positive cell images and corresponding first sample clinical texts; for each first image-text sample pair, input the first sample image features of the plurality of first sample suspicious positive cell images and the first sample clinical texts in the first image-text sample pair into the image-text alignment network to be trained, so as to extract a plurality of image features related to the first sample clinical texts from the first sample image features of the plurality of first sample suspicious positive cell images by an image processing unit in the image-text alignment network to be trained, obtain second sample image features corresponding to the first image-text sample pair, extract features of the first sample clinical texts by a text processing unit in the image-text alignment network to be trained, and obtain sample text features corresponding to the first image-text sample pair; determine a first loss value according to the second sample image features and the sample text features corresponding to each of the plurality of first image-text sample pairs; and train the image-text alignment network to be trained according to the first loss value, and obtain a trained image-text alignment network.
[0070] In one embodiment, the training module is further configured to obtain category token features; for each first image-text sample pair, interact the category token features with the first sample image features of the first sample suspicious positive cell images in the first image-text sample pair and the sample text features of the corresponding first sample clinical texts, obtain cervical lesion prediction task features, input the cervical lesion prediction task features into a fully connected classifier, and predict first cervical lesion category probabilities corresponding to the first image-text sample pair; determine a second loss value according to the first cervical lesion category probabilities corresponding to each of the plurality of first image-text sample pairs; and train the image-text alignment network to be trained according to the first loss value and the second loss value, and obtain a trained image-text alignment network.
[0071] In an embodiment, the trained image-text alignment network is a preliminary trained image-text alignment network; the training module is further configured to obtain a plurality of second image-text sample pairs; each of the second image-text sample pairs comprises a plurality of second sample suspicious positive cell images and corresponding second sample clinical texts; the second sample suspicious positive cell images are obtained by covering a mask on a central region of the first sample suspicious positive cell images in the first image-text sample pair; the second sample clinical texts are the first sample clinical texts in the corresponding first image-text sample pair; for each of the second image-text sample pairs, the plurality of second sample suspicious positive cell images and the second sample clinical text in the second image-text sample pair are input into the cervical lesion prediction model to be trained, so as to extract third sample image features of the plurality of second sample suspicious positive cell images by the visual network to be trained in the cervical lesion prediction model to be trained, extract a plurality of image features related to the second sample clinical text from the third sample image features of the plurality of second sample suspicious positive cell images by the preliminary trained image-text alignment network to be trained in the cervical lesion prediction model to be trained, obtain fourth sample image features corresponding to the second image-text sample pair, and predict a second cervical lesion category probability corresponding to the second image-text sample pair based on the fourth sample image features and sample text features of the second sample clinical text by the prediction network in the cervical lesion prediction model to be trained; determine a third loss value according to the second cervical lesion category probabilities corresponding to the plurality of second image-text sample pairs; for each of the first image-text sample pairs, the plurality of first sample suspicious positive cell images and the first sample clinical text in the first image-text sample pair are input into the cervical lesion prediction model to be trained, so as to extract fifth sample image features of the plurality of first sample suspicious positive cell images by the visual network to be trained in the cervical lesion prediction model to be trained, extract a plurality of image features related to the first sample clinical text from the fifth sample image features of the plurality of first sample suspicious positive cell images by the preliminary trained image-text alignment network to be trained in the cervical lesion prediction model to be trained, obtain sixth sample image features corresponding to the first image-text sample pair, and predict a third cervical lesion category probability corresponding to the first image-text sample pair based on the sixth sample image features and sample text features of the first sample clinical text by the prediction network in the cervical lesion prediction model to be trained; determine a fourth loss value according to the third cervical lesion category probabilities corresponding to the plurality of first image-text sample pairs; and train the cervical lesion prediction model to be trained according to the third loss value and the fourth loss value, to obtain the trained cervical lesion prediction model.
[0072] In an embodiment, the detection module 404 is further configured to input the plurality of patch images into the trained cell detection model to determine, by the cell detection model, bounding boxes of positive cells in the plurality of patch images and confidence thereof; and select, from the bounding boxes of the positive cells detected from the plurality of patch images, regions corresponding to a preset number of bounding boxes of positive cells with higher confidence as the plurality of suspicious positive cell images.
[0073] The image language fusion prediction learning device for the imbalanced data of cervical cytology described above, by acquiring a cervical cytology whole slide image and cutting the cervical cytology whole slide image into a plurality of patch images, detecting the plurality of patch images to screen a plurality of suspicious positive cell images from the plurality of patch images, acquiring clinical text corresponding to the cervical cytology whole slide image, extracting a plurality of second image features related to the clinical text from first image features of the plurality of suspicious positive cell images, fusing the plurality of second image features and text features of the clinical text to obtain image language fusion features, and predicting a cervical lesion category to which the cervical cytology whole slide image belongs based on the image language fusion features. Compared with the traditional cervical lesion category prediction based on only the cervical cytology whole slide image or the simple fusion of image features of the cervical cytology whole slide image and text features of the corresponding text, the present application can fully utilize the complementarity and mutual relationship between the image and the text, avoid information loss or confusion, and improve the accuracy of the cervical lesion category prediction.
[0074] Each module in the image language fusion prediction learning device for the imbalanced data of cervical cytology described above can be realized by software, hardware, or a combination thereof, in whole or in part. Each module described above can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to each module.
[0075] In an embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in FIG. 8. Figure 5As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control ability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with the terminal outside through the network connection. The computer program is executed by the processor to realize an image language fusion prediction learning method for cervical cytology unbalanced data.
[0076] In one embodiment, a computer device is provided, which can be a terminal, and its internal structure diagram can be as shown in the figure. Figure 6 As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control ability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with the terminal outside through the network connection. The computer program is executed by the processor to realize an image language fusion prediction learning method for cervical cytology unbalanced data.
[0077] Those skilled in the art can understand that, Figure 5 and Figure 6It should be noted that the structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0078] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above-mentioned method embodiments when executing the computer program.
[0079] In one embodiment, a computer readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the steps in the above-mentioned method embodiments.
[0080] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps in the above-mentioned method embodiments.
[0081] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0082] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned method embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. In the embodiments provided by the present application, any reference to memory, storage, database or other medium can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0083] Any combination of the technical features in the above embodiments can be made, and for the sake of brevity, not all possible combinations are described above, however, as long as the combination of the technical features does not exist in contradiction, it shall be considered within the scope of the present disclosure.
[0084] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it shall not be understood as a limitation on the patent scope of the present application. It shall be pointed out that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these shall be within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An image language fusion prediction learning method for cervical cytology unbalanced data, characterized in that, The method comprises: acquiring a cervical cytology whole slide image and cutting the cervical cytology whole slide image into a plurality of patch images; detecting the plurality of patch images to screen a plurality of suspicious positive cell images from the plurality of patch images; the detection of the plurality of patch images to screen a plurality of suspicious positive cell images from the plurality of patch images comprises: inputting the plurality of patch images into a trained cell detection model to determine the detection frame of a positive cell in the plurality of patch images and its confidence by the cell detection model; selecting a region corresponding to a detection frame of a positive cell with a top confidence from the detection frame of the positive cell detected from the plurality of patch images as a plurality of suspicious positive cell images; acquiring clinical text corresponding to the cervical cytology whole slide image; extracting a plurality of second image features related to the clinical text from first image features of the plurality of suspicious positive cell images; fusing the plurality of second image features and text features of the clinical text to obtain image language fusion features, and predicting a cervical lesion category to which the cervical cytology whole slide image belongs based on the image language fusion features; the cervical lesion category is predicted by a trained cervical lesion prediction model; the cervical lesion prediction model comprises a trained visual network, a trained image text alignment network and a trained prediction network; the first image features are extracted by the visual network in the cervical lesion prediction model; the second image features are extracted by the image text alignment network in the cervical lesion prediction model; and the cervical lesion category is predicted by the prediction network in the cervical lesion prediction model; the image text alignment network comprises an image processing unit and a text processing unit; the image processing unit and the text processing unit share feedforward layer parameters; the method further comprises an image text alignment network training step; the image text alignment network training step comprises: acquiring a plurality of first image-text sample pairs; each first image-text sample pair comprises a plurality of first sample suspicious positive cell images and corresponding first sample clinical text; for each first image-text sample pair, inputting first sample image features of the plurality of first sample suspicious positive cell images and the first sample clinical text in the first image-text sample pair into a to-be-trained image text alignment network to extract a plurality of image features related to the first sample clinical text from the first sample image features of the plurality of first sample suspicious positive cell images by the image processing unit in the to-be-trained image text alignment network, to obtain second sample image features corresponding to the first image-text sample pair, and to extract features of the first sample clinical text by the text processing unit in the to-be-trained image text alignment network, to obtain sample text features corresponding to the first image-text sample pair; determine a first loss value according to the second sample image features and the sample text features corresponding to each of the plurality of first image-text sample pairs; train the image-text alignment network to be trained according to the first loss value, to obtain a trained image-text alignment network; the trained image-text alignment network is a preliminarily trained image-text alignment network; the method further comprises a cervical lesion prediction model training step; the cervical lesion prediction model training step comprises: obtain a plurality of second image-text sample pairs; each of the second image-text sample pairs comprises a plurality of second sample suspicious positive cell images and corresponding second sample clinical text; the second sample suspicious positive cell images are obtained by covering masks on the center regions of the first sample suspicious positive cell images in the first image-text sample pairs; and the second sample clinical text is the first sample clinical text in the corresponding first image-text sample pair; for each of the second image-text sample pairs, input the plurality of second sample suspicious positive cell images and the second sample clinical text in the second image-text sample pair into a cervical lesion prediction model to be trained, to extract third sample image features of the plurality of second sample suspicious positive cell images by a visual network to be trained in the cervical lesion prediction model to be trained, extract a plurality of image features related to the second sample clinical text from the third sample image features of the plurality of second sample suspicious positive cell images by the preliminarily trained image-text alignment network to be trained in the cervical lesion prediction model to be trained, obtain fourth sample image features corresponding to the second image-text sample pair, and predict a second cervical lesion category probability corresponding to the second image-text sample pair based on the fourth sample image features and sample text features of the second sample clinical text by a prediction network in the cervical lesion prediction model to be trained; According to the plurality of second image-text sample pairs, a third loss value is determined according to the respective corresponding second cervical lesion category probability; the third loss value is calculated by the following formula: ; wherein, represents the i-th second image-text sample pair corresponding to the second cervical lesion category probability, represents the third loss value; for each of the first image-text sample pairs, input the plurality of first sample suspicious positive cell images and the first sample clinical text in the first image-text sample pair into a cervical lesion prediction model to be trained, to extract fifth sample image features of the plurality of first sample suspicious positive cell images by a visual network to be trained in the cervical lesion prediction model to be trained, extract a plurality of image features related to the first sample clinical text from the fifth sample image features of the plurality of first sample suspicious positive cell images by the preliminarily trained image-text alignment network to be trained in the cervical lesion prediction model to be trained, obtain sixth sample image features corresponding to the first image-text sample pair, and predict a third cervical lesion category probability corresponding to the first image-text sample pair based on the sixth sample image features and sample text features of the first sample clinical text by a prediction network in the cervical lesion prediction model to be trained; determine a fourth loss value according to the third cervical lesion category probabilities corresponding to each of the plurality of first image-text sample pairs; and train the cervical lesion prediction model to be trained according to the fourth loss value, to obtain a trained cervical lesion prediction model. The third loss value and the fourth loss value are used to train the cervical lesion prediction model to be trained, and a trained cervical lesion prediction model is obtained.
2. The method of claim 1, wherein, The training of the image-text alignment network to be trained according to the first loss value comprises: obtaining category token features; For each first image-text sample pair, the category token features, the first sample image features of the suspicious positive cell image of the first sample in the first image-text sample pair, and the sample text features of the corresponding first sample clinical text are interacted to obtain cervical lesion prediction task features, and the cervical lesion prediction task features are input into a fully connected classifier to predict the first cervical lesion category probability corresponding to the first image-text sample pair; A second loss value is determined according to the first cervical lesion category probability corresponding to each of the plurality of first image-text sample pairs. The first loss value and the second loss value are used to train the image-text alignment network to be trained, and a trained image-text alignment network is obtained.
3. An image language fusion prediction learning device for cervical cytology unbalanced data, characterized in that, The device comprises: An acquisition module is configured to acquire a cervical cytology whole slice image and cut the cervical cytology whole slice image into a plurality of patch images; A detection module is configured to detect the plurality of patch images to screen a plurality of suspicious positive cell images from the plurality of patch images, and further configured to input the plurality of patch images into a trained cell detection model to determine a detection frame of a positive cell in the plurality of patch images and a confidence thereof by the cell detection model, and select a region corresponding to a detection frame of a positive cell with a top confidence from the detection frames of the positive cells detected from the plurality of patch images as a plurality of suspicious positive cell images. The acquisition module is further configured to acquire clinical text corresponding to the cervical cytology whole slice image. An extraction module is configured to extract a plurality of second image features related to the clinical text from first image features of the plurality of suspicious positive cell images. A prediction module is configured to fuse the plurality of second image features and text features of the clinical text to obtain image language fusion features, and predict a cervical lesion category to which the cervical cytology whole slice image belongs based on the image language fusion features. The cervical lesion category is predicted by a trained cervical lesion prediction model; the cervical lesion prediction model comprises a trained visual network, a trained image-text alignment network, and a trained prediction network; the first image features are extracted by the visual network in the cervical lesion prediction model; the second image features are extracted by the image-text alignment network in the cervical lesion prediction model; and the cervical lesion category is predicted by the prediction network in the cervical lesion prediction model. The image-text alignment network comprises an image processing unit and a text processing unit; the image processing unit and the text processing unit share a feedforward layer parameter; the device further comprises a training module configured to obtain a plurality of first image-text sample pairs; each of the first image-text sample pairs comprises a plurality of first sample suspicious positive cell images and corresponding first sample clinical texts; for each of the first image-text sample pairs, the first sample image features of the plurality of first sample suspicious positive cell images and the first sample clinical texts in the first image-text sample pair are input into the image-text alignment network to be trained, so as to extract, by the image processing unit in the image-text alignment network to be trained, a plurality of image features related to the first sample clinical texts from the first sample image features of the plurality of first sample suspicious positive cell images, obtain second sample image features corresponding to the first image-text sample pair, extract, by the text processing unit in the image-text alignment network to be trained, features of the first sample clinical texts, and obtain sample text features corresponding to the first image-text sample pair; a first loss value is determined according to the second sample image features and the sample text features corresponding to each of the plurality of first image-text sample pairs; the image-text alignment network to be trained is trained according to the first loss value, and a trained image-text alignment network is obtained; The trained image-text alignment network is a preliminarily trained image-text alignment network; the training module is further configured to obtain a plurality of second image-text sample pairs; each of the second image-text sample pairs comprises a plurality of second sample suspicious positive cell images and corresponding second sample clinical texts; the second sample suspicious positive cell images are obtained by covering a mask on a central region of the first sample suspicious positive cell images in the first image-text sample pairs; the second sample clinical texts are the first sample clinical texts in the corresponding first image-text sample pairs; for each of the second image-text sample pairs, the plurality of second sample suspicious positive cell images and the second sample clinical texts in the second image-text sample pair are input into a to-be-trained cervical lesion prediction model, so as to extract third sample image features of the plurality of second sample suspicious positive cell images by a to-be-trained visual network in the to-be-trained cervical lesion prediction model, extract a plurality of image features related to the second sample clinical texts from the third sample image features of the plurality of second sample suspicious positive cell images by the preliminarily trained image-text alignment network in the to-be-trained cervical lesion prediction model, obtain fourth sample image features corresponding to the second image-text sample pair, and predict a second cervical lesion category probability corresponding to the second image-text sample pair based on the fourth sample image features and sample text features of the second sample clinical texts by a prediction network in the to-be-trained cervical lesion prediction model; a third loss value is determined according to the second cervical lesion category probabilities corresponding to the plurality of second image-text sample pairs; the third loss value is calculated by the following formula: ; wherein, represents a second cervical lesion category probability corresponding to an i-th second image-text sample pair in n second image-text sample pairs, representing the third loss value; for each first image-text sample pair, input the plurality of first sample suspicious positive cell images and the first sample clinical text in the first image-text sample pair into a cervical lesion prediction model to be trained, to extract fifth sample image features of the plurality of first sample suspicious positive cell images by a visual network to be trained in the cervical lesion prediction model to be trained, extract a plurality of image features related to the first sample clinical text from the fifth sample image features of the plurality of first sample suspicious positive cell images by the preliminary trained image text alignment network to be trained in the cervical lesion prediction model to be trained, obtain sixth sample image features corresponding to the first image-text sample pair, and predict a third cervical lesion category probability corresponding to the first image-text sample pair based on the sixth sample image features and sample text features of the first sample clinical text by a prediction network in the cervical lesion prediction model to be trained; determine a fourth loss value according to the third cervical lesion category probability corresponding to each of the plurality of first image-text sample pairs; and train the cervical lesion prediction model to be trained according to the third loss value and the fourth loss value, to obtain a trained cervical lesion prediction model.
4. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor, when executing the computer program, implements the steps of the method of claim 1 or 2.
5. A computer readable storage medium storing a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of claim 1 or 2.
6. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of claim 1 or 2. The computer program, when executed by the processor, implements the steps of the method of claim 1 or 2.
Citation Information
Patent Citations
Image processing method, device and equipment
CN115497092A
Image processing method and device, equipment and medium
CN117711001A
Medical report generation method, system and terminal based on cross-view semantic alignment
CN117809794A
Image and clinical information fusion processing method and system based on joint attention cross-modal network
CN118366626A