Image Processing Method, Apparatus, Device, and Storage Medium
Through the combination of student model and teacher model, pseudo-labels are generated using fully supervised and unsupervised training, the problem of time-consuming and labor-consuming data annotation in image processing model training is solved, and high-accuracy medical image recognition is achieved.
Patent Information
- Application Number
- CN202111296226.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-03
AI Technical Summary
In the prior art, the training of an image processing model requires a large amount of manually labeled sample data, which results in time and effort being used for data labeling and insufficient model training accuracy.
Through the combination of student model and teacher model, the first sample image is used for full supervision training, and the second sample image is used for unsupervised training, generating pseudo labels, reducing the amount of data annotation, and model training is carried out by mixing the difference information of the annotation image and predicting the annotation image to improve the model accuracy.
It greatly reduces the workload of data annotation, and improves the accuracy of model training, and the generated pseudo-labels are highly accurate, and the model performs excellently in medical image recognition.
Smart Images

Figure CN114332553B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an image processing method, apparatus, device, and storage medium. Background Art
[0002] With the vigorous development of artificial intelligence technology, image processing technology, as a branch of artificial intelligence technology, has a wider and wider range of applications. For example, image processing technology can be applied in medical scenarios. By using an image processing model, various physiological tissue regions in medical images are identified to obtain an identification result, which includes the tissue categories to which each physiological tissue region belongs.
[0003] In related technologies, a large amount of labeled sample data often needs to be obtained, and then an image processing model is trained based on this sample data. Among them, the labeling of sample data requires personnel with professional medical knowledge to manually label various physiological tissue regions in medical images.
[0004] However, the amount of sample data required for model training is large, and the process of labeling sample data is time-consuming and laborious. Therefore, there is an urgent need for an image processing method that can reduce the workload of data labeling while ensuring the accuracy of model training. Summary of the Invention
[0005] Embodiments of this application provide an image processing method, apparatus, device, and storage medium, which can greatly reduce the workload of data labeling and improve the accuracy of model training. The technical solutions are as follows:
[0006] On the one hand, an image processing method is provided, and the method includes:
[0007] Obtain a first predicted labeled image of a first sample image in a first sample set through a student model, and obtain a mixed labeled image of a mixed image of a second sample image pair in a second sample set, where the second sample image pair includes at least two second sample images;
[0008] Obtain a labeled mixed image of the second sample image pair through a teacher model, where the labeled mixed image is obtained based on second predicted labeled images of the at least two second sample images;
[0009] Train the student model and the teacher model based on the first predicted labeled image, the first labeled image of the first sample image, the mixed labeled image, and the labeled mixed image to obtain a first image processing model, where the first image processing model is used to determine a labeled image of an input image, and the first labeled image is used to represent the tissue category to which the physiological tissue region in the first sample image belongs;
[0010] Based on the first image processing model, obtain the second labeled images of multiple second sample images in the second sample set, where the second labeled images are used to represent the predicted tissue categories to which the physiological tissue regions in the second sample images belong;
[0011] Based on the multiple first sample images in the first sample set, the first labeled images of the multiple first sample images, the multiple second sample images, and the multiple second labeled images, train a second image processing model, where the second image processing model is used to determine the labeled image of an input image.
[0012] In some embodiments, the method further includes:
[0013] Based on a test image, the test labeled image of the test image, and the trained second image processing model, determine the model accuracy of the second image processing model, where the test labeled image is obtained by labeling the tissue category to which the physiological tissue region in the test image belongs;
[0014] When the model accuracy meets the conditions, perform the step of re-determining the second labeled images of each of the second sample images based on the trained second image processing model to obtain multiple updated second labeled images.
[0015] On the one hand, an image processing method is provided, and the method includes:
[0016] Receive an image recognition request sent by a terminal, where the image recognition request carries a target image to be recognized, and the target image includes at least one physiological tissue region;
[0017] Input the target image into the second image processing model, and through the second image processing model, output the labeled image of the target image, where the labeled image is used to represent the tissue category to which each physiological tissue region belongs;
[0018] Send the labeled image to the terminal;
[0019] Wherein, the second image processing model is trained by multiple first sample images, the first labeled image of each first sample image, multiple second sample images, and the second labeled image of each second sample image. The first labeled image is obtained by manually labeling the tissue category to which the physiological tissue region in the first sample image belongs, and the second labeled image is obtained by the first image processing model labeling the tissue category to which the physiological tissue region in the second sample image belongs.
[0020] On the one hand, an image processing device is provided, and the device includes:
[0021] A first acquisition module, configured to obtain, through a student model, a first predicted annotation image of a first sample image in a first sample set, and obtain a mixed annotation image of a mixed image of a pair of second sample images in a second sample set, where the pair of second sample images includes at least two second sample images;
[0022] A second acquisition module, configured to obtain, through a teacher model, an annotated mixed image of the pair of second sample images, where the annotated mixed image is obtained based on second predicted annotation images of the at least two second sample images;
[0023] A first training module, configured to train the student model and the teacher model based on the first predicted annotation image, a first annotation image of the first sample image, the mixed annotation image, and the annotated mixed image, so as to obtain a first image processing model, where the first image processing model is used to determine an annotation image of an input image, and the first annotation image is used to represent a tissue category to which a physiological tissue region in the first sample image belongs;
[0024] A third acquisition module, configured to obtain, based on the first image processing model, second annotation images of multiple second sample images in the second sample set, where the second annotation images are used to represent predicted tissue categories to which physiological tissue regions in the second sample images belong;
[0025] A second training module, configured to train a second image processing model based on multiple first sample images in the first sample set, first annotation images of the multiple first sample images, the multiple second sample images, and multiple second annotation images, where the second image processing model is used to determine an annotation image of an input image.
[0026] In some embodiments, the first training module includes:
[0027] A first determination unit, configured to determine first difference information between the first predicted annotation image and the first annotation image;
[0028] A second determination unit, configured to determine second difference information between the mixed annotation image and the annotated mixed image;
[0029] A training unit, configured to train the student model and the teacher model based on the first difference information and the second difference information, so as to obtain the first image processing model.
[0030] In some embodiments, the second determination unit is configured to determine third difference information between the mixed annotation image and the annotated mixed image; and use a ratio between the third difference information and a total number of pixel points included in the mixed annotation image as the second difference information.
[0031] In some embodiments, the second determination unit is configured to determine a plurality of pixel point pairs, and determine the square value of the pixel value difference between a first pixel point and a second pixel point included in each of the pixel point pairs. The first pixel point is a pixel point in the mixed annotation image, and the second pixel point is a pixel point at the corresponding position in the annotated mixed image. The pixel value of any pixel point represents the predicted tissue category to which the pixel point belongs; the sum of the square values of the plurality of pixel point pairs is used as the third difference information.
[0032] In some embodiments, the training unit is configured to perform weighted summation on the first difference information and the second difference information to obtain fourth difference information; based on the fourth difference information, adjust the model parameters of the student model, and based on the adjusted model parameters of the student model, adjust the model parameters of the teacher model; train the student model and the teacher model respectively based on the corresponding adjusted model parameters; determine the first image processing model from the trained student model and the trained teacher model.
[0033] In some embodiments, the apparatus further includes:
[0034] A cut-mix module, configured to obtain the second sample image pair from the second sample set; perform cut-mix on at least two second sample images in the second sample image pair to obtain the mixed image.
[0035] In some embodiments, the at least two second sample images include a third sample image and a fourth sample image; the cut-mix module is configured to determine a cropping window; crop a first image area corresponding to the cropping window in the third sample image; extract a second image area corresponding to the cropping window from the fourth sample image, and fill the second image area into the position where the first image area is located in the third sample image to obtain the mixed image.
[0036] In some embodiments, the second acquisition module is configured to input the second sample image pair into the teacher model, and output a second predicted annotation image of the at least two second sample images through the teacher model; perform cut-mix on at least two of the second predicted annotation images to obtain the annotated mixed image.
[0037] In some embodiments, the second training module is configured to input the multiple first sample images, the multiple first annotation images, the multiple second sample images, and the multiple second annotation images into the second image processing model, and train the second image processing model supervised by the multiple first annotation images and the multiple second annotation images; based on the trained second image processing model, re-determine the second annotation image of each second sample image to obtain multiple updated second annotation images; and based on the multiple first sample images, the multiple first annotation images, the multiple second sample images, and the multiple updated second annotation images, perform updated training on the basis of the trained second image processing model.
[0038] In some embodiments, the apparatus further includes:
[0039] A testing module, configured to determine the model accuracy of the second image processing model based on a test image, a test annotation image of the test image, and the trained second image processing model, where the test annotation image is obtained by annotating the tissue category to which the physiological tissue region in the test image belongs;
[0040] The second training module is further configured to, when the model accuracy meets the condition, re-determine the second annotation image of each second sample image based on the trained second image processing model to obtain the multiple updated second annotation images.
[0041] On the one hand, an image processing apparatus is provided, and the apparatus includes:
[0042] A receiving module, configured to receive an image recognition request sent by a terminal, where the image recognition request carries a target image to be recognized, and the target image includes at least one physiological tissue region;
[0043] An annotation module, configured to input the target image into the second image processing model, and output an annotation image of the target image through the second image processing model, where the annotation image is used to represent the tissue category to which each physiological tissue region belongs;
[0044] A sending module, configured to send the annotation image to the terminal;
[0045] Among them, the second image processing model is trained with a plurality of first sample images, the first annotation image of each first sample image, a plurality of second sample images, and the second annotation image of each second sample image. The first annotation image is obtained by manually annotating the tissue category to which the physiological tissue region in the first sample image belongs, and the second annotation image is obtained by annotating the tissue category to which the physiological tissue region in the second sample image by the first image processing model.
[0046] On the one hand, a computer device is provided. The computer device includes one or more processors and one or more memories. At least one computer program is stored in the one or more memories. The computer program is loaded and executed by the one or more processors to implement the image processing method.
[0047] On the one hand, a computer-readable storage medium is provided. At least one computer program is stored in the computer-readable storage medium. The computer program is loaded and executed by a processor to implement the image processing method.
[0048] On the one hand, a computer program product or a computer program is provided. The computer program product or the computer program includes program code. The program code is stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above-mentioned image processing method.
[0049] In the technical solution provided in the embodiment of the present application, training is performed in the mode of a student model and a teacher model based on the first sample image and the second sample image. The student model is trained in a fully supervised manner based on the first sample image and in an unsupervised manner based on the second sample image, and the teacher model is trained based on the second sample image, so as to obtain a first image processing model with relatively high accuracy. Then, the pseudo-label of the second sample image, that is, the second annotation image, is determined by the first image processing model, so that it is not necessary for technicians to manually annotate the second sample image; further, the second image processing model is trained in a fully supervised manner based on the labeled first sample image and the second sample image. Since the accuracy of the annotation images of the two types of sample images is relatively high, the accuracy of the trained second image processing model is also relatively high. It can be seen that the above solution not only greatly reduces the workload of data annotation, but also improves the accuracy of model training. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0051] Figure 1 It is a schematic diagram of the implementation environment of an image processing method provided by an embodiment of the present application;
[0052] Figure 2 It is a schematic diagram of a medical image provided by an embodiment of the present application;
[0053] Figure 3 It is a schematic diagram of the structure of an image processing model provided by an embodiment of the present application;
[0054] Figure 4 It is a flowchart of an image processing method provided by an embodiment of the present application;
[0055] Figure 5 It is a flowchart of an image processing method provided by an embodiment of the present application;
[0056] Figure 6 It is a schematic diagram of an image processing method provided by an embodiment of the present application;
[0057] Figure 7 It is a schematic diagram of an image processing method provided by an embodiment of the present application;
[0058] Figure 8 It is a schematic diagram of an image processing method provided by an embodiment of the present application;
[0059] Figure 9 It is a schematic diagram of an image processing method provided by an embodiment of the present application;
[0060] Figure 10 It is a schematic diagram of an image processing method provided by an embodiment of the present application;
[0061] Figure 11 It is a schematic diagram of an image processing method provided by an embodiment of the present application;
[0062] Figure 12 It is a flowchart of an image processing method provided by an embodiment of the present application;
[0063] Figure 13 It is a schematic diagram of the structure of an image processing device provided by an embodiment of the present application;
[0064] Figure 14 It is a schematic diagram of the structure of an image processing device provided by an embodiment of the present application;
[0065] Figure 15 It is a schematic structural diagram of a terminal provided by an embodiment of the present application;
[0066] Figure 16 It is a schematic structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0067] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0068] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependence between "first", "second", and "nth", nor are the quantity and execution order limited.
[0069] In the present application, the term "at least one" means one or more, and the meaning of "multiple" means two or more. For example, multiple first sample images mean two or more first sample images.
[0070] To facilitate the understanding of the technical process of the embodiments of the present application, some terms involved in the embodiments of the present application are explained below:
[0071] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0072] Computer Vision Technology (CV) is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further perform graphic processing to make the images processed by the computer more suitable for human eye observation or transmission to instruments for detection.
[0073] Convolutional Neural Networks (CNN) are a class of feedforward neural networks (Feedforward Neural Networks) that contain convolutional calculations and have a deep structure, and are one of the representative algorithms of deep learning.
[0074] Fully Convolutional Networks (FCN) is a network for image semantic segmentation, specifically for image pixel-level classification.
[0075] Semi-Supervised Learning (SSL) is a learning method that combines supervised learning and unsupervised learning. Semi-supervised learning uses a large amount of unlabeled data and also uses labeled data to perform pattern recognition tasks.
[0076] Consistency Regularization means that for an input, even with minor perturbations, its predictions should be consistent. Consistency means minimizing the difference between the prediction and the label, that is, the prediction is consistent with the label.
[0077] The technical solution provided by the embodiments of this application can also be combined with cloud technology. For example, the trained image processing model is deployed on a cloud server. Cloud Technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.
[0078] Among them, the Medical Cloud in cloud technology refers to creating a medical and health service cloud platform using "cloud computing" based on new technologies such as cloud computing, mobile technology, multimedia, 4G communication, big data, and the Internet of Things, combined with medical technology, achieving the sharing of medical resources and the expansion of the medical scope. Due to the application and combination of cloud computing technology, the Medical Cloud improves the efficiency of medical institutions and facilitates residents' medical treatment. For example, the current hospital appointment registration, electronic medical records, medical insurance, etc. are all the products of the combination of cloud computing and the medical field. The Medical Cloud also has the advantages of data security, information sharing, dynamic expansion, and overall layout. Exemplarily, the image processing model provided by the embodiments of this application is deployed on a medical and health service cloud platform.
[0079] Optionally, the computer device provided by the embodiments of this application can be provided as a terminal or a server. The implementation environment composed of the terminal and the server will be introduced below.
[0080] Figure 1 It is a schematic diagram of the implementation environment of an image processing method provided by the embodiments of this application. Refer to Figure 1 , this implementation environment includes an image acquisition device 110, a terminal 120, and a server 130. Among them, the terminal 120 is connected to the image acquisition device 110 and the server 130 through a wireless network or a wired network respectively.
[0081] An image acquisition device 110 is configured to acquire medical images of pathological sections and send the medical images to a terminal 120. Optionally, the image acquisition device 110 is an electron microscope or a slide scanner. For example, a doctor prepares a pathological section and places it on the image acquisition device 110. The image acquisition device 110 acquires an image within the current field of view and sends the acquired image to the terminal 120.
[0082] The terminal 120 is configured to receive the image sent by the image acquisition device 110, forward the image to the server 130, or the terminal 120 directly performs recognition on the image without forwarding it to the server 130. This is not limited in the embodiments of the present application. Optionally, the terminal 120 includes, but is not limited to, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, a smart voice interaction device, a smart home appliance, or a vehicle-mounted terminal, etc. The terminal 120 installs and runs an application program that supports image processing.
[0083] The server 130 is configured to receive the image sent by the terminal 120, perform recognition on the image to obtain a recognition result, and send the recognition result to the terminal 120. Correspondingly, the terminal 120 is further configured to receive the recognition result and display the recognition result. Optionally, the server 130 is an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, a content delivery network (CDN), and a big data and artificial intelligence platform.
[0084] Optionally, the terminal 120 generally refers to one of multiple terminals. Only the terminal 120 is used as an example in the embodiments of the present application. Those skilled in the art can understand that the number of the above-mentioned terminals 120 can be more or less. For example, the above-mentioned terminal 120 is only one, or the above-mentioned terminal 120 is dozens or hundreds, or more. In this case, other terminals are further included in the above-mentioned implementation environment. The embodiments of the present application do not limit the number and device type of the terminals.
[0085] The sample image in the embodiments of the present application can be acquired by the image acquisition device and sent to the server through the terminal, or can be obtained by the server from the database. This is not limited in the embodiments of the present application.
[0086] After introducing the implementation environment of the image processing method provided in the embodiments of the present application, the application scenarios of the image processing method provided in the embodiments of the present application will be described below. It should be noted that the terminal in the following description process is the terminal 120 in the above implementation environment, and the server is the server 130 in the above implementation environment. The image processing method provided in the embodiments of the present application can be applied to a variety of image processing scenarios. For example, it can be applied to the scenario of identifying medical images, or the scenario of identifying other images containing different category regions. The embodiments of the present application do not limit this.
[0087] In the scenario of identifying medical images, for example, identifying the tissue category to which the physiological tissue region in the medical image belongs. Since identifying the tissue category corresponding to the physiological tissue region in a medical image requires relatively rich experience, and the identification of tissue categories is the basis for effective treatment. For some medical staff with insufficient experience, the image processing method provided in the embodiments of the present application can be used to identify medical images, determine the tissue categories to which each physiological tissue region in the medical image belongs, and use this as a reference.
[0088] When it is necessary to identify a medical image, the medical staff uploads the medical image to the server through the terminal. The server executes the image processing method provided in the embodiments of the present application on the received medical image to obtain an annotated image, which is used to represent the tissue categories to which each physiological tissue region in the medical image belongs. The server returns the annotated image to the terminal, and the terminal presents the annotated image to the medical staff. The annotated image can play a reference role for the medical staff, thus assisting the medical staff in formulating a treatment plan. For example, with the help of the annotated image, qualitative or quantitative analysis of the tumor cancer situation of the patient is carried out, so as to evaluate the malignancy level of the patient's tumor, and then according to the analysis results.
[0089] For example, taking the pathological section image of tumor tissue as a medical image, the tissue categories include tumor, lymph, stroma, normal duct or necrotic tissue, etc. In the treatment of cancer (such as breast cancer, lung cancer or liver cancer, etc.), tumor tissue is obtained from the patient by methods such as forceps extraction, resection or incisional biopsy, made into a pathological section through staining, and a pathological section image is obtained through an image acquisition device. See Figure 2, the left figure is the original pathological section image, and the right figure is the annotation image. Among them, the red area in the annotation image represents the tumor, the green area represents the lymph, the yellow area represents the normal duct, the cyan area represents necrosis, and the blue area represents the stroma. Medical staff can more intuitively observe the tumor canceration situation of the patient according to the annotation image, and can also directly perform further quantitative analysis on the ROI (Region Of Interest). For example, calculate the immunohistochemical ki67 nuclear positive index (a cell proliferation index) in the tumor area to further quantify the tumor canceration situation. Among them, the staining method can be the HE (Hematoxylin-Eosin) method, the IHC (Immuno Histo Chemistry) method or other staining methods.
[0090] In the embodiment of the present application, the computer device can implement the image processing method provided by the embodiment of the present application with the help of an image processing model. The structure of the image processing model provided by the embodiment of the present application will be described below. In some embodiments, the image processing model includes a first convolutional layer, a pooling layer, an encoder, a decoder, a first deconvolutional layer, a second convolutional layer, and a second deconvolutional layer. Among them, the encoder includes a plurality of encoding blocks, each encoding block includes a plurality of residual blocks, the decoder includes a plurality of decoding blocks, and each decoding block includes a third convolutional layer, a third deconvolutional layer, and a fourth convolutional layer.
[0091] Among them, each convolutional layer is used to perform convolutional processing on the matrix input to this convolutional layer to obtain a feature map. The pooling layer is used to pool the feature map obtained by the first convolutional layer to further extract features. The encoder is used to perform downsampling on the feature map output by the pooling layer. The decoder is used to perform upsampling on the feature map output by the encoder to obtain a feature map with the same size as the input image.
[0092] Among them, the residual block includes a fifth convolutional layer, a sixth convolutional layer, and a seventh convolutional layer. In the image processing model, a batch normalization layer and a ReLU (Rectified Linear Unit) activation function layer are further included between two adjacent convolutional layers.
[0093] It should be noted that the number of encoders, the number of residual blocks included in each encoder, and the number of decoders can all be set as needed, and the embodiment of the present application does not limit this.
[0094] In some embodiments, the encoding block further includes a Dropout layer. The Dropout layer is used to make the weight or output of each node in the target convolutional layer have a probability of being zero during each training process of the image processing model, that is, "dropout", and this probability is any value greater than or equal to 0 and less than or equal to 1, such as the probability p = 0.2. The target convolutional layer is the previous convolutional layer connected to the Dropout layer. Optionally, the Dropout layer is arranged behind the sixth convolutional layer of the residual block, and the target convolutional layer is the sixth convolutional layer. In the embodiments of the present application, the Dropout layer can be set in each encoding block, or only in one encoding block, and the embodiments of the present application do not limit this. In the embodiments of the present application, by setting the Dropout layer in the encoding block, the problem that the model is overfitted due to the excessive weight of a certain neuron can be avoided.
[0095] For example, taking the LinKNet neural network model as an example of the image processing model, LinKNet is a lightweight convolutional neural network with the characteristics of few parameters, fast training speed, and remarkable effects. See Figure 3 , the encoder in LinKNet consists of 4 encoding blocks (encoder block), namely encoding block A (encoder1), encoding block B (encoder2), encoding block C (encoder3), and encoding block D (encoder4), and the decoder consists of 4 decoding blocks (decoder block), namely decoding block E (decoder1), decoding block F (decoder2), decoding block G (decoder3), and decoding block H (decoder4). At the same time, LinkNet uses the skip connection method, enabling communication between each encoding block and decoding block, thereby improving the accuracy of image segmentation by increasing information sharing. Among them, the parameters of the first convolutional layer are a 7*7 convolutional kernel with a stride of 2, the parameters of the pooling layer are a 3*3 convolutional kernel with a stride of 2, and the parameters of the second convolutional layer are a 3*3 convolutional kernel with a stride of 1. The parameters of the first deconvolutional layer and the second deconvolutional layer are both 3*3 convolutional kernels with a stride of 2. The parameters of the third convolutional layer and the fourth convolutional layer in each decoding block are both 1*1 convolutional kernels with a stride of 1, and the parameters of the third deconvolutional layer are 3*3 convolutional kernels with a stride of 2. Each encoding block includes 2 residual blocks. The parameters of the fifth convolutional layer a in residual block 1 are a 3*3 convolutional kernel with a stride of 2, the parameters of the sixth convolutional layer b are a 3*3 convolutional kernel with a stride of 1, and the parameters of the seventh convolutional layer c are a 1*1 convolutional kernel with a stride of 1; the parameters of the fifth convolutional layer d in residual block 2 are a 3*3 convolutional kernel with a stride of 1, the parameters of the sixth convolutional layer e are a 3*3 convolutional kernel with a stride of 1, and the parameters of the seventh convolutional layer f are a 1*1 convolutional kernel with a stride of 1.
[0096] It should be noted that Figure 3 only the convolutional layers are drawn in []. For each convolutional layer, there corresponds a batch normalization layer and a ReLU activation function layer. That is, the output data of the convolutional layer is first normalized and then activated, and the activated output data is then input into the next convolutional layer.
[0097] It should be noted that the structures of the student model, the teacher model, the first image processing model, and the second image processing model involved in the embodiments of the present application can all refer to the structure of the above image processing model, and the structure of the above image processing model is only an example. In other possible implementation manners, the image processing model can also be other structures, and the embodiments of the present application do not limit this.
[0098] After introducing the implementation environment, application scenarios, and the structure of the image processing model of the embodiments of the present application, the image processing method provided by the embodiments of the present application will be described below. In the embodiments of the present application, the image processing method provided by the embodiments of the present application can be implemented with the server or the terminal as the execution subject, or can be implemented through the interaction between the terminal and the server. Among them, the terminal is the terminal 120 in the above implementation environment, and the server is the server 130 in the above implementation environment. For the interaction between the terminal and the server, the terminal sends the sample data set to the server, and the server trains the image processing model. The server can return the trained image processing model to the terminal, and the terminal performs image recognition through the image processing model; or, the server can also store the trained image processing model, recognize the image sent by the terminal, and return the recognition result to the terminal, thereby saving the workload of the terminal. The embodiments of the present application do not limit the execution subject.
[0099] Figure 4 is a flowchart of an image processing method provided by an embodiment of the present application. Refer to Figure 4 , taking the server as the execution subject as an example in the embodiments of the present application, the method includes:
[0100] 401. The server obtains, through the student model, the first predicted annotation image of the first sample image in the first sample set, and obtains the mixed annotation image of the mixed image of the second sample image pair in the second sample set, where the second sample image pair includes at least two second sample images.
[0101] In the embodiments of the present application, a student model and a teacher model are trained using the Mean teacher method. During the process of training the student model, two types of sample images are used. One type is the first sample image. The student model uses the first sample image as input data and is trained in a fully supervised learning manner with the first annotation image of the first sample image as supervision. The other type is the mixed image of the second sample image pair. The student model uses the mixed image as input data and is trained in an unsupervised learning manner.
[0102] In some embodiments, both the first sample image and the second sample image are medical images. Both the first sample image and the second sample image include at least one physiological tissue region, and each physiological tissue region corresponds to a tissue category. For example, the tissue categories include categories such as tumor, lymph, stroma, normal duct, or necrotic tissue.
[0103] The first sample image in the first sample set is an annotated image, that is, the first sample image has a corresponding first annotation image. The first annotation image is an image obtained by annotating the tissue categories to which each physiological tissue region in the first sample image belongs. Among them, the first annotation image is obtained by manually annotating the first sample image before model training. Therefore, the annotation result is real and can be regarded as a true label. In the embodiments of the present application, the annotation of an image is to annotate the tissue category to which each pixel point in the image belongs. Then, in the annotation image, the pixel value of each pixel point represents the tissue category to which the pixel point belongs. Correspondingly, the first annotation image is represented in the form of a matrix, and each element in the matrix represents the pixel value of the pixel point at that position, and this pixel value represents the tissue category to which the pixel point belongs.
[0104] Among them, the pixel values corresponding to each tissue category are set in advance. For example, the tissue category corresponding to the pixel value 1 is a tumor, and the tissue category corresponding to the pixel value 2 is lymph. Then, if the pixel value of pixel point A is 1, it means that the tissue category of pixel point A is a tumor. In some embodiments, different tissue categories correspond to pixel values of different colors, so that users can clearly distinguish different physiological tissue regions according to the colors when viewing the annotation image. For example, the tissue category "tumor" corresponds to red, and the tissue category "lymph" corresponds to blue. The embodiments of the present application do not limit the setting of the pixel value size or the color type corresponding to the tissue category.
[0105] The first predicted labeled image is a predicted labeled image obtained by the student model labeling the first sample image. In some embodiments, the predicted labeled image is represented in the form of a probability matrix with C channels. Each channel corresponds to a tissue category, and C is the number of tissue categories, such as 5. Each element in the probability matrix includes an array, and each array includes C probabilities. Each element represents the probability that the pixel at that position belongs to each tissue category respectively. The sum of the C probabilities of each pixel is 1. For example, in the probability matrix of the first predicted labeled image, the array of pixel A is (0.1, 0.5, 0.1, 0.2, 0.1). Taking the tissue categories including tumor, lymph, stroma, normal duct, and necrotic tissue as an example, the probability that pixel A belongs to the tumor is 0.1, the probability that it belongs to the lymph is 0.5, the probability that it belongs to the stroma is 0.1, the probability that it belongs to the normal duct is 0.2, and the probability that it belongs to the necrotic tissue is 0.1.
[0106] The mixed image is an image obtained by processing at least two second sample images included in the second sample image pair. In some embodiments, the processing operation is a Cut out and Mixup (CutMix) operation. The mixed labeled image is a predicted labeled image obtained by the student model labeling the tissue category to which the physiological tissue region in the mixed image belongs.
[0107] 402. The server obtains the labeled mixed image of the second sample image pair through the teacher model. The labeled mixed image is obtained based on the second predicted labeled images of the at least two second sample images.
[0108] Among them, the server inputs the second sample images in the second sample set into the teacher model and trains the teacher model to obtain the second predicted labeled images, so as to provide a comparison object for the output of the student model. In the embodiments of the present application, the teacher model labels the tissue category to which the physiological tissue region in each second sample image belongs to obtain the second predicted labeled images. The server processes the second predicted labeled images of at least two second sample images to obtain the labeled mixed image. In some embodiments, the processing operation and the processing operation in step 401 are both CutMix operations.
[0109] 403. The server trains the student model and the teacher model based on the first predicted labeled image, the first labeled image of the first sample image, the mixed labeled image, and the labeled mixed image to obtain a first image processing model. The first image processing model is used to determine the labeled image of the input image, and the first labeled image is used to represent the tissue category to which the physiological tissue region in the first sample image belongs.
[0110] After the training of the student model and the teacher model is completed, the server determines the first image processing model from the student model and the teacher model.
[0111] 404. The server, based on the first image processing model, obtains second labeled images of multiple second sample images in the second sample set, and the second labeled images are used to represent the predicted tissue categories to which the physiological tissue regions in the second sample images belong.
[0112] Since the first image processing model is a trained image processing model, the second sample images can be labeled through the first image processing model to obtain second labeled images, without the need for manual labeling of the second sample images, greatly reducing the data labeling volume. Among them, since the second labeled images are obtained by model prediction and are not real labeled images, they can be regarded as pseudo-labels of the second sample images.
[0113] 405. The server trains a second image processing model based on the multiple first sample images in the first sample set, the first labeled images of the multiple first sample images, the multiple second sample images, and the multiple second labeled images, and the second image processing model is used to determine the labeled image of the input image.
[0114] After the training of the second image processing model is completed, the server can deploy the second image processing model on its own end. When the terminal needs image recognition, it sends the image to be recognized to the server, and the server obtains the labeled image of the image based on the second image processing model, and then returns the labeled image to the terminal; or, the server sends the second image processing model to the terminal, and the terminal performs image recognition through the second image processing model.
[0115] In the technical solution provided by the embodiments of the present application, training is performed in the mode of a student model and a teacher model based on the first sample images and the second sample images. The student model is trained in a fully supervised manner based on the first sample images and in an unsupervised manner based on the second sample images, and the teacher model is trained based on the second sample images, so as to obtain a first image processing model with relatively high accuracy. Then, the pseudo-labels of the second sample images, that is, the second labeled images, are determined through the first image processing model, and there is no need for technicians to manually label the second sample images; further, the second image processing model is trained in a fully supervised manner based on the labeled first sample images and second sample images. Since the accuracies of the labeled images of the two types of sample images are relatively high, the accuracy of the trained second image processing model is also relatively high. It can be seen that the above solution not only greatly reduces the workload of data labeling, but also improves the accuracy of model training.
[0116] In the embodiments of the present application, the training process of the student model and the teacher model includes multiple iteration processes. In the first iteration process, the server inputs the sample images into the corresponding models to obtain the prediction results of the first iteration process; based on the prediction results of the first iteration process, the difference information is determined, and based on the difference information, the model parameters of the models are adjusted; the model parameters adjusted in the first iteration are used as the model parameters of the second iteration, and then the second iteration process is carried out; the above iteration process is repeated multiple times until the training meets the target conditions, and the trained student model and teacher model are obtained.
[0117] In some embodiments, the target condition for the training to be satisfied is that the number of training iterations of the model reaches the target number, and the target number is a preset number of training iterations, such as 1000 times; or, the target condition for the training to be satisfied is that the difference information meets the target threshold condition, such as the difference information is a loss value, and the loss value is less than 0.0001. The embodiments of the present application do not limit the setting of the target conditions. In the embodiments of the present application, the case where the target condition is that the difference information meets the target threshold condition is taken as an example for description.
[0118] Figure 5 It is a flowchart of an image processing method provided by the embodiments of the present application. Refer to Figure 5 , the embodiments of the present application take the i-th iteration process of the student model and the teacher model as an example for description, where i is a positive integer not less than 1, and the method includes:
[0119] 501. In the i-th iteration process, the server obtains a second sample image pair from the second sample set.
[0120] Among them, the second sample images in the second sample set are sample images without labeled tissue categories. The second sample image pair includes at least two second sample images. In the embodiments of the present application, the case where the second sample image pair includes two second sample images is taken as an example for description. In some embodiments, the implementation manner of step 501 includes: the server randomly obtains two second sample images from the second sample set, and these two second sample images are called the second sample image pair.
[0121] 502. The server performs cutmix on at least two second sample images in the second sample image pair to obtain a mixed image.
[0122] In the embodiments of the present application, by performing cutmix on the second sample images without labeled tissue categories, a mixed image is obtained, and the mixed image is used as the input data of the student model, and a high-level noise is added to the input data in the form of cutmix image perturbation, so that the overfitting of the student model can be reduced by means of consistency regularization in the subsequent process, and thus the training accuracy of the model is improved.
[0123] In some embodiments, at least two second sample images in the second sample image pair include a third sample image and a fourth sample image; the implementation manner for the server to perform cutmix on at least two second sample images in the second sample image pair to obtain a mixed image includes: the server determines a cropping window; in the third sample image, a first image region corresponding to the cropping window is cropped off; from the fourth sample image, a second image region corresponding to the cropping window is extracted, and the second image region is filled into the position where the first image region is located in the third sample image to obtain a mixed image.
[0124] Among them, the cropping window is a rectangular window, and the size of the cropping window can be set as needed, and the embodiments of the present application do not limit this. For example, for each second sample image pair, a cropping window with a randomly sized or fixed size is generated.
[0125] In the embodiments of the present application, by replacing the first image region corresponding to the cropping window in the third sample image with the second image region in the fourth sample image, a mixed image is obtained, thereby achieving the effect of adding noise to the image in the cutmix manner, increasing the noise intensity of the input data of the model, and thus increasing the regularization intensity of the model, and further providing data support for preventing overfitting in model training.
[0126] 503. The server obtains a first predicted annotation image of the first sample image in the first sample set through the student model.
[0127] Among them, the server inputs the first sample image into the student model, and the student model makes a prediction on the first sample image to obtain a first predicted annotation image.
[0128] 504. The server obtains a mixed annotation image of the mixed image through the student model.
[0129] Among them, the server inputs the mixed image into the student model, and the student model makes a prediction on the mixed image to obtain a mixed annotation image. It should be noted that in each iteration process, the student model makes two predictions, which are respectively based on the first sample image and the mixed image for prediction.
[0130] In the embodiments of the present application, both the first sample set and the second sample set include a large number of sample images. The server pre-sets the sample quantity (batch size) used in one iteration. In each iteration process, the server respectively obtains the first sample images with the batch size from the first sample set, and obtains the second sample image pairs with the batch size from the second sample set. Based on the mixed images of the obtained first sample images and the second sample image pairs, the student model is trained respectively. For example, the batch size is 10, 20 or 30, etc. In the embodiments of the present application, taking the batch size of 1 as an example for illustration, when the batch size is greater than 1, the training process of the student model is the same as that when the batch size is 1, which will not be elaborated here.
[0131] In some embodiments, the server does not perform cutmix on the second sample image pairs, but performs Cutout on each second sample image. Correspondingly, step 501-step 502 are replaced with the following steps: In the i-th iteration process, the server cuts the second sample images in the second sample set to obtain cut images. Correspondingly, step 504 is replaced with the following steps: The server obtains the cut annotation images of the cut images through the student model.
[0132] Optionally, the implementation manner of the server cutting the second sample images in the second sample set to obtain cut images includes: The server determines a cropping window, and in the second sample image, fills the pixel values of multiple pixel points in the image area corresponding to the cropping window with 0 to obtain a cut image. The difference between the cut image and the second sample image is that there is an image area with pixel values of 0. The setting manner of the cropping window in this step is the same as that in step 501, which will not be elaborated here.
[0133] It should be noted that the embodiments of the present application are described by taking the server performing cutmix on the second sample image pairs as an example.
[0134] 505. The server inputs the second sample image pairs into the teacher model, and through the teacher model, outputs the second predicted annotation images of at least two second sample images.
[0135] Among them, the second sample image set is the input data of the teacher model. The teacher model predicts the second sample images to obtain the second predicted annotation images. In the embodiments of the present application, the student model and the teacher model are trained synchronously. Therefore, step 503-step 504 and step 505-step 506 are performed synchronously.
[0136] 506. The server performs cutmix on at least two second predicted annotation images to obtain an annotation mixed image.
[0137] In the embodiment of the present application, the input data of the teacher model is the unprocessed second sample image, that is, the image without added noise. By performing cutmix on the predicted second predicted annotation image, and the cutmix operation is the same as the cutmix operation performed on the second sample image pair before training the student model, it provides accurate data support for the determination of subsequent difference information.
[0138] In some embodiments, the implementation manner of step 506 is the same as that of step 502, which will not be elaborated here. Optionally, steps 505 - 506 are an implementation manner for the server to obtain the annotation mixed image of the second sample image pair through the teacher model.
[0139] In some embodiments, when the server cuts each second sample image to obtain a cut image and obtains the cut annotation image of the cut image through the student model, steps 505 - 506 are replaced with the following steps: The server inputs the second sample images in the second sample set into the teacher model, and the teacher model outputs the second predicted annotation images of each second sample image; The server cuts the second predicted annotation image to obtain an annotated cut image.
[0140] Among them, the cut operation performed by the server on the second predicted annotation image is the same as the cut operation performed on the second sample image before training the student model.
[0141] 507. The server determines the first difference information between the first predicted annotation image and the first annotation image, where the first annotation image is used to represent the tissue category to which the physiological tissue region in the first sample image belongs.
[0142] Among them, the first difference information is the loss value when the student model is trained according to the first sample image, that is, the full supervision loss value in the process of training the student model in a full supervision learning manner. In some embodiments, the first difference information is the cross - entropy loss value. Correspondingly, the server calculates the cross - entropy loss value between the first predicted annotation image and the first annotation image. Each image includes a plurality of pixel points, and the number of pixel points included in the first predicted annotation image is the same as the number of pixel points included in the first annotation image, both of which are the number of pixel points included in the first sample image. The cross - entropy loss value is calculated on a per - pixel - point basis.
[0143] In the i-th iteration process, if the number of the first sample images input into the student model is 1, the cross-entropy loss value between the first predicted annotation image and the first annotation image of the first sample image is used as the first difference information. If the number of the first sample images input into the student model is multiple, the average value of the cross-entropy loss values of the multiple first sample images is used as the first difference information. Since the first difference information refers to the cross-entropy loss values of each first sample image, the first difference information can better reflect the overall loss situation of the multiple first sample images and has a higher accuracy.
[0144] 508. The server determines the second difference information between the mixed annotation image and the annotated mixed image.
[0145] Among them, since the second sample image is an unannotated sample image, the second difference information is the unsupervised loss value in the process of training the student model in an unsupervised learning manner. In some embodiments, the implementation manner for the server to determine the second difference information between the mixed annotation image and the annotated mixed image includes the following steps (1)-(2):
[0146] (1) The server determines the third difference information between the mixed annotation image and the annotated mixed image.
[0147] In some embodiments, the implementation manner for the server to determine the third difference information between the mixed annotation image and the annotated mixed image includes: The server determines multiple pixel point pairs, determines the square value of the pixel value difference between the first pixel point and the second pixel point included in each pixel point pair, the first pixel point is a pixel point in the mixed annotation image, the second pixel point is a pixel point at the corresponding position in the annotated mixed image, and the pixel value of any pixel point represents the predicted tissue category to which the pixel point belongs; The sum of the square values of the multiple pixel point pairs is used as the third difference information.
[0148] Among them, the number of the multiple pixel point pairs is the same as the number of pixel points included in the mixed annotation image. The mixed annotation image is represented by a first probability matrix with C channels, and the annotated mixed image is represented by a second probability matrix with C channels. For any pixel point pair, the implementation manner for the server to determine the square value of the pixel value difference between the first pixel point and the second pixel point included in each pixel point pair includes: The server determines the square value of the probability difference between the corresponding channels of the first pixel point and the second pixel point, obtains C square values, and uses the sum of the C square values as the square value of the pixel point pair.
[0149] In the embodiments of the present application, since the pixel value can represent the tissue category to which the pixel point belongs, by determining the third difference information based on the pixel values of multiple pixel point pairs, the third difference information can reflect the overall difference between the mixed annotation image and the annotated mixed image, thereby improving the accuracy of the third difference information.
[0150] (2) The server uses the ratio between the third difference information and the total number of pixel points included in the mixed annotation image as the second difference information.
[0151] Among them, the total number of pixel points can be represented by the size of the mixed annotation image, such as height × width. In some embodiments, in the i-th iteration process, the number of mixed images input to the student model and the number of second sample images input to the teacher model are both n, then steps (1)-(2) are implemented by Formula 1:
[0152] Formula 1:
[0153] Among them, d MSE is the second difference information, H is the height of the mixed annotation image, W is the width of the mixed annotation image, k is the k-th image, k = 1, 2, 3,..., n, a is the third sample image, b is the fourth sample image, mix is the cutmix operation, mix(a, b) is the mixed image of the third sample image and the fourth sample image, f θ (mix(a, b)) is the mixed annotation image, is the second predicted annotation image of the third sample image, is the second predicted annotation image of the fourth sample image, is the annotation mixed image, θ is the model parameter of the student model in the i-th iteration process, and φ is the model parameter of the teacher model in the i-th iteration process.
[0154] Among them, each group of mixed annotation images and annotation mixed images corresponds to a third difference information. The ratio of the sum value of the third difference information of n images to the total number of pixel points is used as the second difference information in the i-th iteration process.
[0155] In the embodiments of the present application, the student model and the teacher model can predict the probability that a pixel point belongs to each tissue category. When determining the second difference information, by taking a pixel point as a sample and using the ratio of the third difference information to the total number of pixel points in the image as the second difference information, the second difference information can represent the differences in details between the mixed annotation image and the annotation mixed image, thereby improving the accuracy of the second difference information.
[0156] In some embodiments, when the server performs cutting on each second sample image to obtain a cut image, and obtains the cut annotation image of the cut image through the student model. When obtaining the annotated cut image through the teacher model, step 508 is replaced with the following steps: The server determines the second difference information between the cut annotation image and the annotated cut image.
[0157] Among them, for the implementation manner in which the server determines the second difference information between the clipped annotation image and the annotated clipped image, refer to the above steps (1)-(2), which will not be elaborated in this embodiment of the present application.
[0158] In the embodiment of the present application, by clipping the second sample image, noise is added to the input image of the student model, which can also prevent the student model from overfitting during training. The annotated clipped image is obtained by clipping the second predicted annotation image output by the teacher model, providing a reference for calculating the difference information of the student model. Moreover, since the clipping operation is more time-saving, the model training efficiency is improved.
[0159] 509. The server performs weighted summation on the first difference information and the second difference information to obtain the fourth difference information.
[0160] Among them, the weights of the first difference information and the second difference information can be set as needed, which are not limited in this embodiment of the present application. In some embodiments, the weight of the second difference information is positively correlated with the number of iterations, and the weight of the first difference information is negatively correlated with the number of iterations. That is to say, as the number of iterations increases, the weight of the second difference information increases, and the weight of the first difference information decreases. In the embodiment of the present application, the first difference information is the difference information of full supervision learning, and the second difference information is the difference information of unsupervised learning, then the fourth difference information is the difference information of semi-supervised learning.
[0161] 510. If the fourth difference information meets the target condition, the server stops the iteration and obtains the trained student model and the trained teacher model.
[0162] Among them, that the fourth difference information meets the target condition means that the fourth difference information is small, then the accuracies of the student model and the teacher model are high, and the server can stop the iteration, that is, stop training the student model and the teacher model.
[0163] 511. If the fourth difference information does not meet the target condition, the server performs the (i + 1)-th iteration process on the student model and the teacher model based on the fourth difference information, and repeats the above iteration process multiple times until the training meets the target condition, obtaining the trained student model and the trained teacher model.
[0164] In some embodiments, the server first adjusts the model parameters of the student model and then adjusts the model parameters of the teacher model. That is, the model parameters of the teacher model are updated based on the model parameters of the student model. Accordingly, the implementation manner of the server for performing the (i + 1)-th iteration process on the student model and the teacher model based on the fourth difference information includes: the server adjusts the model parameters of the student model based on the fourth difference information, and adjusts the model parameters of the teacher model based on the adjusted model parameters of the student model; and performs the (i + 1)-th iteration process on the student model and the teacher model respectively based on the corresponding adjusted model parameters.
[0165] In some embodiments, the model parameters of the teacher model are updated by the EMA (Exponential Moving Average) method. Accordingly, step (2) is implemented by the following formula two:
[0166] Formula two: θ′ t = αθ′ t-1 + (1 - α)θ t
[0167] where t is the t-th iteration, θ′ t is the model parameter of the teacher model, θ t is the model parameter of the student model, and α is a smoothing coefficient hyperparameter, and α is generally set to 0.99.
[0168] 512. The server determines a first image processing model from the trained student model and the trained teacher model, and the first image processing model is used to determine the annotated image of the input image.
[0169] In some embodiments, the implementation manner of step 509 includes: the server tests the student model and the teacher model based on the test images of the test set to obtain test results, and in the student model and the teacher model, the model with test results better than the other model is used as the first image processing model. Among them, the test images have corresponding test annotated images, and the test annotated images are obtained by manually annotating the test images. Optionally, the test results are used to represent the accuracy of the model test, and the first image processing model is the model with higher accuracy than the other model in the student model and the teacher model.
[0170] In the embodiments of the present application, through multiple iteration processes on the student model and the teacher model, the student model and the teacher model can better learn the input data according to the first difference information and the second difference information, so that the accuracy of the obtained first image processing model is relatively high.
[0171] Among them, steps 509 - 512 are an implementation manner in which the server trains the student model and the teacher model based on the first difference information and the second difference information to obtain the first image processing model; steps 507 - 512 are an implementation manner in which the server trains the student model and the teacher model based on the first predicted labeled image, the first labeled image of the first sample image, the mixed labeled image, and the labeled mixed image to obtain the first image processing model.
[0172] For example, referring to Figure 6 , input the first sample image (with labeled data) into the student model to determine the first difference information (cross-entropy loss) between the first predicted labeled image and the first labeled image. Cut-mix the second sample image (unlabeled image) A and the second sample image (unlabeled image) B to obtain a mixed image. Input the mixed image into the student model to output a mixed labeled image. Input the second sample image A and the second sample image B into the teacher model to output the second predicted labeled image a and the second predicted labeled image b. Cut-mix a and b to obtain a labeled mixed image. Determine the second difference information (mean squared error loss) between the mixed labeled image and the labeled mixed image. Determine the fourth difference information (unsupervised loss) based on the first difference information and the second difference information. Thus, based on the fourth difference information, adjust the model parameters of the student model, and based on the model parameters of the student model, adjust the model parameters of the teacher model by the EMA method.
[0173] 513. The server obtains the second labeled images of multiple second sample images in the second sample set based on the first image processing model, and the second labeled images are used to represent the predicted tissue categories to which the physiological tissue regions in the second sample images belong.
[0174] Among them, the server inputs the second sample images in the second sample set into the first image processing model, and outputs the second labeled images through the first image processing model. In the embodiments of the present application, the first image processing model is trained by a semi-supervised learning method using the first sample images and the second sample images. Therefore, the first image processing model can accurately label the input images.
[0175] In the embodiments of the present application, through semi-supervised learning, a data set containing a small amount of labeled data and a large amount of unlabeled data is learned. Thus, the first image processing model is used to label the unlabeled data, and further approaches or even achieves the effect of fully labeling the data set, thereby reducing the heavy manual labeling amount in the image segmentation task. For example, in the case of 30% labeled data, the effect close to 100% fully labeled data can be achieved.
[0176] 514. The server trains a second image processing model based on multiple first sample images, first annotation images of the multiple first sample images, multiple second sample images, and multiple second annotation images. The second image processing model is used to determine the annotation image of the input image.
[0177] Among them, the input data of the second image processing model are the first sample images, the first annotation images, the second sample images, and the second annotation images, and the output data are the annotation images obtained by annotating the first sample images and the second sample images respectively. In some embodiments, the implementation manner of the server training the second image processing model based on the multiple first sample images, the first annotation images of the multiple first sample images, the multiple second sample images, and the multiple second annotation images includes the following steps (1)-(3):
[0178] (1) The server inputs the multiple first sample images, the multiple first annotation images, the multiple second sample images, and the multiple second annotation images into the second image processing model, and trains the second image processing model with the multiple first annotation images and the multiple second annotation images as supervision.
[0179] Among them, since the input data of the second image processing model are all sample images with annotated images, the second image processing model is trained in a fully supervised learning manner. In the embodiments of the present application, the training process of the second image processing model includes multiple iteration processes, and the iteration manner is the same as that used in the training processes of the student model and the teacher model, and will not be elaborated here. Among them, the second image processing model annotates the second sample images through the first image processing model, so as to perform self-training in a fully supervised learning manner, improving the accuracy of self-training.
[0180] (2) The server re-determines the second annotation image of each second sample image based on the trained second image processing model, and obtains multiple updated second annotation images.
[0181] Among them, the server inputs the second sample images into the second image processing model, and outputs the updated second annotation images through the second image processing model.
[0182] In some embodiments, the server first tests the accuracy of the second image processing model, and then updates the second annotation image of the second sample image. Before step (2), the image processing method provided in the embodiments of the present application further includes: the server determines the model accuracy of the second image processing model based on the test image, the test annotation image of the test image, and the trained second image processing model. The test annotation image is obtained by annotating the tissue category to which the physiological tissue region in the test image belongs; when the model accuracy meets the conditions, the operation of step (2) is executed.
[0183] Optionally, the server inputs the test image into the trained second image processing model to obtain a predicted labeled image of the test image, compares the predicted labeled image with the test labeled image, and obtains a test result, which can represent the accuracy of the model. This condition is that the accuracy is higher than the first threshold, and the embodiments of the present application do not limit the setting of the first threshold.
[0184] In the embodiments of the present application, by first testing the second image processing model and using the second image processing model to relabel the second sample image only when the accuracy of the model meets the condition, the accuracy of the updated second labeled image is improved.
[0185] (3) The server performs updated training based on multiple first sample images, multiple first labeled images, multiple second sample images, and multiple updated second labeled images on the basis of the trained second image processing model.
[0186] Among them, since the second labeled image is updated, the server can train a new second image processing model based on the updated second labeled image. In the embodiments of the present application, the server repeatedly executes the operations in steps (1)-(3) until the model accuracy of the second image processing model meets the condition. This condition may be that the accuracy no longer improves, or this condition may also be that the accuracy is higher than the second threshold, and the embodiments of the present application do not limit the setting of the second threshold.
[0187] For example, referring to Figure 7 , taking the consistency regularization method of CutMix as an example, first train the first image processing model, predict the second sample image (unlabeled data) through the first image processing model to generate a second labeled image, that is, a pseudo label, and then load the second labeled image and the second sample image together with the first sample image (labeled data) and the first labeled image into the second image processing model for model training. Through multiple iterations, the second labeled image of the second sample image is continuously updated until the accuracy of the second image processing model on the test set no longer improves.
[0188] In the embodiments of the present application, after the second image processing model is trained, the technical personnel compare the performance of the second image processing model with the performance of the image processing models trained by other image processing methods in the related art. For example, the consistency regularization method, the self-training method, or the fully supervised learning method. Referring to Figure 8, the model trained through full_supervision, self-training, and consistency regularization (such as CutMix) is compared with the first image processing model (proposed) and the second image processing model trained through the image processing method of the embodiments of the present application. The results show that when the dataset includes 30% first sample images and 70% second sample images, the performance of the second image processing model trained by the image processing method provided by the embodiments of the present application is much higher than that of the conventional consistency regularization method and self-training method. In addition, referring to Figure 9 , when the proportion of the first sample images in the dataset is different (that is, the proportion of labeled data is different), the performance of the second image processing model trained by the image processing method provided by the embodiments of the present application is still much higher than that of the conventional consistency regularization method and self-training method. Among them, the evaluation index can be the accuracy of the model.
[0189] In the technical solution provided by the embodiments of the present application, training is performed in the mode of a student model and a teacher model based on the first sample images and the second sample images. The student model is trained in a fully supervised manner based on the first sample images and in an unsupervised manner based on the second sample images, and the teacher model is trained based on the second sample images, so as to obtain a first image processing model with relatively high accuracy. Then, the pseudo-labels of the second sample images, that is, the second labeled images, are determined through the first image processing model, and there is no need for technicians to manually label the second sample images; further, the second image processing model is trained in a fully supervised manner based on the labeled first sample images and the second sample images. Since the accuracy of the labeled images of the two types of sample images is relatively high, the accuracy of the trained second image processing model is also relatively high. It can be seen that the above solution not only greatly reduces the workload of data annotation, but also improves the accuracy of model training.
[0190] Figure 10 is a flowchart of an image processing method provided by the embodiments of the present application. Referring to Figure 10 , the embodiments of the present application will be described by taking the execution entity as a server as an example. The method includes:
[0191] 1001. The server receives an image recognition request sent by the terminal. The image recognition request carries the target image to be recognized, and the target image includes at least one physiological tissue region.
[0192] Among them, when the user needs to perform image recognition, the terminal is triggered to send an image recognition request to the server. The image recognition request may carry one or more target images to be recognized. The target image includes medical images.
[0193] 1002. The server inputs the target image into the second image processing model, and through the second image processing model, an annotated image of the target image is output. The annotated image is used to represent the tissue category to which each physiological tissue region belongs.
[0194] Among them, the second image processing model is trained with multiple first sample images, the first annotated image of each first sample image, multiple second sample images, and the second annotated image of each second sample image. The first annotated image is obtained by manually annotating the tissue category to which the physiological tissue region in the first sample image belongs, and the second annotated image is obtained by the first image processing model annotating the tissue category to which the physiological tissue region in the second sample image belongs.
[0195] 1003. The server sends the annotated image to the terminal.
[0196] In the embodiment of the present application, since the server has previously trained the second image processing model, the terminal can determine a relatively accurate annotated image by relying on the server to determine the annotated image of the target image to be recognized, thereby improving the accuracy of the annotated image.
[0197] See Figure 11 , in the embodiment of the present application, the second image processing model is first trained offline on the server side. After the training is completed, the second image processing model is deployed to the medical and health service cloud platform. Thus, the user logs in to the cloud platform through the terminal, uploads a medical image (such as a pathological section image), and the terminal calls the second image processing model with the help of the server, thereby determining the annotated image of the medical image.
[0198] Figure 12 is a flowchart of an image processing method provided by an embodiment of the present application. See Figure 12 , the embodiment of the present application takes the data interaction between the terminal and the server as an example for illustration. The method includes:
[0199] 1201. The terminal acquires a target image to be recognized, and the target image contains at least one physiological tissue region.
[0200] In some embodiments, the terminal acquires the target image to be recognized through an image acquisition device.
[0201] 1202. The terminal sends an image recognition request to the server, and the image recognition request carries the target image to be recognized.
[0202] For example, taking the target image as a pathological section image, the user captures the pathological section image through an image acquisition device, and the image acquisition device sends the pathological section image to the terminal. The terminal generates an image recognition request in response to the image recognition operation triggered based on the pathological section image. For example, the image recognition operation is a submission operation or other triggering operation for the pathological section image.
[0203] 1203. The server receives the image recognition request sent by the terminal.
[0204] After receiving the image recognition request, the server obtains the target image from the image recognition request.
[0205] 1204. The server inputs the target image into the second image processing model, and through the second image processing model, outputs an annotated image of the target image. The annotated image is used to represent the tissue category to which each physiological tissue region belongs.
[0206] The second image processing model performs image recognition on the target image to obtain the annotated image. In some embodiments, for the training process of the second image processing model, refer to the embodiments shown Figure 5 and will not be elaborated here.
[0207] 1205. The server sends the annotated image to the terminal.
[0208] After obtaining the annotated image, the server sends the annotated image to the terminal so that the user can view the annotated image through the terminal.
[0209] 1206. The terminal receives the annotated image and displays the annotated image.
[0210] Among them, the terminal displays the annotated image on the display interface. The terminal can display or hide at least one physiological tissue region in the annotated image. For any physiological tissue region, in the case of display, the physiological tissue region is displayed in the style of the annotated image, and in the case of hiding, the physiological tissue region is displayed in the style of the target image. In some embodiments, the terminal responds to the hiding operation on any physiological tissue region in the annotated image and displays the physiological tissue region in the style of the target image. Among them, the present application embodiment does not limit the setting of the hiding operation. For example, the hiding operation is a click operation. After receiving the annotated image, the terminal defaults to displaying the at least one physiological tissue region in the style of the annotated image. In the embodiments of the present application, by displaying or hiding the tissue category to which the physiological tissue region belongs, the user can conveniently view the annotated image and the target image, improving the convenience of image viewing.
[0211] In the embodiments of the present application, since the server has pre-trained the second image processing model, the terminal can determine a relatively accurate labeled image by determining the labeled image of the target image to be recognized with the help of the server, thereby improving the accuracy of the labeled image.
[0212] Figure 13 It is a schematic structural diagram of an image processing device provided by an embodiment of the present application. Refer to Figure 13 , the device includes:
[0213] The first acquisition module 1301 is configured to obtain, through the student model, the first predicted labeled image of the first sample image in the first sample set, and obtain the mixed labeled image of the mixed image of the second sample image pair in the second sample set, where the second sample image pair includes at least two second sample images;
[0214] The second acquisition module 1302 is configured to obtain, through the teacher model, the labeled mixed image of the second sample image pair, where the labeled mixed image is obtained based on the second predicted labeled images of at least two second sample images;
[0215] The first training module 1303 is configured to train the student model and the teacher model based on the first predicted labeled image, the first labeled image of the first sample image, the mixed labeled image, and the labeled mixed image, so as to obtain the first image processing model, where the first image processing model is used to determine the labeled image of the input image, and the first labeled image is used to represent the tissue category to which the physiological tissue region in the first sample image belongs;
[0216] The third acquisition module 1304 is configured to obtain, based on the first image processing model, the second labeled images of multiple second sample images in the second sample set, where the second labeled images are used to represent the predicted tissue categories to which the physiological tissue regions in the second sample images belong;
[0217] The second training module 1305 is configured to train the second image processing model based on multiple first sample images in the first sample set, the first labeled images of the multiple first sample images, multiple second sample images, and multiple second labeled images, where the second image processing model is used to determine the labeled image of the input image.
[0218] In some embodiments, the first training module 1303 includes:
[0219] The first determination unit is configured to determine the first difference information between the first predicted labeled image and the first labeled image;
[0220] The second determination unit is configured to determine the second difference information between the mixed labeled image and the labeled mixed image;
[0221] A training unit for training a student model and a teacher model based on first difference information and second difference information to obtain a first image processing model.
[0222] In some embodiments, a second determination unit is configured to determine third difference information between a mixed labeled image and a labeled mixed image; and use the ratio between the third difference information and the total number of pixel points included in the mixed labeled image as the second difference information.
[0223] In some embodiments, the second determination unit is configured to determine a plurality of pixel point pairs, and determine the squared value of the pixel value difference between a first pixel point and a second pixel point included in each pixel point pair, where the first pixel point is a pixel point in the mixed labeled image, the second pixel point is a pixel point at the corresponding position in the labeled mixed image, and the pixel value of any pixel point represents the predicted tissue category to which the pixel point belongs; and use the sum of the squared values of the plurality of pixel point pairs as the third difference information.
[0224] In some embodiments, the training unit is configured to perform weighted summation on the first difference information and the second difference information to obtain fourth difference information; adjust the model parameters of the student model based on the fourth difference information, and adjust the model parameters of the teacher model based on the adjusted model parameters of the student model; train the student model and the teacher model respectively based on the corresponding adjusted model parameters; and determine the first image processing model from the trained student model and the trained teacher model.
[0225] In some embodiments, the apparatus further includes:
[0226] A cutmix module for obtaining a second sample image pair from a second sample set; and performing cutmix on at least two second sample images in the second sample image pair to obtain a mixed image.
[0227] In some embodiments, the at least two second sample images include a third sample image and a fourth sample image; a cutmix unit is configured to determine a cropping window; crop a first image region corresponding to the cropping window in the third sample image; extract a second image region corresponding to the cropping window from the fourth sample image, and fill the second image region to the position where the first image region is located in the third sample image to obtain a mixed image.
[0228] In some embodiments, a second acquisition module 1302 is configured to input the second sample image pair into the teacher model, and output second predicted labeled images of at least two second sample images through the teacher model; and perform cutmix on the at least two second predicted labeled images to obtain a labeled mixed image.
[0229] In some embodiments, the second training module 1305 is configured to input a plurality of first sample images, a plurality of first annotation images, a plurality of second sample images, and a plurality of second annotation images into a second image processing model, and train the second image processing model supervised by the plurality of first annotation images and the plurality of second annotation images; based on the trained second image processing model, re-determine the second annotation image of each second sample image to obtain a plurality of updated second annotation images; and perform updated training on the basis of the trained second image processing model based on the plurality of first sample images, the plurality of first annotation images, the plurality of second sample images, and the plurality of updated second annotation images.
[0230] In some embodiments, the apparatus further includes:
[0231] A testing module, configured to determine the model accuracy of the second image processing model based on a test image, a test annotation image of the test image, and the trained second image processing model, where the test annotation image is obtained by annotating the tissue category to which the physiological tissue region in the test image belongs;
[0232] The second training module 1305 is further configured to, when the model accuracy meets the conditions, re-determine the second annotation image of each second sample image based on the trained second image processing model to obtain a plurality of updated second annotation images.
[0233] In the technical solution provided by the embodiments of the present application, training is performed in the mode of a student model and a teacher model based on the first sample image and the second sample image. By training the student model in a fully supervised manner based on the first sample image and in an unsupervised manner based on the second sample image, and training the teacher model based on the second sample image, a first image processing model with relatively high accuracy is obtained. Then, the pseudo-label of the second sample image, that is, the second annotation image, is determined through the first image processing model, so that there is no need for technicians to manually annotate the second sample image; further, by training the second image processing model in a fully supervised manner based on the annotated first sample image and the second sample image, since the accuracy of the annotation images of the two types of sample images is relatively high, the accuracy of the trained second image processing model is also relatively high. It can be seen that the above solution not only greatly reduces the workload of data annotation, but also improves the accuracy of model training.
[0234] Figure 14 is a schematic structural diagram of an image processing apparatus provided by an embodiment of the present application. Refer to Figure 14 The apparatus includes:
[0235] A receiving module 1401, configured to receive an image recognition request sent by a terminal, where the image recognition request carries a target image to be recognized, and the target image includes at least one physiological tissue region;
[0236] A labeling module 1402, configured to input a target image into a second image processing model, and output a labeled image of the target image through the second image processing model, where the labeled image is used to represent the tissue category to which each physiological tissue region belongs;
[0237] A sending module 1403, configured to send the labeled image to a terminal;
[0238] Wherein, the second image processing model is trained by multiple first sample images, the first labeled image of each first sample image, multiple second sample images, and the second labeled image of each second sample image. The first labeled image is obtained by manually labeling the tissue category to which the physiological tissue region in the first sample image belongs, and the second labeled image is obtained by the first image processing model labeling the tissue category to which the physiological tissue region in the second sample image belongs.
[0239] In the embodiments of the present application, since the server has pre-trained the second image processing model, the terminal can determine a relatively accurate labeled image by relying on the server to determine the labeled image of the target image to be recognized, thereby improving the accuracy of the labeled image.
[0240] It should be noted that: when the image processing device provided in the above embodiments performs image processing, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the image processing device provided in the above embodiments and the embodiments of the image processing method belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0241] The embodiments of the present application provide a computer device for executing the above image processing method. The computer device can be provided as a terminal. The structure of the terminal will be introduced below:
[0242] Figure 15 It is a schematic structural diagram of a terminal provided in the embodiments of the present application. The terminal 1500 can be: a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch, etc., but is not limited thereto.
[0243] Generally, the terminal 1500 includes: one or more processors 1501 and one or more memories 1502.
[0244] The processor 1501 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1501 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1501 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0245] The memory 1502 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1502 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1502 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 1501 to implement the image processing method provided in the method embodiments of the present application.
[0246] In some embodiments, the terminal 1500 may further optionally include: a peripheral device interface 1503 and at least one peripheral device. The processor 1501, the memory 1502, and the peripheral device interface 1503 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1503 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1504, a display screen 1505, a camera assembly 1506, an audio circuit 1507, and a power supply 1509.
[0247] The peripheral device interface 1503 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1501 and the memory 1502. In some embodiments, the processor 1501, the memory 1502, and the peripheral device interface 1503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1501, the memory 1502, and the peripheral device interface 1503 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0248] The radio frequency circuit 1504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1504 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1504 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on.
[0249] The display screen 1505 is used to display a UI (User Interface). The UI can include graphics, text, icons, videos, and any combination thereof. When the display screen 1505 is a touch display screen, the display screen 1505 also has the ability to collect touch signals on or above the surface of the display screen 1505. The touch signal can be input to the processor 1501 as a control signal for processing. At this time, the display screen 1505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard.
[0250] The camera assembly 1506 is used to collect images or videos. Optionally, the camera assembly 1506 includes a front camera and a rear camera. Usually, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal.
[0251] The audio circuit 1507 can include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1501 for processing, or input to the radio frequency circuit 1504 to achieve voice communication.
[0252] The power supply 1509 is used to supply power to each component in the terminal 1500. The power supply 1509 can be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0253] In some embodiments, the terminal 1500 further includes one or more sensors 1510. The one or more sensors 1510 include but are not limited to: an acceleration sensor 1511, a gyroscope sensor 1512, a pressure sensor 1513, an optical sensor 1515, and a proximity sensor 1516.
[0254] The acceleration sensor 1511 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 1500.
[0255] The gyroscope sensor 1512 can detect the body direction and rotation angle of the terminal 1500. The gyroscope sensor 1512 can cooperate with the acceleration sensor 1511 to collect the 3D actions of the user on the terminal 1500.
[0256] The pressure sensor 1513 can be disposed on the side frame of the terminal 1500 and / or the lower layer of the display screen 1505. When the pressure sensor 1513 is disposed on the side frame of the terminal 1500, it can detect the holding signal of the user on the terminal 1500, and the processor 1501 can perform left - hand / right - hand recognition or quick operations according to the holding signal collected by the pressure sensor 1513. When the pressure sensor 1513 is disposed on the lower layer of the display screen 1505, the processor 1501 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1505.
[0257] The optical sensor 1515 is used to collect the ambient light intensity. In one embodiment, the processor 1501 can control the display brightness of the display screen 1505 according to the ambient light intensity collected by the optical sensor 1515.
[0258] The proximity sensor 1516 is used to collect the distance between the user and the front of the terminal 1500.
[0259] Those skilled in the art can understand that Figure 15 the structure shown in does not constitute a limitation on the terminal 1500, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.
[0260] The above - mentioned computer device can also be provided as a server. The structure of the server will be introduced below:
[0261] Figure 16FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1600 may vary greatly due to different configurations or performances, and may include one or more Central Processing Units (CPUs) 1601 and one or more memories 1602. Among them, at least one computer program is stored in the one or more memories 1602, and the at least one computer program is loaded and executed by the one or more processors 1601 to implement the image processing methods provided by the above-mentioned method embodiments. Of course, the server 1600 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server 1600 may also include other components for implementing the functions of the device, which will not be elaborated here.
[0262] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program. The above computer program can be executed by a processor to complete the image processing method in the above embodiment. For example, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0263] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes program code, and the program code is stored in a computer-readable storage medium. The processor of the computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above image processing method.
[0264] In some embodiments, the computer program involved in the embodiments of the present application may be deployed to be executed on a single computer device, or on multiple computer devices located at one location. Or, on multiple computer devices distributed at multiple locations and interconnected through a communication network. The multiple computer devices distributed at multiple locations and interconnected through a communication network may form a blockchain system.
[0265] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a magnetic disk, or an optical disc, etc.
[0266] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining, by a student model, a first predicted annotation image of a first sample image in a first sample set, and obtaining a mixed annotation image of a mixed image of a second sample image pair in a second sample set, where the second sample image pair includes at least two second sample images; Obtaining, by a teacher model, an annotated mixed image of the second sample image pair, where the annotated mixed image is obtained based on second predicted annotation images of the at least two second sample images; Training the student model and the teacher model based on the first predicted annotation image, a first annotation image of the first sample image, the mixed annotation image, and the annotated mixed image; testing the trained student model and teacher model with test images in a test set to obtain a test result, where the test result is used to represent the accuracy of model testing; in the trained student model and teacher model, using the model with a better test result than the other model as a first image processing model, where the first image processing model is used to determine an annotation image of an input image, and the first annotation image is used to represent the tissue category to which the physiological tissue region in the first sample image belongs; Based on the first image processing model, obtaining second annotation images of multiple second sample images in the second sample set, where the second annotation images are used to represent the predicted tissue categories to which the physiological tissue regions in the second sample images belong; Training a second image processing model based on the multiple first sample images in the first sample set, the first annotation images of the multiple first sample images, the multiple second sample images, and the multiple second annotation images, where the second image processing model is used to determine an annotation image of an input image; Based on the trained second image processing model, re - determining the second annotation images of each second sample image to obtain multiple updated second annotation images; Based on the multiple first sample images, the multiple first annotation images, the multiple second sample images, and the multiple updated second annotation images, performing updated training on the basis of the trained second image processing model.
2. The method according to claim 1, wherein The training the student model and the teacher model based on the first predicted annotation image, the first annotation image of the first sample image, the mixed annotation image, and the annotated mixed image includes: Determining first difference information between the first predicted annotation image and the first annotation image; Determining second difference information between the mixed annotation image and the annotated mixed image; Training the student model and the teacher model based on the first difference information and the second difference information.
3. The method according to claim 2, wherein The determining the second difference information between the mixed annotation image and the annotated mixed image includes: Determining third difference information between the mixed annotation image and the annotated mixed image; Using the ratio between the third difference information and the total number of pixel points included in the mixed annotation image as the second difference information.
4. The method according to claim 3, wherein The determining the third difference information between the mixed annotation image and the annotated mixed image includes: Determine multiple pairs of pixel points, and determine the square value of the pixel value difference between the first pixel point and the second pixel point included in each of the pairs of pixel points. The first pixel point is a pixel point in the hybrid annotation image, and the second pixel point is a pixel point at the corresponding position in the annotation hybrid image. The pixel value of any pixel point represents the predicted tissue category to which the pixel point belongs; Use the sum of the square values of the multiple pairs of pixel points as the third difference information.
5. The method according to claim 2, wherein The training of the student model and the teacher model based on the first difference information and the second difference information includes: Perform weighted summation on the first difference information and the second difference information to obtain fourth difference information; Based on the fourth difference information, adjust the model parameters of the student model, and based on the adjusted model parameters of the student model, adjust the model parameters of the teacher model; Based on the corresponding adjusted model parameters, train the student model and the teacher model respectively; Determine the first image processing model from the trained student model and the trained teacher model.
6. The method according to claim 1, wherein The method further includes: Obtain the second sample image pair from the second sample set; Perform cutmix on at least two second sample images in the second sample image pair to obtain the hybrid image.
7. The method according to claim 6, wherein The at least two second sample images include a third sample image and a fourth sample image; the performing cutmix on at least two second sample images in the second sample image pair to obtain the hybrid image includes: Determine a cropping window; In the third sample image, crop the first image area corresponding to the cropping window; Extract the second image area corresponding to the cropping window from the fourth sample image, and fill the second image area into the position where the first image area is located in the third sample image to obtain the hybrid image.
8. The method according to claim 1, wherein The obtaining the annotation hybrid image of the second sample image pair through the teacher model includes: Input the second sample image pair into the teacher model, and through the teacher model, output the second predicted annotation images of the at least two second sample images; Perform cutmix on at least two of the second predicted annotation images to obtain the annotation hybrid image.
9. The method according to claim 1, characterized in that The training of the second image processing model based on the multiple first sample images in the first sample set, the first annotation images of the multiple first sample images, the multiple second sample images, and the multiple second annotation images includes: Input the multiple first sample images, the multiple first annotation images, the multiple second sample images, and the multiple second annotation images into the second image processing model, and train the second image processing model with the multiple first annotation images and the multiple second annotation images as supervision.
10. The method according to claim 1, characterized in that, The method further includes: Based on the test image, the test annotation image of the test image, and the trained second image processing model, determine the model accuracy of the second image processing model. The test annotation image is obtained by annotating the tissue category to which the physiological tissue area in the test image belongs; When the accuracy of the model meets the conditions, perform the step of re-determining the second labeled image of each second sample image based on the trained second image processing model to obtain a plurality of updated second labeled images.
11. An image processing method, characterized in that, The method includes: Receiving an image recognition request sent by a terminal, where the image recognition request carries a target image to be recognized, and the target image includes at least one physiological tissue region; Inputting the target image into a second image processing model, and outputting a labeled image of the target image through the second image processing model, where the labeled image is used to represent the tissue category to which each physiological tissue region belongs; Sending the labeled image to the terminal; Among them, the second image processing model is trained by a plurality of first sample images, the first labeled image of each first sample image, a plurality of second sample images, and the second labeled image of each second sample image. The first labeled image is obtained by manually labeling the tissue category to which the physiological tissue region in the first sample image belongs. The second labeled image is obtained by a first image processing model labeling the tissue category to which the physiological tissue region in the second sample image belongs. The first image processing model is the model with better test results among the trained student model and teacher model. The test result is obtained by testing the trained student model and teacher model based on the test images in the test set, and the test result is used to represent the accuracy of the model test; The trained second image processing model is used to re-determine the second labeled image of each second sample image to obtain a plurality of updated second labeled images; the plurality of first sample images, the plurality of first labeled images, the plurality of second sample images, and the plurality of updated second labeled images are used for updated training based on the trained second image processing model.
12. An image processing apparatus, characterized in that, The device includes: A first acquisition module, configured to obtain, through a student model, a first predicted labeled image of a first sample image in a first sample set, and obtain a mixed labeled image of a mixed image of a second sample image pair in a second sample set, where the second sample image pair includes at least two second sample images; A second acquisition module, configured to obtain, through a teacher model, a labeled mixed image of the second sample image pair, where the labeled mixed image is obtained based on the second predicted labeled images of the at least two second sample images; The first training module is used to train the student model and the teacher model based on the first predicted annotation image, the first annotation image of the first sample image, the mixed annotation image, and the annotated mixed image; test the trained student model and teacher model based on the test images of the test set to obtain test results, and the test results are used to represent the accuracy of the model test; among the trained student model and teacher model, the model with better test results than the other model is used as the first image processing model, and the first image processing model is used to determine the annotation image of the input image, and the first annotation image is used to represent the tissue category to which the physiological tissue region in the first sample image belongs; The third acquisition module is used to obtain the second annotation images of multiple second sample images in the second sample set based on the first image processing model, and the second annotation images are used to represent the predicted tissue categories to which the physiological tissue regions in the second sample images belong; The second training module is used to train the second image processing model based on multiple first sample images in the first sample set, the first annotation images of the multiple first sample images, the multiple second sample images, and multiple second annotation images, and the second image processing model is used to determine the annotation image of the input image; based on the trained second image processing model, re-determine the second annotation images of each second sample image to obtain multiple updated second annotation images; based on the multiple first sample images, the multiple first annotation images, the multiple second sample images, and the multiple updated second annotation images, perform updated training on the basis of the trained second image processing model.
13. The device according to claim 12, characterized in that, The first training module includes: The first determination unit is used to determine the first difference information between the first predicted annotation image and the first annotation image; The second determination unit is used to determine the second difference information between the mixed annotation image and the annotated mixed image; The training unit is used to train the student model and the teacher model based on the first difference information and the second difference information.
14. The device according to claim 13, characterized in that The second determination unit is used to: Determine the third difference information between the mixed annotation image and the annotated mixed image; Use the ratio between the third difference information and the total number of pixel points included in the mixed annotation image as the second difference information.
15. The device according to claim 14, characterized in that, The second determination unit is used to: Determine multiple pixel point pairs, and determine the squared value of the pixel value difference between the first pixel point and the second pixel point included in each pixel point pair. The first pixel point is a pixel point in the mixed annotation image, and the second pixel point is a pixel point at the corresponding position in the annotated mixed image. The pixel value of any pixel point represents the predicted tissue category to which the pixel point belongs; Use the sum of the squared values of the multiple pixel point pairs as the third difference information.
16. The device according to claim 13, characterized in that, The training unit is used to: Perform weighted summation on the first difference information and the second difference information to obtain the fourth difference information; Adjust the model parameters of the student model based on the fourth difference information, and adjust the model parameters of the teacher model based on the adjusted model parameters of the student model; Train the student model and the teacher model respectively based on the corresponding adjusted model parameters; Determine the first image processing model from the trained student model and the trained teacher model.
17. The device according to claim 12, characterized in that, The apparatus further includes a cutmix module, configured to: Obtain the second sample image pair from the second sample set; Perform cutmix on at least two second sample images in the second sample image pair to obtain the mixed image.
18. The device according to claim 17, characterized in that, The at least two second sample images include a third sample image and a fourth sample image; the cutmix module is configured to: Determine a cropping window; In the third sample image, crop the first image region corresponding to the cropping window; Extract the second image region corresponding to the cropping window from the fourth sample image, and fill the second image region into the position of the first image region in the third sample image to obtain the mixed image.
19. The device according to claim 12, characterized in that, The second acquisition module is configured to: Input the second sample image pair into the teacher model, and output, through the teacher model, the second predicted annotation images of the at least two second sample images; Perform cutmix on at least two of the second predicted annotation images to obtain the annotation mixed image.
20. The device according to claim 12, characterized in that, The second training module is configured to: Input the multiple first sample images, the multiple first annotation images, the multiple second sample images, and the multiple second annotation images into the second image processing model, and train the second image processing model with the multiple first annotation images and the multiple second annotation images as supervision.
21. The device according to claim 12, wherein, The apparatus further includes: A test module, configured to determine the model accuracy of the second image processing model based on a test image, a test annotation image of the test image, and the trained second image processing model, where the test annotation image is obtained by annotating the tissue category to which the physiological tissue region in the test image belongs; The second training module is further configured to, when the model accuracy meets the condition, perform the step of re-determining the second annotation image of each second sample image based on the trained second image processing model to obtain multiple updated second annotation images.
22. An image processing apparatus, characterized in that, The apparatus includes: A receiving module, configured to receive an image recognition request sent by a terminal, where the image recognition request carries a target image to be recognized, and the target image includes at least one physiological tissue region; A labeling module, configured to input the target image into the second image processing model, and output, through the second image processing model, a labeling image of the target image, where the labeling image is used to represent the tissue category to which each physiological tissue region belongs; A sending module, configured to send the labeling image to the terminal; Among them, the second image processing model is trained with a plurality of first sample images, the first annotation image of each first sample image, a plurality of second sample images, and the second annotation image of each second sample image. The first annotation image is obtained by manually annotating the tissue category to which the physiological tissue region in the first sample image belongs. The second annotation image is obtained by the first image processing model annotating the tissue category to which the physiological tissue region in the second sample image belongs. The first image processing model is the model with better test results among the trained student model and teacher model. The test results are obtained by testing the trained student model and teacher model based on the test images in the test set. The test results are used to represent the accuracy of the model test; The trained second image processing model is used to re-determine the second annotation image of each second sample image to obtain a plurality of updated second annotation images; the plurality of first sample images, the plurality of first annotation images, the plurality of second sample images, and the plurality of updated second annotation images are used to perform updated training based on the trained second image processing model.
23. A computer device, characterized in that, The computer device includes one or more processors and one or more memories. At least one computer program is stored in the one or more memories. The computer program is loaded and executed by the one or more processors to implement the image processing method according to any one of claims 1 to 11.
24. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium. The computer program is loaded and executed by a processor to implement the image processing method according to any one of claims 1 to 11.
25. A computer program product, the computer program product includes program code, the program code is stored in a computer-readable storage medium, a processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code so that the computer device executes to implement the image processing method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Semi-supervised learning method and system based on target segmentation field self-learning
CN112381098A
Remote sensing image deep network semi-supervised semantic segmentation method based on transformation consistency regularization
CN113378736A