A training method for an image segmentation model and related devices

Through the methods of data augmentation and pseudo-label generation of medical images, the problems of poor prediction accuracy and insufficient generalization caused by grayscale distribution drift in medical images are solved, and the performance and generalization of image segmentation models are improved.

CN113989501BActive Publication Date: 2025-06-03SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111236105.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-06-03
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

There is a problem of grayscale distribution drift in medical images, resulting in poor accuracy and lack of generalization in intelligent models when predicting.

Method used

By augmenting the test images data, a pre-trained segmentation model is used to determine the predicted probability map, a master probability map is generated, the maximum joint mask map is divided, the pseudo label is determined, and the segmentation model is fine-tuned based on the pseudo label until the fine-tuning end condition is met.

Benefits of technology

The performance and generalization of image segmentation models are improved, the model is adjusted online through self-supervised learning, and the model is driven online learning using data augmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989501B_ABST
    Figure CN113989501B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for an image segmentation model and related devices. The method includes performing data augmentation on test images in a test image set to obtain a number of augmented test images; determining a predicted probability map of the test images and predicted probability maps of the augmented test images through a pre-segmentation model, and determining a master probability map based on the predicted probability maps; determining a number of maximum joint mask maps based on the master probability map and determining pseudo-labels based on the maximum joint mask maps; fine-tuning the model parameters of the segmentation model based on the pseudo-labels to obtain an image segmentation model. The present application online adjusts the segmentation model through self-supervised learning, uses test-time data augmentation to generate reliable pseudo-labels to drive the online learning of the segmentation model, and dynamically stops learning according to the consistency degree of the pseudo-labels, thereby improving the model performance and generalization degree of the image segmentation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of medical image processing, and particularly relates to a training method for an image segmentation model and related devices. Background Art

[0002] Data-driven learning algorithms have made significant breakthroughs in various critical and challenging tasks, but their premise is that the data used for model training and the data during testing are independently sampled from the same distribution. However, due to the influence of factors such as imaging devices, imaging parameters, and even different operators and imaging time points, there is a typical gray-scale distribution drift problem in medical images, which makes the intelligent model learned based on the training data have poor prediction accuracy and lack of generalization when dealing with medical images.

[0003] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0004] The technical problem to be solved by the present application is to provide a training method for an image segmentation model and related devices in view of the deficiencies of the existing technology.

[0005] To solve the above technical problem, in the first aspect of the embodiments of the present application, a training method for an image segmentation model is provided. The training method includes:

[0006] Performing data augmentation on the test images in the test image set to obtain a number of augmented test images;

[0007] Determining the prediction probability map of the test image and the prediction probability maps of the augmented test images through a pre-trained segmentation model, and determining a master probability map based on the determined prediction probability maps;

[0008] Respectively dividing the master probability map based on each of a number of preset thresholds to obtain a number of maximum union mask maps;

[0009] Determining the pseudo-labels of the test image and the augmented test images based on the number of maximum union mask maps;

[0010] Fine-tuning the segmentation model based on each pseudo-label, and continuing to perform the step of performing data augmentation on the test images in the test image set until the fine-tuning end condition is met to obtain an image segmentation model.

[0011] In the training method of the image segmentation model, wherein the determining a master probability map based on the determined prediction probability maps specifically includes:

[0012] Performing inverse augmentation operations on the prediction probability maps corresponding to the augmented test images respectively to obtain candidate test probability maps corresponding to the augmented test images respectively;

[0013] Add the candidate test probability maps and the predicted probability map corresponding to the test image to obtain a master probability map.

[0014] The training method of the image segmentation model, wherein the step of respectively dividing the master probability map based on each of a plurality of preset thresholds to obtain a plurality of maximum union mask maps specifically includes:

[0015] Select a plurality of preset thresholds, wherein the plurality of preset thresholds include all integers less than the number of predicted probability maps;

[0016] For each preset threshold, obtain the maximum difference among the differences between the pixel values of each pixel in the master probability map and the preset threshold, and generate the maximum union mask map corresponding to the preset threshold based on the minimum value of the maximum difference and 1, so as to obtain a plurality of maximum union mask maps.

[0017] The training method of the image segmentation model, wherein the step of determining the pseudo-labels of the test image and each enhanced test image based on the plurality of maximum union mask maps specifically includes:

[0018] Select a preset number of maximum union mask maps from the plurality of maximum union mask maps to form a pseudo-label set;

[0019] Adopt a distance similarity index to select the pseudo-label corresponding to each predicted probability map in the pseudo-label set, so as to obtain the pseudo-labels of the test image and each enhanced test image.

[0020] The training method of the image segmentation model, wherein the step of fine-tuning the segmentation model based on each pseudo-label and continuing to perform the step of data augmentation on the test images in the test image set until the fine-tuning end condition is met to obtain an image segmentation model specifically includes:

[0021] Correct the model parameters of the segmentation model based on each pseudo-label;

[0022] Select the first maximum union mask map corresponding to the maximum preset threshold and the second maximum union mask map corresponding to the minimum preset threshold, and calculate the similarity coefficient between the first maximum union mask map and the second maximum union mask map;

[0023] Judge whether the segmentation model meets the fine-tuning end condition based on the similarity coefficient and the number of fine-tuning times of the segmentation model, wherein the fine-tuning end condition is that the similarity coefficient is less than a preset coefficient, or the number of fine-tuning times is equal to a preset number threshold;

[0024] When the fine-tuning end condition is met, use the fine-tuned segmentation model as the image segmentation model;

[0025] When the fine-tuning end condition is not met, continue to perform the step of data augmentation on the test images in the test image set.

[0026] The training method of the image segmentation model, wherein the segmentation model includes at least one style conversion module; the specific steps of determining the predicted probability map of the test image and the predicted probability maps of the augmented test images by the pre-trained segmentation model are as follows:

[0027] For each reference test image in the test image group composed of the test image and the augmented test images, input the reference test image and the source domain image into the segmentation model respectively, wherein the source domain image is a training image in a preset training sample set for training the segmentation model;

[0028] Control the network layer of the segmentation model before the style conversion module to determine the first feature map corresponding to the reference test image and the second feature map corresponding to the source domain image;

[0029] Control the style conversion module to adjust the image gray level of the first feature map based on the image gray level distribution of the second feature map to obtain the adjusted first feature map;

[0030] Control the adjusted first feature map and the second feature map to pass through the network layer of the segmentation model after the style conversion module to obtain the predicted probability map of the reference test image, so as to obtain the predicted probability map of the test image and the predicted probability maps of the augmented test images.

[0031] The training method of the image segmentation model, wherein the specific steps of controlling the style conversion module to adjust the image gray level of the first feature map based on the image gray level distribution of the second feature map to obtain the adjusted first feature map are as follows:

[0032] For each channel of the first feature map and each channel of the second feature map, determine the sequence number of each pixel value in the channel among all pixels in the channel, wherein the sequence number is formed according to the order of pixel values from small to large;

[0033] For each channel in the first feature map, select the corresponding candidate channel in the second feature map, and select the corresponding candidate pixels for each pixel in the channel based on the sequence numbers of the pixels in the channel, and use the pixel values of the candidate pixels to replace the pixel values of the corresponding pixels to obtain the adjusted first feature map.

[0034] The training method of the image segmentation model, wherein the control style conversion module adjusts the image gray level of the first feature map based on the image gray level distribution of the second feature map to obtain the adjusted first feature map, specifically including:

[0035] For each channel of the first feature map and each channel of the second feature map, the channel is divided into a plurality of sliding windows in a sliding window manner, and the sequence number of each pixel value in all pixels of the channel window is determined for each pixel in each sliding window, where the sequence number is formed according to the order of pixel values from small to large;

[0036] For each sliding window in the first feature map, a candidate sliding window corresponding to the sliding window is selected in the second feature map, and candidate pixels corresponding to each pixel are selected in the candidate sliding window based on the sequence numbers corresponding to each pixel in the sliding window, and the pixel values of the candidate pixels are used to replace the pixel values of the pixels corresponding to the candidate pixels respectively to obtain the adjusted sliding window corresponding to the sliding window;

[0037] For each pixel in each channel of the first feature map, the adjusted pixel values of the pixel in each adjusted sliding window including the pixel are obtained, and the average value of all the obtained adjusted pixel values is used as the pixel value of the pixel to obtain the adjusted first feature map.

[0038] A second aspect of the embodiments of the present application provides a computer-readable storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the training method of the image segmentation model as described in any one of the above.

[0039] A third aspect of the embodiments of the present application provides a terminal device, which includes: a processor, a memory, and a communication bus; a computer-readable program executable by the processor is stored on the memory;

[0040] The communication bus realizes the connection and communication between the processor and the memory;

[0041] When the processor executes the computer-readable program, the steps in the training method of the image segmentation model as described in any one of the above are implemented.

[0042] Beneficial effects: Compared with the prior art, the present application provides a method for training an image segmentation model and related devices. The method includes performing data augmentation on test images in a test image set to obtain a plurality of augmented test images; determining a predicted probability map of the test images and predicted probability maps of the augmented test images through a pre-trained segmentation model, and determining a master probability map based on the determined predicted probability maps; determining a plurality of maximum joint mask maps based on the master probability map; determining pseudo-labels of the test images and the augmented test images based on the plurality of maximum joint mask maps; fine-tuning the segmentation model based on each pseudo-label, and continuing to perform the step of performing data augmentation on the test images in the test image set until a fine-tuning end condition is met, so as to obtain an image segmentation model. The present application online adjusts the segmentation model through self-supervised learning, uses test-time data augmentation to generate reliable pseudo-labels to drive the online learning of the segmentation model, and dynamically stops learning according to the consistency degree of the pseudo-labels, thereby improving the model performance and generalization degree of the image segmentation model. Brief Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative labor, other drawings can be obtained based on these drawings.

[0044] Figure 1 It is a flowchart of the method for training the image segmentation model provided by the present application.

[0045] Figure 2 It is a flowchart of an embodiment of the method for training the image segmentation model provided by the present application.

[0046] Figure 3 It is a schematic diagram of the order statistic alignment operation in the method for training the image segmentation model provided by the present application.

[0047] Figure 4 It is a schematic diagram of the order statistic alignment based on a sliding window in the method for training the image segmentation model provided by the present application.

[0048] Figure 5 It is a schematic diagram of the pseudo-label generation strategy framework in the method for training the image segmentation model provided by the present application.

[0049] Figure 6 It is a schematic diagram of the structure of the terminal device provided by the present application. Detailed Embodiments

[0050] This application provides a training method and related device for an image segmentation model. To make the objectives, technical solutions, and effects of this application clearer and more explicit, the following further describes this application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain this application and are not used to limit this application.

[0051] Those skilled in the art of this technology can understand that unless specifically stated, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of this application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0052] Those skilled in the art of this technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0053] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0054] The inventors have found through research that data-driven learning algorithms have made significant breakthroughs in various key and challenging tasks, but its assumption is that the data used for model training and the data during testing are independently sampled from the same distribution. However, due to the influence of factors such as imaging devices, imaging parameters, and even different operators and imaging time points, there will be a typical gray-scale distribution drift problem in medical images, resulting in poor prediction accuracy and lack of generalization of the intelligent model learned based on the training data when applied to medical images.

[0055] The most direct way to solve this problem is to continuously increase the amount of training data, enabling the model to learn more complex possibilities of gray-scale transformation drift during the training process. However, in reality, the mixture of numerous imaging factors will cause each test image to exhibit a unique gray-scale distribution drift, and the possibilities of the distribution are infinite. Moreover, meticulously delineating annotation information for the data is both time-consuming and costly, especially in the medical field that requires professional knowledge, making it very challenging to apply in clinical practice. On the other hand, after deploying the well-trained model to medical devices for daily use, due to limitations in machine computing resources, such as some ultrasound machines only equipped with a CPU and limited memory, it is very difficult to implement a process of substantial re-learning for the network model. Therefore, how to overcome the risks brought by this unpredictable gray-scale distribution drift of images and largely avoid the dilemma of the model facing a large amount of repeated learning under different imaging conditions.

[0056] Many scholars have proposed various solutions to this problem. According to whether target data and its corresponding labels are required, these solutions can be divided into supervised learning and unsupervised learning:

[0057] (1) Supervised learning: The network is supervised with target data. At this time, not only sufficient target data is required, but also its corresponding annotation information. This method has been proven that in multi-center prostate segmentation, by using at least 8 annotated images from unknown centers to learn the pre-trained network, the problem of the model's generalization ability can be solved. However, similarly, annotating each new image is time-consuming and laborious, and requires certain professional knowledge, making it less practical in clinical practice.

[0058] (2) Unsupervised learning: For the target domain, only its images are used without corresponding annotation information, thus saving the annotation process for the target images. The two main methods are discriminative adversarial networks and image-to-image translation networks. The core idea of discriminative adversarial networks is to make the network learn the common invariant features of different image appearances through adversarial learning, while ignoring the drift of different appearance distributions. The image-to-image translation method can be further divided into generative adversarial networks and feed-forward approximate style transfer. Their essence is to convert the images in the source domain into images with the style features of the target domain images, and then train the target task network on the converted source domain data to achieve the purpose of model generalization. However, under clinical conditions, there are countless possible gray-scale distribution drifts, and not all drifts are known or available. Therefore, the premise of obtaining a sufficient amount of target domain image data does not always hold.

[0059] Based on this, in the embodiments of the present application, data augmentation is performed on the test images in the test image set to obtain a number of enhanced test images; the prediction probability map of the test image and the prediction probability maps of the enhanced test images are determined by a pre-trained segmentation model, and a master probability map is determined based on the determined prediction probability maps; a number of maximum joint mask maps are determined based on the master probability map; pseudo-labels of the test image and the enhanced test images are determined based on the number of maximum joint mask maps; the segmentation model is fine-tuned based on each pseudo-label, and the step of performing data augmentation on the test images in the test image set is continued until the fine-tuning end condition is met, so as to obtain an image segmentation model. The present application online adjusts the segmentation model through self-supervised learning, uses the method of test-time data augmentation to generate reliable pseudo-labels to drive the online learning of the segmentation model, and dynamically stops learning according to the consistency degree of the pseudo-labels, thereby improving the model performance and generalization degree of the image segmentation model.

[0060] The following further illustrates the application content by describing the embodiments in conjunction with the accompanying drawings.

[0061] This embodiment provides a training method for an image segmentation model, as Figure 1 and Figure 2 shown, the method includes:

[0062] S10. Perform data augmentation on the test images in the test image set to obtain a number of enhanced test images.

[0063] Specifically, the test image set includes a number of test images, and each test image in the number of test images does not carry annotation information. In this embodiment, each test image in the number of test images is a medical image. For example, each test image is an ultrasound image, or is a magnetic resonance image, etc. The test image set is used to test a pre-trained segmentation model, where the segmentation model is used for medical image segmentation. For example, the test image is a thyroid ultrasound image, and the segmentation model is used to segment the thyroid ultrasound image to obtain a nodule area in the thyroid ultrasound image, etc.

[0064] Each of the several enhanced test images is obtained by performing data augmentation on the same test image. It can be understood that for a test image in the test image set, data augmentation is performed on the test image several times to obtain several enhanced test images. Among them, the augmentation methods or augmentation intensities of each of the several enhanced test images are different. That is to say, for any two enhanced test images among the several enhanced test images, denoted as the first enhanced test image and the second enhanced test image respectively, the augmentation method and augmentation intensity of the first enhanced test image are not exactly the same as those of the second enhanced test image. For example, the augmentation method of the first enhanced test image is rotation, and the augmentation method of the second enhanced test image is horizontal mirroring, or the augmentation methods of the first enhanced test image and the second enhanced test image are both rotation, the rotation angle of the first enhanced test image is 30 degrees, and the rotation angle of the second enhanced test image is 40 degrees, etc. In a specific implementation, the several enhanced test images are 4 enhanced test images. One of the 4 enhanced test images is enhanced by the horizontal mirroring method, and the other three enhanced test images are all enhanced by the rotation method and the rotation angles of each enhanced test image are different.

[0065] S20. Determine the prediction probability map of the test image and the prediction probability maps of the enhanced test images through a pre-trained segmentation model, and determine the master probability map based on the determined prediction probability maps.

[0066] Specifically, the segmentation model is a pre-trained network model. The preset training sample set for training the segmentation model includes several training images, and each training image in the several training images carries annotation information. In addition, before training the segmentation model using the preset training sample set, data preprocessing can be performed on the preset training sample set. Among them, the data preprocessing at least includes standardization. In addition, it can also include operations such as normalization, scaling, data augmentation, and unbalanced sample processing. The standardization is specifically to subtract the mean of the pixels from all the pixels of the image and then divide by the standard deviation of the pixels, so that the gray distribution of the image satisfies a distribution with a mean of 0 and a variance of 1. In this embodiment, by normalizing the preset training sample set, the robustness and generalization ability of the segmentation model when facing test images with different gray distributions can be improved. In addition, in most cases, the appearance differences between the pre-training sample sets are large and the distributions are unbalanced. If the data is not preprocessed, it may have a certain impact on the subsequent segmentation model training process, such as restricting its accuracy, convergence speed, and generalization ability, etc. Therefore, through data preprocessing in this embodiment, the differences between the data can be reduced, the data imaged by different devices can be adapted, the generalization ability of the model can be enhanced, and at the same time, it is more conducive to training optimization.

[0067] In one implementation of this embodiment, the segmentation model includes at least one style conversion module; the specific process of determining the prediction probability map of the test image and the prediction probability maps of the enhanced test images through the pre-trained segmentation model is as follows:

[0068] For each reference test image in the test image group composed of the test image and the enhanced test images, input the reference test image and the source domain image into the segmentation model respectively, where the source domain image is a training image in a preset training sample set used for training the segmentation model;

[0069] Control the network layer of the segmentation model before the style conversion module to determine the first feature map corresponding to the reference test image and the second feature map corresponding to the source domain image;

[0070] Control the style conversion module to adjust the image gray level of the first feature map based on the image gray level distribution of the second feature map to obtain an adjusted first feature map;

[0071] Control the adjusted first feature map and the second feature map to pass through the network layer of the segmentation model after the style conversion module to obtain the prediction probability map of the reference test image, so as to obtain the prediction probability map of the test image and the prediction probability maps of the enhanced test images.

[0072] Specifically, the style conversion module is used to change the appearance gray level distribution of the test image according to the appearance gray level distribution of the source domain image to adjust the image gray level style of the test image. Among them, the style conversion module can adopt a plug-and-play framework and is configured in the segmentation model during the test process of the segmentation model. In addition, during the training process of the segmentation model, the segmentation model may not include a style conversion module, but the style conversion module is inserted into the segmentation model after training is completed, and the segmentation model with the inserted style conversion module is used as the pre-trained segmentation model; or, the style conversion module is configured during the training process of the segmentation model. When the training image passes through the network layer before the style conversion module, it skips the style conversion module and is input into the network layer after the style conversion module. Or, the style conversion module is configured during the training process of the segmentation model, and the input item of the style conversion module remains unchanged, that is, the input item and the output item of the style conversion module are the same.

[0073] In an implementation of this embodiment, the segmentation model may include an encoder and a decoder. The style conversion module is located inside the encoder, and the encoder may include one style conversion module or multiple style conversion modules. When there are multiple style conversion modules, their positions in the encoder are different. When the functions of the style conversion modules are the same, they are all used to adjust the image grayscale of the first feature map based on the image grayscale distribution of the second feature map. Herein, the first feature map refers to the one generated by the network layer before the style conversion module based on the test image, and the second feature map refers to the one generated by the network layer before the style conversion module based on the training image. It can be understood that when there are multiple style conversion modules, the image grayscale style of the test image is adjusted multiple times based on the image grayscale style of the source domain image. Here, an example where the encoder includes one style conversion module is used for illustration.

[0074] The reference test image includes the test image and the enhanced test images that form a test image group. That is to say, the reference test image can be the test image or an enhanced test image. The source domain image is the training image in the preset training sample set used to train the segmentation model. That is to say, when testing the segmentation model with the test image, a training image is selected from the preset training sample set used to train the segmentation model as the source domain image corresponding to the test image. Then, the test image and the source domain image are respectively input into the segmentation model, and the first feature map of the test image and the second feature map of the source domain image are respectively obtained through the segmentation model, so as to facilitate subsequent adjustment of the image grayscale of the first feature map based on the image grayscale distribution of the second feature map.

[0075] In an implementation of this embodiment, the specific steps for the control style conversion module to adjust the image grayscale of the first feature map based on the image grayscale distribution of the second feature map to obtain the adjusted first feature map include:

[0076] For each channel of the first feature map and each channel of the second feature map, determine the sequence number of each pixel value in all pixels of the channel;

[0077] For each channel in the first feature map, select the corresponding candidate channel in the second feature map, and based on the sequence numbers corresponding to each pixel in the channel, select the corresponding candidate pixels in the candidate channel, and use the pixel values of the candidate pixels to replace the pixel values of the corresponding pixels to obtain the adjusted first feature map.

[0078] Specifically, the image grayscale is used to reflect the image style features, and the image style features may include contrast, brightness, texture, and fine noise, etc. The sequence number is formed according to the ascending order of pixel values. It can be understood that the pixel values of each pixel in the channel are sorted in ascending order to obtain a pixel sequence, and the position number of each pixel in this pixel sequence is the sequence number of the pixel. For example, as Figure 3 shown in the channel F c [n], the pixel values of each pixel are sorted in ascending order to obtain Figure 3 the sequence number diagram of F c [n] shown in. Among them, the sequence number corresponding to -1.32 is 0, indicating that the pixel value of the pixel in the first column of the third row is ranked 0 among all pixel values; the sequence number corresponding to 2.57 is 8, indicating that the pixel value of the pixel in the third column of the third row is ranked 8 among all pixel values. Thus, each channel in the first feature map and each channel in the second feature map can form a sequence number diagram, and the sequence number at each position in this sequence number diagram identifies the sequence number of the pixel value of the pixel at this position in the corresponding channel among all pixel value rankings.

[0079] After obtaining the sequence number of each pixel in each channel, the channels in the first feature map are matched with the channels in the second feature map according to the channel numbers to obtain the channels in the second feature map corresponding to each channel in the first feature map, denoted as the candidate channels corresponding to each channel. Among them, the channel numbers of each channel are the same as the corresponding candidate channel numbers. This is because both the first feature map and the second feature map are output by the network layer before the style conversion module, so the image scale of the first feature map is the same as that of the second feature map, and thus each channel in the first feature map can select a candidate channel with the same channel number in the second feature map.

[0080] After obtaining the candidate channels, for each pixel in the channels of the first feature map, a candidate pixel with the same sequence number as the sequence number of this pixel is selected in the candidate channels, and the pixel value of the candidate pixel is used as the pixel value of the pixel corresponding to the candidate pixel. For example, as Figure 3 shown, for the pixel with sequence number 0 in the channel F c [n], the pixel with sequence number 0 in the candidate channel F s [n] is selected as the candidate pixel. If the pixel value of the pixel with sequence number 0 in F s [n] is -0.74, then -1.32 of the pixel with sequence number 0 in the channel F c [n] is replaced by -0.74. Thus, the channel F c [n] can be adjusted to obtain F′ c[n], that is, for each pixel value in F c [n], replace it with the pixel value in F s [n] that has the same sequence number as this value in F c [n]. When the pixel values in F c [n] are replaced to obtain F' c [n], the values between F' c [n] and F s [n] are all the same. In this way, the characteristic statistics representing the image style information, such as the mean, standard deviation, covariance, etc., can be migrated. At the same time, the ranking of the values of F' c [n] after replacement and F c [n] before replacement is consistent, that is, its order statistic characteristics are retained. In this way, the style conversion module can change the pixel values in the first feature map of the test image based on the pixel values of the pixels in the second feature map of the source domain image, so that the appearance distribution of the test image is transformed into the appearance distribution of the source domain image, and at the same time, the semantic structure information of the test image can be retained by maintaining the order statistic characteristics of the values of the feature map, so that the test image with distribution drift can be robustly segmented by the existing segmentation model.

[0081] In another implementation manner of this embodiment, since medical images generally have uneven gray-scale distributions, especially in ultrasonic images, the frequency of uneven gray-scale distributions is higher. Therefore, when adjusting the image gray scale through the style conversion module, the channels can be divided into several sliding windows, and then the image gray scale of each sliding window is adjusted to further solve the problem that medical images generally have uneven gray-scale distributions. Based on this, the control style conversion module adjusts the image gray scale of the first feature map based on the image gray-scale distribution of the second feature map to obtain the adjusted first feature map, which specifically includes:

[0082] For each channel of the first feature map and each channel of the second feature map, divide the channel into several sliding windows in a sliding window manner, and determine the sequence numbers of the pixel values of each pixel in each sliding window among all the pixels in the channel window, where the sequence numbers are formed according to the order of pixel values from small to large;

[0083] For each sliding window in the first feature map, select the corresponding candidate sliding window in the second feature map, and select the corresponding candidate pixels for each pixel based on the sequence numbers corresponding to each pixel in the sliding window in the candidate sliding window, and use the pixel values of the candidate pixels to replace the pixel values of the corresponding pixels for each pixel to obtain the adjusted sliding window corresponding to this sliding window;

[0084] For each pixel in each channel of the first feature map, obtain the adjusted pixel values of the pixel in each adjusted sliding window including the pixel, and take the average of all the obtained adjusted pixel values as the pixel value of the pixel, so as to obtain the adjusted first feature map.

[0085] Specifically, the sliding window is obtained by means of a sliding window method, and the window sizes of the sliding windows are the same. For example, they are all 3*3, etc., so that the number of sliding windows divided from each channel in the first feature map is the same as the number of sliding windows divided from each channel in the second feature map. Thus, the number of sliding windows corresponding to the first feature map is the same as the number of sliding windows corresponding to the second feature map, and the sliding windows corresponding to the first feature map and the sliding windows corresponding to the second feature map are in one-to-one correspondence. In addition, after obtaining a number of sliding windows corresponding to the first feature map and a number of sliding windows corresponding to the second feature map, perform image gray-scale adjustment on the obtained number of sliding windows corresponding to the first feature map. Among them, the process of performing image gray-scale adjustment on each sliding window is the same as the process of performing image gray-scale adjustment on each channel in the above embodiment, which will not be elaborated here, and can specifically refer to the above description.

[0086] Further, after obtaining the adjusted sliding windows corresponding to the sliding windows, since there may be a situation where two sliding windows include the same pixel. For example, if the sliding window is 3*3 and the step size is 2, then the pixel located in the third column of the first row will be included in two sliding windows. Therefore, when determining the adjusted first feature map after obtaining the adjusted sliding windows, for each pixel in each channel of the first feature map, obtain the pixel position of the pixel, and select the sliding window including the pixel position, read the pixel value at the pixel position in the adjusted sliding window corresponding to the sliding window, and take the average of all the read pixel values as the adjusted pixel value of the pixel, so as to obtain the adjusted first feature map. For example, as Figure 4 shown, the channel includes four sliding windows, which are respectively denoted as and Each sliding window performs order statistic alignment to obtain and Then, based on and form the adjusted channel, where the adjusted pixel value of each pixel included in multiple sliding windows is the average of the pixel values of the pixel in the adjusted sliding windows including the pixel, and the order statistic alignment is the process of replacing pixels based on the order number as described above.

[0087] In an implementation manner of this embodiment, the determining the master probability map based on the determined prediction probability maps specifically includes:

[0088] Inverse enhancement operations are respectively performed on the prediction probability maps corresponding to the respective enhanced test images to obtain the candidate test probability maps corresponding to the respective enhanced test images;

[0089] The candidate test probability maps and the prediction probability map corresponding to the test image are added together to obtain a master probability map.

[0090] Specifically, the master probability map is obtained by adding the candidate test probability maps and the prediction probability map corresponding to the test image. Here, "proximity" means adding the pixel values at the corresponding pixel positions of the candidate prediction probability maps and the prediction probability map. The image scale of the master probability map is the same as that of the prediction probability map, and the pixel value at each pixel position in the master probability map is equal to the sum of the pixel values at this pixel position in the candidate test probability maps and the pixel value at this pixel position in the prediction probability map. Among them, the value range of the pixel value at each pixel position in the master probability map is [0, N], and N is equal to the number of several enhanced prediction images plus 1. For example, as Figure 5 shown, if the number of several enhanced prediction images is 4, then the value range of the pixel value at each pixel position in the master probability map is [0, 5]. In addition, the inverse enhancement operation is used to restore the prediction results of each pixel in the prediction probability map corresponding to the enhanced test image to the original position to obtain the candidate test probability maps corresponding to the respective enhanced test images.

[0091] S30: The master probability map is respectively divided based on each of several preset thresholds to obtain several maximum union mask maps.

[0092] Specifically, the several preset thresholds are set in advance. Each of the several preset thresholds is an integer, and the number of the several preset thresholds is less than or equal to the number of several enhanced test images plus 1. In a typical implementation, the number of the several preset thresholds is equal to the number of several enhanced test images plus 1. For example, as Figure 5 shown, if the number of several enhanced test images is 4 and the number of the several preset thresholds is 5. The number of the several maximum union mask maps is equal to the number of the several preset thresholds. That is to say, the master probability map is divided into a maximum union mask map based on each preset threshold. Each maximum union mask map is obtained by dividing the master probability map, and the preset thresholds corresponding to the maximum union mask maps are different.

[0093] In an implementation manner of this embodiment, the step of respectively dividing the master probability map based on each of several preset thresholds to obtain several maximum union mask maps specifically includes:

[0094] Select several preset thresholds;

[0095] For each preset threshold, obtain the maximum difference among the differences between the pixel values of each pixel in the master probability map and the preset threshold, and generate the maximum joint mask map corresponding to the preset threshold based on the minimum value of the maximum difference and 1, so as to obtain a plurality of maximum joint mask maps.

[0096] Specifically, each preset threshold among a plurality of preset thresholds is an integer less than the number of prediction probability maps. Since the prediction probability maps include the prediction probability map corresponding to the test image and the prediction probability maps corresponding to each enhanced test image, the number of prediction probability maps is equal to the number of enhanced test images plus 1, and the number of a plurality of preset thresholds is equal to the number of enhanced test images plus 1. Therefore, each integer less than the number of prediction probability maps is a preset threshold. Thus, a plurality of preset thresholds can be 0, 1, 2,..., N - 1, where N is equal to the number of enhanced prediction images plus 1, that is, N is equal to the number of prediction probability maps. For example, if the number of preset probability maps is 5, then a plurality of preset thresholds include 0, 1, 2, 3, and 4. Thus, specifically selecting a plurality of preset thresholds can include: obtaining the number of prediction probability maps, selecting all integers less than the number of prediction probability maps, and finally taking all the selected integers as a plurality of preset thresholds.

[0097] After a plurality of preset thresholds are selected, divide the master probability map based on each preset threshold. Among them, the determination process of the maximum joint mask map can be expressed as:

[0098] p i,i+1 = min{max{P - i}, 1}, {i is an integer and i < N}

[0099] where i represents the preset threshold, P represents the master probability map, and p i,i+1 represents the maximum joint mask map, and N is equal to the number of enhanced prediction images plus 1.

[0100] In this embodiment, by using a plurality of adjacent integers to divide the master probability map into corresponding neighborhood intervals, a plurality of maximum joint mask maps with different confidence levels can be obtained. For example, as Figure 5 shown, a plurality of preset thresholds include 0, 1, 2, 3, and 4. Then, 5 maximum joint mask maps with different confidence levels can be obtained, which are p 0,1 , p 1,2 , p 2,3 , p 3,4 , p 4,5 , where p 0,1 can be regarded as the maximum joint union map of five prediction probability maps, which retains all positions with non-zero probabilities, p 1,2 , p 2,3 , p 3,4 are joint maps where the positions with at least 1, 2, 3 predictions are shown as foreground, p4,5 It can also be regarded as the intersection graph of five predicted probability graphs. In this way, the maximum joint mask graph with different confidence levels can be used as supplementary information for the predicted probability graph. The pseudo-labels determined based on the maximum joint mask graph and the predicted probability graph can give a more reasonable foreground indication, thereby improving the accuracy of fine-tuning the segmentation model based on the test image.

[0101] S40. Determine the pseudo-labels of the test image and each enhanced test image based on a plurality of maximum joint mask graphs.

[0102] Specifically, the pseudo-labels are used as annotation information. After obtaining the pseudo-labels of the test image and each enhanced test image, the segmentation model is fine-tuned based on the pseudo-labels. The pseudo-labels of the test image and each enhanced test image are both based on one maximum joint mask graph among a plurality of maximum joint mask graphs, and the pseudo-labels of the test image and each enhanced test image are determined based on the distance similarity index between their respective predicted probability graphs and each maximum joint mask graph.

[0103] Based on this, in an implementation manner of this embodiment, determining the pseudo-labels of the test image and each enhanced test image based on a plurality of maximum joint mask graphs specifically includes:

[0104] Select a preset number of maximum joint mask graphs from the plurality of maximum joint mask graphs to form a pseudo-label set;

[0105] Use the distance similarity index to select the pseudo-label corresponding to each predicted probability graph in the pseudo-label set to obtain the pseudo-labels of the test image and each enhanced test image.

[0106] Specifically, the pseudo-label set includes a preset number of maximum joint mask graphs, and the number of maximum joint mask graphs included in the pseudo-label set is less than or equal to the number of the plurality of maximum joint mask graphs. That is, the pseudo-label set may include some or all of the maximum joint mask graphs among the plurality of maximum joint mask graphs. In a specific implementation manner, since the maximum joint mask generated with a lower threshold can contain more supplementary information of the prediction result of the test image, the preset number of maximum joint mask graphs is selected in ascending order of the preset threshold corresponding to the maximum joint mask graph, and the preset number is less than the number of the plurality of maximum joint mask graphs. For example, the plurality of maximum joint mask graphs include p 0,1 , p 1,2 , p 2,3 , p 3,4 and p 4,5 . Select p 0,1 , p 1,2 , p 2,3 to form the pseudo-label set, that is, the pseudo-label set includes p 0,1, p 1,2 , p 2,3 。

[0107] After obtaining the pseudo-label set, as Figure 5 shown, using the Earth Mover’s Distance (EMD) similarity metric, for each predicted probability map, select the most similar one of the maximum joint mask maps in the pseudo-label set Pset as its corresponding pseudo-label p' i , to obtain the pseudo-labels of the test image and each enhanced test image. In this embodiment, selecting the probability map with the smallest distribution distance as the corresponding pseudo-label can, to a large extent, combine its own predicted probability map, thus avoiding the collapse that may be caused by large-scale adjustments during the self-learning process of the model.

[0108] S50. Fine-tune the segmentation model based on each pseudo-label, and continue to perform the step of data augmentation on the test images in the test image set until the fine-tuning end condition is met, to obtain an image segmentation model.

[0109] Specifically, after obtaining each pseudo-label, modify the segmentation model based on the predicted probability maps and pseudo-labels corresponding to the test image and each enhanced test image respectively, to fine-tune the segmentation model. In addition, since the test image is annotation information, it is impossible to determine whether the segmentation model is optimized in the correct direction to produce reasonable results and when to stop fine-tuning the segmentation model. Therefore, this embodiment gives a fine-tuning end condition, where the fine-tuning end condition is that the similarity coefficient D cur is less than the preset coefficient D pre , or the number of fine-tuning times k is equal to the preset number threshold k max . When the similarity coefficient corresponding to the test image is less than the preset coefficient, or the number of fine-tuning times is equal to the preset number threshold, it means that the segmentation model meets the fine-tuning end condition, and the fine-tuning process of the segmentation model is ended.

[0110] Based on this, the step of fine-tuning the segmentation model based on each pseudo-label and continuing to perform the step of data augmentation on the test images in the test image set until the fine-tuning end condition is met to obtain an image segmentation model specifically includes:

[0111] Correct the model parameters of the segmentation model based on each pseudo-label;

[0112] Select the first maximum joint mask map corresponding to the maximum preset threshold and the second maximum joint mask map corresponding to the minimum preset threshold, and calculate the similarity coefficient between the first maximum joint mask map and the second maximum joint mask map;

[0113] Determine whether the segmentation model meets the fine-tuning end condition based on the similarity coefficient and the number of fine-tuning times of the segmentation model, where the fine-tuning end condition is that the similarity coefficient is less than a preset coefficient, or the number of fine-tuning times is equal to a preset number threshold;

[0114] When the fine-tuning end condition is met, use the fine-tuned segmentation model as the image segmentation model;

[0115] When the fine-tuning end condition is not met, continue to perform the step of data augmentation on the test images in the test image set.

[0116] Specifically, the similarity coefficient uses the Dice coefficient. Calculate the shape consistency of the first maximum union mask map corresponding to the maximum preset threshold and the second maximum union mask map corresponding to the minimum preset threshold through the Dice coefficient, and use the calculated shape consistency as the similarity coefficient. When the similarity coefficient between the first maximum union mask map and the second maximum union mask map is less than the preset coefficient, it indicates that the maximum union masks at different confidence levels should tend to be consistent with each other, indicating that the prediction results of the segmentation model for the test image and each enhanced test image corresponding to the test image will be stable, thus indicating that the segmentation model is correctly optimized, and thus the fine-tuning of the segmentation model can be stopped. In addition, adding the number of fine-tuning times equal to the preset number threshold to the fine-tuning end condition can avoid entering an infinite loop during the fine-tuning process.

[0117] In summary, this embodiment provides a training method for an image segmentation model. The method includes performing data augmentation on the test images in the test image set to obtain a number of enhanced test images; determining the prediction probability map of the test image and the prediction probability maps of the enhanced test images through a pre-trained segmentation model, and determining the master probability map based on the determined prediction probability maps; determining a number of maximum union mask maps based on the master probability map; determining the pseudo-labels of the test image and each enhanced test image based on the number of maximum union mask maps; fine-tuning the segmentation model based on each pseudo-label, and continuing to perform the step of data augmentation on the test images in the test image set until the fine-tuning end condition is met to obtain an image segmentation model. This application online adjusts the segmentation model through self-supervised learning, uses the method of test-time data augmentation to generate reliable pseudo-labels to drive the online learning of the segmentation model, and dynamically stops learning according to the consistency degree of the pseudo-labels, thereby improving the model performance and generalization degree of the image segmentation model.

[0118] Based on the above training method of the image segmentation model, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the training method of the image segmentation model as described in the above embodiment.

[0119] Based on the above training method of the image segmentation model, the present application also provides a medical image segmentation method. An image segmentation model trained by using the training method of the image segmentation model provided in the above embodiments is applied. The medical image segmentation method specifically includes:

[0120] Input the medical image to be segmented into the image segmentation model, and determine the segmentation region corresponding to the medical image through the image segmentation model.

[0121] Based on the above training method of the image segmentation model, the present application also provides a terminal device, as Figure 6 shown. It includes at least one processor 20; a display screen 21; and a memory 22. It may also include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiments.

[0122] In addition, when the logical instructions in the above-mentioned memory 22 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0123] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the methods in the above embodiments.

[0124] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes can also be transient storage media.

[0125] In addition, the specific processes of loading and executing multiple instructions by the instruction processor in the above storage medium and terminal device have been described in detail in the above method, and will not be elaborated here one by one.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A training method for an image segmentation model, characterized in that, the training method includes: Performing data augmentation on the test images in the test image set to obtain a number of augmented test images; Determining the prediction probability map of the test image and the prediction probability maps of the augmented test images through a pre-trained segmentation model; Performing inverse augmentation operations on the prediction probability maps corresponding to the augmented test images respectively to obtain the candidate test probability maps corresponding to the augmented test images respectively; Adding the candidate test probability maps and the prediction probability map corresponding to the test image to obtain a master probability map; Selecting a number of preset thresholds, where the number of preset thresholds includes all integers less than the number of prediction probability maps; For each preset threshold, obtaining the maximum difference among the differences between the pixel values of the pixels in the master probability map and the preset threshold, and generating the maximum joint mask map corresponding to the preset threshold based on the minimum value of the maximum difference and 1, so as to obtain a number of maximum joint mask maps; Determining the pseudo-labels of the test image and the augmented test images based on the number of maximum joint mask maps; Fine-tuning the segmentation model based on each pseudo-label, and continuing to perform the step of performing data augmentation on the test images in the test image set until the fine-tuning end condition is met, so as to obtain an image segmentation model.

2. The training method for an image segmentation model according to claim 1, characterized in that, the determining the pseudo-labels of the test image and the augmented test images based on the number of maximum joint mask maps specifically includes: Selecting a preset number of maximum joint mask maps from the number of maximum joint mask maps to form a pseudo-label set; Selecting the pseudo-labels corresponding to the prediction probability maps respectively in the pseudo-label set by using a distance similarity index, so as to obtain the pseudo-labels of the test image and the augmented test images.

3. The training method for an image segmentation model according to claim 1, characterized in that, the fine-tuning the segmentation model based on each pseudo-label, and continuing to perform the step of performing data augmentation on the test images in the test image set until the fine-tuning end condition is met, so as to obtain an image segmentation model specifically includes: Correcting the model parameters of the segmentation model based on each pseudo-label; Selecting the first maximum joint mask map corresponding to the maximum preset threshold and the second maximum joint mask map corresponding to the minimum preset threshold, and calculating the similarity coefficient between the first maximum joint mask map and the second maximum joint mask map; Judging whether the segmentation model meets the fine-tuning end condition based on the similarity coefficient and the fine-tuning times of the segmentation model, where the fine-tuning end condition is that the similarity coefficient is less than a preset coefficient, or the fine-tuning times is equal to a preset number threshold; When the fine-tuning end condition is met, taking the fine-tuned segmentation model as the image segmentation model; When the fine-tuning end condition is not met, continuing to perform the step of performing data augmentation on the test images in the test image set.

4. The training method for an image segmentation model according to any one of claims 1-3, characterized in that, The segmentation model includes at least one style conversion module; the specific steps of determining the prediction probability map of the test image and the prediction probability maps of the augmented test images through the pre-trained segmentation model are as follows: For each reference test image in the test image group composed of the test image and the augmented test images, input the reference test image and the source domain image into the segmentation model respectively, where the source domain image is a training image in a preset training sample set for training the segmentation model; Control the network layer of the segmentation model before the style conversion module to determine the first feature map corresponding to the reference test image and the second feature map corresponding to the source domain image; Control the style conversion module to adjust the image gray level of the first feature map based on the image gray level distribution of the second feature map to obtain the adjusted first feature map; Control the adjusted first feature map and the second feature map to pass through the network layer of the segmentation model after the style conversion module to obtain the prediction probability map of the reference test image, so as to obtain the prediction probability map of the test image and the prediction probability maps of the augmented test images.

5. The training method of the image segmentation model according to claim 4, characterized in that, The specific steps of controlling the style conversion module to adjust the image gray level of the first feature map based on the image gray level distribution of the second feature map to obtain the adjusted first feature map are as follows: For each channel of the first feature map and each channel of the second feature map, determine the sequence number of each pixel value in the channel among all the pixels in the channel, where the sequence number is formed according to the order of pixel values from small to large; For each channel in the first feature map, select the corresponding candidate channel in the second feature map, and select the corresponding candidate pixel for each pixel in the candidate channel based on the sequence number corresponding to each pixel in the channel, and use the pixel value of each candidate pixel to replace the pixel value of the pixel corresponding to each candidate pixel to obtain the adjusted first feature map.

6. The training method of the image segmentation model according to claim 4, characterized in that, The specific steps of controlling the style conversion module to adjust the image gray level of the first feature map based on the image gray level distribution of the second feature map to obtain the adjusted first feature map are as follows: For each channel of the first feature map and each channel of the second feature map, divide the channel into several sliding windows in a sliding window manner, and determine the sequence number of each pixel value in the sliding window among all the pixels in the channel window, where the sequence number is formed according to the order of pixel values from small to large; For each sliding window in the first feature map, select the corresponding candidate sliding window in the second feature map, and select the corresponding candidate pixel for each pixel in the candidate sliding window based on the sequence number corresponding to each pixel in the sliding window, and use the pixel value of each candidate pixel to replace the pixel value of the pixel corresponding to each candidate pixel to obtain the adjusted sliding window corresponding to the sliding window. For each pixel in each channel of the first feature map, obtain the adjusted pixel values of the pixel in each adjusted sliding window including the pixel, and use the average value of all the obtained adjusted pixel values as the pixel value of the pixel, so as to obtain the adjusted first feature map.

7. A computer-readable storage medium, characterized in that the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the training method of the image segmentation model according to any one of claims 1-6.

8. A terminal device, characterized in that it includes: a processor, a memory and a communication bus; the memory stores a computer-readable program executable by the processor; the communication bus realizes the connection and communication between the processor and the memory; when the processor executes the computer-readable program, the steps in the training method of the image segmentation model according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Unsupervised cross-domain self-adaptive medical image segmentation method based on deep adversarial learning

    AU2020103905A4

  • A man-machine cooperative image segmentation and labeling method

    CN109741332A