Model training, image generation and lesion detection method, device and electronic equipment
By using two independent diffusion model branches with a shared time-encoded neural network in the medical detection model, the problem of insufficient training data is solved, realistic lesion images are generated, and the model's generalization performance is improved. This fulfills the requirement for labeled data, generates the lesion detection and segmentation effects of the model, improves the model's generalization performance, generates the lesion detection effect of the model application, and enhances the model's segmentation ability.
Patent Information
- Application Number
- CN202410052654.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-01-12
AI Technical Summary
Existing medical detection models lack a large amount of labeled data during training, resulting in low generalization and poor robustness, especially poor detection performance under new data distributions.
A neural network with two independent diffusion model branches is used for image generation and target segmentation tasks, respectively. It is trained on fully supervised or unsupervised datasets and shares temporal encoding to ensure semantic consistency, generating realistic lesion images and performing target segmentation.
It effectively alleviates the need for labeled data, improves the detection performance of the model, generates images without artifacts and noise, improves segmentation accuracy, and enhances the generalization performance of the model.
Smart Images

Figure CN117852615B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the field of medical image processing, and in particular to a model training, image generation and lesion detection method and device and electronic equipment. BACKGROUND
[0002] At the current stage, medical detection models are often used for computer-aided diagnosis. These medical models complete the positioning and lesion segmentation tasks of specific tissue parts for input medical images. When training the model, a large amount of labeled data is required. If a small amount of labeled data is used for training, the obtained model often has low generalization and poor robustness, and the detection performance of the sample under the new data distribution is not good. However, the lack of a large amount of labeled training samples is a major pain point currently faced by medical detection.
[0003] Therefore, it is urgent to solve the problem of obtaining a large amount of labeled data at low cost, and on the other hand, it is also an urgent need to obtain a detection model with better performance using a small amount of labeled data. SUMMARY
[0004] To solve the problems in the related art, the present application provides a model training, image generation, lesion detection method, device, electronic equipment and medium, which can obtain a large amount of labeled data through a neural network model, and can also perform target detection and segmentation on the image to be detected. Not only can it effectively alleviate the demand for labeled data, but also can improve the performance of target detection, and has significantly improved the effectiveness and reliability of the model.
[0005] One aspect of the present application provides a model training method, comprising:
[0006] obtaining a training data set;
[0007] training a neural network model based on the training data set to update network parameters;
[0008] The neural network model includes two branches of diffusion models that are independent of each other, and are respectively used to implement an image generation task and a target segmentation task.
[0009] According to the method of the present application, the training of the neural network model based on the training data set to update the network parameters comprises:
[0010] If the training data set is labeled, the network parameters are updated based on the loss function of the two branches;
[0011] If the training data set is not labeled, the network parameters are updated based on the loss function of the branch implementing the image generation task.
[0012] According to the method of the application, the two branch diffusion models share time coding.
[0013] Another aspect of the application provides an image generation method, comprising:
[0014] Obtaining image data of a first modality;
[0015] Processing the image data of the first modality by a neural network model to generate image data of a second modality; wherein,
[0016] If the image data of the first modality is a contour of a first target and a contour of a second target, the generated image data of the second modality is an image containing the first target and the second target, and the second target is located in an internal region of the first target;
[0017] If the image data of the first modality is an image containing the first target and the second target, and the second target is located in an internal region of the first target, the generated image data of the second modality is a segmented contour of the first target and a contour of the second target;
[0018] And the neural network model is obtained by the model training method as described above.
[0019] According to the method of the application, the method further comprises:
[0020] If the image data of the first modality is a contour of the first target, the image data of the second modality is an image containing the first target;
[0021] If the image data of the first modality is an image containing the first target, the image data of the second modality is a segmented contour of the first target.
[0022] Another aspect of the application provides a chest digital image generation method, comprising:
[0023] Obtaining a contour of a lung region and a contour of a lesion region;
[0024] Processing the contour of the lung region and the contour of the lesion region by a neural network model to generate a chest radiograph image of a patient with the lesion;
[0025] Wherein, the neural network model is obtained by the model training method as described above.
[0026] Another aspect of the application provides a lesion detection method based on a chest digital image, comprising:
[0027] Obtaining a chest digital image;
[0028] processing the chest digital image through a neural network model to segment a lung region and a lesion region from the chest digital image;
[0029] The neural network model is obtained through the model training method as described above.
[0030] Another aspect of the present application also provides a model training device, comprising:
[0031] A training data acquisition module configured to acquire a training data set;
[0032] A parameter updating module configured to train a neural network model based on the training data set to update network parameters.
[0033] The neural network model comprises two branches of diffusion models independent of each other, respectively used to implement an image generation task and a target segmentation task.
[0034] Another aspect of the present application also provides an image generation device, comprising:
[0035] An input data acquisition module configured to acquire image data of a first modality;
[0036] An image generation module configured to process the image data of the first modality through a neural network model to generate image data of a second modality.
[0037] If the image data of the first modality is an outline of a first target and an outline of a second target, the generated image data of the second modality is an image containing the first target and the second target, and the second target is located in an internal region of the first target.
[0038] If the image data of the first modality is an image containing the first target and the second target, and the second target is located in an internal region of the first target, the generated image data of the second modality is the outline of the first target and the outline of the second target.
[0039] The neural network model is obtained through the model training method as described above.
[0040] Another aspect of the present application also provides an electronic device, comprising at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the processor to implement the method as described above.
[0041] Another aspect of the present application also provides a computer readable storage medium having stored thereon computer readable instructions which, when executed by a processor, cause the processor to implement the method as described above.
[0042] Another aspect of the present application also provides a computer program which, when executed by a processor, cause the processor to implement the method as described above.
[0043] The technical scheme of the present application brings the following beneficial effects:
[0044] 1. The neural network model constructed by using two branch independent diffusion models can respectively complete the tasks of image generation and target segmentation, and the neural network model can complete training by using a supervised data set or an unsupervised data set, thereby reducing the requirement for training samples.
[0045] 2. When the model is applied in the medical detection field, a specific human tissue region mask and a lesion region mask can be used as input to generate a realistic image with a lesion, the digital image of the specific human tissue can be expanded to increase the number of samples, and the model is convenient to use in the training process of other detection models, thereby enhancing the generalization performance of these models.
[0046] 3. The model can also use the digital image of the specific human tissue as input to segment the corresponding tissue region and lesion region therefrom, thereby realizing the function of lesion detection.
[0047] 4. In the practice of applying the model to the chest X-ray image generation task, compared with the generative adversarial network (GAN), the image generated by the model has no obvious artifacts and noise, and is closer to the real chest X-ray image.
[0048] 5. In the practice of applying the model to the lung tuberculosis lesion segmentation task based on the chest X-ray image, compared with the traditional convolution-based and Transformer-based image segmentation method, the model can obtain more accurate segmentation results. According to experimental verification, the average accuracy index dice of the model for lesion segmentation reaches 86.5%, which is 3 percentage points higher than the 83.4% of the traditional Unet model. BRIEF DESCRIPTION OF DRAWINGS
[0049] Other features, objects and advantages of the present application will become more apparent from the following detailed description of the non-limiting embodiments, taken in conjunction with the accompanying drawings.
[0050] Figure 1 A system architecture schematic diagram to which the method of the embodiment of the present application is applied is schematically shown;
[0051] Figure 2A flowchart of a model training method according to an embodiment of the present application is schematically shown;
[0052] Figure 3 A structural schematic diagram of a neural network model in an embodiment of the present application is schematically shown;
[0053] Figure 4 A flowchart of an image generation method according to an embodiment of the present application is schematically shown;
[0054] Figure 5 A flowchart of a chest digital image generation method according to an embodiment of the present application is schematically shown;
[0055] Figure 6 Experimental results of a chest radiograph image generated by applying the method of an embodiment of the present application;
[0056] Figure 7 A flowchart of a detection method based on a chest digital image according to an embodiment of the present application is schematically shown;
[0057] Figure 8 A block diagram of a model training device according to an embodiment of the present application is schematically shown;
[0058] Figure 9 A block diagram of an image generation device according to an embodiment of the present application is schematically shown;
[0059] Figure 10 A schematic diagram of a processing procedure for detecting a chest DR image according to the prior art;
[0060] Figure 11 A structural schematic diagram of a computer system suitable for implementing the model training method and the image generation method of an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0061] Hereinafter, exemplary embodiments of the present application will be described in detail with reference to the accompanying drawings so as to make them more readily understood by those skilled in the art. Further, portions irrelevant to the description of the exemplary embodiments have been omitted in the accompanying drawings for the sake of clarity.
[0062] In the present application, it should be understood that terms such as "include" or "have" are intended to indicate that there are features, numbers, steps, actions, components, parts or combinations thereof disclosed in the specification, and do not exclude the possibility that one or more other features, numbers, steps, actions, components, parts or combinations thereof exist or are added.
[0063] It should also be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0064] It should be noted that the acquisition or display of data in the present application is authorized, confirmed or actively selected by the user.
[0065] At the present stage, the combination of computer-aided diagnosis target detection model usually uses detection model (such as MaskRCNN) and classification model (such as ResNet50+MLP) for processing. By using a pre-labeled data set to train a neural network, the target region in the digital image is located and numbered. There is also a way of using transfer learning to pre-train a model (such as ImageNet) on a large-scale image data set. After fine-tuning the model, it is applied to the detection task of a specific target. In some methods, the diversity of the detection target at different scales is also considered, and multi-scale detection and segmentation technology is used, which can improve the detection and positioning ability of the model for multi-scale targets. For example, for chest DR (Digital Radiography) images, the positioning and segmentation of TB (Tuberculosis) in the lungs are one of the detection tasks.
[0066] Taking a chest DR image as an example, Figure 10 The processing process of the prior art for the chest DR image to be detected is shown.
[0067] These existing detection methods all rely on models trained using a large amount of labeled data. If a small amount of labeled data is used, the model trained will often have low generalization and poor robustness, especially for images in new data distribution. However, a large amount of labeled data is difficult to obtain in some fields. For example, in the medical field, medical images need to be labeled by doctors with professional experience, but the time of doctors is uncontrollable, and it is a big pain to obtain a large amount of labeled training data.
[0068] To solve these problems, the inventors propose a neural network model that can generate sample images for training and complete target detection from input images. The same model can meet both needs. When the model is applied in the medical detection field, it can well solve the problem of difficulty in obtaining a large amount of labeled data, and can also complete the detection task of the lesion area at the same time. For example, in chest DR image detection, using the neural network model of the present application, realistic chest X-ray images with TB and masks corresponding to the lesion can be generated, and the TB area can also be segmented from the input chest X-ray image, which can complete both sample generation and lesion segmentation tasks.
[0069] The technical solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0070] Figure 1A schematic diagram of a system architecture of a model training, image generation and lesion detection method according to an embodiment of the present application is shown.
[0071] As shown in Figure 1 The system architecture 100 can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or fiber optic cables, etc.
[0072] The terminal devices 101, 102, 103 interact with the server 105 through the network 104 to receive or send messages, etc. Various client applications can be installed on the terminal devices 101, 102, 103. For example, a special application program with functions such as medical image display, lesion display and editing, report generation, etc.
[0073] The terminal devices 101, 102, 103 can be hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be various special-purpose or general-purpose electronic devices, including but not limited to smartphones, tablet computers, laptop computers and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they can be installed in the above-mentioned electronic devices. They can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or as a single software or software module.
[0074] The server 105 can be a server that provides various services, such as a backend server that provides services for client applications installed on the terminal devices 101, 102, 103. For example, the server can train and apply an image generation model to implement a multi-modal image conversion function to display the final result on the terminal devices 101, 102, 103.
[0075] The server 105 can be hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server 105 is software, it can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or as a single software or software module.
[0076] The model training, image generation and detection method provided by the embodiments of the present application can be executed by the server 105 or the terminal devices 101, 102, 103. Alternatively, the various methods of the embodiments of the present application can be partially executed by the terminal devices 101, 102, 103 and partially executed by the server 105.
[0077] It should be understood that Figure 1 the number of terminal devices, networks and servers in the above-mentioned system is only illustrative. Any number of terminal devices, networks and servers can be provided according to the implementation needs.
[0078] The following describes an embodiment of a model training method for implementing the present application. Figure 2
[0079] Figure 2 A flowchart of the model training method according to an embodiment of the present application is schematically shown.
[0080] As Figure 2 shown, the model training method includes operations S210-S240.
[0081] In operation S210, a training data set is obtained.
[0082] In operation S220, a neural network model is trained based on the training data set to update network parameters, wherein the neural network model includes two branches of diffusion models independent of each other, respectively used to implement an image generation task and a target segmentation task.
[0083] First, the neural network model in the embodiment of the present application is described. The structure of the neural network model includes two branches of diffusion models. A diffusion model is a model that learns the distribution of training data and then outputs pictures similar to the distribution of training data. Common diffusion generative models include DDPM (Denoising Diffusion Probabilistic Models) and DDIM (Denoising Diffusion Implicit Models). The core architecture of the diffusion model is a UNet model. The basic idea is to gradually add noise to an image at a large time step, and to reconstruct the UNet model by adding time encoding in the process. The diffusion model includes two processes: a forward diffusion process and a reverse generation process. The forward diffusion process is a process of gradually adding Gaussian noise to an image until it becomes random noise. The reverse generation process is a denoising process that starts from a random noise and gradually denoises until an image is generated.
[0084] The following describes the structure of the neural network model in the embodiment of the present application, taking a chest X-ray image and a tuberculosis detection task as an example. Figure 3
[0085] Figure 3 The structure of the neural network model in the embodiment of the present application is schematically shown.
[0086] As Figure 3 As shown, this neural network model includes two branches. The upper branch generates a chest X-ray image, while the lower branch segments the lung field mask and the tuberculosis lesion mask. The upper branch adds noise by converting the chest X-ray image to Gaussian noise and removes noise by converting the Gaussian noise back to the chest X-ray image. The lower branch adds noise by converting the two-channel image (stitched from the lung field mask and lesion mask) to Gaussian noise and removes noise by converting the Gaussian noise back to the lung field mask and lesion mask.
[0087] According to an embodiment of the present invention, in this neural network model, the diffusion models of the upper and lower branches share temporal encoding, that is, the noise added at the same time step during the forward diffusion process of the two branches is the same, to ensure semantic consistency of image generation and target segmentation tasks. Furthermore, the diffusion models of the upper and lower branches may also have the same initialization noise and noise scheduling.
[0088] like Figure 3 As shown, for an input sample x, the sample with added noise at time t is represented as x t x t-1 It is a sample from the previous time step, β t It is the variance of the added noise, which follows a normal distribution. In addition, there is a noise scheduler in the diffusion model, which is used to control the parameters of the noise added during the forward diffusion process.
[0089] The diffusion process involves gradually adding Gaussian noise to the data until it becomes random noise. For the original data, the diffusion process consists of T steps, each step adding Gaussian noise to the data obtained in the previous step as follows:
[0090]
[0091]
[0092] noise β t ∈(0,1) normal distribution.
[0093] When the diffusion step size T is large enough, the intermediate process variable x can be considered as adding noise. t It also follows a Gaussian distribution.
[0094] When a posterior distribution q(x) t-1 |x t When x is known, use t The distribution follows a standard Gaussian distribution N(0,1), which yields q(x0). Therefore, the reverse process can be expressed as:
[0095]
[0096]
[0097] p θ is the estimation of q corresponding to each time step t, and p and q constitute the components of the variational autoencoder in the diffusion model. Correspondingly, p θ is responsible for generating chest X-ray images, is responsible for generating lung field mask and pulmonary tuberculosis lesion mask.
[0098] The final loss function L can be expressed as the negative lower bound of the log-likelihood, and is formally rewritten as the sum of the loss of each step.
[0099] L v1b : = L0 +... + L t-1 +... + L T
[0100] where
[0101] L0: = -logp θ (x0|x1)
[0102] L t-1 : = D KL (q(x t-1 |x t , x0) || p θ (x t-1 |x t ))
[0103] L T : = D KL (q(x T |x0) || p(x T ))
[0104] The neural network model in the embodiment of the application, the diffusion models of the two branches are independent of each other, that is, the diffusion models of the two branches have independent model parameters according to the tasks implemented; the noise added by the diffusion models of the two branches at different time steps is shared by the two branches, that is, the diffusion models of the two branches add the same noise at the same time step in the diffusion process.
[0105] Further, the diffusion model in the branch can select the image generation branch when initializing the noise. Since the diffusion models of the two branches share the time encoding, not only is the model training simplified, but also the semantic consistency of the two branch tasks is ensured.
[0106] The neural network model in the embodiment of the application can use a fully supervised data set and / or an unsupervised data set to complete training. If the training data set is labeled, the network parameters are updated based on the loss function of the two branches. If the training data set is not labeled, the network parameters are updated based on the loss function of the image generation task branch.
[0107] Specifically, the data in the full-supervised data set is labeled, for example, the chest X-ray image contains both the lung field mask and the lesion mask, when training the model using this part of data, both upper and lower branches can provide loss functions for updating the parameters of the entire model. The data in the unsupervised data set is not labeled, for example, only contains chest X-ray images, and does not contain the corresponding lung field mask and lesion region mask, when training the model using this part of data, the loss function provided by the branch (the upper branch in the figure) that completes the image generation is used to update the model parameters. Figure 3
[0108] The neural network model in the embodiment of the application can also use the full-supervised data set and the unsupervised data set to complete the training, and during the training, the parameters of the network model are updated according to different training data.
[0109] The model trained by this method can synthesize a large number of semantically consistent paired images in the application, for example, obtain images containing pulmonary tuberculosis and pulmonary tuberculosis lesion region data, which can expand the training samples and enhance the robustness of the trained model when training other detection models.
[0110] Based on the same inventive concept, the application also provides an image generation method, which will be described below with reference to Figure 4 .
[0111] Figure 4 The image generation method of the embodiment of the application is schematically shown. The method comprises operations S410-S420.
[0112] In operation S410, image data of a first modality is acquired.
[0113] In operation S420, the image data of the first modality is processed by a neural network model to generate image data of a second modality; if the image data of the first modality is the contour of a first target and the contour of a second target, the generated image data of the second modality is an image containing the first target and the second target, and the second target is located in the internal region of the first target; if the image data of the first modality is an image containing the first target and the second target, and the second target is located in the internal region of the first target, the generated image data of the second modality is the contour of the first target and the contour of the second target.
[0114] According to an embodiment of the present application, the neural network model is obtained by the foregoing model training method. The neural network model includes two branches of diffusion generative models that are independent of each other, and are respectively used to implement an image generation task and a target segmentation task. The specific construction and training method of the neural network model can refer to the foregoing model training method.
[0115] According to an embodiment of the present application, a modality represents a form of an image, for example, a mask of a target object and a real or near-real image containing the target object are two modalities, and for another example, images taken by different devices are also images of different modalities.
[0116] By the image generation method according to an embodiment of the present application, the image generation and target segmentation tasks suitable for various scenarios can be completed, for example, inputting a lung region mask and a pulmonary tuberculosis lesion mask to generate an abnormal chest film of a patient with pulmonary tuberculosis, and vice versa. When a large number of images with labeled data are needed to train other detection models, the method according to the present application can meet this demand.
[0117] Further, if the image data of the first modality is a contour of a first target, the image data of the second modality is an image containing the first target; if the image data of the first modality is an image containing the first target, the image data of the second modality is a segmented contour of the first target. In this way, simpler image generation and target segmentation tasks can be completed, for example, inputting a lung region mask to generate a normal chest film of a healthy person, and vice versa. In some application scenarios, a large number of normal images are also needed to train a model, and the method according to the present application can meet this demand.
[0118] According to an embodiment of the present application, the pre-trained neural network model is applied to realize pairing of multi-modal images, thereby meeting the demand for images of different modalities in reality.
[0119] Taking medical detection as an example, the image data of the first modality can be a mask of a lesion, and the corresponding image data of the second modality is an image of a patient with the lesion. For example, assuming that the first modality is a pulmonary tuberculosis lesion mask and a lung field mask, the second modality is a chest X-ray image of a patient with pulmonary tuberculosis, and assuming that the first modality is a lung field mask, the second modality is a chest X-ray image of a healthy person. For another example, assuming that the first modality is a breast nodule lesion mask, the second modality is a mammography image of a patient with a breast nodule. In this way, the model can be used to expand the required data samples at low cost, including abnormal images of patients with diseases and normal images of healthy persons.
[0120] Still taking medical detection as an example, the image data of the first modality can be a medical image to be detected, and the image data of the corresponding second modality is a lesion mask detected and segmented from the medical image. For example, assuming that the first modality is a chest X-ray, the second modality can be a detected tuberculosis lesion mask and a lung field mask. If the patient is healthy, only the lung field mask is present. For another example, assuming that the first modality is a mammogram of a patient with a breast nodule, the second modality is a detected breast mask and a breast nodule lesion mask. In this way, the model can be used to complete the detection and segmentation of the target object. Moreover, because the model guarantees semantic consistency between the image generation and the target segmentation task through the shared time encoding during training, the segmentation accuracy of the model can also be improved. Experimental verification shows that the model can provide better matching accuracy compared with traditional convolution-based and Transformer-based image segmentation methods.
[0121] It should be noted that the neural network model, the training method thereof, and the image generation method provided by the present application can also be applied in other image processing fields, such as object detection and pattern detection. The above medical detection examples do not limit the application of the present application in other fields.
[0122] The following describes the embodiments of the present application in the medical detection field. Figure 5 and Figure 6 The following describes the embodiments of the present application in the medical detection field.
[0123] Figure 5 A flowchart of a chest digital image generation method according to an embodiment of the present application is schematically shown. The method includes operations S510-S520.
[0124] In operation S510, the contour of the lung region and the contour of the lesion region are obtained.
[0125] In operation S520, the contour of the lung region and the contour of the lesion region are processed by a neural network model to generate a chest X-ray image with the lesion, wherein the neural network model includes two branches of independent diffusion models respectively used to implement a chest X-ray image generation task and a lesion segmentation task, and is obtained through the above-mentioned model training method.
[0126] According to an embodiment of the present application, the contour mask of the lung region and the contour mask of the lesion region can come from existing other medical data, such as from a CT image, or from an artificially drawn contour. The generated chest X-ray image can be a chest X-ray image with the lesion. It can be understood that when only the mask of the lung region is input, the chest X-ray image generated will not contain the lesion, so that a normal chest X-ray image of a healthy person can be generated. The lung lesion can be tuberculosis, lung tumor, or lung disease.
[0127] Figure 6 The experimental results of the chest radiograph images generated by applying the method of the embodiments of the present application are given. The top row is the real chest radiograph, the bottom row is the chest radiograph generated by the method of the embodiments of the present application, and the middle row is the chest radiograph obtained by applying the generative adversarial network (GAN). It can be seen that the chest radiograph images obtained by applying the method of the embodiments of the present application have no obvious artifacts and noise, and the generated images are closer to the real chest X-ray images.
[0128] Figure 7 A flowchart of the lesion detection method based on chest digital images according to the embodiments of the present application is schematically shown. The method includes operations S710-S720.
[0129] In operation S710, a chest digital image is acquired.
[0130] In operation S720, the chest digital image is processed by a neural network model to segment a lung region and a lesion region therefrom; wherein the neural network model includes two branches of diffusion generative models independent of each other, respectively used to implement an image generation task and a lesion segmentation task, and the neural network model is obtained by the aforementioned model training method.
[0131] According to the embodiments of the present application, the same neural network model is applied to perform lesion detection on the input chest radiograph. Through experimental verification, the method of the embodiments of the present application can obtain more accurate segmentation results compared with the traditional convolution-based and Transformer-based image segmentation methods. The experimental comparison results are shown in the following table.
[0132] Model name Evaluation index Dice Unet 0.834 DeepLabV3 0.847 TransUnet 0.851 The present application 0.865
[0133] It can be seen that the embodiments of the present application adopt two independent diffusion model branches, and during training, the shared temporal coding is used to give semantic consistency to the image generation task and the image segmentation task, which improves the segmentation accuracy of the model. This improvement effect is not possessed by the traditional image segmentation methods and the single diffusion model.
[0134] Based on the same inventive concept, the present application also provides a model training device and an image generation device, which will be described below with reference to Figure 8 and Figure 9 .
[0135] Figure 8 A block diagram of the model training device 800 of the embodiments of the present application is schematically shown. The device 800 can be realized by software, hardware or a combination of both as part or all of an electronic device.
[0136] As Figure 8As shown, the model training apparatus 800 comprises a training data acquisition module 810 and a parameter updating module 820. The model training apparatus 800 can perform the model training method described above.
[0137] The training data acquisition module 810 is configured to acquire a training data set.
[0138] The parameter updating module 820 is configured to train a neural network model based on the training data set to update network parameters; wherein the neural network model comprises two branches of diffusion generative models independent of each other, respectively used to implement an image generation task and a target segmentation task. The diffusion models of the two branches have the same initialization noise and noise scheduling, and share the same temporal coding.
[0139] During training, if the training data set is labeled, the network parameters are updated based on the loss functions of the two branches; if the training data set is not labeled, the network parameters are updated based on the loss function of the branch implementing the image generation task.
[0140] According to the embodiment of the present application, the above-mentioned model training apparatus 800 can be used to train a fully supervised and / or unsupervised data set, and the obtained model can complete the tasks of image generation and target segmentation.
[0141] Figure 9 A block diagram of an image generation apparatus 900 according to an embodiment of the present application is schematically shown. The apparatus 900 can be realized by software, hardware or a combination of both to become part or all of an electronic device.
[0142] As shown, the image generation apparatus 900 comprises an input data acquisition module 910 and an image generation module 920. The image generation apparatus 900 can perform the various image generation methods described above. Figure 9
[0143] The input data acquisition module 910 is configured to acquire image data of a first modality.
[0144] The image generation module 920 is configured to process the image data of the first modality by a neural network model to generate image data of a second modality; if the image data of the first modality is the contour of a first target and the contour of a second target, the generated image data of the second modality is an image containing the first target and the second target, and the second target is located in the internal region of the first target; if the image data of the first modality is an image containing the first target and the second target, and the second target is located in the internal region of the first target, the generated image data of the second modality is the contour of the first target and the contour of the second target segmented out; wherein the neural network model is obtained by the aforementioned model training method.
[0145] According to the embodiments of the present application, by inputting images of different modalities, images of the required modality are generated correspondingly, so that the same neural network model can meet the requirements of generating a large number of labeled images or detecting target objects from the images to be detected.
[0146] Figure 11 The structural schematic diagram of a computer system suitable for implementing the model training method and the image generation method according to the embodiments of the present application is shown schematically.
[0147] As shown in Figure 11 The computer system 1100 includes a processing unit 1101, which can execute various processes in the above embodiments according to programs stored in a read-only memory (ROM) 1102 or loaded into a random access memory (RAM) 1103 from a storage portion 1108. Various programs and data required for the operation of the system 1100 are also stored in the RAM 1103. The processing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0148] The following components are connected to the I / O interface 1105: an input portion 1106 including a keyboard, a mouse, and the like; an output portion 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage portion 1108 including a hard disk, and the like; and a communication portion 1109 including a network interface card such as a LAN card, a modem, and the like. The communication portion 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as needed. A removable medium 1111 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is mounted on the drive 1110 as needed, so that a computer program read therefrom is installed in the storage portion 1108 as needed. Among them, the processing unit 1101 can be implemented as a CPU, a GPU, a TPU, a FPGA, an NPU, and the like.
[0149] In particular, the method described above can be implemented as a computer software program according to the embodiments of the present application. For example, the embodiments of the present application include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program containing program code for executing the above method. In such embodiments, the computer program can be downloaded and installed from a network by the communication portion 1109, and / or installed from the removable medium 1111.
[0150] The computer program product of the present application can be a computer program embodied on a non-transitory computer readable medium. When the program is executed by a computer, it can achieve the above-described functions of the present application. The computer program can be stored and / or distributed on media, such as CD-ROM, DVD, USB, flash memory, RAM, ROM, etc. The computer program can also be distributed online, through a network, such as the Internet, Intranet, Extranet, etc. The computer program can exist in a whole program or any one of parts as independent codes, with remote procedures, functions, objects and / or data structures that can use, communicate or operate between different computers via the network. The computer program can be in source code, object code, encrypted code, etc.
[0151] The units or modules described in the embodiments of the present application can be implemented by software, or by programmable hardware. The described units or modules can also be arranged in a processor, and the name of the unit or module does not constitute a limitation on the unit or module itself in some cases.
[0152] As another aspect, the present application also provides a computer readable storage medium, which can be the computer readable storage medium contained in the electronic device or computer system in the above embodiments, or can exist separately and not be assembled into the device. The computer readable storage medium stores one or more programs, which are used by one or more processors to execute the method of the embodiments of the present application.
[0153] The above description is merely preferred embodiments of the present application and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application described in the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features can be replaced with technical features disclosed in the present application (but not limited to) having similar functions to form technical solutions.
Claims
1. A model training method, characterized in that, The method comprises: obtaining a training data set, wherein the training data set comprises a fully supervised data set and / or an unsupervised data set; training a neural network model based on the training data set to update network parameters; wherein the neural network model comprises two branches of diffusion models independent of each other, respectively used to implement an image generation task and a target segmentation task, the diffusion models of the two branches have the same initial noise and noise scheduling, and share the same time coding; and the neural network model synthesizes a pair of images with consistent semantics. The training of the neural network model to update the network parameters comprises: if the training data set is labeled, updating the network parameters based on a loss function of the two branches; and if the training data set is not labeled, updating the network parameters based on a loss function of the branch implementing the image generation task.
2. An image generation method characterized by, The method comprises: obtaining image data of a first modality; processing the image data of the first modality through a neural network model to generate image data of a second modality; wherein if the image data of the first modality is an outline of a first target and an outline of a second target, the generated image data of the second modality is an image containing the first target and the second target, and the second target is located in an internal region of the first target; if the image data of the first modality is an image containing the first target and the second target, and the second target is located in an internal region of the first target, the generated image data of the second modality is the outline of the first target and the outline of the second target segmented out; and the neural network model is obtained by the model training method of claim 1.
3. The method of claim 2, wherein, The method further comprises: if the image data of the first modality is the outline of the first target, the image data of the second modality is an image containing the first target; if the image data of the first modality is an image containing the first target, the image data of the second modality is the outline of the first target segmented out.
4. A method of generating a digital image of a breast, characterized in that, The method comprises: obtaining an outline of a lung region and an outline of a lesion region; processing the outline of the lung region and the outline of the lesion region through a neural network model to generate a chest radiograph image with the lesion; wherein the neural network model is obtained by the model training method of claim 1.
5. A method of lesion detection based on chest digitized images, characterized by, The method comprises: obtaining a chest digital image; processing the chest digital image through a neural network model to segment out a lung region and a lesion region therefrom; wherein the neural network model is obtained by the model training method of claim 1.
6. A model training apparatus characterized by comprising: The method comprises: a training data acquisition module configured to obtain a training data set, wherein the training data set comprises a fully supervised data set and / or an unsupervised data set; a parameter updating module configured to train a neural network model based on the training data set to update network parameters; The neural network model includes two branches of diffusion models independent of each other, respectively used for implementing an image generation task and a target segmentation task, the diffusion models of the two branches have the same initial noise and noise scheduling, and share the same time coding; the neural network model synthesizes pairs of images with consistent semantics; the training of the neural network model to update network parameters includes: if the training data set is labeled, updating network parameters based on loss functions of the two branches; if the training data set is unlabeled, updating network parameters based on a loss function of the branch implementing the image generation task.
7. An image generation apparatus characterized by comprising: Comprising: an input data acquisition module configured to acquire image data of a first modality; an image generation module configured to generate image data of a second modality by processing the image data of the first modality through a neural network model; if the image data of the first modality is an outline of a first target and an outline of a second target, the generated image data of the second modality is an image containing the first target and the second target, and the second target is located in an internal region of the first target; if the image data of the first modality is an image containing the first target and the second target, and the second target is located in an internal region of the first target, the generated image data of the second modality is the outline of the first target and the outline of the second target segmented out; wherein the neural network model is obtained through the model training method of claim 1.
8. An electronic device, comprising: Comprising: at least one processor; and, a memory in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.
Citation Information
Patent Citations
Cervical vertebra MRI image self-supervised segmentation method using diffusion model to generate data
CN117036386A