Fetal face completion method, training method, device, equipment, medium and product
By combining the generative adversarial network with the stable diffusion generation mode, the details distortion and feature mismatch in fetal facial image completion are solved, and high-precision fetal facial image completion is achieved, which helps doctors to make accurate diagnosis.
Patent Information
- Application Number
- CN202510310492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
Existing image processing methods are difficult to accurately capture the complex features and structures of the fetal face. The completed image has distortion of detail and mismatch in features, which cannot meet the requirements of auxiliary diagnosis.
Generative adversarial network technology is adopted, combined with the stable diffusion generation mode, and through the game between the generator and the discriminator, a complete fetal facial image that is natural and in line with medical common sense is generated. The loss calculation is used by a three-dimensional segmenter, and the parameters of the generator and discriminator are optimized to achieve high-precision completion.
The generated fetal facial image has high detail reduction and good feature matching, which meets the requirements of auxiliary diagnosis, simplifies the model and reduces calculation overhead, and improves the accuracy of fetal ultrasound examination.
Smart Images

Figure CN120259106A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical assistance technologies, and particularly to a method for fetal face completion, a training method, a device, equipment, a medium, and a product. Background Art
[0002] Currently, when doctors use ultrasound instruments to examine a fetus in a pregnant woman's abdomen, there may be missing parts of the fetal face, which poses a great challenge to the doctors' diagnostic work. Due to the principle of ultrasound imaging and the variable positions of the fetus in the mother's body, some facial structures may not be clearly presented. If key parts such as eyes, nose, and mouth are blocked or the imaging is blurred, it is difficult for doctors to comprehensively evaluate whether the fetal facial development is normal based on this.
[0003] Traditional image processing methods, such as simple image enhancement and restoration techniques for facial feature completion, are inadequate when dealing with complex medical images. They cannot accurately capture the complex features and structures of the fetal face, and the completed images often have problems such as detail distortion and feature mismatch, making it difficult to meet the requirements of auxiliary diagnosis. Summary of the Invention
[0004] This application proposes a method for fetal face completion, a training method, a device, equipment, a medium, and a product, which can solve one of the problems in the background art.
[0005] To achieve the above object, this application adopts the following technical solutions:
[0006] In a first aspect, a training method for a fetal face completion model is provided. The method includes:
[0007] Obtaining training data, where the training data includes: real images and images to be completed; and
[0008] Using the training data to train the model,
[0009] The model uses a generative adversarial network, and the generative adversarial network includes: a generator for completing the image to be completed, and a discriminator for discriminating between real images or generated images. The generator adopts the noise addition process in the stable diffusion generation mode.
[0010] Based on the above technical solution, using generative adversarial network technology, GAN consists of a generator and a discriminator, and the two play against each other. Among them, the generator can complete the fetal ultrasound image with missing parts of the face into a complete image that looks natural and conforms to medical common sense. In this way, the complex features and structures of the fetal face can be accurately captured, and the completed image has a high degree of detail restoration and feature matching, meeting the requirements of auxiliary diagnosis. And only the noise addition process in the stable diffusion generation mode is adopted, without the need to reverse the denoising, thus simplifying the model and reducing the computational overhead.
[0011] In a possible design of the first aspect, the generative adversarial network further includes: a three-dimensional segmenter for segmenting the generated image of the generator, and the intersection over union (IoU) obtained by processing of the three-dimensional segmenter is used for the loss calculation of the discriminator.
[0012] In a possible design of the first aspect, the generator includes: a three-dimensional Gaussian noise generator, a first convolutional network for feature extraction of the image to be completed, and a second convolutional network for repeatedly superimposing three-dimensional Gaussian noise on the extracted feature map.
[0013] In a possible design of the first aspect, the generative adversarial network updates parameters by optimizing an objective function, the objective function adopts a loss function, and the loss function includes: a generator loss, a discriminator loss, and a three-dimensional segmenter loss.
[0014] In a second aspect, a fetal face completion method is provided, and the completion method includes:
[0015] Obtaining an image to be processed; and
[0016] Using the trained model as described above to process the image to be processed to obtain a fetal face completion result.
[0017] In a third aspect, a training device for a fetal face completion model is provided, and the training device includes:
[0018] A first acquisition unit for acquiring training data, where the training data includes: real images and images to be completed; and
[0019] A training unit for training the model using the training data,
[0020] The model adopts a generative adversarial network, and the generative adversarial network includes: a generator for completing the image to be completed, and a discriminator for discriminating real images or generated images, and the generator adopts a noise addition process in a stable diffusion generation mode.
[0021] In a fourth aspect, a fetal face completion device is provided, and the completion device includes:
[0022] A second acquisition unit for acquiring an image to be processed; and
[0023] A processing unit for using the trained model as described above to process the image to be processed to obtain a fetal face completion result.
[0024] Fifth aspect, there is provided an electronic device, the electronic device comprising: a processor, and a memory coupled to the processor, the memory for storing a computer program; the processor for executing the computer program stored in the memory so that the electronic device executes the training method according to any possible implementation manner in the first aspect, or executes the completion method according to the second aspect.
[0025] Sixth aspect, there is provided a computer-readable storage medium, comprising a computer program or instruction, when the computer program or instruction runs on a computer, causing the computer to execute the training method according to any possible implementation manner in the first aspect, or execute the completion method according to the second aspect.
[0026] Seventh aspect, there is provided a computer program product, comprising: a computer program or instruction, when the computer program or instruction runs on a computer, causing the computer to execute the training method according to any possible implementation manner in the first aspect, or execute the completion method according to the second aspect. Description of the Drawings
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a schematic diagram of a fetal face completion model provided by an embodiment of the present application;
[0029] Figure 2 It is a schematic diagram of the Generator provided by an embodiment of the present application. Detailed Embodiments
[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0031] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description and claims of the specification and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application pertains. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.
[0033] Embodiments of this application combine the diffusion model and the generative adversarial network (GAN) technology to obtain a fetal face completion model, opening up a brand-new path for completing missing fetal face information. The diffusion model destroys the image by gradually adding noise, and this characteristic has a unique complementarity with the adversarial generation mechanism of GAN. The specific structure and process are as Figure 1 shown:
[0034] 1. "Fake data" (data to be completed) is input into the "generator", and after processing, "generated data" is obtained.
[0035] 2. The "generated data" and "true data" are sent into the "discriminator" together, and the discriminator makes a judgment and outputs the result.
[0036] 3. According to the output result of the discriminator, the "generator loss" and the "discriminator loss" are calculated respectively.
[0037] 4. The "generated data" is input into the "3D U-Net" (3D segmenter, which is a semantic segmentation algorithm in the figure) network to obtain "label data" (inference label).
[0038] The mean squared error (MSE) is calculated for the "label data" and the "true label", thereby obtaining the "label loss".
[0039] The 3D segmenter performs semantic segmentation on the generated data. The intersection over union (IOU) value is calculated between the obtained inference label and the true label, and finally the calculated IOU value is used in the loss calculation of the discriminator.
[0040] 5. Calculate the total loss value, that is, "total loss" = "generator loss" + "discriminator loss" + "label loss".
[0041] 6. Update the parameters of the generator and the discriminator according to the calculated total loss.
[0042] After multiple rounds of parameter iterative updates, the model reaches a convergence state, and finally an available generation model is obtained.
[0043] Among them, the specific content in the Generator is as follows Figure 2 shown:
[0044] The essence of the generator is a forward stable diffusion process:
[0045] 1. The fake data passes through a 3D U-net network to obtain generated data1 (the first data).
[0046] 2. The 3D Gaussian noise passes through a 3D U-net network to obtain generated data2 (the second data).
[0047] 3. The generated data1 and generated data2 are fed into the image conditional unet to obtain an output of 64x64x64 (feature map).
[0048] 4. The whole process is repeated n times to obtain the final generated data. Among them, the generated data1 is only used once, and the generated data2 is used N times;
[0049] In practical applications, first use the Diffusion model to process a large number of complete fetal facial ultrasound images, add noise in the forward direction to simulate interference to master the feature change law, and lay a foundation for facial completion; when processing fetal facial ultrasound images with missing parts, input them into the combined model of Diffusion and GAN. The generator roughly completes the contour, and the generator of GAN further refines the generation to make the completed part fuse with the surrounding features. The discriminator judges and feeds back the generated image, and the generator optimizes accordingly. After multiple iterative trainings, the combined model generates high-quality completed images to assist doctors in prenatal diagnosis, and is expected to improve the accuracy of fetal ultrasound examination and ensure fetal health.
[0050] The innovation points of this embodiment mainly include:
[0051] 1. Innovative Technical Fusion Architecture: A unique combined model architecture of Diffusion and GAN is constructed. The noise addition process of the Diffusion model is organically combined with the generative adversarial mechanism of GAN to form a new fetal face completion model system. In this architecture, the Diffusion model is responsible for learning the feature distribution of fetal face images at different noise levels, providing prior knowledge for subsequent completion; GAN, based on the preliminary results generated by Diffusion and the visible part of the original image, conducts refined generation and discrimination. The two work together, and this innovative architecture design lays the foundation for achieving high-precision fetal face completion.
[0052] 2. Feature Pre-generation Based on Diffusion: The Diffusion model is used to perform forward noise addition training on a large number of complete fetal face ultrasound images, enabling the model to master the variation rules of fetal face features under different interference degrees. When facing fetal face images with missing parts, the Diffusion model can process the initial noise of the missing parts according to the learned feature distribution, generating a preliminary completion contour containing key facial feature information. This pre-generation process provides important guidance and basic information for the subsequent refined generation of GAN.
[0053] 3. Refined Generation and Discrimination Mechanism of GAN: The generator of the GAN part adopts a multi-layer convolutional and transposed convolutional structure. Taking the preliminary contour generated by Diffusion and the visible part of the original image as inputs, it conducts refined generation on the missing areas of the fetal face, continuously optimizing the pixel values to achieve seamless fusion of the completed part with the surrounding facial features. The discriminator strictly discriminates the differences between the generated image and the real complete fetal face image from both the global and local levels, as well as the coherence and rationality between the completed part and the remaining part of the original image. By the feedback of the discriminator, the generator is continuously optimized to ensure the high quality and authenticity of the generated image.
[0054] The beneficial effects brought by this embodiment mainly include:
[0055] 1. High-precision and Real Completion: The Diffusion model fully learns the variation rules of fetal face features, providing prior information for the generator. Based on this and the visible part of the original image, the generator finely generates the missing areas. The discriminator strictly checks to ensure that the completed part is highly integrated with the whole, with real details, texture, and color transitions, and conforms to the medical development characteristics of the fetal face, greatly improving the completion accuracy and authenticity.
[0056] 2. Strong generalization and adaptation ability: Fetal ultrasound images are complex and diverse. The Diffusion model is trained with a large number of images under different conditions to master a wide range of feature patterns. The combined model can thus handle various fetal facial deficiencies, including complex deficiencies caused by body positions, equipment, etc., and can effectively complete unknown deficiency scenarios based on learned experience. Its generalization ability far exceeds that of a single model.
[0057] 3. Strongly assist medical research and diagnosis: The clear images after completion help doctors comprehensively observe the fetal face, accurately diagnose deformity problems such as cleft lip and palate, and improve the accuracy of prenatal diagnosis. At the same time, it provides rich data for medical research, facilitating in-depth analysis of the fetal facial development rules and influencing factors, and promoting the development of fetal medicine.
[0058] The embodiment of the present application also provides a training device for a fetal facial completion model, and the training device includes:
[0059] A first acquisition unit, configured to acquire training data, where the training data includes: real images and images to be completed; and
[0060] A training unit, configured to train the model using the training data,
[0061] The model uses a generative adversarial network, and the generative adversarial network includes: a generator for completing the image to be completed, and a discriminator for discriminating between real images or generated images. The generator adopts a stable diffusion generation mode.
[0062] The embodiment of the present application also provides a fetal facial completion device, and the completion device includes:
[0063] A second acquisition unit, configured to acquire an image to be processed; and
[0064] A processing unit, configured to process the image to be processed using the trained model as described above to obtain a fetal facial completion result.
[0065] The embodiment of the present application also provides an electronic device, including: a processor, and a memory coupled to the processor. The memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in any one of the above embodiments.
[0066] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory.
[0067] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire device through various interfaces and circuits.
[0068] The memory can be used to store the computer program. By running or executing the computer program stored in the memory and calling the data stored in the memory, the processor realizes various functions of the electronic device.
[0069] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0070] The embodiments of the present application also provide a storage medium. The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, Read-Only Memory (ROM), Random Access Memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0071] The embodiment of the present application also provides a computer program product, including: a computer program or instruction, when the computer program or instruction runs on a computer, enabling the computer to execute the method of any one of the above possible implementation manners.
[0072] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made, and these improvements and modifications are also regarded as the protection scope of the present application.
Claims
1. A training method for a fetal facial completion model, characterized in that, The method includes: Obtaining training data, where the training data includes: real images and images to be completed; and Using the training data to train the model, The model adopts a generative adversarial network, and the generative adversarial network includes: a generator for completing the image to be completed, and a discriminator for discriminating real images or generated images. The generator adopts a noise addition process in the stable diffusion generation mode.
2. The training method according to claim 1, wherein The generative adversarial network further includes: a three-dimensional segmenter for segmenting the generated image of the generator. The intersection over union (IoU) obtained by the processing of the three-dimensional segmenter is used for the loss calculation of the discriminator.
3. The training method according to claim 1, wherein The generator includes: A three-dimensional Gaussian noise generator, a first convolutional network for feature extraction of the image to be completed, and a second convolutional network for repeatedly superimposing three-dimensional Gaussian noise on the extracted feature map.
4. The training method according to claim 2, characterized in that The generative adversarial network updates parameters by optimizing the objective function, and the objective function adopts a loss function, and the loss function includes: generator loss, discriminator loss, and three-dimensional segmenter loss.
5. A method for fetal face completion, characterized in that, The completion method includes: Obtaining an image to be processed; and Using the model trained as described in any one of claims 1-4 to process the image to be processed to obtain a fetal face completion result.
6. A training device for a fetal facial completion model, characterized in that, The training device includes: A first acquisition unit for obtaining training data, where the training data includes: real images and images to be completed; and A training unit for using the training data to train the model, The model adopts a generative adversarial network, and the generative adversarial network includes: a generator for completing the image to be completed, and a discriminator for discriminating real images or generated images. The generator adopts a noise addition process in the stable diffusion generation mode.
7. A fetal face completion device, characterized in that, The completion device includes: A second acquisition unit for obtaining an image to be processed; and A processing unit for using the model trained as described in any one of claims 1-4 to process the image to be processed to obtain a fetal face completion result.
8. An electronic device, characterized in that, The electronic device includes: a processor and a memory coupled to the processor, The memory is used for storing a computer program; and The processor is used for executing the computer program stored in the memory so that the electronic device executes the training method described in any one of claims 1-4, or executes the completion method described in claim 5.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the training method described in any one of claims 1-4, or execute the completion method described in claim 5.
10. A computer program product, characterized in that, The computer program product includes: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the training method described in any one of claims 1-4, or execute the completion method described in claim 5.