A method and related apparatus for determining an image synthesis model.
By constructing an image synthesis model and optimizing the convolutional layer parameters using training sample pairs and a distribution loss function, the problem of inaccurate multimodal image synthesis is solved, and more efficient cross-modal image acquisition is achieved.
Patent Information
- Application Number
- CN202211063145.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2042-08-31
AI Technical Summary
Existing technologies for multimodal image synthesis are inaccurate, making it difficult to effectively reduce the consumption of human and material resources, and failing to meet the image acquisition needs in multimodal scenarios.
By constructing an image synthesis model, optimizing convolutional layer parameters using training sample pairs and a distribution loss function, learning modal differences and improving feature extraction accuracy, cross-modal image synthesis is achieved.
It improves the efficiency of image acquisition in multimodal scenarios and synthesizes target modal images that are closer to reality.
Smart Images

Figure CN117011199B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular to a method and apparatus for determining an image synthesis model. Background Technology
[0002] A target object can have images in multiple modalities. Images in different modalities can reflect different aspects of the target object's characteristics, so images in multiple modalities can provide a more comprehensive description of the target object.
[0003] However, acquiring multimodal images often requires multiple methods, thus consuming significant human and material resources. For example, magnetic resonance imaging (MRI), with its modal diversity, can obtain multimodal medical images. This imaging method significantly improves the productivity of routine diagnosis and advanced research; however, the high variability between devices and the high cost of examinations pose challenges to the acquisition and utilization of multimodal images.
[0004] In multimodal scenarios, synthesizing images of missing modalities from images of existing modalities is an effective way to reduce the consumption of human and material resources. However, the images of missing modalities synthesized in related technologies are not accurate and cannot meet the requirements. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a method and related apparatus for determining an image synthesis model, which can synthesize images that more closely resemble the real target modality, thereby improving image acquisition efficiency in multimodal scenarios.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] On one hand, embodiments of this application provide a method for determining an image synthesis model, the method comprising:
[0008] Obtain training sample pairs, wherein the training sample pairs include a first sample image of a first modality and a second sample image of a second modality;
[0009] The training sample pairs are input into the initial synthesis model. The initial synthesis model includes a first feature extraction sub-model for inputting the first sample image and a second feature extraction sub-model for inputting the second sample image. The first feature extraction sub-model includes N first convolutional layers, and the second feature extraction model includes N second convolutional layers, where N≥1.
[0010] Determine the distribution difference in feature distribution between the first output feature of the first convolutional layer of the i-th layer and the second output feature of the second convolutional layer of the i-th layer, where i is a positive integer less than or equal to N;
[0011] The i-th distribution loss function is constructed by the first restoration difference of the first output feature relative to the first sample image, the second restoration difference of the second output feature relative to the second sample image, and the distribution difference;
[0012] Based on the i-th distribution loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the first restoration difference, the second restoration difference, and the distribution difference. The initial synthesis model is then trained to obtain an image synthesis model, which is used for cross-modal image synthesis between the first modality and the second modality.
[0013] On the other hand, embodiments of this application provide an apparatus for determining an image synthesis model, the apparatus comprising an acquisition unit, an input unit, a determination unit, a construction unit, and a training unit:
[0014] The acquisition unit is used to acquire training sample pairs, the training sample pairs including a first sample image of a first modality and a second sample image of a second modality;
[0015] The input unit is used to input the training sample pairs into the initial synthesis model. The initial synthesis model includes a first feature extraction sub-model for inputting the first sample image and a second feature extraction sub-model for inputting the second sample image. The first feature extraction sub-model includes N first convolutional layers, and the second feature extraction model includes N second convolutional layers, where N≥1.
[0016] The determining unit is used to determine the distribution difference in feature distribution between the first output feature of the i-th first convolutional layer and the second output feature of the i-th second convolutional layer, where i is a positive integer less than or equal to N;
[0017] The construction unit is used to construct the i-th distribution loss function by the first restoration difference of the first output feature relative to the first sample image, the second restoration difference of the second output feature relative to the second sample image, and the distribution difference;
[0018] The training unit is used to adjust the layer parameters of the first convolutional layer and the second convolutional layer of the i-th layer according to the i-th distribution loss function by minimizing the first restoration difference, the second restoration difference and the distribution difference optimization objective, and train the initial synthesis model to obtain an image synthesis model. The image synthesis model is used to perform cross-modal image synthesis between the first modality and the second modality.
[0019] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory:
[0020] The memory is used to store program code and transmit the program code to the processor;
[0021] The processor is used to execute the methods described above according to the instructions in the program code.
[0022] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program for performing the methods described above.
[0023] On the other hand, embodiments of this application provide a computer program product including instructions that, when run on a computer, cause the computer to perform the methods described above.
[0024] As can be seen from the above technical solution, after obtaining training sample pairs including the first sample image of the first modality and the second sample image of the second modality, the training sample pairs can be input into the initial synthesis model. The initial synthesis model includes a first feature extraction sub-model for inputting the first sample image and a second feature extraction sub-model for inputting the second sample image. The first feature extraction sub-model includes N first convolutional layers, and the second feature extraction sub-model includes N second convolutional layers. The i-th first convolutional layer can output a first output feature determined based on the first sample image, and the i-th second convolutional layer can output a second output feature determined based on the second sample image. The distribution difference between the first and second output features and the reconstruction difference corresponding to each output feature are determined. Since the distribution difference can identify the modal difference between the first and second modes, the larger the difference, the more unfavorable it is for subsequent cross-modal image synthesis between the first and second modes through the image synthesis model. The reconstruction difference can identify the difference between the reconstructed image obtained by the corresponding output feature and the sample image. The larger the difference, the more difficult it is for the output features used by the initial synthesis model for cross-modal image synthesis to reflect the accurate information of the source mode, and instead carry too much noise information. Therefore, the i-th distribution loss function is constructed through the distribution difference and each restoration difference. Based on the i-th distribution loss function, the layer parameters of the first convolutional layer and the second convolutional layer of the i-th layer are adjusted by minimizing the optimization objective of each restoration difference and the distribution difference, thereby training the initial synthesis model to obtain the image synthesis model. Since the image synthesis model learns the modal differences between the first and second modalities and improves the feature extraction accuracy through the above training, the transferability of effective features in cross-modal synthesis is enhanced. This enables the image synthesis model to accurately convert the image features of the target modal image based on the output features of the source modal image, synthesize an image that is closer to the real target modal, and improve the image acquisition efficiency in multimodal scenarios. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 A schematic diagram illustrating the scene determination by the image synthesis model provided in the embodiments of this application;
[0027] Figure 2 A flowchart illustrating a method for determining an image synthesis model, as provided in an embodiment of this application;
[0028] Figure 3 One of the schematic diagrams for model training provided in the embodiments of this application;
[0029] Figure 4 A second schematic diagram illustrating model training provided in an embodiment of this application;
[0030] Figure 5 The third schematic diagram of model training provided for the embodiments of this application;
[0031] Figure 6 Fourth schematic diagram of model training provided for embodiments of this application;
[0032] Figure 7 A schematic diagram illustrating loss function determination from the perspective of feature space, provided for embodiments of this application;
[0033] Figure 8 This application provides an illustration of image synthesis.
[0034] Figure 9 A schematic diagram of model verification provided for an embodiment of this application;
[0035] Figure 10 This is another schematic diagram of model verification provided in an embodiment of this application;
[0036] Figure 11 A device structure diagram of an image synthesis model determination apparatus provided in an embodiment of this application;
[0037] Figure 12 A structural diagram of a terminal device provided in an embodiment of this application;
[0038] Figure 13 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation
[0039] The embodiments of this application will now be described with reference to the accompanying drawings.
[0040] In multimodal scenarios, synthesizing images of missing modalities from images of existing modalities is an effective way to reduce the consumption of human and material resources. However, the images of missing modalities synthesized in related technologies are not accurate and cannot meet the requirements.
[0041] To this end, this application provides a method and related apparatus for determining an image synthesis model. The resulting image synthesis model can synthesize a more realistic image, thereby improving the efficiency of image acquisition in multimodal scenarios.
[0042] The method for determining the image synthesis model provided in this application embodiment can be implemented by a computer device, which can be a terminal device or a server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0043] Terminal devices include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, and medical imaging. Terminal devices and servers can be directly or indirectly connected via wired or wireless communication methods; this application does not impose any limitations on this.
[0044] This application can be applied to the field of artificial intelligence (AI). AI is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0045] The embodiments of this application mainly relate to computer vision technology and machine learning.
[0046] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0047] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0048] For example, in the embodiments of this application, various types of information included in the image can be identified based on computer vision technology, such as distribution differences, structural differences, offsets in subspaces, and correlation differences. Furthermore, an initial synthesis model can be trained based on machine learning to obtain an image synthesis model. This enables the image synthesis model to perform cross-modal image synthesis, such as synthesizing a second-modal image based on a target image of the first modality, thereby synthesizing a more realistic synthesized image and improving the image acquisition efficiency in multimodal scenarios.
[0049] It is understood that in the specific embodiments of this application, the image samples and target images used may involve data related to user facial features. When the above embodiments of this application are applied to specific products or technologies, each item requires separate user permission or consent, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0050] To facilitate understanding of the technical solutions provided in this application, the following section will introduce a method for determining an image synthesis model provided in an embodiment of this application, in conjunction with a practical application scenario.
[0051] Figure 1 A schematic diagram of the scene determination using the image synthesis model provided in this application embodiment is shown, wherein server 100 is used as an example of the aforementioned computer device for illustration.
[0052] First, training sample pairs 20 and an initial synthetic model 10 are obtained for training the model. The training sample pairs 20 include a first sample image of the first modality and a second sample image of the second modality. The initial synthetic model 10 includes a first feature extraction sub-model 11 and a second feature extraction sub-model 12. The first feature extraction sub-model 11 includes N first convolutional layers, and the second feature extraction sub-model 12 includes N second convolutional layers. The first convolutional layers are arranged in order according to the direction of data flow and are numbered 1 to N according to the order of arrangement. The second convolutional layers are arranged in order according to the direction of data flow and are numbered 1 to N according to the order of arrangement.
[0053] Server 100 can input training sample pairs 20 into the initial synthetic model 10. Specifically, the first sample image is input into the first feature extraction sub-model 11, and the second sample image is input into the second feature extraction sub-model 12. The i-th first convolutional layer can output a first output feature determined based on the first sample image, and the i-th second convolutional layer can output a second output feature determined based on the second sample image. N is a positive integer greater than or equal to 1, i is a positive integer less than or equal to N, the i-th first convolutional layer is one of N first convolutional layers, and the i-th second convolutional layer is one of N second convolutional layers. Figure 1 The example below uses N greater than 3 as an example for illustration.
[0054] Server 100 can determine the distribution difference in feature distribution between the first output feature and the second output feature, as well as the reconstruction difference corresponding to each output feature. Since this distribution difference can identify the modal difference between the first modality and the second modality, the larger the difference, the more unfavorable it is for subsequent cross-modal image synthesis between the first modality and the second modality through the image synthesis model. The reconstruction difference can identify the difference between the reconstruction image obtained from the corresponding output feature and the sample image. The larger the difference, the more difficult it is for the output features used by the initial synthesis model for cross-modal image synthesis to reflect the accurate information of the source modality, and instead carry too much noise information. Therefore, the i-th distribution loss function is constructed through this distribution difference and each reconstruction difference. By minimizing the optimization objective of each reconstruction difference and the distribution difference, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted, thereby training the initial synthesis model 10 to obtain the image synthesis model 30.
[0055] Because the above training enables the image synthesis model to learn the modal differences between the first and second modalities and improve the accuracy of feature extraction, it enhances the transferability of effective features in cross-modal synthesis. This allows the image synthesis model to accurately convert the image features of the target modal image based on the output features of the source modal image, synthesize an image that is closer to the real target modal, and improve the efficiency of image acquisition in multimodal scenarios.
[0056] Figure 2 This application provides a flowchart of a method for determining an image synthesis model. In this embodiment, a server is used as the aforementioned computer device for description. The method includes:
[0057] S101, Obtain training sample pairs.
[0058] In this embodiment, training sample pairs are used to train an initial synthesis model. The initial synthesis model includes a first feature extraction sub-model and a second feature extraction sub-model. The first feature extraction sub-model is used to extract features from images of a first modality, and the second feature extraction sub-model is used to extract features from images of a second modality. For this initial synthesis model, the training sample pairs include first sample images of the first modality and second sample images of the second modality. The first sample images are input into the first feature extraction sub-model, and the second sample images are input into the second feature extraction sub-model. Specifically, the first sample image can be denoted as X, and the second sample image can be denoted as Y.
[0059] Training sample pairs can originate from a training sample set, which may include multiple training sample pairs. These multiple training sample pairs can be sequentially input into an initial synthetic model to train the initial synthetic model. Each training sample pair includes two sample images from different modalities. In this embodiment, a training sample pair including a first sample image and a second sample image is used as an example. The images of different modalities of the target object are used to reflect different aspects of the target object's features. Images from multiple modalities can more comprehensively describe the target object. The target object can be a person, a scene, an object to be detected, etc. The images of different modalities of the target object can be, for example, images of different image styles of the target object, or images used to obtain images reflecting different measurement parameters of the target object. A modality can be considered a domain, and image transformation between different modalities can also be called image transformation between different domains.
[0060] For example, magnetic resonance imaging (MRI) is a medical imaging technique used for medical diagnosis. MRI can obtain multimodal images of a target object. Based on these multimodal images, various feature data of the target object can be obtained. Different modal images can be obtained through different image generation devices, which are also detection devices. The target object can be the brain, spinal cord, heart and major blood vessels, joints and bones, soft tissues, and pelvis, etc. Taking the brain as an example, the image of modality 1 is used to obtain the spin relaxation time T2 of the brain, the image of modality 2 is used to obtain the proton density (PD) of the brain, the image of modality 3 is used to obtain the fluid attenuated inversion recovery (FLAIR) sequence of the brain, and the image of modality 4 is used to obtain the lattice relaxation time T1 of the brain. In addition, there are other modes used to obtain physical properties such as diffusion coefficient, magnetization coefficient, and chemical shift.
[0061] However, acquiring multimodal images often requires multiple methods, thus consuming significant human and material resources. In multimodal scenarios, synthesizing images of missing modalities from existing modal images is an effective way to reduce these resources. However, the images of missing modalities synthesized in related technologies are often inaccurate and fail to meet requirements. For example, the synthesized images of missing modalities may not accurately represent the target object, rendering the synthesized image unusable and failing to achieve the goal of reducing human and material resources. The initial synthesis model provided in this application can learn the modal differences between the first and second modalities and improve feature extraction accuracy, thereby enhancing the transferability of effective features in cross-modal synthesis. The trained image synthesis model can synthesize near-realistic images, improving image acquisition efficiency in multimodal scenarios.
[0062] Specifically, the training sample set for training sample pairs can include sample images generated by different image generation devices, as well as sample images using different institutional standards. By diversifying the sample images, the initial synthetic model can effectively learn the synthetic knowledge of sample images from different sources during the training process, avoiding overfitting caused by too many sample images and improving the generalization of the initial synthetic model.
[0063] For example, the training sample set includes images of multiple modality 1, which come from device A and device B, respectively. The training sample set also includes images of multiple modality 2, which come from hospital C, hospital D, and research institute E, respectively.
[0064] In one possible implementation, the training sample set, including training sample pairs, may include images of target objects of the same category in different states. The category of the target object can be determined based on the function of the image synthesis model. Target objects of the same category can share the same image synthesis model. Examples of target objects of the same category include those belonging to people, animals, or landscapes, as well as those with the same detection regions. The state of the target object can be determined based on the function of the image synthesis model. Images of the target object in different states can share the same image synthesis model. Examples of target objects in different states include people in different standing postures, and the same detection regions in different states. The same detection regions in different states include, for example, a healthy brain and a brain with a brain tumor. As an example, the training sample set includes a first subset and a second subset. The first subset includes images of a healthy brain, and the second subset includes images of a brain with a brain tumor. Images of a healthy brain can come from the publicly available IXI dataset, and images of a brain with a brain tumor can come from the publicly available BraTS dataset. Using brain MRI as a sample image as an example, it can cover both healthy brain MRI images and brain MRI images with pathology, making the model's learning scope more comprehensive.
[0065] When the first and second sample images are brain magnetic resonance imaging based on the cranium, the brain magnetic resonance imaging can be divided into multiple two-dimensional slice data. The first sample image can be the image corresponding to the two-dimensional slice data of the last two bottleneck layers, and the second sample image can also be the image corresponding to the two-dimensional slice data of the last two bottleneck layers.
[0066] In practice, the first and second sample images belonging to the same training sample pair are generated based on the same type of target object. These images can be generated based on different target objects, the same target object, or different target objects. Images of different modalities generated based on the same target object are more likely to have their modality-independent deep information extracted, which helps improve the accuracy of image synthesis in the image synthesis model. For example, the first and second sample images could be images of different modalities obtained from multiple detections of the same cranium. However, since generating multimodal images based on the same target object is challenging, in practice, multimodal images based on the same type of target object (but belonging to different target objects) are often used as the first and second sample images of the same training sample pair. This poses a challenge to improving model accuracy. The embodiments of this application, through subsequent training based on inter-domain differences and feature extraction accuracy, can effectively solve the problem of limited or no paired data for the sample images in the first and second modalities, i.e., the problem caused by the sample images belonging to different target objects in the sample image pair.
[0067] S102, input the training sample pairs into the initial synthetic model.
[0068] In this embodiment, the initial synthesis model can be used to implement the convolutional sparse coding algorithm. Sparse coding (SC) is an unsupervised image reconstruction technique that uses a set of basis vectors to more efficiently represent sample data. The goal of sparse coding is to find a set of over-complete basis vectors (also called a dictionary) D such that the input vector X... i This can be represented as a linear combination of these basis vectors, expressed through the sparse coefficient d. i The sparse coefficients can represent the combination of basis vectors corresponding to each input vector. These coefficients can be determined through training and are reflected in the coefficients of the sparse coding algorithm. The input vector can be represented as: X i =D*d i , where X i ∈R d×1 That is, a matrix with d rows and 1 column, where D∈R d×m That is, a matrix with d rows and m columns, d i ∈R m×1 , which is an m-row, one-column matrix. In sparse coding algorithms, by training two basis vectors corresponding to the source and target domain images respectively, and making the sparse representations of the source and target domains equal, the target domain image can be synthesized by replacing the basis vectors corresponding to the source domain with the basis vectors corresponding to the target domain after sparse representation of the source domain image.
[0069] Convolutional Sparse Coding (CSC) is a sparse coding algorithm implemented through convolutional layers. Compared to other sparse coding algorithms, CSC can combine image features and handle global data, making it widely used for learning translation-invariant dictionaries in image processing. CSC uses a sample-adaptive dictionary, where each filter is a linear combination of a set of basic filters learned from the data. This increases the flexibility of data computation, allowing the capture of a large number of sample-dependent patterns, making it particularly effective when dealing with large or high-dimensional datasets. Models based on CSC can learn to achieve lower computational time and space complexity. In CSC, an image with N pixels is used as the input vector X. CSC is implemented through filters D in the model, which include m basic filters. The parameters of the m basic filters represent the sparsity coefficients. The i-th basic filter d... iFeature extraction is performed on a small image n that matches the convolution kernel, with basis vector Γ. i The i-th feature map is an image with the same size as the input vector X and corresponding to the i-th basic filter. The input vector can be represented as:
[0070]
[0071] Similarly, in the convolutional sparse coding algorithm, by training two basis vectors corresponding to the source domain and the target domain images respectively, and making the sparse representations of the source domain and the target domain equal, the target domain image can be synthesized by replacing the basis vectors corresponding to the source domain with the basis vectors corresponding to the target domain after sparse representation of the source domain image.
[0072] In the convolutional sparse coding algorithm, the filter D can be implemented through convolutional layers. The convolutional layers act as convolutional filters. The initial synthesis model based on the convolutional sparse coding algorithm can include a first feature extraction sub-model and a second feature extraction sub-model. The first feature extraction sub-model is used to extract features of the image of the first modality, and the second feature extraction sub-model is used to extract features of the image of the second modality. The first feature extraction sub-model includes N first convolutional layers, and the second feature extraction model includes N second convolutional layers, where N≥1, that is, the first feature extraction sub-model and the second feature extraction sub-model have the same number of convolutional layers.
[0073] Where N is 1, the first feature extraction sub-model includes one first convolutional layer, and the second feature extraction sub-model includes one second convolutional layer. Features of the first modality image can be extracted through the first convolutional layer, and features of the second modality image can be extracted through the second convolutional layer. When N is greater than 1, the first feature extraction sub-model includes multiple first convolutional layers, and the second feature extraction sub-model includes multiple second convolutional layers. The first convolutional layers have an order based on the data flow and are assigned layer numbers according to this order. The second convolutional layers also have an order based on the data flow and are assigned layer numbers according to this order. Both the layer numbers of the first and second convolutional layers are positive integers less than or equal to N. The output feature of the i-th first convolutional layer is then used as the i-th feature. The input features of the +1th layer of the first convolutional layer and the output features of the i-th layer of the second convolutional layer are used as the input features of the (i+1)th layer of the second convolutional layer, where i is a positive integer less than or equal to N. The feature size of the output features of the (i+1)th layer of the first convolutional layer is smaller than that of the i-th layer of the first convolutional layer, and the feature size of the output features of the (i+1)th layer of the second convolutional layer is smaller than that of the i-th layer of the second convolutional layer. The hierarchical structure of the first convolutional layers enables multiple first convolutional layers to extract deeper features of the first modality of the image, and the hierarchical structure of the second convolutional layers enables multiple second convolutional layers to extract deeper features of the second modality of the image. Multiple first convolutional layers and multiple second convolutional layers can act as convolutional filters, leveraging the advantage of depth to convey information in an increasingly compact manner.
[0074] For example, N can be 9. A batch normalization layer can be added after the first convolutional layer to promote convergence, and another batch normalization layer can be added after the second convolutional layer to promote convergence. A global spatial average pooling layer can be set after the last first convolutional layer, and a global spatial average pooling layer can be set after the last second convolutional layer. The sampling stride of the first and second convolutional layers can be 2, achieving spatial subsampling of the first and second sample images. The initial synthetic model can be set to 200 epochs, the learning rate can be set to 0.0002, the batch size can be 32, and the balancing parameters can be set to λ = 0.2, α = 0.15, and γ = 1. A multi-kernel maximum mean discretization (MK-MMD) model with Gaussian kernels is used, and the bandwidth is configured as median pairwise squared distance.
[0075] In this embodiment, training sample pairs can be input into the initial synthesis model. Specifically, a first sample image can be input into a first feature extraction sub-model, and a second sample image can be input into a second feature extraction sub-model. The i-th first convolutional layer can output a first output feature determined based on the first sample image, and the i-th second convolutional layer can output a second output feature determined based on the second sample image. The first output feature can be denoted as Z. x,i The second output feature can be denoted as Z. y,i When i is 1, the first output feature can be denoted as Z. x,1 The second output feature can be denoted as Z. y,1 Taking the first convolutional layer of the i-th layer as the feature extractor, its operation function is represented by f, and the output of the filter is represented by F. Then the first output feature can be expressed as: Z x,i =f(X,F x,i-1 The second output feature can be expressed as: Z (λ). y,i =f(Y,F y,i-1 ,λ), where λ is a constant, X in f is used to identify that the input features come from the first mode, and Y is used to identify that the input features come from the second mode. The feature representing the tensor attribute with height h and width w. The feature represents the tensor attribute with height h and width w.
[0076] When the sample image is generated by an image synthesis device, even within the same modality, different device manufacturers and device parameter settings can cause errors in the generated image within the same modality.
[0077] In this embodiment of the application, the first output feature and the second output feature may be processed by Intra-Domain Standardization (IDS), see [link to relevant documentation]. Figure 3 This is a schematic diagram of model training provided in an embodiment of this application. Intra-domain standardization can reduce the differences in intra-domain features.
[0078] Within a single modality, intra-domain variability can arise from differences in the image generation device's source and physical parameter settings. Intra-domain feature variability is generally detrimental to effective feature learning; excessive learning of variably occurring intra-domain features can easily lead to model overfitting. Reducing intra-domain feature variability can prevent overfitting. Intra-domain feature variability can be caused by different image generation devices or different image sources. For example, different hospitals using different scanning parameters may produce different images, resulting in feature differences that interfere with the effective features.
[0079] Specifically, we can obtain the first initial output features of the first convolutional layer (i-th layer) and the second initial output features of the second convolutional layer (i-th layer). Through global normalization, we map the first initial output features to a first image size to obtain the first output features. Similarly, we map the second initial output features to a second image size through global normalization to obtain the second output features. The first image size can be the size of a unit sphere, and the second image size can also be the size of a unit sphere. This unifies the maximum norm of the first and second output features, eliminating scaling ambiguity between them within their respective domains.
[0080] S103, determine the distribution difference in feature distribution between the first output feature of the first convolutional layer of the i-th layer and the second output feature of the second convolutional layer of the i-th layer.
[0081] S104, construct the i-th distribution loss function by the first restoration difference of the first output feature relative to the first sample image, the second restoration difference of the second output feature relative to the second sample image, and the distribution difference.
[0082] In this embodiment of the application, after obtaining that the first convolutional layer of the i-th layer can output the first output feature determined based on the first sample image, and the second convolutional layer of the i-th layer can output the second output feature determined based on the second sample image, the distribution difference between the first output feature and the second output feature in the feature distribution can be determined. Since this distribution difference can identify the modal difference between the first mode and the second mode, the larger the difference, the more unfavorable it is for subsequent cross-modal image synthesis between the first mode and the second mode through the image synthesis model.
[0083] Specifically, the distributional differences in the feature distributions can be determined by mapping the first and second output features to the same feature space, and then the i-th distributional loss function can be constructed based on these differences. The feature space can be a Hilbert space, and the distributional differences in the feature distributions can be determined using two-sample testing methods such as MK-MMD or mean embedding test.
[0084] Reducing the distribution difference between the first and second output features, as an inter-domain adaptation process, essentially demonstrates cross-modal invariant structure and effectively learns relevant sparse features from different domains. Through the inter-domain adaptation process, the initial synthetic model can learn inter-domain differences in an unsupervised manner to correct misalignments and adjust the model to better generalize different datasets, thus achieving the goal of adaptation.
[0085] The first restoration difference mentioned in S104 is used to identify the difference between the restored image obtained through the first output feature and the first sample image, and the second restoration difference is used to identify the difference between the restored image obtained through the second output feature and the second sample image.
[0086] The reconstructed image can be obtained based on a first output feature and filter parameters used to generate the first output feature. The filter parameters could be, for example, the aforementioned F... x,i-1 When the accuracy of the first output feature is high and the noise is low, the difference between the restored image and the first sample image is small, and vice versa. In the scenario of obtaining a second modality image through cross-domain image synthesis using the image of the first modality, the first output feature is the basis for cross-domain image synthesis. When the feature extraction accuracy is insufficient, the low-quality first output feature will directly affect the effect of cross-domain image synthesis.
[0087] The reconstructed image can be obtained based on a second output feature and filter parameters used to generate the second output feature. The filter parameters could be, for example, the aforementioned F... y,i-1 When the accuracy of the second output feature is high and the noise is low, the difference between the restored image and the second sample image is small, and vice versa. In scenarios where the first modality image is obtained through cross-domain image synthesis using the second modality image, the second output feature is the foundation for cross-domain image synthesis. When the feature extraction accuracy is insufficient, the low-quality second output feature will directly affect the effect of cross-domain image synthesis.
[0088] Therefore, the reconstruction difference can identify the difference between the reconstructed image obtained from the corresponding output features and the sample image. The larger the difference, the less likely the output features used by the initial synthesis model for cross-modal image synthesis are to accurately reflect the source modality information, and instead carry too much noise information. Therefore, the i-th distribution loss function is constructed through this distribution difference and each reconstruction difference, and the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the optimization objective of each reconstruction difference and the distribution difference.
[0089] Therefore, the i-th distribution loss function is constructed based on this distribution difference. The distribution loss function characterizes the magnitude of the distribution difference; the larger the value of the loss function, the larger the distribution difference, and vice versa. In one possible implementation, the first and second output features are features that have undergone intra-domain difference elimination processing. See [link to relevant documentation]. Figure 3 .
[0090] The i-th distribution loss function, constructed using the distribution difference, the first restored difference, and the second restored difference, can serve as a domain discrepancy metric (DDM). To express, specifically:
[0091]
[0092] in, For the first reduction difference, For the second reduction difference, The expected value of the feature distribution of the first output feature. The expected value of the feature distribution of the second output feature. The difference in feature distribution between the first and second output features.
[0093] In one possible implementation, the first and second output features can be regularized using the MK-MMD hierarchical regularizer to determine the distribution difference between the first and second output features in the feature distribution. In this case, the above equation... It can be adjusted to To express.
[0094] The regularization method described above makes the regularization results more suitable for performing unbiased estimation across domains. Furthermore, it can improve the proximity of unpaired cross-domain data within the same portion when the target objects involved in the first and second sample images do not belong to the same target object.
[0095] S105, according to the i-th distribution loss function, adjust the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer by minimizing the first restoration difference, the second restoration difference and the distribution difference, and train the initial synthesis model to obtain the image synthesis model.
[0096] Both sparse coding and convolutional sparse coding in related technologies can achieve image synthesis. However, because sparse coding and convolutional sparse coding only focus on image features within their respective domains, they do not consider the adaptive extraction of features between domains or the accuracy of features within a domain. This results in poor transferability of image features between different domains. Based on this, after constructing the i-th distribution loss function using the first restoration difference corresponding to the first output feature, the second restoration difference corresponding to the second output feature, and the distribution difference, the layer parameters of the i-th convolutional layer and the i-th convolutional layer can be adjusted according to the i-th distribution loss function by minimizing the optimization objectives of the first restoration difference, the second restoration difference, and the distribution difference. This trains the initial synthesis model into an image synthesis model, which can be used for cross-modal image synthesis between the first and second modalities.
[0097] Thus, the transferability of convolutional sparse features can be enhanced and quantized by the distributional differences in feature distributions, such as MMD, and the transferability of convolutional sparse features can be improved by minimizing the distributional differences, thereby generalizing the heterogeneous representation of features across modal images.
[0098] Therefore, since the distribution difference can identify the modal difference between the first mode and the second mode, the larger the difference, the more unfavorable it is for subsequent cross-modal image synthesis between the first mode and the second mode through the image synthesis model. The reconstruction difference can identify the difference between the reconstructed image obtained by the corresponding output features and the sample image. The larger the difference, the more difficult it is for the output features used by the initial synthesis model for cross-modal image synthesis to reflect the accurate information of the source mode, and instead carry too much noise information. Therefore, the i-th distribution loss function is constructed through the distribution difference and each restoration difference. Based on the i-th distribution loss function, the layer parameters of the first convolutional layer and the second convolutional layer of the i-th layer are adjusted by minimizing the optimization objective of each restoration difference and the distribution difference, thereby training the initial synthesis model to obtain the image synthesis model. Since the image synthesis model learns the modal differences between the first and second modalities and improves the feature extraction accuracy through the above training, the transferability of effective features in cross-modal synthesis is enhanced. This enables the image synthesis model to accurately convert the image features of the target modal image based on the output features of the source modal image, synthesize an image that is closer to the real target modal, and improve the image acquisition efficiency in multimodal scenarios.
[0099] The above explanation uses the i-th first convolutional layer, the i-th second convolutional layer, and the i-th distributed loss function as an example. This means that the i-th convolutional layer of the two feature extraction sub-models in the initial synthesis model shares a single distributed loss function for adjusting the layer parameters. When the number of convolutional layers N in each of the two feature extraction sub-models is greater than 1, multiple distributed loss functions can be calculated. The number of distributed loss functions is consistent with the number of convolutional layers N; that is, the number of distributed loss functions can be equal to the number of the first and second convolutional layers N. Each first and second convolutional layer can correspond to one distributed loss function. Therefore, the layer parameters of each first and second convolutional layer can be adjusted using the corresponding distributed loss function, resulting in higher model accuracy.
[0100] See Figure 4This is a schematic diagram of another model training method provided in an embodiment of this application. The first convolutional layer and the first second convolutional layer correspond to the first distributed loss function, the ith first convolutional layer and the ith second convolutional layer correspond to the ith distributed loss function, and the Nth first convolutional layer and the Nth second convolutional layer correspond to the Nth distributed loss function. The first output feature and the second output feature can also be features processed by IDS. The number of distributed loss functions can also be less than the number of first convolutional layers N, that is, some first convolutional layers can correspond to distributed loss functions. Then the layer parameters of multiple first convolutional layers and multiple second convolutional layers can be adjusted to make the model have higher accuracy.
[0101] When the number of distributed loss functions equals the number of first convolutional layers N and N is greater than 1, there is a one-to-one correspondence between the N first convolutional layers and the N second convolutional layers. The i-th distributed loss function corresponds to the i-th first convolutional layer and the i-th second convolutional layer. Based on the i-th distributed loss function, the layer parameters of the i-th first convolutional layer and the i-th convolutional layer are adjusted by minimizing the optimization objectives of the first restoration difference, the second restoration difference, and the distribution difference. Specifically, based on the N distributed loss functions, the layer parameters of the corresponding first and second convolutional layers are adjusted by minimizing the optimization objectives of the first restoration difference, the second restoration difference, and the distribution difference, thereby training the initial synthesis model to obtain the image synthesis model.
[0102] During the training of the initial synthesis model to obtain the image synthesis model, the layer parameters of the i-th convolutional layer and the i-th convolutional layer can be adjusted through the mapping relationship between the first and second output features. Specifically, the first output feature can be transformed into a feature corresponding to the second modality through the corresponding mapping matrix, serving as the transformed feature corresponding to the i-th convolutional layer. Based on the feature difference between this transformed feature and the second output feature, the i-th association loss function is constructed. Based on the i-th association loss function, the layer parameters of the i-th convolutional layer and the i-th convolutional layer are adjusted by minimizing the feature difference and the optimization objective of the mapping matrix. See also Figure 5 This diagram illustrates another model training method provided in this application. The i-th association loss function can be constructed using the projection matrix, the first output feature, and the second output feature. The layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are then adjusted based on the i-th association loss function. In one possible implementation, the first and second output features are features processed by the IDS (Integrated Device Analyzer).
[0103] The feature difference between the transformed feature and the second output feature reflects the feature correlation between the first output feature and the second output feature. The greater the feature difference, the worse the feature correlation, which is less conducive to cross-modal image synthesis between the first and second modalities through the image synthesis model. Since the above training improves the correlation of effective features in cross-modal synthesis, improves the consistency of spatial relationship between the first and second output features, and enhances the transferability of effective features in cross-modal synthesis, the image synthesis model can accurately transform the image features of the synthesized image based on the output features of the target image, thereby synthesizing a more realistic synthesized image.
[0104] The mapping matrix can be represented as P, and the mapping matrix corresponding to the first convolutional layer of the i-th layer can be represented as P. i As discussed above, the first output feature can be represented as: Z x,i =f(X,F x,i-1 The second output feature can be expressed as: Z (λ). y,i =f(Y,F y,i-1 Since two independently learned sub-models (i, λ) produce unrelated features, concatenating them can lead to the transformation between these unrelated features, potentially resulting in the loss of effective features and the increase of ineffective features. Therefore, a correlation loss function can be constructed to achieve joint optimization of the two sub-models. The constructed i-th correlation loss function can be expressed as:
[0105]
[0106] Where α is a constant, P i Z x,i To transform features, the mapping matrix is an additional parameter that needs to be minimized to reduce its impact on the model.
[0107] The above explanation uses the first convolutional layer of the i-th layer, the second convolutional layer of the i-th layer, and the i-th association loss function as an example. In actual operation, the layer numbers i of the first and second convolutional layers that need to be adjusted can be preset. Based on the i-th association loss function, the layer parameters of the first convolutional layer of the i-th layer and the i-th convolutional layer can be adjusted by minimizing the feature difference and the optimization objective of the mapping matrix.
[0108] In practice, multiple association loss functions can be set, corresponding to multiple first convolutional layers and multiple second convolutional layers. The number of association loss functions can be equal to the number of first convolutional layers N, where N is greater than 1. This means each first convolutional layer can have one association loss function, allowing the layer parameters of both first and second convolutional layers to be adjusted, resulting in higher model accuracy. Alternatively, the number of association loss functions can be less than the number of first convolutional layers N, allowing some first convolutional layers to have association loss functions. This allows the layer parameters of multiple first and second convolutional layers to be adjusted, also resulting in higher model accuracy.
[0109] When the number of association loss functions equals the number of first convolutional layers N and N is greater than 1, the N first convolutional layers and the N second convolutional layers each have a one-to-one correspondence with the N association loss functions. Among them, the i-th association loss function in the N association loss functions corresponds to the i-th first convolutional layer and the i-th second convolutional layer. Then, according to the i-th association loss function, the layer parameters of the i-th first convolutional layer and the i-th convolutional layer are adjusted by minimizing the feature difference and the optimization objective of the mapping matrix. Specifically, according to the N association loss functions, the layer parameters of the corresponding first convolutional layer and the second convolutional layer are adjusted by minimizing the distribution difference, thereby training the initial synthesis model to obtain the image synthesis model.
[0110] In cross-domain synthesis, domain-invariant features within a single modality are crucial for expressing important information within that modality, representing important low-level details reflecting specific domain information. For example, in an image involving a target object within a particular modality, the object's geometric structure within that modality is vital information. If this type of geometrically relevant information is neglected during cross-domain synthesis, the synthesized image may have visual significance but lack practical value in meeting the requirements of neuroimaging.
[0111] In order to preserve the correct geometry in a single modality, embodiments of this application provide a training method for the geometry during the training process.
[0112] First, let's clarify the application scenario. The first sample image is generated for a first target object, and the first sample image includes p-layer image frames for the first target object. The second sample image is generated for a second target object, and the second sample image includes p-layer image frames for the second target object. The first target object and the second target object have the same object type.
[0113] When the target object is a three-dimensional object, the sample image generated for the target object is a three-dimensional image composed of p-layer image frames. Although the first target object and the second target object have the same object type, even if the first target object and the second target object are the same target object, the geometric structure they exhibit in the first mode and the second mode will be different.
[0114] In the process of training the initial synthesis model to obtain the image synthesis model, the method further includes:
[0115] S11: Based on the first geometric structure of the first target object in the first mode, and the first sub-features in the first output features that correspond to the p-layer image frames respectively, determine the first structural difference between the first sub-features for the first geometric structure;
[0116] S12: Based on the second geometric structure of the second target object in the second modality, and the second sub-features in the second output features that correspond to the p-layer image frames respectively, determine the second structural differences between the second sub-features for the second geometric structure;
[0117] S13: Construct the i-th structural difference loss function using the first structural difference and the second structural difference;
[0118] S14: Based on the i-th structural difference loss function, adjust the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer by minimizing the optimization objective of the first structural difference and the second structural difference.
[0119] Since the first structural difference can identify the degree of difference between the geometric structure reflected in the first output feature and the geometric structure of the target object in the first modality, the greater the difference, the greater the loss of important geometric structures when extracting the first output feature. The more difficult it is for the first output feature to accurately reflect the actual geometric structure of the first target object in the first modality. In the scenario of obtaining the second modality image through cross-domain image synthesis using the image of the first modality, the first output feature is the basis for cross-domain image synthesis. When the extracted features cannot accurately reflect the geometric structure of the first target object in the first modality, it will directly affect the effect of cross-domain image synthesis.
[0120] Similarly, the difference in the second structure can identify the degree of difference between the geometric structure reflected in the second output feature and the geometric structure of the target object in the second modality. The greater the difference, the greater the loss of important geometric structures when extracting the second output feature. The more difficult it is for the second output feature to accurately reflect the actual geometric structure of the second target object in the second modality. In the scenario of obtaining the first modality image through cross-domain image synthesis using the image of the second modality, the second output feature is the basis for cross-domain image synthesis. When the extracted features cannot accurately reflect the geometric structure of the second target object in the second modality, it will directly affect the effect of cross-domain image synthesis.
[0121] Adjusting the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer according to the i-th structural difference loss function allows the i-th first convolutional layer and the i-th second convolutional layer to learn how to avoid losing important geometric structural information when extracting the first output feature, thereby improving the accuracy of important information related to geometric structure in the first output feature and the second output feature.
[0122] In one possible implementation, it can be achieved through manifold learning of Laplacian co-regularization (LCR), thereby preserving complementary properties in the geometry.
[0123] S11 includes:
[0124] The first sub-features in the first output features that correspond to the p-layer image frames are regularized to obtain the first regularized sub-features that correspond to the p-layer image frames respectively.
[0125] Based on the first geometric structure of the first target object in the first modality and the first regularized sub-features corresponding to the p-layer image frames respectively, the first structural differences between the first regularized sub-features for the first geometric structure are determined.
[0126] S12 includes:
[0127] The second sub-features in the second output features that correspond to the p-layer image frames are regularized to obtain the second regularized sub-features that correspond to the p-layer image frames respectively.
[0128] Based on the second geometric structure of the second target object in the second modality, and the second regularized sub-features corresponding to the p-layer image frames respectively, the second structural differences between the second regularized sub-features for the second geometric structure are determined.
[0129] After processing the first and second output features using LCR, the geometrically related features carried in the first and second output features can be effectively highlighted, which facilitates the improvement of the accuracy of the determined structural differences and structural difference loss function.
[0130] Structural difference loss function L G The specific formula is as follows:
[0131]
[0132] Where W is the weight, and in order to make it clear in this embodiment, l is used instead of i as the layer index of the first convolutional layer or the second convolutional layer of N layers. Used to identify the first structural difference. Used to identify the second structural difference.
[0133] For the N first and second convolutional layers of the entire initial synthesis model, the above formula can be expressed as:
[0134]
[0135] To further reduce the impact of intra-domain errors and better preserve the geometric structure information in the output features, in one possible implementation, the method further includes, during the training of the initial synthesis model to obtain the image synthesis model:
[0136] S21: Determine the first hierarchical orthogonal matrix corresponding to the first output feature and the second hierarchical orthogonal matrix corresponding to the second output feature based on singular value decomposition;
[0137] S22: Based on the first hierarchical orthogonal matrix and the second hierarchical orthogonal matrix, determine the first offset of the vertex of the geometric structure identified by the first output feature relative to the reference angle, and the second offset of the vertex of the geometric structure identified by the second geometric structure feature relative to the reference angle;
[0138] S23: Construct the i-th subspace mismatch loss function based on the first offset and the second offset;
[0139] S24: Based on the i-th subspace mismatch loss function, adjust the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer by minimizing the optimization objective of the first offset and the second offset.
[0140] Given the heterogeneity of sample images acquired from different manufacturers and image synthesis devices set with different physical parameters, all of these heterogeneities can lead to feature conflicts and inconsistencies, resulting in oversmoothing of cross-domain synthesis and potential intra-domain scaling-based mismatches.
[0141] To reduce generalization error and better preserve the geometric parameters carried by sample images in cross-domain synthesis tasks, this application proposes a subspace mismatch regularizer to constrain in detail the basis of true similarity in the subspace.
[0142] Considering that the characteristic matrix decomposed into singular values can achieve such an effect, the embodiments of this application can use general singular value decomposition (SVD) to obtain hierarchical orthogonal matrices. For example, the first hierarchical orthogonal matrix and the second hierarchical orthogonal matrix in the above steps.
[0143] By using a hierarchical orthogonal matrix, the vertex offset of the geometric structure in the output features can be accurately extracted, such as the first offset or the second offset. This offset can be based on the rotated sample image, or the rotated target object, or it can be based on error or noise.
[0144] The mismatch loss function L for the i-th subspace constructed by the first and second offsets S It can be used as a subspace mismatch penalization (SMP) to enable the initial synthetic model to learn some rotational errors in the domain when trained based on the i-th subspace mismatch loss function, and to better preserve the accurate geometric structure in the output features.
[0145] In summary, in one possible implementation, the embodiments of this application can combine the aforementioned distributed loss function. Correlation loss function Structural difference loss function L G and subspace mismatch loss function L S Effective training of the initial synthesis model, based on the aforementioned loss function, enables it to accurately eliminate intra-domain and cross-domain errors, thereby improving the accuracy of cross-domain synthesis. The overall loss function can be expressed as follows:
[0146]
[0147] The resulting image synthesis model can be named a Transferable Convolutional Sparse Coding Network (TransCSCN). The learned image synthesis model is then applied to a cross-domain synthesis task, given a test image X of a source modality. t The synthesized image of the relevant target mode can be obtained through Y. t =F y Z ty The calculation yields Z, where Z ty ≈PZ tx Ztx =f(X) t ,λ), P is the aforementioned projection matrix, and F is... y Filter parameters.
[0148] like Figure 6 As shown, the first detection layer and the Nth convolutional layer of the first and second feature extraction sub-models in the initial synthesis model are used as examples for illustration.
[0149] exist Figure 6 The example shows three sets of modules, namely the DDM module, which is used to generate the distributed loss function. The LCR module is used to generate the structural difference loss function L. G The SMP module is used to generate the subspace mismatch loss function L. S The correlation loss function is obtained through... As shown.
[0150] The output features corresponding to the first and second convolutional layers of the first layer are processed by the DDM module, LCR module, and SMP module to obtain the first distribution loss function. The first structural difference loss function L G and the first subspace mismatch loss function L S And obtain the first association loss function. Therefore, the layer parameters of the first convolutional layer and the second convolutional layer can be adjusted using the above formula.
[0151] Correspondingly, the output features corresponding to the first and second convolutional layers of the Nth layer are processed by the DDM module, LCR module, and SMP module to obtain the Nth distribution loss function. The Nth structural difference loss function L G and the mismatch loss function L of the Nth subspace S And obtain the Nth association loss function. Therefore, the layer parameters of the first convolutional layer and the second convolutional layer of the Nth layer can be adjusted using the above formula.
[0152] like Figure 7 As shown, the aforementioned distribution loss function is illustrated from the perspective of the feature space. Structural difference loss function L G Subspace mismatch loss function L S The forms of expression.
[0153] exist Figure 7In the feature space on the left, white squares represent feature data from the first output feature corresponding to the first modality, and gray squares represent feature data from the second output feature corresponding to the second modality. The closer the white and gray squares are, the more they indicate that they belong to paired data under the same distribution. For example, in this feature space, the white and gray squares on the right are not paired data, while the other three pairs of squares are paired data. This is achieved through a distributional loss function. Training can make the cross-domain features output by the convolutional layers in the initial synthetic model, which belong to the same distribution, similar to each other.
[0154] exist Figure 7 The lower right portion is used to identify the structural difference loss function L. G The method for determining the structure difference loss function L involves using white dots to represent vertices of the geometric structure in the first mode and gray dots to represent vertices of the geometric structure in the second mode. The structural difference loss function L can be determined by analyzing the geometric structures identified in image frames from different layers. G .
[0155] exist Figure 7 The upper right portion is used to identify the subspace mismatch loss function L. S The determination method, whereby, for the first mode, is exemplarily represented by a white dot, which represents a first offset (as shown by the offset angle) from the reference angle in the subspace; and exemplarily represented by a gray dot, which represents a second offset (as shown by the offset angle) from the reference angle in the subspace. The subspace mismatch loss function L can be determined using the first and second offsets. S .
[0156] In one possible implementation, embodiments of this application provide a method for image synthesis using an image synthesis model. The method includes:
[0157] S201: Obtain the target image of the source modality of the target object.
[0158] S202: Based on the target image, synthesize a synthetic image of the target modality using an image synthesis model.
[0159] In this embodiment, the trained image synthesis model has the ability to synthesize an image of the second modality using an image of the first modality, and also has the ability to synthesize an image of the first modality using an image of the second modality. This image synthesis model is trained from an initial synthesis model using the training method described in the foregoing embodiments.
[0160] Then, the target image of the source modality of the target object can be obtained, and the synthesized image of the target modality can be obtained through image synthesis model. See [link to relevant documentation]. Figure 8This is a schematic diagram of image synthesis provided in an embodiment of this application. The image synthesis model 30 can be configured in a server. The target image can be acquired through a terminal connected to the server, and the synthesized image can be sent to the terminal connected to the server. When the source modality is a first modality, the target modality is a second modality. That is, based on the target image of the first modality, a synthesized image of the second modality can be synthesized through the image synthesis model. Correspondingly, when the source modality is the second modality, the target modality is the first modality. That is, based on the target image of the second modality, a synthesized image of the first modality can be synthesized through the image synthesis model.
[0161] Based on the target image, a synthesized image of the target modality is obtained through an image synthesis model. Specifically, the target image is input into the feature extraction sub-model corresponding to the source modality in the image synthesis model. The target output features of the Nth convolutional layer in the feature extraction sub-model corresponding to the source modality are transformed into the features to be synthesized in the target modality. Then, based on the features to be synthesized, the synthesized image is obtained by synthesizing the Nth convolutional layer in the feature extraction sub-model corresponding to the target modality through the layer parameters.
[0162] Specifically, when the source mode is the first mode and the target mode is the second mode, the target image is input into the first feature extraction sub-model corresponding to the first mode in the image synthesis model. The target output features of the first convolutional layer of the Nth layer in the first feature extraction sub-model are transformed into the features to be synthesized in the second mode. Then, based on the features to be synthesized, the synthesized image is obtained by synthesizing the image through the layer parameters of the second convolutional layer of the Nth layer in the second feature extraction sub-model corresponding to the second mode. For example, if the first mode is mode 4 used to obtain the lattice relaxation time T1 of the brain, and the second mode is mode 1 used to obtain the spin relaxation time T2 of the brain, then the image of mode 1 can be synthesized from the image of mode 4.
[0163] Specifically, when the source mode is the second mode and the target mode is the first mode, the target image is input into the second feature extraction sub-model corresponding to the second mode in the image synthesis model. The target output features of the second convolutional layer of the Nth layer in the second feature extraction sub-model are transformed into the features to be synthesized in the first mode. Then, based on the synthesized features, the synthesized image is obtained by synthesizing the Nth layer of the first convolutional layer in the first feature extraction sub-model corresponding to the first mode.
[0164] Before synthesizing the target modality image through the image synthesis model, the image synthesis model can be validated. The validation of the image synthesis model can be achieved through test sample pairs, which include the third sample image of the first modality and the fourth sample image of the second modality. The third and fourth sample images are generated based on the same target object and are well aligned. The third and fourth sample images are the basis facts and can be used to verify the quality of the synthesis result of the image synthesis model.
[0165] The test sample pairs can be derived from the test sample set, which may include sample images generated by different image generation devices, sample images from different institutions, and images of the same type of target object in different states.
[0166] For example, validation can be performed on two public multimodal brain datasets: IXI and BraTS. Specifically, proton density-weighted (PDw) and T2w MRI scans from the IXI dataset (with significant differences) and T1w and fluid attenuation inversion recovery (FLAIR) acquisitions from the BraTS dataset (with significant differences) were used. Physically, PDw data identified body fluids and fat; T2w data reflected medium-bright fat and bright fluid; T1w data provided good contrast between gray matter (GM) and white matter (WM); FLAIR data showed that GM was brighter than WM, and cerebrospinal fluid (CSF) was dark, not bright. The evaluation was conducted in two parts:
[0167] T2-w images are generated from PD-w data acquired on the IXI dataset, and vice versa;
[0168] FLAIR data is synthesized from T1w images on the BraTS dataset, and vice versa.
[0169] We fixed the number of test cases: 80 for IXI and 45 for BraTS, selecting 60 samples from IXI and 20 samples from BraTS for validation. After discarding half of the data pairs, we constructed fully unsupervised training data using 219 unpaired PDw and T2w MRIs for IXI and 80 unpaired T2w and FLAIR MRIs for BraTS.
[0170] The hyperparameters of TransCSCN were tuned on a validation set. In addition to visual efforts, anatomical accuracy also requires equal attention. To this end, segmentation results from synthetic data were computed and compared to their ground truth.
[0171] Both real scans and synthetic results were input into a segmentation tool, the FMRIB software library (FSL4
[16] ), to segment the main brain tissues into categories (GM, WM, and CSF), and the results were averaged for each brain volume. Tissue prior probability templates were based on average multiple automatic segmentations in the standard spaces from the IXI and BraTS datasets, respectively. Evaluation criteria included peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and Dice score to quantitatively assess the quality of the synthetic results.
[0172] The test and training sample sets can have the same source, and their data do not overlap. For example, from 578 healthy brain images, 80 can be extracted as test samples, and the rest as training samples; from 225 brain images with brain tumors, 45 can be extracted as test samples, and the rest as training samples. Alternatively, from 578 healthy brain images, 60 can be extracted as test samples, and the rest as training samples; from 225 brain images with brain tumors, 20 can be extracted as test samples, and the rest as training samples.
[0173] Using healthy cranial images, the transformation from Mode 2 images (used to obtain PD of the cranium) to Mode 1 images (used to obtain T2 of the cranium), and the transformation from Mode 1 images to Mode 2 images, are examples. The results are compared using PSNR and SSIM. See [link to documentation]. Figure 9 This is a schematic diagram of model verification provided in an embodiment of this application, wherein... Figure 9 A and Figure 9 C shows a test sample pair. Figure 9 A is the input image for mode 1. Figure 9 C is the real image of mode 2 corresponding to the input image. Figure 9 B represents the image of mode 2 synthesized by the image synthesis model. As can be seen from the figure, the image synthesized by the image synthesis model has a high similarity to the real image.
[0174] Of course, this can also be verified by converting images from FLAIR modality 3 (used to obtain images of the brain) to T1 modality 4 (used to obtain images of the brain) in cranial images of a brain tumor state. See [link to relevant documentation]. Figure 10 This is another model verification diagram provided in an embodiment of this application, wherein... Figure 10 A and Figure 10 C shows a test sample pair. Figure 10 A is the input image for modality 3. Figure 10 C is the real image of mode 4 corresponding to the input image. Figure 10 B represents the image of mode 4 synthesized by the image synthesis model. As can be seen from the figure, the image synthesized by the image synthesis model has a high similarity to the real image.
[0175] See Table 1 for an example of a test result provided in an embodiment of this application.
[0176] Table 1. An example of a test result.
[0177]
[0178] The primary evaluation focuses on the visual quality and segmentation performance of synthetic data, presenting quantitative results alongside other findings. The generalizability of TransCSCN is explored through testing on numerous tasks distributed across two independent datasets with consistent properties.
[0179] Specifically, in Figures 9-10 The visual results are presented, including those synthesized using different methods and the corresponding metric measurements.
[0180] Visual measurements are shown as averages of PSNR and SSIM synthesis performance. Observations reveal that the method in this application produces more realistic results with good approximation and better quantitative results. Table 1 presents the summarized performance of TransCSCN and other comparative methods on different datasets across different tasks. The last column of Table 1 shows the performance improvement compared to the worst and best comparative results, respectively. In particular, TransCSCN consistently outperforms all advanced methods and significantly improves performance in PSNR, SSIM, and Dice scores.
[0181] In the foregoing Figures 1-10 Based on the corresponding embodiments, this application provides a device structure diagram for determining an image synthesis model, as shown in the following figure. Figure 11 As shown, the image synthesis model determination device 1100 includes an acquisition unit 1101, an input unit 1102, a determination unit 1103, a construction unit 1104, and a training unit 1105.
[0182] The acquisition unit 1101 is used to acquire training sample pairs, the training sample pairs including a first sample image of a first modality and a second sample image of a second modality;
[0183] The input unit 1102 is used to input the training sample pair into an initial synthesis model. The initial synthesis model includes a first feature extraction sub-model for inputting the first sample image and a second feature extraction sub-model for inputting the second sample image. The first feature extraction sub-model includes N first convolutional layers, and the second feature extraction model includes N second convolutional layers, where N≥1.
[0184] The determining unit 1103 is used to determine the distribution difference in feature distribution between the first output feature of the i-th first convolutional layer and the second output feature of the i-th second convolutional layer, where i is a positive integer less than or equal to N;
[0185] The construction unit 1104 is used to construct the i-th distribution loss function by the first restoration difference of the first output feature relative to the first sample image, the second restoration difference of the second output feature relative to the second sample image, and the distribution difference;
[0186] The training unit 1105 is used to adjust the layer parameters of the first convolutional layer and the second convolutional layer of the i-th layer according to the i-th distribution loss function by minimizing the first restoration difference, the second restoration difference and the distribution difference optimization objective, and train the initial synthesis model to obtain an image synthesis model. The image synthesis model is used to perform cross-modal image synthesis between the first modality and the second modality.
[0187] In one possible implementation, the acquisition unit is further configured to:
[0188] Obtain the first initial output feature of the first convolutional layer of the i-th layer and the second initial output feature of the second convolutional layer of the i-th layer;
[0189] The first initial output feature is mapped to the first image size through global normalization to obtain the first output feature; the second initial output feature is mapped to the second image size through global normalization to obtain the second output feature.
[0190] In one possible implementation, the first sample image is generated for a first target object, the first sample image includes p-layer image frames for the first target object, the second sample image is generated for a second target object, the second sample image includes p-layer image frames for the second target object, and the first target object and the second target object have the same object type;
[0191] The training unit is also used for:
[0192] Based on the first geometric structure of the first target object in the first modality, and the first sub-features in the first output features that correspond to the p-layer image frames respectively, a first structural difference between the first sub-features for the first geometric structure is determined;
[0193] Based on the second geometric structure of the second target object in the second modality, and the second sub-features in the second output features that correspond to the p-layer image frames respectively, the second structural differences between the second sub-features for the second geometric structure are determined;
[0194] The i-th structural difference loss function is constructed using the first structural difference and the second structural difference;
[0195] Based on the i-th structural difference loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the optimization objective of the first structural difference and the second structural difference.
[0196] In one possible implementation, the training unit is further used for:
[0197] The first sub-features in the first output features that correspond to the p-layer image frames are regularized to obtain the first regularized sub-features that correspond to the p-layer image frames respectively.
[0198] Based on the first geometric structure of the first target object in the first modality and the first regularized sub-features corresponding to the p-layer image frames respectively, the first structural differences between the first regularized sub-features for the first geometric structure are determined.
[0199] The second sub-features in the second output features that correspond to the p-layer image frames are regularized to obtain the second regularized sub-features that correspond to the p-layer image frames respectively.
[0200] Based on the second geometric structure of the second target object in the second modality, and the second regularized sub-features corresponding to the p-layer image frames respectively, the second structural differences between the second regularized sub-features for the second geometric structure are determined.
[0201] In one possible implementation, the training unit is further used for:
[0202] The first hierarchical orthogonal matrix corresponding to the first output feature and the second hierarchical orthogonal matrix corresponding to the second output feature are determined based on singular value decomposition.
[0203] Based on the first hierarchical orthogonal matrix and the second hierarchical orthogonal matrix, determine the first offset of the vertices of the geometric structure identified by the first output feature relative to the reference angle, and the second offset of the vertices of the geometric structure identified by the second geometric feature relative to the reference angle;
[0204] Construct the i-th subspace mismatch loss function based on the first offset and the second offset;
[0205] Based on the i-th subspace mismatch loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the optimization objectives of the first offset and the second offset.
[0206] In one possible implementation, when N≥2, the first output feature is the input feature of the (i+1)th first convolutional layer, and the second output feature is the input feature of the (i+1)th second convolutional layer. The feature size of the output feature of the (i+1)th first convolutional layer is smaller than the feature size of the first output feature, and the feature size of the output feature of the (i+1)th second convolutional layer is smaller than the feature size of the second output feature.
[0207] In one possible implementation, the N first convolutional layers and the N second convolutional layers each have a one-to-one correspondence with N distributed loss functions, wherein the i-th distributed loss function among the N distributed loss functions corresponds to the i-th first convolutional layer and the i-th second convolutional layer.
[0208] The training unit is further configured to adjust the layer parameters of the corresponding first and second convolutional layers according to the N distribution loss functions by minimizing the optimization objective of the first restoration difference, the second restoration difference and the distribution difference, so as to train the initial synthesis model to obtain an image synthesis model.
[0209] In one possible implementation, the training unit is further used for:
[0210] Based on the feature difference between the transformation feature corresponding to the first convolutional layer of the i-th layer and the second output feature, an i-th association loss function is constructed, wherein the transformation feature is the feature corresponding to the second modality converted from the first output feature through the corresponding mapping matrix;
[0211] Based on the i-th association loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the feature difference and the optimization objective of the mapping matrix.
[0212] In one possible implementation, the training sample pairs are derived from a training sample set, which includes sample images generated by different image generation devices for the same object type, and the training images in the same modality are generated by at least two image generation devices.
[0213] In one possible implementation, the first sample image and the second sample image are generated based on different target objects.
[0214] In one possible implementation, the apparatus further includes a synthesis unit:
[0215] The acquisition unit is also used to acquire the target image of the source modality of the target object;
[0216] The synthesis unit is used to synthesize a synthetic image of the target modality based on the target image and through the image synthesis model.
[0217] Wherein, when the source mode is the first mode, the target mode is the second mode, or when the source mode is the second mode, the target mode is the first mode.
[0218] In one possible implementation, the synthesis unit is further configured to:
[0219] The target image is input into the feature extraction sub-model corresponding to the source modality in the image synthesis model;
[0220] The target output features of the Nth convolutional layer in the feature extraction sub-model corresponding to the source modality are transformed into the features to be synthesized in the target modality.
[0221] The synthesized image is obtained by synthesizing the Nth convolutional layer in the feature extraction sub-model corresponding to the target modality based on the features to be synthesized.
[0222] Therefore, since the distribution difference can identify the modal difference between the first mode and the second mode, the larger the difference, the more unfavorable it is for subsequent cross-modal image synthesis between the first mode and the second mode through the image synthesis model. On the other hand, the reconstruction difference can identify the difference between the reconstructed image obtained by the corresponding output features and the sample image. The larger the difference, the more difficult it is for the output features used by the initial synthesis model for cross-modal image synthesis to reflect the accurate information of the source mode, and instead carry too much noise information. Therefore, the i-th distribution loss function is constructed through the distribution difference and each restoration difference. Based on the i-th distribution loss function, the layer parameters of the first convolutional layer and the second convolutional layer of the i-th layer are adjusted by minimizing the optimization objective of each restoration difference and the distribution difference, thereby training the initial synthesis model to obtain the image synthesis model. Since the image synthesis model learns the modal differences between the first and second modalities and improves the feature extraction accuracy through the above training, the transferability of effective features in cross-modal synthesis is enhanced. This enables the image synthesis model to accurately convert the image features of the target modal image based on the output features of the source modal image, synthesize an image that is closer to the real target modal, and improve the image acquisition efficiency in multimodal scenarios.
[0223] This application also provides a computer device, which is the computer device described above, and may include a terminal device or a server. The aforementioned image synthesis model determination device may be configured in this computer device. The computer device will now be described in conjunction with the accompanying drawings.
[0224] If the computer device is a terminal device, please refer to Figure 12 As shown, this application provides a terminal device, taking a mobile phone as an example:
[0225] Figure 12 This diagram illustrates a partial structural representation of a mobile phone related to the terminal device provided in this embodiment. (Reference) Figure 12 The mobile phone includes components such as a radio frequency (RF) circuit 1410, a memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a Wi-Fi module 1470, a processor 1480, and a power supply 1490. Those skilled in the art will understand that... Figure 12 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0226] The following is combined with Figure 12 A detailed introduction to each component of a mobile phone:
[0227] The RF circuit 1410 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1480; in addition, it transmits uplink data to the base station.
[0228] The memory 1420 can be used to store software programs and modules, and the processor 1480 executes various functions and data processing of the mobile phone by running the software programs and modules stored in the memory 1420.
[0229] The input unit 1430 may include a touch panel 1431 and other input devices 1432.
[0230] The display unit 1440 may include a display panel 1441.
[0231] The mobile phone may also include at least one sensor 1450, such as a light sensor, a motion sensor, and other sensors.
[0232] Audio circuitry 1460, speaker 1461, and microphone 1462 provide an audio interface between the user and the mobile phone.
[0233] The processor 1480 is the control center of the mobile phone. It connects to various parts of the mobile phone through various interfaces and lines. It performs various functions of the mobile phone and processes data by running or executing software programs and / or modules stored in the memory 1420 and calling data stored in the memory 1420.
[0234] In this embodiment, the processor 1480 included in the terminal device also has the following functions:
[0235] Obtain training sample pairs, wherein the training sample pairs include a first sample image of a first modality and a second sample image of a second modality;
[0236] The training sample pairs are input into the initial synthesis model. The initial synthesis model includes a first feature extraction sub-model for inputting the first sample image and a second feature extraction sub-model for inputting the second sample image. The first feature extraction sub-model includes N first convolutional layers, and the second feature extraction model includes N second convolutional layers, where N≥1.
[0237] Determine the distribution difference in feature distribution between the first output feature of the first convolutional layer of the i-th layer and the second output feature of the second convolutional layer of the i-th layer, wherein the first output feature is determined based on the first sample image and the second output feature is determined based on the second sample image, and i is a positive integer less than or equal to N;
[0238] The first restoration difference corresponding to the first output feature, the second restoration difference corresponding to the second output feature, and the distribution difference are used to construct the i-th distribution loss function. The first restoration difference is used to identify the difference between the restored image obtained by the first output feature and the first sample image, and the second restoration difference is used to identify the difference between the restored image obtained by the second output feature and the second sample image.
[0239] Based on the i-th distribution loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the first restoration difference, the second restoration difference, and the distribution difference. The initial synthesis model is then trained to obtain an image synthesis model, which is used for cross-modal image synthesis between the first modality and the second modality.
[0240] If the computer device is a server, this application embodiment also provides a server; please refer to [link to relevant documentation]. Figure 13 As shown, Figure 13This is a structural diagram of a server 1500 provided in an embodiment of this application. The server 1500 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1522 (e.g., one or more processors) and a memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. The memory 1532 and storage media 1530 can be temporary or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 1522 may be configured to communicate with the storage media 1530 and execute the series of instruction operations in the storage media 1530 on the server 1500.
[0241] Server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0242] The steps performed by the server in the above embodiments can be based on Figure 13 The server structure shown.
[0243] In addition, this application embodiment also provides a storage medium for storing a computer program for executing the method provided in the above embodiment.
[0244] This application also provides a computer program product including instructions that, when run on a computer, cause the computer to perform the methods provided in the above embodiments.
[0245] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, or optical disk, etc., and other media capable of storing program code.
[0246] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Moreover, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for determining an image synthesis model, characterized in that, The method includes: Obtain training sample pairs, wherein the training sample pairs include a first sample image of a first modality and a second sample image of a second modality; The training sample pairs are input into the initial synthesis model. The initial synthesis model includes a first feature extraction sub-model for inputting the first sample image and a second feature extraction sub-model for inputting the second sample image. The first feature extraction sub-model includes N first convolutional layers, and the second feature extraction model includes N second convolutional layers, where N≥1. Determine the distribution difference in feature distribution between the first output feature of the first convolutional layer of the i-th layer and the second output feature of the second convolutional layer of the i-th layer, where i is a positive integer less than or equal to N; The i-th distribution loss function is constructed by the first restoration difference of the first output feature relative to the first sample image, the second restoration difference of the second output feature relative to the second sample image, and the distribution difference; Based on the i-th distribution loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the first restoration difference, the second restoration difference, and the distribution difference. The initial synthesis model is then trained to obtain an image synthesis model, which is used for cross-modal image synthesis between the first modality and the second modality.
2. The method according to claim 1, characterized in that, Before determining the distribution difference in feature distribution between the first output feature of the i-th first convolutional layer and the second output feature of the i-th second convolutional layer, the method further includes: Obtain the first initial output feature of the first convolutional layer of the i-th layer and the second initial output feature of the second convolutional layer of the i-th layer; The first initial output feature is mapped to the first image size through global normalization to obtain the first output feature; the second initial output feature is mapped to the second image size through global normalization to obtain the second output feature.
3. The method according to claim 1, characterized in that, The first sample image is generated for a first target object, and the first sample image includes p-layer image frames for the first target object. The second sample image is generated for a second target object, and the second sample image includes p-layer image frames for the second target object. The first target object and the second target object have the same object type. In the process of training the initial synthesis model to obtain the image synthesis model, the method further includes: Based on the first geometric structure of the first target object in the first modality, and the first sub-features in the first output features that correspond to the p-layer image frames respectively, a first structural difference between the first sub-features for the first geometric structure is determined; Based on the second geometric structure of the second target object in the second modality, and the second sub-features in the second output features that correspond to the p-layer image frames respectively, the second structural differences between the second sub-features for the second geometric structure are determined; The i-th structural difference loss function is constructed using the first structural difference and the second structural difference; Based on the i-th structural difference loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the optimization objective of the first structural difference and the second structural difference.
4. The method according to claim 3, characterized in that, The step of determining the first structural difference between the first sub-features for the first geometric structure based on the first geometric structure of the first target object in the first modality and the first sub-features in the first output features corresponding to the p-layer image frames respectively includes: The first sub-features in the first output features that correspond to the p-layer image frames are regularized to obtain the first regularized sub-features that correspond to the p-layer image frames respectively. Based on the first geometric structure of the first target object in the first modality and the first regularized sub-features corresponding to the p-layer image frames respectively, the first structural differences between the first regularized sub-features for the first geometric structure are determined. The step of determining the second structural differences between the second sub-features for the second geometric structure based on the second geometric structure of the second target object in the second modality and the second sub-features in the second output features corresponding to the p-layer image frames respectively includes: The second sub-features in the second output features that correspond to the p-layer image frames are regularized to obtain the second regularized sub-features that correspond to the p-layer image frames respectively. Based on the second geometric structure of the second target object in the second modality, and the second regularized sub-features corresponding to the p-layer image frames respectively, the second structural differences between the second regularized sub-features for the second geometric structure are determined.
5. The method according to claim 1, characterized in that, In the process of training the initial synthesis model to obtain the image synthesis model, the method further includes: The first hierarchical orthogonal matrix corresponding to the first output feature and the second hierarchical orthogonal matrix corresponding to the second output feature are determined based on singular value decomposition. Based on the first hierarchical orthogonal matrix and the second hierarchical orthogonal matrix, determine the first offset of the vertices of the geometric structure identified by the first output feature relative to the reference angle, and the second offset of the vertices of the geometric structure identified by the second output feature relative to the reference angle; Construct the i-th subspace mismatch loss function based on the first offset and the second offset; Based on the i-th subspace mismatch loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the optimization objectives of the first offset and the second offset.
6. The method according to claim 1, characterized in that, When N≥2, the first output feature is the input feature of the (i+1)th first convolutional layer, and the second output feature is the input feature of the (i+1)th second convolutional layer. The feature size of the output feature of the (i+1)th first convolutional layer is smaller than the feature size of the first output feature, and the feature size of the output feature of the (i+1)th second convolutional layer is smaller than the feature size of the second output feature.
7. The method according to claim 6, characterized in that, The N first convolutional layers and the N second convolutional layers each have a one-to-one correspondence with N distributed loss functions, wherein the i-th distributed loss function among the N distributed loss functions corresponds to the i-th first convolutional layer and the i-th second convolutional layer; The step of adjusting the layer parameters of the first and second convolutional layers of the i-th layer according to the i-th distribution loss function by minimizing the first restoration difference, the second restoration difference, and the distribution difference, and training the initial synthesis model to obtain the image synthesis model, includes: Based on the N distribution loss functions, the layer parameters of the corresponding first and second convolutional layers are adjusted by minimizing the optimization objective of the first restoration difference, the second restoration difference, and the distribution difference, and the initial synthesis model is trained to obtain the image synthesis model.
8. The method according to any one of claims 1-7, characterized in that, In the process of training the initial synthesis model to obtain the image synthesis model, the method further includes: Based on the feature difference between the transformation feature corresponding to the first convolutional layer of the i-th layer and the second output feature, an i-th association loss function is constructed, wherein the transformation feature is the feature corresponding to the second modality converted from the first output feature through the corresponding mapping matrix; Based on the i-th association loss function, the layer parameters of the i-th first convolutional layer and the i-th second convolutional layer are adjusted by minimizing the feature difference and the optimization objective of the mapping matrix.
9. The method according to claim 1, characterized in that, The method further includes: Obtain the target image of the source modality of the target object; Based on the target image, a synthesized image of the target modality is obtained through the image synthesis model; Wherein, when the source mode is the first mode, the target mode is the second mode, or when the source mode is the second mode, the target mode is the first mode.
10. A device for determining an image synthesis model, characterized in that, The device includes an acquisition unit, an input unit, a determination unit, a construction unit, and a training unit: The acquisition unit is used to acquire training sample pairs, the training sample pairs including a first sample image of a first modality and a second sample image of a second modality; The input unit is used to input the training sample pairs into the initial synthesis model. The initial synthesis model includes a first feature extraction sub-model for inputting the first sample image and a second feature extraction sub-model for inputting the second sample image. The first feature extraction sub-model includes N first convolutional layers, and the second feature extraction model includes N second convolutional layers, where N≥1. The determining unit is used to determine the distribution difference in feature distribution between the first output feature of the i-th first convolutional layer and the second output feature of the i-th second convolutional layer, where i is a positive integer less than or equal to N; The construction unit is used to construct the i-th distribution loss function by the first restoration difference of the first output feature relative to the first sample image, the second restoration difference of the second output feature relative to the second sample image, and the distribution difference; The training unit is used to adjust the layer parameters of the first convolutional layer and the second convolutional layer of the i-th layer according to the i-th distribution loss function by minimizing the first restoration difference, the second restoration difference and the distribution difference optimization objective, and train the initial synthesis model to obtain an image synthesis model. The image synthesis model is used to perform cross-modal image synthesis between the first modality and the second modality.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the method according to any one of claims 1-9.
12. A computer program product comprising instructions that, when run on a computer, cause the computer to perform the method of any one of claims 1-9.
Citation Information
Patent Citations
Cross-modal image synthesis method
CN113012086A
Medical image cross-modal synthesis system and method based on multi-source confrontation strategy
CN114387481A