Training method of noise determination model, and medical image generation method and device
By training a noise determination model and performing multiple noise additions using knowledge from multiple tissues and medical images, target medical images from different perspectives are generated. This solves the problem of difficulty in acquiring multi-view medical images in existing technologies and improves the accuracy of analysis results.
Patent Information
- Application Number
- CN202410631961.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-11-21
AI Technical Summary
Current technologies struggle to acquire multi-view medical images, impacting the accuracy of disease diagnosis and treatment planning.
By training a noise determination model, utilizing multiple tissue knowledge and medical images, and performing multiple noise addition processes, the denoised data is determined based on sample knowledge and the feature data after noise addition. The neural network model is then trained to generate target medical images from different perspectives.
It improves the visual richness of medical images, enhances the accuracy of analysis results, and improves the accuracy of disease diagnosis and treatment planning.
Smart Images

Figure CN120997070A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a training method for a noise determination model, a medical image generation method, and an apparatus. Background Technology
[0002] Medical images are images acquired using techniques such as computed tomography (CT) or magnetic resonance imaging (MRI) to visualize the internal structure of an organism or a part of an organism. Because medical images provide rich visual information, they are often analyzed using electronic devices to aid doctors in disease diagnosis and treatment planning. The accuracy of the analysis results is closely related to the richness of the medical images' perspectives; therefore, acquiring multi-view medical images has become a pressing issue. Summary of the Invention
[0003] This application provides a method for training a noise determination model, a method for generating medical images, and an apparatus for doing so. The method can train a noise determination model, accurately predict the noise data that needs to be removed, and then remove the noise data to obtain a target medical image with a different perspective from the source medical image. The technical solution includes the following:
[0004] Firstly, a training method for a noise determination model is provided, the method comprising:
[0005] Acquire multiple tissue knowledge, a first medical image, and a second medical image. Each tissue knowledge represents the visual information of a tissue structure from multiple perspectives, and the second medical image is an image of the tissue structure in the first medical image from different perspectives.
[0006] Based on the first medical image, sample knowledge is determined from the plurality of tissue knowledge, wherein the sample knowledge characterizes the visual information of the tissue structure in the first medical image from multiple perspectives;
[0007] At least one noise addition is performed on the image features of the second medical image to obtain the feature data after each noise addition.
[0008] Based on the sample knowledge and the feature data after each noise addition, the neural network model determines the first denoised data corresponding to each noise addition, and the first denoised data corresponding to any noise addition is used to denoise the feature data after any noise addition.
[0009] The neural network model is trained based on the first denoised data corresponding to each denoising step to obtain a noise determination model. The noise determination model is used to determine the target noise data, and the target noise data is used to determine the target medical image of the tissue structure in the source medical image from different perspectives.
[0010] Secondly, a method for generating medical images is provided, the method comprising:
[0011] Acquire multiple tissue knowledge and source medical images, where any tissue knowledge represents the visual information of an tissue structure from multiple perspectives;
[0012] Source knowledge is determined from the plurality of tissue knowledge based on the source medical image, and the source knowledge represents the visual information of the tissue structure in the source medical image from multiple perspectives;
[0013] The target noise data is determined by a noise determination model based on the source knowledge, and the noise determination model is trained according to the method shown in the first aspect.
[0014] The target medical image is obtained by denoising the reference noise data based on the target noise data. The target medical image is an image of the tissue structure in the source medical image from different perspectives.
[0015] Thirdly, a training device for a noise determination model is provided, the device comprising:
[0016] The acquisition module is used to acquire multiple tissue knowledge, a first medical image, and a second medical image. Each tissue knowledge represents the visual information of a tissue structure from multiple perspectives, and the second medical image is an image of the tissue structure in the first medical image from different perspectives.
[0017] A determination module is used to determine sample knowledge from the plurality of tissue knowledge based on the first medical image, wherein the sample knowledge characterizes the visual information of the tissue structure in the first medical image from multiple perspectives;
[0018] The noise-adding module is used to perform at least one noise-adding operation on the image features of the second medical image to obtain the feature data after each noise-adding operation.
[0019] The determining module is further configured to determine the first denoised data corresponding to each denoising step based on the sample knowledge and the feature data after each denoising step using a neural network model, wherein the first denoised data corresponding to any denoising step is used to denoise the feature data after any denoising step.
[0020] The training module is used to train the neural network model based on the first denoised data corresponding to each denoising step to obtain a noise determination model. The noise determination model is used to determine the target noise data, and the target noise data is used to determine the target medical image of the tissue structure in the source medical image from different perspectives.
[0021] In one possible implementation, the acquisition module is configured to acquire multiple image sets, each image set including reference medical images of the tissue structure of the same object from multiple perspectives, the reference medical images including multiple reference image blocks; cluster the features of each reference image block to obtain multiple clusters, each cluster including features of at least one reference image block; and determine each tissue knowledge based on each cluster.
[0022] In one possible implementation, the features of the reference image patch are extracted via a first extraction network; the apparatus further includes:
[0023] The extraction module is used to extract features from the third medical image through the first neural network;
[0024] A reconstruction module is used to reconstruct a fourth medical image based on the features of the third medical image;
[0025] The training module is further configured to train the first neural network to obtain the first extraction network based on at least one of a first loss, a second loss, or a third loss.
[0026] Wherein, the first loss characterizes the pixel difference between the third medical image and the fourth medical image, the second loss characterizes the distribution difference between the third medical image and the fourth medical image, and the third loss characterizes the feature difference between the third medical image and the fourth medical image.
[0027] In one possible implementation, the first medical image includes a plurality of first image blocks;
[0028] The determining module is configured to, for any first image block, determine the distance between the features of the first image block and each organizational knowledge, and take the organizational knowledge corresponding to the distance that satisfies the distance condition as the organizational knowledge of the first image block; and determine the sample knowledge based on the organizational knowledge of each first image block.
[0029] In one possible implementation, noise is added multiple times;
[0030] The noise-adding module is used to add noise to the image features of the second medical image for the first noise addition, to obtain the feature data after the first noise addition; and to add noise to the feature data after the previous noise addition for non-first noise addition, to obtain the feature data after non-first noise addition.
[0031] In one possible implementation, the determining module is configured to acquire at least one of difference information, task description, or sample medical report, wherein the difference information characterizes the difference between the sample medical report and the first medical image, the task description describes the task of generating medical images from different perspectives based on the first medical image, and the sample medical report describes the diagnostic content of the first medical image and the second medical image; and the first denoised data corresponding to each denoising step is determined by a neural network model based on the difference information, the task description, or at least one of the sample medical report, the sample knowledge, and the feature data after each denoising step.
[0032] In one possible implementation, the determining module is configured to determine a first probability that the sample medical report belongs to multiple categories and a second probability that the first medical image belongs to the multiple categories; for any category, determine the difference between the first probability and the second probability of the any category; determine a target category whose difference is not less than a reference difference; and determine the difference information based on the target category.
[0033] In one possible implementation, the neural network model includes a first network part and a second network part;
[0034] The determining module is configured to, for any instance of noise addition, determine first information based on the first medical image, the sample knowledge, and the feature data after any instance of noise addition by the first network portion, wherein the first information represents information after the fusion of the first medical image, the sample knowledge, and the feature data after any instance of noise addition; and determine first denoised data corresponding to the instance of noise addition by the second network portion based on the first information and the feature data after any instance of noise addition.
[0035] In one possible implementation, the determining module is configured to determine second information based on the feature data after any one instance of noise addition by the second network portion, the second information representing the feature data after any one instance of noise addition; and determine first denoised data corresponding to the any one instance of noise addition by the second network portion based on the first information and the second information.
[0036] In one possible implementation, the determining module is further configured to determine, through the third network part, a second denoised data corresponding to any one denoising based on the feature data after any one denoising, and the second denoised data corresponding to any one denoising is used to denoise the feature data after any one denoising.
[0037] The training module is also used to train the third network part based on the second denoised data corresponding to each noise addition, so as to obtain the second network part.
[0038] In one possible implementation, the training module is used to determine the loss of any noise addition based on the first denoised data and the added noise data corresponding to the noise addition, wherein the added noise data corresponding to the noise addition is used to add noise to obtain the feature data after the noise addition; and to train the neural network model based on the loss of each noise addition to obtain a noise determination model.
[0039] Fourthly, a medical image generation apparatus is provided, the apparatus comprising:
[0040] The acquisition module is used to acquire multiple tissue knowledge and source medical images, where any tissue knowledge represents the visual information of an tissue structure from multiple perspectives.
[0041] A determination module is used to determine source knowledge from the plurality of tissue knowledge based on the source medical image, wherein the source knowledge characterizes the visual information of the tissue structure in the source medical image from multiple perspectives;
[0042] The determining module is further configured to determine the target noise data based on the source knowledge using a noise determination model, wherein the noise determination model is trained according to the method shown in the first aspect;
[0043] A denoising module is used to denoise the reference noise data based on the target noise data to obtain a target medical image, wherein the target medical image is an image of the tissue structure in the source medical image from different perspectives.
[0044] In one possible implementation, the noise reduction process is repeated multiple times.
[0045] The determining module is used to determine the target noise data for the first denoising based on the source knowledge and the reference noise data using a noise determination model for the first denoising.
[0046] The denoising module is used to denoise the reference noise data based on the target noise data of the first denoising, and obtain the feature data after the first denoising.
[0047] The determining module is used to determine the target noise data for non-first-time denoising based on the source knowledge and the feature data after the previous denoising of the non-first-time denoising, using the noise determination model.
[0048] The denoising module is used to denoise the previously denoised feature data based on the target noise data after the non-first denoising, to obtain the non-first denoised feature data; and to determine the target medical image based on the feature data after the last denoising.
[0049] Fifthly, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to implement the methods described in the first or second aspect above.
[0050] In a sixth aspect, a computer-readable storage medium is also provided, wherein at least one computer program is stored therein, the at least one computer program being loaded and executed by a processor to enable an electronic device to implement the methods shown in the first or second aspect above.
[0051] In a seventh aspect, a computer program is also provided, said computer program being at least one, which is loaded and executed by a processor to enable an electronic device to implement the methods shown in the first or second aspect above.
[0052] Eighthly, a computer program product is also provided, wherein at least one computer program is stored therein, the at least one computer program being loaded and executed by a processor to enable an electronic device to implement the methods shown in the first or second aspect above.
[0053] The technical solution provided in this application brings at least the following beneficial effects:
[0054] The technical solution provided in this application achieves the acquisition of multi-view visual information of various tissue structures in the first medical image by determining sample knowledge of the first medical image from multiple tissue knowledge sources. Since the first and second medical images have different viewpoints, the image features of the second medical image are denoised multiple times. Based on the sample knowledge and the feature data after each denoising step, the first denoised data corresponding to that denoising step is determined. This enables the model to accurately predict the noise added to the image features at a certain viewpoint based on multi-view visual information, improving the accuracy of the first denoised data. Training the neural network model with the first denoised data corresponding to each denoising step improves the training effect and accuracy of the model. This allows the trained noise determination model to accurately predict the noise data that needs to be removed. By removing this noise data, medical images from different viewpoints are obtained, increasing the viewpoint richness of the medical images and thus improving the accuracy of the analysis results obtained from analyzing the medical images. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 This is a schematic diagram of a computer system for a noise determination model training method or a medical image generation method provided in an embodiment of this application;
[0057] Figure 2 This is a flowchart of a training method for a noise determination model provided in an embodiment of this application;
[0058] Figure 3 This is a schematic diagram illustrating the determination of multiple organizational knowledge provided in an embodiment of this application;
[0059] Figure 4 This is a schematic diagram illustrating the determination of sample knowledge provided in an embodiment of this application;
[0060] Figure 5 This is a schematic diagram illustrating the determination of text content provided in an embodiment of this application;
[0061] Figure 6 This is a schematic diagram illustrating the determination of first denoised data provided in an embodiment of this application;
[0062] Figure 7 This is a schematic diagram of the structure of a first network portion and a second network portion provided in an embodiment of this application;
[0063] Figure 8 This is a flowchart of a medical image generation method provided in an embodiment of this application;
[0064] Figure 9 This is a schematic diagram illustrating the generation of a target medical image provided in an embodiment of this application;
[0065] Figure 10 This is a schematic diagram illustrating disease diagnosis based on source medical images and target medical images, provided in an embodiment of this application.
[0066] Figure 11 This is a flowchart illustrating the training process of a noise determination model provided in an embodiment of this application.
[0067] Figure 12 This is a training diagram of another noise determination model provided in an embodiment of this application;
[0068] Figure 13 This is a schematic diagram of model training provided in an embodiment of this application;
[0069] Figure 14 This is a schematic diagram illustrating the generation of a target X-ray image provided in an embodiment of this application;
[0070] Figure 15This is a comparative schematic diagram of a source X-ray film, a target X-ray film, and X-ray films generated by various models provided in an embodiment of this application;
[0071] Figure 16 This is a schematic diagram of the structure of a training device for a noise determination model provided in an embodiment of this application;
[0072] Figure 17 This is a schematic diagram of the structure of a medical image generation device provided in an embodiment of this application;
[0073] Figure 18 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;
[0074] Figure 19 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0075] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0076] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0077] Medical images are images acquired using technologies such as CT imaging or MRI to visualize the internal structure of an organism or a part of an organism. Because medical images provide rich visual information, they are often analyzed using electronic devices to assist doctors in disease diagnosis and treatment planning. In the medical field, the perspectives of medical images include AP (Anterior Posterio View), PA (Posteroanterior View), LAT (Lateral View), LL (Left Lateral View), RL (Right Lateral View), LPO (Left Anterior Oblique View), and RPO (Right Anterior Oblique View), and the accuracy of the analysis results is closely related to the richness of the medical image's perspectives.
[0078] Based on this, embodiments of this application provide a training method for a noise determination model or a medical image generation method. The noise determination model can accurately predict the noise data that needs to be removed. By removing the noise data, a target medical image with a different perspective from the source medical image is obtained. By enriching the perspective of the medical image, the accuracy of the analysis results is improved.
[0079] like Figure 1 As shown, Figure 1 This is a schematic diagram of a computer system for a noise determination model training method or a medical image generation method provided in this application embodiment. The computer system includes a terminal device 101 and a server 102. The noise determination model training method or medical image generation method provided in this application embodiment can be executed by the terminal device 101, by the server 102, or by both the terminal device 101 and the server 102. This application embodiment does not limit the execution of this method.
[0080] Terminal device 101 has a client installed and running for displaying medical images, and server 102 provides background services for this client. In one possible implementation, server 102 performs the primary computational work, and terminal device 101 performs secondary computational work. Alternatively, server 102 performs secondary computational work, and terminal device 101 performs the primary computational work. Or, terminal device 101 and server 102 collaborate on computation using a distributed computing architecture.
[0081] Optionally, the terminal device 101 can be any electronic device product capable of human-computer interaction with the user through one or more methods such as a keyboard, touchpad, remote control, voice interaction, or handwriting device. For example, the terminal device 101 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, PC (Personal Computer), mobile phone, PDA (Personal Digital Assistant), wearable device, PPC (Pocket PC), smart car system, smart TV, etc.
[0082] Terminal device 101 can refer to one of a plurality of terminal devices. This embodiment uses terminal device 101 as an example. Those skilled in the art will know that the number of terminal devices 101 can be more or less. For example, there may be only one terminal device 101, or there may be dozens or hundreds of terminal devices 101, or more. This application embodiment does not limit the number or type of terminal devices 101.
[0083] Server 102 can be a single server, a server cluster consisting of multiple servers, or any of the following: a cloud computing platform or a virtualization center. This embodiment of the application does not limit this. Server 102 communicates directly or indirectly with terminal device 101 via a wired or wireless network. Server 102 has data receiving, data processing, and data sending functions. Of course, server 102 may also have other functions, which are not limited in this embodiment of the application.
[0084] In an exemplary embodiment of this application, the target object 103 can interact with the terminal device 101 via an interactive interface. In one possible implementation, the target object 103 inputs the source medical image 104 into the terminal device 101 through the interactive interface, and the terminal device 101 sends the source medical image 104 to the server 102 via a wireless network or a wired network. The server 102 has pre-configured the interface according to the specified parameters. Figure 2 An embodiment of the training method for the relevant noise determination model is provided, in which the noise determination model is trained based on a first medical image and a second medical image. Figure 1 The diagram shows a first medical image from a PA (Parallel Perspective) viewpoint and a second medical image from a LAT (Local Atmosphere) viewpoint. In practical applications, the first medical image can be any viewpoint, and the second medical image can be any viewpoint different from the first medical image, as long as the viewpoints of the first and second medical images are different. After the server 102 receives the source medical image 104, it determines the noise determination model based on the source medical image 104, according to... Figure 8An embodiment of a related medical image generation method identifies and outputs target noise data, and then determines a target medical image based on the target noise data. Figure 1 The image shows the source medical image from a PA (Parallel View) perspective and the target medical image from a LAT (Local Atmosphere View) perspective. In practical applications, the source medical image can be any viewpoint, and the target medical image can be any viewpoint different from the source medical image. Server 102 sends the target medical image to terminal device 101 via a wireless or wired network. Terminal device 101 displays the source medical image 104 and the target medical image 105 through an interactive interface.
[0085] Those skilled in the art should understand that the terminal device 101 and server 102 described above are merely illustrative examples. Other existing or future terminal devices or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0086] The various optional embodiments of this application can be implemented based on artificial intelligence (AI) technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making functions.
[0087] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0088] like Figure 2 As shown, Figure 2 This is a flowchart illustrating a training method for a noise determination model provided in this application embodiment. The terminal device 101 or server 102 executing the method of this application embodiment can be collectively referred to as an electronic device; that is, the method of this application embodiment can be executed by an electronic device. Figure 2 As shown, the method includes the following steps.
[0089] Step 201: Obtain multiple tissue knowledge, a first medical image, and a second medical image. Each tissue knowledge represents the visual information of an tissue structure from multiple perspectives, and the second medical image is an image of the tissue structure in the first medical image from different perspectives.
[0090] In this embodiment, the multiple perspectives of any tissue structure include multiple aspects of the AP, PA, LAT, LL, RL, LPO, and RPO perspectives mentioned above. The first medical image is an image of at least one tissue structure of an object from any one of the multiple perspectives, and the second medical image is an image of at least one tissue structure of the same object from another perspective among the multiple perspectives. That is, the first medical image and the second medical image are medical images of the same object's tissue structure from different perspectives, and can be considered as an image pair. In practical applications, there are multiple image pairs, and any two image pairs can correspond to the same or different objects.
[0091] The electronic device can acquire a first medical image and a second medical image, and the acquisition method is not limited herein. For example, the electronic device is connected to an acquisition device for acquiring medical images, which acquires at least one of the first or second medical images in real time and transmits it to the electronic device. Alternatively, the electronic device can acquire at least one of the first or second medical images input by a user. Alternatively, the electronic device can acquire at least one of the first or second medical images via the Internet.
[0092] Furthermore, electronic devices can acquire multiple tissue knowledge sets. Each tissue knowledge set corresponds to a specific tissue structure, and any two tissue knowledge sets correspond to different tissue structures. A tissue structure is a part of an organism divided according to its structure, function, etc. Any tissue structure can be any of the following: thoracic cage, ribs, lungs, heart, pleura, liver, kidneys, thyroid gland, spleen, spine, etc. In practical applications, tissue structures can be parts divided into finer-grained or coarser-grained sets. For example, a tissue structure can include the thoracic cage and pleura; similarly, the left and right atria in the heart can be considered two separate tissue structures. Each tissue knowledge set is a set of feature vectors used to represent the visual information of a tissue structure from multiple perspectives. This visual information includes location, texture, color, size, etc. For example, lung tissue knowledge includes feature vectors from three RGB (Red-Green-Blue) channels, which represent the color of the lung from multiple perspectives.
[0093] The method by which the electronic device acquires organizational knowledge is not limited here. For example, the electronic device can acquire any organizational knowledge input by the user. Alternatively, the electronic device can acquire any organizational knowledge via the Internet. Or, the electronic device can acquire images of any organizational structure from multiple perspectives, and obtain the corresponding organizational knowledge by extracting features from these images. Alternatively, multiple pieces of organizational knowledge can be acquired by following steps 2011 to 2013 (not shown in the figure).
[0094] Step 2011: Obtain multiple image sets, each image set including reference medical images of the tissue structure of the same object from multiple perspectives, and each reference medical image including multiple reference image blocks.
[0095] In this embodiment, any image set includes multiple reference medical images, and any two reference medical images correspond to the tissue structure of the same object from different perspectives. There are multiple image sets, and any two image sets can correspond to the same or different objects, and any two image sets can include the same or different numbers of reference medical images. The acquisition methods for the reference medical images are similar to those for the first and second medical images, and will not be repeated here. Any two reference medical images belonging to the same image set can be used as the first medical image and the second medical image, respectively. Each reference medical image includes multiple reference image blocks, and any two reference image blocks can have the same or different sizes.
[0096] Step 2012: Cluster the features of each reference image block to obtain multiple clusters, and each cluster includes features of at least one reference image block.
[0097] In this embodiment, any reference medical image can be input into a first extraction network to extract features from the reference medical image, and the features of the reference medical image can be divided into features of multiple reference image blocks. Alternatively, any reference medical image can be divided into multiple reference image blocks, and each reference image block can be input into the first extraction network to extract features from each reference image block.
[0098] Optionally, the first extraction network f is used. e Mapping the reference medical image x to the latent space yields the features z of the reference medical image. z = f e (x)∈R h×w×c In this matrix, h, w, and c represent the height, width, and number of channels of the feature, respectively, and R represents a real number. That is, z is an h×w×c matrix composed of real numbers. Next, z is divided into multiple h′×w′×c feature blocks, where each feature block represents a feature of a reference image block.
[0099] The first extraction network includes at least one of the following: linear layers, convolutional layers, normalization layers, feedforward layers, and attention layers. For example, the first extraction network is the encoder in a VAE (Variational Auto Encoder) or AE (Auto Encoder). Understandably, different structures of the first extraction network result in different feature extraction methods, which will not be elaborated upon here.
[0100] Since the features of the reference image patch are extracted through the first extraction network, the electronic device needs to train the first extraction network first. In this case, before step 2022, the method further includes: extracting features of the third medical image through the first neural network; reconstructing the fourth medical image based on the features of the third medical image; and training the first neural network based on at least one of the first loss, the second loss, or the third loss to obtain the first extraction network.
[0101] In this embodiment, there are multiple third medical images. Any one of the third medical images can be a first medical image, a second medical image, a reference medical image, or other medical images. The method of acquiring the third medical image is similar to that of acquiring the first and second medical images, and will not be described again here. The third medical image can be input into a first neural network. The structure and function of the first neural network are similar to those of the first extraction network. Therefore, features of the third medical image can be extracted through the first neural network. The method of feature extraction will not be described again here.
[0102] The features of a third medical image can be input into a second neural network, which then reconstructs a fourth medical image based on these features. The second extraction network includes at least one of the following: linear layers, convolutional layers, normalization layers, feedforward layers, and attention layers. For example, the second extraction network can be a decoder in a VAE (Variational Auto Encoder) or AE (Auto Encoder). Understandably, different structures of the second extraction network will result in different image reconstruction methods, which will not be elaborated upon here.
[0103] The first loss can be determined based on the third and fourth medical images using formulas such as mean absolute error (also called L1 loss) or root mean square error. The first loss characterizes the pixel differences between the third and fourth medical images. Optionally, the third and fourth medical images are of the same size, each including multiple rows and columns of pixels. For any pixel in any row and column, the difference between the pixel value in the third medical image and the pixel value in the fourth medical image can be calculated, and the first loss is determined based on the differences between each pixel.
[0104] Alternatively, a second loss can be determined based on the third and fourth medical images using the KL loss calculation formula. This second loss characterizes the distribution differences between the third and fourth medical images. Optionally, the distribution satisfied by the third medical image is determined based on the pixel values of each pixel in the third medical image. Similarly, the distribution satisfied by the fourth medical image is determined based on the pixel values of each pixel in the fourth medical image. The second loss is then determined based on the distributions satisfied by the third and fourth medical images.
[0105] Furthermore, a third loss can be determined based on the third and fourth medical images using formulas such as cosine similarity, Manhattan distance, or perceptual loss. This third loss characterizes the feature differences between the third and fourth medical images. Optionally, features from both the third and fourth medical images can be extracted, and the third loss can be determined based on these features.
[0106] Next, the loss of the first neural network is obtained by averaging or weighting at least one of the first loss, second loss, or third loss. Based on the loss of the first neural network, the first neural network is trained once to obtain the trained first neural network. If the trained first neural network meets the first termination condition, it is used as the first extraction network. If the trained first neural network does not meet the first termination condition, it is used as the first neural network for the next training iteration, and the training continues until the trained first neural network meets the first termination condition, thus obtaining the first extraction network.
[0107] This application does not limit the content of the first termination condition being met by the first trained neural network. For example, the first termination condition being met by the first trained neural network includes: the number of training iterations of the first trained neural network reaching the first number, or the performance index of the first trained neural network being not less than a first index, etc. Here, the first number or the first index can be a value preset based on human experience, or it can be a value input by the target object.
[0108] Since the first loss, second loss, and third loss are all determined based on the third and fourth medical images, and the fourth medical image is obtained by reconstructing features from the third medical image, training the first neural network with at least one of the first, second, or third losses optimizes the network in the direction that the extracted features can reconstruct the original image, thereby improving the accuracy of the network's feature extraction and improving the accuracy of subsequent clustering results.
[0109] After obtaining the first extraction network, the features of each reference image patch can be extracted through the first extraction network. Then, the features of each reference image patch are clustered according to any one of the following methods: K-Means clustering, mean-shift clustering, or expectation-maximum clustering based on Gaussian mixture model, to obtain multiple clusters. The clustering method is not limited here.
[0110] Step 2013: Determine the organizational knowledge of each cluster.
[0111] In this embodiment, for any cluster, the cluster includes features of at least one reference image block. If a cluster includes features of one reference image block, those features are considered as organizational knowledge. If a cluster includes features of multiple reference image blocks, the mean of the features of each reference image block is determined, and this mean is considered as organizational knowledge. Optionally, if the number of features of the reference image blocks included in a cluster is not greater than a reference number, the cluster is filtered out. If the number of features of the reference image blocks included in a cluster is greater than the reference number, since the features of the reference image blocks include elements in multiple rows and columns, for any element in any row and column, the mean of the elements in that row and column of the features of each reference image block is calculated to obtain the element in that row and column of an organizational knowledge. In this way, the mean of the features of each reference image block is determined, and an organizational knowledge is obtained. The reference number can be preset based on human experience or can be data input by the target object. In this way, multiple pieces of organizational knowledge can be determined.
[0112] like Figure 3 As shown, the reference medical image is input into an image feature extraction network (i.e., the first extraction network) to obtain the features of each reference image patch in the reference medical image. A clustering algorithm is used to group the features of each reference image patch into multiple clusters, and tissue knowledge is determined based on each cluster. Optionally, the feature z of the reference medical image includes multiple h′×w′×c feature blocks, where each feature block represents the features of a reference image patch. The K-means clustering method is used to cluster these feature blocks into N... com The prototype is composed of several clusters, the centroids of each cluster (i.e., the mean values of features of each reference image patch in the cluster). Wherein, prototype O is N composed of real numbers com A matrix of size ×h′×w′×c, including N com Each organizational knowledge is a feature of h′×w′×c.
[0113] Since image regions belonging to the same tissue structure in different reference medical images are similar in texture, color, and location, and since reference image patches are image regions of tissue structures, clustering is achieved at the tissue structure level by clustering the features of each reference image patch. This allows image regions belonging to the same tissue structure to be grouped into one cluster, and image regions belonging to different tissue structures to be grouped into different clusters. Furthermore, since the image set includes reference medical images of the same object's tissue structure from multiple perspectives, clustering based on multiple image sets ensures that the features of each reference image patch in the cluster represent the visual information of the same tissue structure from multiple perspectives. This allows the tissue knowledge determined based on each cluster to represent the visual information of the corresponding tissue structure from multiple perspectives. Subsequently, when generating medical images based on multiple tissue knowledge, the tissue knowledge can provide rich prior knowledge of multi-view visual information, which is beneficial for guiding the generation of medical images from different perspectives and improving the perspective richness of medical images.
[0114] Step 202: Determine sample knowledge from multiple tissue knowledge based on the first medical image. The sample knowledge represents the visual information of the tissue structure in the first medical image from multiple perspectives.
[0115] The first medical image involves some or all of multiple tissue structures, and each tissue structure corresponds to a specific tissue knowledge. Based on this, the electronic device can determine the sample knowledge corresponding to the tissue structure involved in the first medical image from multiple tissue knowledge sources. By determining the sample knowledge, multiple perspective visual information of the tissue structure involved in the first medical image is obtained, which is beneficial for subsequent neural network models to generate medical images from different perspectives based on the sample knowledge.
[0116] In one possible implementation, the first medical image includes multiple first image blocks. Step 202 includes: for any first image block, determining the distance between the features of the first image block and the tissue knowledge of each first image block, and taking the tissue knowledge corresponding to the distance that satisfies the distance condition as the tissue knowledge of the first image block; and determining sample knowledge based on the tissue knowledge of each first image block.
[0117] In this embodiment, a first medical image can be input into a first extraction network to extract features from the first medical image, and these features can be divided into features of multiple first image blocks. Alternatively, any first medical image can be divided into multiple first image blocks, and each first image block can be input into the first extraction network to extract features from each first image block. Any two first image blocks may have the same or different sizes. Optionally, the first extraction network f is used. e x sMapping to the latent space yields the features z of the first medical image. s =f e (x s Next, z s The image is divided into multiple feature blocks, and each feature block is a feature of a first image block.
[0118] For any first image patch, the distance between the features of the first image patch and each piece of organizational knowledge is calculated using distance calculation formulas such as Euclidean distance, cosine similarity, or Manhattan distance. For any distance, if the distance satisfies a distance condition, the organizational knowledge corresponding to that distance is taken as the organizational knowledge of the first image patch. Wherein, satisfying the distance condition includes: the distance being the minimum distance among the distances between the features of the first image patch and each piece of organizational knowledge, or the distance not exceeding a value preset based on human experience.
[0119] Next, based on the positions of each first image block in the first medical image, the tissue knowledge of each first image block is stitched together to obtain sample knowledge. This ensures that the position of the tissue knowledge of any first image block within the sample knowledge corresponds to its position in the first medical image. In the following text, the sample knowledge may use the parameter z. com To characterize.
[0120] like Figure 4 As shown, for any first image patch in the first medical image, the Nearest Neighbors (NN) algorithm is used to determine the tissue knowledge with the closest distance to the features of the first image patch from multiple tissue knowledge sources. The tissue knowledge of each first image patch is then concatenated to obtain sample knowledge. By calculating the distance between the features of the first image patch and each piece of tissue knowledge, and selecting the tissue knowledge corresponding to the distance that satisfies the distance condition, fine-grained acquisition of multi-view visual information of the tissue structures involved in the local region of the first medical image is achieved, thereby improving the accuracy of the sample knowledge.
[0121] Step 203: Perform noise addition on the image features of the second medical image at least once to obtain the feature data after each noise addition.
[0122] A second medical image can be input into a first extraction network to extract image features from the second medical image. Then, noise is added to the image features of the second medical image at least once to obtain noisy feature data. Each noisy feature data represents the noisy second medical image. The noise includes, but is not limited to, at least one of Gaussian noise, Poisson noise, and salt-and-pepper noise, and is a type of noisy data. Any two added noises can be of the same or different types, and if any two added noises are of the same type, the distributions of the added noises can be the same or different.
[0123] In an exemplary embodiment, the noise is added multiple times; step 203 includes: for the first noise addition, the image features of the second medical image are denoised to obtain the feature data after the first noise addition; for non-first noise addition, the feature data after the previous noise addition is denoised to obtain the feature data after the non-first noise addition.
[0124] In this embodiment, the electronic device first determines the noise-added data for the first denoising step, and then adds the noise-added data for the first denoising step to the image features of the second medical image, obtaining the feature data after the first denoising step. This feature data represents the second medical image with the noise added for the first time. Next, the electronic device determines the noise-added data for the second denoising step, and adds the noise-added data for the second denoising step to the feature data after the first denoising step, obtaining the feature data after the second denoising step. This feature data represents the second medical image with noise added twice. Then, the electronic device determines the noise-added data for the third denoising step, and adds the noise-added data for the third denoising step to the feature data after the second denoising step, obtaining the feature data after the third denoising step. This feature data represents the second medical image with noise added three times. This process continues until the feature data after the final denoising step is obtained, which represents the second medical image with noise added multiple times.
[0125] Optionally, for the t-th (t is a positive integer) noise addition, z is used t-1 The feature data representing the data before the t-th noise addition, i.e., z t-1 For the image features of the second medical image or the feature data after adding noise for the (t-1)th time, use z t The feature data after the t-th noise addition is represented. Where, δ t It is a hyperparameter for the t-th noise addition, used to determine the scale of the noisy data, ∈ t Characterizes Gaussian noise.
[0126] The process of adding noise to image features multiple times is also called the forward pass. In the forward pass, noisy data is gradually added to the image features of the second medical image, progressively destroying the image features until the final noisy feature data is used to represent the noise. Subsequently, a neural network model is trained so that the denoised data determined by the model approximates the noisy data from each subsequent addition. This allows for the gradual denoising of the noisy data based on these individual denoised data, enabling the recovery of the image representation and improving the model's accuracy.
[0127] Step 204: Based on sample knowledge and feature data after each noise addition, the neural network model determines the first denoised data corresponding to each noise addition. The first denoised data corresponding to any noise addition is used to denoise the feature data after any noise addition.
[0128] Sample knowledge and any noisy feature data can be input into a neural network model. The neural network model predicts the first denoised data corresponding to that noisy addition. This first denoised data is a feature used to characterize the noise. The first denoised data can be subtracted from any noisy feature data to denoise the noisy image features and restore the image features. The neural network model includes at least one of the following network layers: convolutional layer, normalization layer, feedforward layer, activation layer, linear layer, and attention layer. Different neural network models have different structures, and the methods for predicting the first denoised data also differ, which will not be elaborated further here.
[0129] In one possible implementation, step 204 includes steps A1 to A2 (not shown in the figure).
[0130] Step A1: Obtain at least one of the following: difference information, task description, or sample medical report. The difference information characterizes the difference between the sample medical report and the first medical image. The task description describes the task of generating medical images from different perspectives based on the first medical image. The sample medical report describes the diagnostic content of the first medical image and the second medical image.
[0131] As mentioned above, the first and second medical images are medical images of the same object from different perspectives. Typically, a medical image includes at least one tissue structure. By analyzing the first and second medical images, diagnostic information for each tissue structure can be obtained. The diagnostic information for any tissue structure includes, but is not limited to, its pathological condition and location. In other words, the sample medical report describes the diagnostic information of the tissue structures involved in the first and second medical images. Electronic devices can access the sample medical report, and the method of access is not limited here. For example, an electronic device can access a sample medical report input by a user, or it can access the sample medical report via the internet.
[0132] Because the sample medical report describes the diagnostic content of the first and second medical images, there may be differences in content between the sample medical report and the first medical image. For example, the sample medical report and the first medical image may differ in terms of tissue structures such as the heart, lungs, and liver, or in terms of abnormal information such as masses, fractures, and inflammation. For instance, if the sample medical report mentions two abnormalities, but the first medical image only shows one, it can be reasonably inferred that the other abnormality is likely shown in the second medical image. Based on this, electronic devices can acquire the difference information used to characterize the differences between the sample medical report and the first medical image. For example, the electronic device can acquire the difference information input by the user, or it can read the difference information via the internet.
[0133] In an exemplary embodiment, step A1, “obtaining difference information”, includes: determining a first probability that the sample medical report belongs to multiple categories and a second probability that the first medical image belongs to multiple categories; for any category, determining the difference between the first probability and the second probability of any category; determining a target category whose difference is not less than a reference difference; and determining difference information based on the target category.
[0134] In this embodiment, the electronic device can input a sample medical report into a text classification network, extract features from the sample medical report through the text classification network, and these features can characterize the semantics of the sample medical report. The text classification network then classifies the features of the sample medical report to obtain a text classification result, which includes a first probability that the sample medical report belongs to multiple categories.
[0135] These categories can be related to organizational structure, such as including categories for heart, lungs, and liver, or they can be categories related to abnormal information, such as including categories for lumps, fractures, and inflammation. The higher the probability of a category being the primary category, the more likely the sample medical report belongs to that category. In other words, if a sample medical report describes the content of a particular category, then the sample medical report belongs to that category.
[0136] On the other hand, the electronic device can input the first medical image into an image classification network, which extracts features from the first medical image. These features characterize the semantics of the first medical image. The image classification network then classifies the features of the first medical image to obtain an image classification result. The image classification result includes a second probability that the first medical image belongs to multiple categories. The higher the second probability of any category, the more likely the first medical image belongs to that category.
[0137] For any category, subtract the second probability of the category from the first probability of the category, or subtract the first probability of the category from the second probability of the category, to obtain a positive difference. If the difference is not less than a reference difference, then that category is taken as the target category. The reference difference is a value pre-set based on human experience.
[0138] Optionally, the sample medical report r can be input into a text classification network. Obtain image classification results Wherein, K represents the number of categories. This is used to characterize the first probability that a sample medical report belongs to the k-th category (k takes any value from 1 to K), where k and K are positive integers. The first medical image x... s Input Image Classification Network Obtain image classification results in, The second probability is used to characterize the first medical image belonging to the k-th category. Next, the difference between each category, Δp = |p|, is calculated. x -p r | and |·| are the absolute value symbols. The set is determined based on the differences between the categories. Where, Δp k The difference represents the k-th category, and th represents the reference difference.
[0139] Next, the difference information is determined using information such as the name, first probability, second probability, or difference of each target category. For example, if the target category is the anomalous information "edema" and the corresponding difference is 0.57, the difference information can be "Different: 'edema': 0.57".
[0140] By determining the difference between the first probability of the category to which the sample medical report belongs and the second probability of the category to which the first medical image belongs, and by determining the target category whose difference is not less than the reference difference, the target category with a larger probability difference is selected. This can reduce the error of the classification results to a certain extent and improve the accuracy of the difference information determined based on the target category.
[0141] Furthermore, since the first and second medical images have different perspectives, and the embodiments of this application expect the neural network model to generate the second medical image based on the first medical image, the electronic device can also acquire a task description to prompt the neural network model to generate images from different perspectives based on the first medical image. Optionally, the generated medical image has the same perspective as the second medical image. For example, if the first medical image is from a PA perspective and the second medical image is from a LAT perspective, the task description could be "Based on sample medical reports and difference information, combined with medical images from a PA perspective, generate medical images from a LAT perspective."
[0142] Step A2: Using a neural network model, determine the first denoised data corresponding to each denoising step based on at least one of the following: difference information, task description, or sample medical report, sample knowledge, and feature data after each denoising step.
[0143] In this embodiment, the text content includes at least one of the following: difference information, task description, or sample medical report. Optionally, the text content D satisfies: Among them, class_name k The name representing the target category, Δp k The difference that represents the target category. The representation includes a set of target categories and differences. T represents the task description, and r represents the sample medical report. The text content, sample knowledge, and feature data after any initial denoising are input into a neural network model, which then predicts the first denoised data corresponding to that denoising iteration.
[0144] like Figure 5 As shown, on one hand, the sample medical report is input into a text classification network to obtain text classification results, which include the first probability that the sample medical report belongs to multiple categories. On the other hand, the first medical image is input into an image classification network to obtain image classification results, which include the second probability that the first medical image belongs to multiple categories. The target category with the largest difference between the first and second probabilities is selected from the multiple categories, and difference information is determined based on the target category and the corresponding difference. The difference information, task description, and sample medical report are concatenated end-to-end to obtain the text content. For example, the task description is concatenated after the difference information to obtain concatenated text, and the sample medical report is concatenated after the concatenated text to obtain the text content.
[0145] Since the text content can reflect the content differences between the sample medical report and the first medical image, as well as the perspective differences between the first medical image and the generated medical image, the text content can guide the neural network model to predict denoised data. This allows the image features obtained after removing noise from the denoised data to characterize medical images that are related to the sample medical report and the first medical image, but have different perspectives and content than the first medical image, thus improving the accuracy of the model's prediction of denoised data.
[0146] In another possible implementation, the neural network model includes a first network part and a second network part. Step 204 includes steps B1 to B2 (not shown in the figure).
[0147] Step B1: For any noise addition, the first network part determines the first information based on the first medical image, sample knowledge and feature data after any noise addition. The first information represents the information after the fusion of the first medical image, sample knowledge and feature data after any noise addition.
[0148] In this embodiment, the features of a first medical image can be determined first. The features of the first medical image, sample knowledge, and feature data after any one instance of noise addition are then concatenated to obtain concatenated features. Optionally, the concatenated features include more information, such as at least one of the features of the concatenated text content or the number of noise additions, as well as the features of the first medical image, sample knowledge, and feature data after any one instance of noise addition. The text content features represent the text content, and the number of noise additions features represent the number of noise additions corresponding to any one instance of noise addition. The concatenated features are input into a first network component, which performs feature extraction on the concatenated features to fuse the semantics of the various information involved in the concatenated features, thus obtaining first information.
[0149] like Figure 6 As shown, the text content includes difference information, task description, and sample medical reports. The text content can be input into a text feature extraction network to obtain text content features, which are then input into the first network part. Alternatively, the number of noise additions (t) can be input into a time-based feature extraction network to obtain noise-added features, which are then input into the first network part. The structure, type, and feature extraction method of the text feature extraction network (also called a text encoder) or the time-based feature extraction network (also called a time encoder) are not limited here. Features of the first medical image can also be input into the first network part. Furthermore, the image features of the second medical image are subjected to t noise additions to obtain t-fold noise-added feature data. This t-fold noise-added feature data is then concatenated with sample knowledge and input into the first network part. The first network part determines and outputs the first information based on each input.
[0150] The first network component includes at least one of the following: linear layers, convolutional layers, normalization layers, feedforward layers, and attention layers. A possible structure for the first network component is shown below.
[0151] In this example, the first network part includes an encoder and intermediate layers. The encoder includes at least one concatenated encoding block, each encoding block comprising multiple convolutional structures and pooling layers. The intermediate layers include convolutional structures. The convolutional structures include convolutional layers and activation layers, used to perform convolutional and activation processing on the features, thereby fusing the semantics of various information involved in the features and improving the feature representation ability. The pooling layers are used to downsample the features, achieving dimensionality reduction, reducing feature complexity, and facilitating model understanding of the features.
[0152] After inputting the concatenated features into the first network part, convolution, activation, and downsampling are performed on the concatenated features through the first encoding block. This integrates the semantics of information such as the differential information involved in the concatenated features, task description, sample medical report, first medical image, sample knowledge, and noisy second medical image represented by feature data after any noise addition, thereby reducing the complexity of the concatenated features and obtaining the output features of the first encoding block. Next, convolution, activation, and downsampling are performed on the output features of the previous encoding block through non-first encoding blocks to further integrate the semantics of various information and reduce feature complexity, obtaining the output features of non-first encoding blocks. Convolution and activation are then performed on the output features of the last encoding block through intermediate layers to further integrate the semantics of various information, obtaining the output features of the intermediate layers. The output features of the intermediate layers can be used as the first information, or the output features of each encoding block and the output features of the intermediate layers can be used as the first information.
[0153] Step B2: Based on the first information and the feature data after any noise addition, the second network part determines the first denoised data corresponding to any noise addition.
[0154] In this embodiment, the first information and the feature data after any one instance of noise addition can be input into the second network part. In practical applications, more information can be input. For example, at least one of the features of the text content or the features of the number of noise additions, along with the first information and the feature data after any one instance of noise addition, can be input into the second network part, and the second network part can predict the first denoised data corresponding to any one instance of noise addition based on the input.
[0155] like Figure 6As shown, the text content is input into a text feature extraction network to obtain the text content features, and these features are then input into the second network part. The number of noise additions, *t*, is input into a noise addition feature extraction network to obtain the features of the number of noise additions, and these features are then input into the second network part. The image features of the second medical image are subjected to *t* noise additions to obtain the t-noise-added feature data, which is then input into the second network part. Furthermore, the first information output from the first network part is also input into the second network part. Based on the various inputs, the second network part determines the first denoised data corresponding to the *t* noise additions.
[0156] The first network component determines the first information, achieving semantic integration of differential information, task description, sample medical reports, the first medical image, sample knowledge, and a noisy second medical image. This reduces the complexity of the first information, improves its representational ability, and facilitates model understanding. The first information can provide multi-view visual information about the various tissue structures involved in the first medical image, and may even provide prior information such as the content and perspective of the medical image to be generated. This allows the first denoised data predicted based on the first information to have features related to the first medical image after noise removal, while also exhibiting differences in content and perspective, thus facilitating the generation of medical images from different perspectives. In other words, the accuracy of the first denoised data predicted based on the first information is high, which is beneficial for generating high-quality medical images from different perspectives.
[0157] In an exemplary embodiment, a second network part can be trained first, and then the first denoised data can be predicted using the second network part. In this case, before step B2, the method further includes: determining the second denoised data corresponding to any denoising based on the feature data after any denoising using a third network part, wherein the second denoised data corresponding to any denoising is used to denoise the feature data after any denoising; and training the third network part based on the second denoised data corresponding to each denoising to obtain the second network part.
[0158] In this embodiment, the feature data after any one instance of noise addition is input into the third network part. In practical applications, more information can be input. For example, at least one of the features of the text content or the features of the number of noise additions, as well as the feature data after any one instance of noise addition, can be input into the third network part, and the third network part can predict the second denoised data corresponding to any one instance of noise addition based on the input.
[0159] The third network component includes at least one of the following: linear layers, convolutional layers, normalization layers, feedforward layers, and attention layers. A possible structure for the third network component is shown below.
[0160] In this example, the third network component includes an encoder, intermediate layers, and a decoder. The encoder includes at least one cascaded encoding block, each encoding block comprising multiple convolutional structures and pooling layers. The intermediate layers include convolutional structures. The decoder includes at least one cascaded decoding block, each decoding block comprising a deconvolutional layer and multiple convolutional structures. The encoder and decoder each include the same number of encoding blocks. The encoder and intermediate layers have been described in detail above and will not be repeated here. The deconvolutional layer is used to deconvolve the features, thereby increasing the dimensionality of the features.
[0161] The first coded block is used to perform convolution, activation, and downsampling on the input of the third network part to obtain the output features of the first coded block. Similarly, non-first coded blocks are used to perform convolution, activation, and downsampling on the output features of the previous coded block to obtain the output features of the non-first coded block. Intermediate layers are used to perform convolution and activation on the output features of the last coded block to obtain the output features of the intermediate layer. Since the number of coded and decoded blocks is the same, there is a one-to-one correspondence between them. For the first decoded block, the output features of the intermediate layer are deconvolutionally processed, the output features of the corresponding coded block are concatenated with the deconvolution result, and the concatenated result is then convolutionally and activated to obtain the output features of the first decoded block. For non-first decoded blocks, the output features of the previous decoded block are deconvolutionally processed, the output features of the corresponding coded block are concatenated with the deconvolution result, and the concatenated result is then convolutionally and activated to obtain the output features of the non-first decoded block. By concatenating the output features of the coded block with the deconvolution result, information loss due to feature processing can be reduced, improving the representational ability of the concatenated result. Convolution and activation processing of the concatenated result allows for the extraction of features representing noise. Therefore, the output feature of the last decoded block is the second denoised data corresponding to any denoising step, used to represent the noise. Denoising the feature data is achieved by subtracting the corresponding second denoised data from the feature data after any denoising step, thus restoring the image features used to represent the medical image.
[0162] The loss for each denoising iteration can be determined based on the second denoised data and the denoised data corresponding to that iteration. The loss of the third network component is then determined based on the losses from each denoising iteration. Optionally, the loss L of the third network component... SD satisfy: Where E represents the sign of the mean. z represents the second medical image. V represents the difference information and task description. r represents the sample medical report. ∈ represents the noisy data added in the t-th iteration. N(0,1) represents a normal distribution with a mean of 0 and a variance of 1 (i.e., Gaussian noise). f SD (zt ,t,τ([V;r])) represents the second denoised data corresponding to the t-th denoising, z t The feature data after the t-th noise addition is represented. τ([V;r]) represents the spliced features after combining the features of the difference information, the features of the task description, and the features of the sample medical report, and τ represents the text encoder. The square of the L2 norm.
[0163] The third network part can be trained once based on its loss to obtain the trained third network part. If the trained third network part meets the second termination condition, it is used as the second network part. If the trained third network part does not meet the second termination condition, it is used as the third network part for the next training iteration, and the training continues until the trained third network part meets the second termination condition, thus obtaining the second network part.
[0164] This application does not limit the content of the second termination condition for the trained third network part. For example, the training iterations of the trained third network part meeting the second termination condition include: the training iterations of the trained third network part reaching the second number, or the performance index of the trained third network part not being less than the second index, etc. The second number or the second index can be a value preset based on human experience, or it can be a value input by the target object.
[0165] The third network part is trained by using the second denoised data and the denoised data corresponding to each denoising step. This optimizes the network so that the predicted denoised data is closer to the denoised data, improving the accuracy of the network's prediction of denoised data. This is beneficial for improving the effect of subsequent training and thus improving the accuracy of the noise determination model.
[0166] After training to obtain the second network part, the denoised data can be predicted using the second network part. Optionally, step B2 includes: determining second information based on the feature data after any denoising process using the second network part, wherein the second information represents the feature data after any denoising process; and determining the first denoised data corresponding to any denoising process using the second network part based on the first information and the second information.
[0167] Since the second network part is based on the prediction of the third network part, the second network part and the third network part have similar structures and functions, and the only difference between them is in the network parameters.
[0168] For example, the second network portion includes an encoder, an intermediate layer, and a decoder, such as Figure 6As shown. The feature data after any one round of noise addition can be input into the second network part, or at least one of the text content features or the number of noise additions, along with the feature data after any one round of noise addition, can be input into the second network part. After processing the input through the encoder and intermediate layers, the output features of each coding block and the output features of the intermediate layers are obtained. The processing method is described above and will not be repeated here. The output features of the intermediate layers can be used as the second information, or the output features of each coding block and the output features of the intermediate layers can be used as the second information.
[0169] In this example, the number of decoded blocks in the second network portion is equal to the number of encoded blocks in the second network portion, and also equal to the number of encoded blocks in the first network portion. In other words, there is a one-to-one correspondence between the decoded blocks, the encoded blocks in the second network portion, and the encoded blocks in the first network portion.
[0170] For the first decoding block, deconvolution is performed on the output features of the intermediate layers of the second and first network parts. The deconvolution result is then concatenated with the output features of the corresponding coding blocks in the second and first network parts. This concatenation is followed by convolution and activation to obtain the output features of the first decoding block. For subsequent decoding blocks, deconvolution is performed on the output features of the previous decoding block. The deconvolution result is then concatenated with the corresponding coding blocks in the second and first network parts. This concatenation is followed by convolution and activation to obtain the output features of the subsequent decoding blocks. The output features of the last decoding block are the first denoised data corresponding to any noise addition, used to characterize the noise. By subtracting the first denoised data corresponding to any noise addition from the feature data after any noise addition, denoising is achieved, thereby restoring the image features used to characterize the medical image.
[0171] like Figure 7 As shown, the first network part includes an encoder and an intermediate layer M'. The encoder includes four encoding blocks, namely encoding blocks A' to D'. The second network part includes an encoder, an intermediate layer M, and a decoder. The encoder includes four encoding blocks, namely encoding blocks A to D, and the decoder includes four decoding blocks, namely decoding blocks D to A.
[0172] The first information is obtained by inputting at least one of the following into the first network part: the features of the text content, the features of the number of times noise is added, the features of the first medical image, the feature data after any noise addition, or sample knowledge. The first information includes the output features of the coding blocks A' to D' and the output features of the intermediate layer M'.
[0173] Similarly, at least one of the features of the text content, the features of the number of times noise was added, or the feature data after any one noise addition was input into the second network part to obtain the second information, which includes the output features of coding blocks A to D and the output features of intermediate layer M.
[0174] For decoded block D, the output features of intermediate layer M' are zero-convolved, and the output features of encoded block D' are also zero-convolved. Here, the zero-convolved layer is a 1×1 convolutional layer with zero network parameters, used to convert the size of one feature to be concatenated into the size of another feature to facilitate feature concatenation. For decoded block D, the output features of intermediate layer M' after zero-convolution and the concatenated output features of intermediate layer M are first deconvolved to obtain the deconvolution result. Then, the deconvolution result, the output features of encoded block D, and the output features of encoded block D' after zero-convolution are concatenated, and the concatenated features are subjected to convolution and activation processing to obtain the output features of decoded block D.
[0175] For decoded block C, the output features of encoded block C' are subjected to zero convolution. Then, the output features of decoded block D are deconvolved using decoded block C. The deconvolution result, the output features of encoded block C, and the zero-convolutioned output features of encoded block C' are concatenated. Finally, convolution and activation processing are applied to the concatenated features to obtain the output features of decoded block C.
[0176] The processing principles of decoding blocks B and A are similar to those of decoding block C, and will not be elaborated upon here. The second network part can determine the first denoised data corresponding to any noise addition based on the output characteristics of decoding block A.
[0177] The second network component determines the second information, achieving semantic integration of differential information, task description, sample medical reports, and noisy second medical images. This reduces the complexity of the second information, improves its representational ability, and facilitates model understanding. The first information can provide multi-view visual information of various tissue structures involved in the first medical image. The first and second information can provide prior information such as the content and perspective of the medical image to be generated. This ensures that the first denoised data predicted based on the first and second information has features related to the first medical image after noise removal, but also differs from the first medical image in content and perspective, which is beneficial for generating medical images from different perspectives.
[0178] Step 205: Train a neural network model based on the first denoised data corresponding to each denoising step to obtain a noise determination model. The noise determination model is used to determine the target noise data, and the target noise data is used to determine the target medical image of the tissue structure in the source medical image from different perspectives.
[0179] In this embodiment, the loss of the neural network model can be determined based on the first denoised data corresponding to each denoising iteration. The neural network model is then trained once based on its loss to obtain a trained neural network model. If the trained neural network model satisfies the third termination condition, it is used as the noise determination model. If the trained neural network model does not satisfy the third termination condition, it is used as the neural network model for the next training iteration, and training continues until the trained neural network model satisfies the third termination condition, thus obtaining the noise determination model.
[0180] This application does not limit the content of the trained neural network model satisfying the third termination condition. For example, the trained neural network model satisfying the third termination condition includes: the trained neural network model reaching a third training iteration, or the performance index of the trained neural network model being not less than a third metric, etc. The third training iteration or the third metric can be a value preset based on human experience, or it can be a value input by the target object.
[0181] In one possible implementation, step 205 includes: for any noise addition, determining the loss of any noise addition based on the first denoised data and the noise-added data corresponding to any noise addition, using the noise-added data corresponding to any noise addition to add noise to obtain the feature data after any noise addition; training a neural network model based on the loss of each noise addition to obtain a noise determination model.
[0182] In this embodiment, the loss for each denoising iteration can be determined based on the difference between the first denoised data and the denoised data corresponding to any given denoising iteration. The loss of the neural network model is obtained by averaging and summing the losses for each denoising iteration.
[0183] Optionally, the loss L of the neural network model satisfies: Where E represents the sign of the mean. z represents the second medical image. t ′=[z t [zcom],z t zcom represents the feature data after the t-th noise addition, z represents the sample knowledge, z t ′representation splicing z t Features following zcom. D-representation includes differential information, task description, and textual content from at least one of the sample medical reports, x s The first medical image is represented. ∈ represents the noisy data from the t-th addition. N(0,1) represents a normal distribution with a mean of 0 and a variance of 1 (i.e., Gaussian noise). f SD (z t ,t,τ(D),fCon (z t ′,t,τ(D),x s )) represents the first denoised data corresponding to the t-th denoising. τ(D) represents the features of the text content, and τ represents the text encoder. f Con (z t ′,t,τ(D),x s ) represents the first information. The square of the L2 norm.
[0184] Next, the neural network model is trained using a loss function based on the neural network model to obtain a noise determination model. The neural network model is trained using the first denoised data and the denoised data corresponding to each denoising iteration. This optimizes the model to make the predicted denoised data approximate the denoised data, improving the accuracy of the model's prediction of the denoised data. After removing noise from the denoised data predicted by the model, feature data representing the medical image can be obtained, enabling the model to generate high-quality medical images.
[0185] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions. For example, the first medical image, the second medical image, and the sample medical report involved in this application were all obtained with full authorization.
[0186] In the above method, sample knowledge of the first medical image is determined from multiple tissue knowledge sources, thereby acquiring multi-view visual information of various tissue structures in the first medical image. Since the first and second medical images have different perspectives, multiple noise additions are applied to the image features of the second medical image. Based on sample knowledge and the feature data after each noise addition, the first denoised data corresponding to that noise addition is determined. This allows the model to accurately predict the noise added to the image features at a certain perspective based on multi-view visual information, improving the accuracy of the first denoised data. Training the neural network model with the first denoised data corresponding to each noise addition improves the model's training effect and accuracy. This enables the trained noise determination model to accurately predict the noise data that needs to be removed. By removing this noise data, medical images from different perspectives are obtained, increasing the perspective richness of medical images and thus improving the accuracy of the analysis results obtained from analyzing medical images.
[0187] like Figure 8 As shown, Figure 8This is a flowchart illustrating a medical image generation method provided in this application embodiment. The terminal device 101 or server 102 executing the method of this application embodiment can be collectively referred to as an electronic device; that is, the method of this application embodiment can be executed by an electronic device. Figure 8 As shown, the method includes the following steps.
[0188] Step 801: Acquire multiple tissue knowledge and source medical images, where any tissue knowledge represents the visual information of an tissue structure from multiple perspectives.
[0189] In this embodiment, the content of multiple tissue knowledge has been described in step 201, and the content of the source medical image is similar to the content of the first medical image and the second medical image involved in step 201. Therefore, the implementation of step 801 can be found in the description of step 201, and will not be repeated here.
[0190] Step 802: Determine source knowledge from multiple tissue knowledge sources based on the source medical image. The source knowledge represents the visual information of the tissue structure in the source medical image from multiple perspectives.
[0191] The content of the source knowledge is similar to the content of the sample knowledge involved in step 202. Therefore, the implementation method of step 802 can be found in the description of step 202, and will not be repeated here.
[0192] Step 803: Determine the target noise data based on source knowledge using the noise determination model.
[0193] Among them, the noise determination model is based on the same principle as... Figure 2 The relevant noise determination model is trained using an embodiment of the training method. The electronic device inputs source knowledge into the noise determination model, and the model predicts the target noise data according to the implementation principle of step 204.
[0194] Step 804: Denoise the reference noise data based on the target noise data to obtain the target medical image. The target medical image is an image of the tissue structure in the source medical image from different perspectives.
[0195] In this embodiment, the electronic device can randomly generate reference noise data, which is used to characterize at least one of Gaussian noise, Poisson noise, salt-and-pepper noise, etc. By removing target noise data from the reference noise data, denoising is achieved to obtain target features. These target features characterize the target medical image, and can be mapped to the target medical image.
[0196] Since the source knowledge includes visual information from multiple perspectives of various tissue structures in the source medical image, by using the noise determination model to predict the target noise data based on the source knowledge and then removing the target noise data, a target medical image that is related to the source medical image but has a different perspective from the source medical image can be obtained.
[0197] In an exemplary embodiment, the denoising is performed multiple times; steps 803 and 804 include: for the first denoising, the noise determination model determines the target noise data for the first denoising based on source knowledge and reference noise data, and the reference noise data is denoised based on the target noise data for the first denoising to obtain the feature data after the first denoising; for subsequent denoising, the noise determination model determines the target noise data for subsequent denoising based on source knowledge and the feature data after the previous denoising, and the feature data after the previous denoising is denoised based on the target noise data for subsequent denoising to obtain the feature data after the previous denoising; and the target medical image is determined based on the feature data after the last denoising.
[0198] In this embodiment, source knowledge and reference noise data can be input into the noise determination model. Following the implementation principle shown in step 204, the noise determination model determines the target noise data for the first denoising step. The target noise data for the first denoising step is then removed from the reference noise data to obtain the feature data after the first denoising step. Next, the source knowledge and the feature data after the first denoising step are input into the noise determination model. Following the implementation principle shown in step 204, the noise determination model determines the target noise data for the second denoising step. The target noise data for the second denoising step is then removed from the feature data after the first denoising step to obtain the feature data after the second denoising step. Afterward, the source knowledge and the feature data after the second denoising step are input into the noise determination model. Following the implementation principle shown in step 204, the noise determination model determines the target noise data for the third denoising step. The target noise data for the third denoising step is then removed from the feature data after the second denoising step to obtain the feature data after the third denoising step. This process continues until the final denoised feature data is obtained, which is the target feature. The noise determination model maps the target feature to the target medical image.
[0199] Since the noise determination model is trained from a neural network model, similar to the neural network model, its input can include more information besides source knowledge and reference noise data (or the feature data after the previous denoising). For example, it can include at least one of the following: target difference information, target task description, or target medical report. The noise determination model processes each piece of information in a similar way to the neural network model, and will not be elaborated further here. Specifically, the target medical report describes the diagnostic content of the source medical image. Target difference information characterizes the differences between the target medical report and the source medical image. The target task description describes the generation of a medical image from a target perspective based on the source medical image; the target perspective differs from the perspective of the source medical image.
[0200] like Figure 9 As shown, a noise determination model can be used to determine target noise data based on the source medical image and the corresponding target medical report. The target medical image is then generated based on the target noise data, and the source and target medical images have different perspectives, thus improving the visual richness of the medical images.
[0201] Subsequently, based on at least one of the source medical image, target medical image, and target medical report, objectives such as disease diagnosis and treatment planning can be achieved. Figure 10 As shown, the source medical image and the target medical image are input into the disease diagnosis model to obtain the disease diagnosis result. The structure and processing principles of the disease diagnosis model are not limited here.
[0202] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions. For example, the source medical images and multiple tissue knowledge involved in this application were obtained with full authorization.
[0203] The above method achieves the acquisition of multi-view visual information of various tissue structures in the source medical image by determining the source knowledge of the source medical image from multiple tissue knowledge sources. The noise determination model determines the target noise data based on the source knowledge, enabling the model to identify information unrelated to the source medical image and the target viewpoint. This ensures that the target medical image obtained after denoising the reference noise data based on the target noise data is relevant to the source medical image and represents the target viewpoint. Since the source knowledge contains multi-view visual information, the target viewpoint can differ from that of the source medical image. In other words, the noise determination model can generate high-quality medical images with different viewpoints, improving the viewpoint richness of medical images and thus enhancing the accuracy of the analysis results obtained from analyzing medical images.
[0204] The above describes the training method of the noise determination model and the medical image generation method of the embodiments of this application from the perspective of method steps. The following is a detailed description in conjunction with a scenario. The method of the embodiments of this application is applicable to the medical field. In practical applications, a noise determination model can be trained based on medical images such as X-ray films, CT scans, MRI scans, and ultrasound films, and corresponding medical reports, according to the method of the embodiments of this application. The noise determination model determines target noise data based on source medical images such as X-ray films, CT scans, MRI scans, and ultrasound films, and corresponding target medical reports, and generates target medical images such as X-ray films, CT scans, MRI scans, and ultrasound films from different perspectives based on the target noise data. For ease of description, this example uses a scenario of X-ray films (such as chest X-rays or mammograms) and corresponding medical reports to illustrate the method of the embodiments of this application.
[0205] like Figure 11 As shown, in this example, a noise determination model can be trained through a three-stage training process. These three training stages include the first stage, the second stage, and the third stage. The first stage includes steps 1101 to 1102.
[0206] Step 1101: Extract the features of the first X-ray film through the first neural network, and reconstruct the second X-ray film based on the features of the first X-ray film.
[0207] Step 1102: Based on the first X-ray film and the second X-ray film, train the first neural network to obtain the image feature extraction network.
[0208] In this context, the first X-ray corresponds to the third medical image mentioned above, the second X-ray corresponds to the fourth medical image mentioned above, and the image feature extraction network corresponds to the first extraction network mentioned above. Based on this, the first-stage training process can be found in the section above regarding training the first extraction network; their implementation principles are similar and will not be repeated here. After completing the first-stage training process, the second-stage training process is executed, which includes steps 1103 to 1106.
[0209] Step 1103 involves determining the image features of the target X-ray image through an image feature extraction network. Noise is gradually added to the image features of the target X-ray image based on multiple rounds of noise addition, resulting in noisy feature data for each round. The target X-ray image corresponds to the second medical image mentioned above. The implementation principle of step 1103 can be found in the description of step 203, and will not be repeated here.
[0210] Step 1104: Determine the image features of the source X-ray image using an image feature extraction network. Based on the image features of the source X-ray image and the medical report, determine the text content. The source X-ray image corresponds to the first medical image mentioned above. The implementation principle of step 1104 can be found in the description of step A1, and will not be repeated here.
[0211] Step 1105: For any noise addition, the diffusion network determines the denoised data based on the feature data and text content after the noise addition, and determines the loss of the noise addition based on the denoised data and the noise-added data.
[0212] Step 1106: Train the diffusion network based on the loss from each noise addition.
[0213] The pre-training diffusion network is the third network part mentioned above, and the post-training diffusion network is the second network part mentioned above. Based on this, the implementation principle of steps 1105 and 1106 can be seen from the content of the trained second network part. The implementation principles of the two are similar and will not be repeated here.
[0214] The second phase of training process is as follows: Figure 12 As shown. On one hand, based on the source X-ray film and medical report, the text content is determined through a text content generation module. The implementation principle of the text content generation module is as follows: Figure 5As shown, details will not be repeated here. The text content is input into the text feature extraction network to obtain the text content features, and these features are then input into the diffusion network. On the other hand, the number of times noise is added (t) is input into the number of noise additions feature extraction network to obtain the features for that number of noise additions, and these features are then input into the diffusion network. Furthermore, based on the target X-ray image, the feature data after t noise additions is determined, and this feature data is input into the diffusion network. The diffusion network, based on each input, determines the denoised data corresponding to each t noise addition. The diffusion network is trained using the denoised data and the noise-added data corresponding to each noise addition. It is understood that in the second stage, the text feature extraction network or the number of noise additions feature extraction network may or may not be trained.
[0215] The second stage of training completes the training of the diffusion network, enabling it to accurately predict denoised data. After removing noise from the denoised data, the image features of the target X-ray film can be restored. To improve model performance, the third stage of training is then performed, comprising steps 1107 to 1109.
[0216] Step 1107: Based on the image features of the source X-ray film, relevant knowledge is determined from multiple tissue knowledge sources. This relevant knowledge corresponds to the sample knowledge mentioned above; therefore, the implementation principle of step 1107 can be found in the description of step 202, and will not be repeated here.
[0217] Step 1108: For any noise addition, the control network determines the control information based on the noise-added feature data, text content, features of the source X-ray film, and related knowledge. The diffusion network determines the denoised data based on the control information, the noise-added feature data, and text content. The loss for this noise addition is determined based on the denoised data and the noise-added data.
[0218] Step 1109: Based on the loss of each noise addition, train the control network and the diffusion network to obtain the noise determination model.
[0219] The control network before training is the first network part mentioned above, and the control information is the first information mentioned above. Based on this, the implementation principle of steps 1108 and 1109 can be found in the description of steps B1 and B2. The implementation principles of the two are similar and will not be repeated here.
[0220] The third stage of training process is as follows: Figure 12As shown. First, the nearest neighbor algorithm is used to determine the relevant knowledge of the source X-ray image based on multiple organizational knowledge bases. This relevant knowledge is then concatenated with the feature data after t levels of noise addition. The concatenated data, along with the features of the text content, the features after t levels of noise addition, and the image features of the source X-ray image, is input into the control network. The control network determines the control information. Next, the features of the text content, the features after t levels of noise addition, the feature data after t levels of noise addition, and the control information are input into the diffusion network. The diffusion network determines the denoised data corresponding to t levels of noise addition based on each input. The diffusion network and the control network are trained using the denoised data and the noise-added data corresponding to each level of noise addition. In the third stage, the text feature extraction network or the level of noise addition feature extraction network may or may not be trained. Optionally, in the third stage, the encoders and intermediate layers in the text feature extraction network, the level of noise addition feature extraction network, and the diffusion network are frozen, and only the decoders in the control network and the diffusion network are trained.
[0221] The training processes in the second and third stages described above involve adding noise to the target X-ray multiple times to obtain noisy feature data, and then using a neural network model to predict the denoised data from each noise addition. It can be understood that by denoising the noisy feature data based on this denoised data, a predicted X-ray that approximates the target X-ray can be obtained. In other words, as... Figure 13 As shown, the embodiments of this application are equivalent to: using a target X-ray film, supervising a neural network model to generate a predicted X-ray film based on the source X-ray film and a medical report.
[0222] In practical applications, such as Figure 14 As shown, the noise determination model can be deployed on the backend. After frontend A obtains the source X-ray image, it can send the source X-ray image to the backend. The backend uses the noise determination model to determine the target denoising data based on the source X-ray image, generates the target X-ray image based on the target denoising data, and sends the target X-ray image to frontend B. Frontend A and frontend B can be the same or different frontends.
[0223] Furthermore, in the training environment of the noise determination model in this application embodiment, several related models were also trained. These models can also be used to generate target X-rays based on source X-rays. Optionally, the training environment is as follows: a publicly available dataset is obtained via the Internet, which includes 377,110 chest X-rays and 227,835 corresponding medical reports. 100,882 samples are selected from the dataset, of which 84,170 samples contain X-rays with two views and 16,712 samples contain X-rays with three views.
[0224] Alternatively, the dataset D can be represented as: Where N is the number of samples, X i Representing the i-th sample, r iThe medical report representing the i-th sample. The i-th sample X i satisfy: Where, x i,j The X-ray image representing the j-th viewpoint in the i-th sample, N i The number of viewpoints for the i-th sample. 2≤N i ≤3.
[0225] From 100,882 samples, 99,087 samples can be selected as training data, 757 samples as validation data, and 1,038 samples as test data. The training data is used to adjust the network parameters during the training phase, and the validation data is used to verify whether the model after adjusting the network parameters has problems such as overfitting or underfitting.
[0226] When performing the above three-stage training using the training data: For the first stage, the training includes an image feature extraction network with four layers, each with 64, 128, 128, and 128 channels respectively. The image features extracted by the image feature extraction network are represented as 64×64×3. During training, the Adam optimizer is used at a 5×10⁻⁶ kcal / s ratio. -5 The learning rate was set to 100 training iterations, with each training iteration consisting of 8 batches of input to the first X-ray image, which had a size of 512×512. For the third stage, the AdamW optimizer was used at a learning rate of 2.5×10⁻⁶. -5 Training is performed using a learning rate of N. com =1024, h′×w′×c=8×8×3 to determine relevant knowledge, and determine text content according to 11 categories and reference difference th=0.5.
[0227] The test data was used to test the noise determination model (i.e., model 6) and related models (i.e., models 1 to 5) of the embodiments of this application. The test results are shown in Table 1 below.
[0228] Table 1
[0229]
[0230]
[0231] The similarity between the predicted X-ray and the target X-ray was measured using MAE (Mean Absolute Error), SSIM (Structural Similarity), MS-SSIM (Multi-Scale Structural Similarity), and FID (Frechet Inception). The performance of the predicted X-ray on the image classification task was measured using F1 and AUC (Area Under Curve). The performance of the predicted X-ray on the report generation task was measured using BLEU-4 (Bilingual Evaluation Understudy-4). In Table 1, "↓" indicates that lower values indicate better model performance, and conversely, "↑" indicates that higher values indicate better model performance. For MAE, SSIM, and MS-SSIM, values within parentheses represent the standard deviation, and values outside parentheses represent the mean. The values for FID, F1, AUC, and BLEU-4 are the mean values.
[0232] As can be seen from Table 1, Model 6 achieved the best performance across all metrics, exceeding Model 5 by 2.20 in FID, 0.0279 in AUC, and 0.0094 in BLEU-4. These results demonstrate that the method described in this application can train a high-performance noise determination model.
[0233] like Figure 15 As shown, this application embodiment also provides three sets of X-ray images. Each set of X-ray images includes a source X-ray image, a target X-ray image, an X-ray image generated by model 2, and X-ray images generated by models 4 to 6. Model 6 is the noise determination model of this application embodiment, and models 2, 4, and 5 are correlation models. Figure 15 It is evident that the X-ray film generated by Model 6 has similar details to the target X-ray film, indicating that the noise determination model of this application embodiment can generate higher quality X-ray films with different perspectives.
[0234] In addition, the embodiments of this application also tested the performance impact of real single-view X-rays, synthetic multi-view X-rays (including real source view X-rays and synthetic target view X-rays) and real multi-view X-rays on downstream tasks (such as image classification tasks and report generation tasks), as shown in Table 2.
[0235] Table 2
[0236]
[0237]
[0238] As shown in Table 2, real multi-view X-ray images can maximize the performance improvement of downstream tasks. Furthermore, the performance improvement of synthesized multi-view X-ray images on downstream tasks is greater than that of real single-view X-ray images, and the performance improvement of synthesized multi-view X-ray images on downstream tasks is close to that of real multi-view X-ray images. This demonstrates that the target X-ray images with different perspectives from the source X-ray images generated by the noise determination model in this application can improve the performance of downstream tasks, thereby increasing the accuracy of X-ray image analysis results.
[0239] Figure 16 The diagram shown is a structural schematic of a training device for a noise determination model provided in an embodiment of this application. Figure 16 As shown, the device includes:
[0240] The acquisition module 1601 is used to acquire multiple tissue knowledge, a first medical image and a second medical image. Each tissue knowledge represents the visual information of a tissue structure from multiple perspectives, and the second medical image is an image of the tissue structure in the first medical image from different perspectives.
[0241] The determination module 1602 is used to determine sample knowledge from multiple tissue knowledge based on a first medical image, wherein the sample knowledge characterizes the visual information of the tissue structure in the first medical image from multiple perspectives.
[0242] The noise addition module 1603 is used to perform at least one noise addition on the image features of the second medical image to obtain the feature data after each noise addition.
[0243] The determination module 1602 is also used to determine the first denoised data corresponding to each denoising based on sample knowledge and feature data after each denoising through a neural network model, and the first denoised data corresponding to any denoising is used to denoise the feature data after any denoising.
[0244] Training module 1604 is used to train a neural network model based on the first denoised data corresponding to each denoising step to obtain a noise determination model. The noise determination model is used to determine the target noise data, and the target noise data is used to determine the target medical image of the tissue structure in the source medical image from different perspectives.
[0245] In one possible implementation, the acquisition module 1601 is used to acquire multiple image sets, each image set including reference medical images of the tissue structure of the same object from multiple perspectives, the reference medical images including multiple reference image blocks; clustering the features of each reference image block to obtain multiple clusters, each cluster including the features of at least one reference image block; and determining each tissue knowledge based on each cluster.
[0246] In one possible implementation, features of the reference image patch are extracted via a first extraction network; the apparatus further includes:
[0247] The extraction module is used to extract features from the third medical image through the first neural network;
[0248] The reconstruction module is used to reconstruct a fourth medical image based on the features of the third medical image.
[0249] Training module 1604 is also used to train a first neural network to obtain a first extraction network based on at least one of a first loss, a second loss, or a third loss;
[0250] Among them, the first loss represents the pixel difference between the third and fourth medical images, the second loss represents the distribution difference between the third and fourth medical images, and the third loss represents the feature difference between the third and fourth medical images.
[0251] In one possible implementation, the first medical image comprises multiple first image blocks;
[0252] The determination module 1602 is used to determine the distance between the features of any first image block and each organizational knowledge for any first image block, and to take the organizational knowledge corresponding to the distance that satisfies the distance condition as the organizational knowledge of any first image block; and to determine the sample knowledge based on the organizational knowledge of each first image block.
[0253] In one possible implementation, noise is added multiple times;
[0254] The noise-adding module 1603 is used to add noise to the image features of the second medical image for the first noise addition, to obtain the feature data after the first noise addition; and to add noise to the feature data after the previous noise addition for non-first noise addition, to obtain the feature data after non-first noise addition.
[0255] In one possible implementation, the determining module 1602 is used to obtain at least one of the following: difference information, task description, or sample medical report. The difference information characterizes the difference between the sample medical report and the first medical image. The task description describes the task of generating medical images from different perspectives based on the first medical image. The sample medical report describes the diagnostic content of the first medical image and the second medical image. The first denoised data corresponding to each denoising step is determined by a neural network model based on at least one of the difference information, task description, or sample medical report, sample knowledge, and feature data after each denoising step.
[0256] In one possible implementation, the determining module 1602 is used to determine a first probability that the sample medical report belongs to multiple categories and a second probability that the first medical image belongs to multiple categories; for any category, determine the difference between the first probability and the second probability of any category; determine a target category whose difference is not less than a reference difference; and determine difference information based on the target category.
[0257] In one possible implementation, the neural network model includes a first network part and a second network part;
[0258] The determination module 1602 is used to determine, for any noise addition, first information based on the first medical image, sample knowledge and feature data after any noise addition by the first network part, the first information representing the information after the fusion of the first medical image, sample knowledge and feature data after any noise addition; and first denoised data corresponding to any noise addition by the second network part based on the first information and feature data after any noise addition.
[0259] In one possible implementation, the determining module 1602 is used to determine second information based on the feature data after any noise addition by the second network part, the second information representing the feature data after any noise addition; and to determine the first denoised data corresponding to any noise addition by the second network part based on the first information and the second information.
[0260] In one possible implementation, the determining module 1602 is further configured to determine the second denoised data corresponding to any noise addition based on the feature data after any noise addition by the third network part, and the second denoised data corresponding to any noise addition is used to denoise the feature data after any noise addition.
[0261] Training module 1604 is also used to train the third network part based on the second denoised data corresponding to each noise addition, thus obtaining the second network part.
[0262] In one possible implementation, the training module 1604 is used to determine the loss of any noise addition based on the first denoised data and the added data corresponding to any noise addition, and the added data corresponding to any noise addition is used to add noise to obtain the feature data after any noise addition; and to train a neural network model based on the loss of each noise addition to obtain a noise determination model.
[0263] In the aforementioned device, by determining sample knowledge of the first medical image from multiple tissue knowledge sources, multi-view visual information of various tissue structures in the first medical image is obtained. Since the first and second medical images have different perspectives, by adding noise multiple times to the image features of the second medical image, and determining the first denoised data corresponding to each noise addition based on sample knowledge and the feature data after each noise addition, the device guides the model to accurately predict the noise added to the image features at a certain perspective based on multi-view visual information, thus improving the accuracy of the first denoised data. Training the neural network model with the first denoised data corresponding to each noise addition improves the model's training effect and accuracy, enabling the trained noise determination model to accurately predict the noise data that needs to be removed. By removing noise data, medical images from different perspectives are obtained, increasing the perspective richness of medical images and thus improving the accuracy of the analysis results obtained from analyzing medical images.
[0264] It should be understood that the above Figure 16 The provided device, in implementing its functions, is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.
[0265] Figure 17 The diagram shown is a structural schematic of a medical image generation device provided in an embodiment of this application. Figure 17 As shown, the device includes:
[0266] The acquisition module 1701 is used to acquire multiple tissue knowledge and source medical images, where any tissue knowledge represents the visual information of an tissue structure from multiple perspectives.
[0267] The determination module 1702 is used to determine source knowledge from multiple tissue knowledge based on the source medical image, wherein the source knowledge represents the visual information of the tissue structure in the source medical image from multiple perspectives.
[0268] Module 1702 is further used to determine target noise data based on source knowledge using a noise determination model, wherein the noise determination model follows the same rules as... Figure 2 The relevant method implementation examples were trained;
[0269] The denoising module 1703 is used to denoise the reference noise data based on the target noise data to obtain the target medical image, which is an image of the tissue structure in the source medical image from different perspectives.
[0270] In one possible implementation, the noise reduction process is repeated multiple times.
[0271] The determination module 1702 is used to determine the target noise data for the first denoising by using a noise determination model based on source knowledge and reference noise data for the first denoising.
[0272] Denoising module 1703 is used to denoise reference noise data based on the target noise data of the first denoising, and obtain feature data after the first denoising.
[0273] The determination module 1702 is used to determine the target noise data for non-first denoising by using the noise determination model based on source knowledge and the feature data after the previous denoising of non-first denoising.
[0274] The denoising module 1703 is used to denoise the feature data after the previous denoising based on the target noise data after the previous denoising, to obtain the feature data after the previous denoising; and to determine the target medical image based on the feature data after the last denoising.
[0275] In the aforementioned device, by determining the source knowledge of the source medical image from multiple tissue knowledge sources, multi-view visual information of various tissue structures in the source medical image is obtained. The noise determination model determines the target noise data based on the source knowledge, enabling the model to identify information unrelated to the source medical image and the target viewpoint. This ensures that the target medical image obtained after denoising the reference noise data based on the target noise data is relevant to the source medical image and represents the target viewpoint. Since the source knowledge contains multi-view visual information, the target viewpoint can differ from that of the source medical image. In other words, the noise determination model can generate high-quality medical images with different viewpoints, increasing the viewpoint richness of medical images and thus improving the accuracy of the analysis results obtained from analyzing medical images.
[0276] It should be understood that the above Figure 17 The provided device, in implementing its functions, is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.
[0277] Figure 18 A structural block diagram of a terminal device 1800 provided in an exemplary embodiment of this application is shown. The terminal device 1800 includes a processor 1801 and a memory 1802.
[0278] Processor 1801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1801 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1801 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0279] The memory 1802 may include one or more computer-readable storage media, which may be non-transitory. The memory 1802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1802 are used to store at least one computer program, which is executed by the processor 1801 to implement the noise determination model training method or medical image generation method provided in the method embodiments of this application.
[0280] In some embodiments, the terminal device 1800 may also optionally include a peripheral device interface 1803 and at least one peripheral device. The processor 1801, memory 1802, and peripheral device interface 1803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1803 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: radio frequency circuitry 1804, display screen 1805, camera assembly 1806, audio circuitry 1807, and power supply 1808.
[0281] Peripheral device interface 1803 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1801 and memory 1802. In some embodiments, processor 1801, memory 1802 and peripheral device interface 1803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1801, memory 1802 and peripheral device interface 1803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0282] The radio frequency (RF) circuit 1804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1804 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1804 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0283] Display screen 1805 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1805 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1801 for processing. In this case, display screen 1805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1805, disposed on the front panel of terminal device 1800; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal device 1800 or in a folded design; in still other embodiments, display screen 1805 may be a flexible display screen, disposed on a curved or folded surface of terminal device 1800. Furthermore, display screen 1805 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1805 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0284] The camera assembly 1806 is used to acquire images or videos. Optionally, the camera assembly 1806 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0285] The audio circuit 1807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 1801 for processing, or to the radio frequency circuit 1804 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal device 1800. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1801 or the radio frequency circuit 1804 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1807 may also include a headphone jack.
[0286] Power supply 1808 is used to power the various components in terminal device 1800. Power supply 1808 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1808 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0287] In some embodiments, the terminal device 1800 further includes one or more sensors 1809. The one or more sensors 1809 include, but are not limited to: an acceleration sensor 1811, a gyroscope sensor 1812, a pressure sensor 1813, an optical sensor 1814, and a proximity sensor 1815.
[0288] Accelerometer 1811 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal device 1800. For example, accelerometer 1811 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1801 can control display screen 1805 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1811. Accelerometer 1811 can also be used for games or for acquiring user motion data.
[0289] The gyroscope sensor 1812 can detect the orientation and rotation angle of the terminal device 1800. The gyroscope sensor 1812 can work in conjunction with the accelerometer sensor 1811 to acquire the user's 3D movements on the terminal device 1800. Based on the data acquired by the gyroscope sensor 1812, the processor 1801 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0290] The pressure sensor 1813 can be disposed on the side bezel of the terminal device 1800 and / or on the lower layer of the display screen 1805. When the pressure sensor 1813 is disposed on the side bezel of the terminal device 1800, it can detect the user's grip signal on the terminal device 1800, and the processor 1801 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1813. When the pressure sensor 1813 is disposed on the lower layer of the display screen 1805, the processor 1801 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0291] An optical sensor 1814 is used to collect ambient light intensity. In one embodiment, the processor 1801 can control the display brightness of the display screen 1805 based on the ambient light intensity collected by the optical sensor 1814. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1805 is increased; when the ambient light intensity is low, the display brightness of the display screen 1805 is decreased. In another embodiment, the processor 1801 can also dynamically adjust the shooting parameters of the camera assembly 1806 based on the ambient light intensity collected by the optical sensor 1814.
[0292] The proximity sensor 1815, also known as a distance sensor, is typically located on the front panel of the terminal device 1800. The proximity sensor 1815 is used to detect the distance between the user and the front of the terminal device 1800. In one embodiment, when the proximity sensor 1815 detects that the distance between the user and the front of the terminal device 1800 is gradually decreasing, the processor 1801 controls the display screen 1805 to switch from a screen-on state to a screen-off state; when the proximity sensor 1815 detects that the distance between the user and the front of the terminal device 1800 is gradually increasing, the processor 1801 controls the display screen 1805 to switch from a screen-off state to a screen-on state.
[0293] Those skilled in the art will understand that Figure 18 The structure shown does not constitute a limitation on the terminal device 1800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0294] Figure 19This is a schematic diagram of the server structure provided in the embodiments of this application. The server 1900 can vary considerably due to different configurations or performance. It may include one or more processors 1901 and one or more memories 1902. The one or more memories 1902 store at least one computer program, which is loaded and executed by the one or more processors 1901 to implement the noise determination model training method or medical image generation method provided in the above-described method embodiments. For example, the processor 1901 is a CPU. Of course, the server 1900 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1900 may also include other components for implementing device functions, which will not be elaborated here.
[0295] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one computer program that is loaded and executed by a processor to enable an electronic device to implement any of the above-described methods for training a noise determination model or for generating medical images.
[0296] Optionally, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0297] In an exemplary embodiment, a computer program is also provided, which is at least one such computer program, loaded and executed by a processor, to enable an electronic device to implement any of the above-described noise determination model training methods or medical image generation methods.
[0298] In an exemplary embodiment, a computer program product is also provided, which stores at least one computer program that is loaded and executed by a processor to enable an electronic device to implement any of the above-described noise determination model training methods or medical image generation methods.
[0299] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0300] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0301] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A training method for a noise determination model, characterized in that, The method includes: Acquire multiple tissue knowledge, a first medical image, and a second medical image. Each tissue knowledge represents the visual information of a tissue structure from multiple perspectives, and the second medical image is an image of the tissue structure in the first medical image from different perspectives. Based on the first medical image, sample knowledge is determined from the plurality of tissue knowledge, wherein the sample knowledge characterizes the visual information of the tissue structure in the first medical image from multiple perspectives; At least one noise addition is performed on the image features of the second medical image to obtain the feature data after each noise addition. Based on the sample knowledge and the feature data after each noise addition, the neural network model determines the first denoised data corresponding to each noise addition, and the first denoised data corresponding to any noise addition is used to denoise the feature data after any noise addition. The neural network model is trained based on the first denoised data corresponding to each denoising step to obtain a noise determination model. The noise determination model is used to determine the target noise data, and the target noise data is used to determine the target medical image of the tissue structure in the source medical image from different perspectives.
2. The method according to claim 1, characterized in that, The acquisition of multiple organizational knowledge includes: Acquire multiple image sets, any image set including reference medical images of the tissue structure of the same object from multiple perspectives, the reference medical images including multiple reference image blocks; Clustering the features of each reference image patch yields multiple clusters, and each cluster includes features from at least one reference image patch. Knowledge of each organization is determined based on each cluster.
3. The method according to claim 2, characterized in that, The features of the reference image blocks are extracted through a first extraction network; before clustering the features of each reference image block to obtain multiple clusters, the method further includes: Features of the third medical image are extracted using the first neural network; Based on the features of the third medical image, a fourth medical image is reconstructed; The first extraction network is obtained by training the first neural network based on at least one of the first loss, the second loss, or the third loss. Wherein, the first loss represents the pixel difference between the third medical image and the fourth medical image, the second loss represents the distribution difference between the third medical image and the fourth medical image, and the third loss represents the feature difference between the third medical image and the fourth medical image.
4. The method according to claim 1, characterized in that, The first medical image includes multiple first image patches; the step of determining sample knowledge from the multiple tissue knowledge patches based on the first medical image includes: For any first image block, determine the distance between the features of the first image block and each organizational knowledge, and take the organizational knowledge corresponding to the distance that satisfies the distance condition as the organizational knowledge of the first image block. The sample knowledge is determined based on the organizational knowledge of each first image block.
5. The method according to claim 1, characterized in that, The noise is added multiple times; the image features of the second medical image are subjected to noise at least once to obtain the feature data after each noise addition, including: For the first noise addition, noise is added to the image features of the second medical image to obtain the feature data after the first noise addition; For non-first-time noise addition, noise is added to the feature data after the previous noise addition, to obtain the feature data after non-first-time noise addition.
6. The method according to any one of claims 1 to 5, characterized in that, The step of determining the first denoised data corresponding to each denoising step using a neural network model based on the sample knowledge and the feature data after each denoising step includes: Obtain at least one of the following: difference information, task description, or sample medical report; the difference information characterizes the difference between the sample medical report and the first medical image; the task description describes the task of generating medical images from different perspectives based on the first medical image; and the sample medical report describes the diagnostic content of the first medical image and the second medical image. The neural network model determines the first denoised data corresponding to each denoising step based on the difference information, the task description, or at least one of the sample medical reports, the sample knowledge, and the feature data after each denoising step.
7. The method according to claim 6, characterized in that, The acquisition of difference information includes: Determine a first probability that the sample medical report belongs to multiple categories and a second probability that the first medical image belongs to the multiple categories; For any category, determine the difference between the first probability and the second probability of that category; Determine the target category whose difference is not less than the reference difference; The difference information is determined based on the target category.
8. The method according to any one of claims 1 to 5, characterized in that, The neural network model includes a first network part and a second network part; the step of determining the first denoised data corresponding to each denoising step based on the sample knowledge and the feature data after each denoising step using the neural network model includes: For any noise addition, the first network part determines first information based on the first medical image, the sample knowledge, and the feature data after any noise addition. The first information represents the information after the fusion of the first medical image, the sample knowledge, and the feature data after any noise addition. The second network portion determines the first denoised data corresponding to any one instance of denoising based on the first information and the feature data after any one instance of denoising.
9. The method according to claim 8, characterized in that, The step of determining the first denoised data corresponding to any one denoising step through the second network part based on the first information and the feature data after any one denoising step includes: The second network part determines second information based on the feature data after any one noise addition, and the second information represents the feature data after any one noise addition. The second network component determines the first denoised data corresponding to any one instance of noise addition based on the first information and the second information.
10. The method according to claim 8, characterized in that, Before determining the first denoised data corresponding to any one denoising step by the second network part based on the first information and the feature data after any one denoising step, the method further includes: Based on the feature data after any noise addition, the third network part determines the second denoised data corresponding to any noise addition, and the second denoised data corresponding to any noise addition is used to denoise the feature data after any noise addition. The third network part is trained based on the second denoised data corresponding to each denoising step, thus obtaining the second network part.
11. The method according to any one of claims 1 to 5, characterized in that, The process of training the neural network model based on the first denoised data corresponding to each denoising iteration to obtain a noise determination model includes: For any denoising, based on the first denoised data and the denoised data corresponding to the denoising, the loss of the denoising is determined, and the denoised data corresponding to the denoising is used to add noise to obtain the feature data after the denoising. The neural network model is trained based on the loss of each noise addition to obtain a noise determination model.
12. A method for generating medical images, characterized in that, The method includes: Acquire multiple tissue knowledge and source medical images, where any tissue knowledge represents the visual information of an tissue structure from multiple perspectives; Source knowledge is determined from the plurality of tissue knowledge based on the source medical image, and the source knowledge represents the visual information of the tissue structure in the source medical image from multiple perspectives; The noise determination model determines the target noise data based on the source knowledge, and the noise determination model is trained according to the method described in any one of claims 1 to 11. The target medical image is obtained by denoising the reference noise data based on the target noise data. The target medical image is an image of the tissue structure in the source medical image from different perspectives.
13. The method according to claim 12, characterized in that, The noise reduction process is repeated multiple times. The step of determining target noise data based on the source knowledge using a noise determination model, and denoising reference noise data based on the target noise data to obtain the target medical image includes: For the first denoising, the noise determination model determines the target noise data for the first denoising based on the source knowledge and the reference noise data, and the reference noise data is denoised based on the target noise data for the first denoising to obtain the feature data after the first denoising. For non-first-time denoising, the noise determination model determines the target noise data for non-first-time denoising based on the source knowledge and the feature data after the previous denoising of the non-first-time denoising. Based on the target noise data for non-first-time denoising, the feature data after the previous denoising is denoised to obtain the feature data after non-first-time denoising. The target medical image is determined based on the feature data after the last denoising step.
14. A training device for a noise determination model, characterized in that, The device includes: The acquisition module is used to acquire multiple tissue knowledge, a first medical image, and a second medical image. Each tissue knowledge represents the visual information of a tissue structure from multiple perspectives, and the second medical image is an image of the tissue structure in the first medical image from different perspectives. A determination module is used to determine sample knowledge from the plurality of tissue knowledge based on the first medical image, wherein the sample knowledge characterizes the visual information of the tissue structure in the first medical image from multiple perspectives; The noise-adding module is used to perform at least one noise-adding operation on the image features of the second medical image to obtain the feature data after each noise-adding operation. The determining module is further configured to determine the first denoised data corresponding to each denoising step based on the sample knowledge and the feature data after each denoising step using a neural network model, wherein the first denoised data corresponding to any denoising step is used to denoise the feature data after any denoising step. The training module is used to train the neural network model based on the first denoised data corresponding to each denoising step to obtain a noise determination model. The noise determination model is used to determine the target noise data, and the target noise data is used to determine the target medical image of the tissue structure in the source medical image from different perspectives.
15. A medical image generation device, characterized in that, The device includes: The acquisition module is used to acquire multiple tissue knowledge and source medical images, where any tissue knowledge represents the visual information of an tissue structure from multiple perspectives. A determination module is used to determine source knowledge from the plurality of tissue knowledge based on the source medical image, wherein the source knowledge characterizes the visual information of the tissue structure in the source medical image from multiple perspectives; The determining module is further configured to determine target noise data based on the source knowledge using a noise determination model, wherein the noise determination model is trained according to the method described in any one of claims 1 to 11; A denoising module is used to denoise reference noise data based on the target noise data to obtain a target medical image, wherein the target medical image is an image of the tissue structure in the source medical image from different perspectives.
16. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to enable the electronic device to perform the method as described in any one of claims 1 to 13.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the electronic device to perform the method as described in any one of claims 1 to 13.
18. A computer program product, characterized in that, The computer program product stores at least one computer program, which is loaded and executed by a processor to enable the electronic device to perform the method as described in any one of claims 1 to 13.
Citation Information
Cited By
Classification method, device and equipment for fallopian tube images and storage medium
CN121639702A