Tongue diagnosis image generation method and device, equipment and storage medium
By using reference tongue coating status information and noise images in the diffusion model for denoising processing, a tongue diagnosis image with high authenticity and diversity is generated, and the problem of insufficient diversity of tongue diagnosis images in the prior art is solved.
Patent Information
- Application Number
- CN202510069647.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-03
AI Technical Summary
Due to the limited number of true acquisitions of tongue diagnosis images, the prior art generates new images by scaling, rotating or affine methods of the original tongue diagnosis image. However, the newly generated images are too similar to the original images, resulting in less diversity in tongue diagnosis images used for training models.
By obtaining the reference tongue coating status information and randomly generated noise images, inputting them into the diffusion model for denoising processing, and generating the target tongue diagnosis image. This method uses the denoising process of the diffusion model combined with reference information to ensure the authenticity and diversity of the generated images.
While maintaining the quality of the target tongue diagnosis image, it improves its diversity and ensures that the generated images have sufficient authenticity and diversity, thereby improving the prediction accuracy of the training model.
Smart Images

Figure CN120088597A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of medical imaging technology, and particularly relates to a method, device, equipment, and storage medium for generating tongue diagnosis images. Background Art
[0002] Tongue diagnosis images are image data that record the appearance characteristics of the tongue in traditional Chinese medicine inspection. In order to improve the prediction accuracy of the tongue image classification model, it is necessary to ensure that the data scale of the tongue diagnosis images used for training the model is large enough.
[0003] Currently, due to the limited number of real acquisitions of tongue diagnosis images, in order to increase the number of tongue diagnosis images, the related technology is to scale, rotate, or perform affine transformation on the original tongue diagnosis images to generate new tongue diagnosis images. However, the newly generated tongue diagnosis images are too similar to the original tongue diagnosis images, resulting in low diversity of the tongue diagnosis images used for training the model. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of this application provide a method, device, equipment, and storage medium for generating tongue diagnosis images, which can improve the diversity of the target tongue diagnosis images while ensuring the sufficient authenticity of the generated target tongue diagnosis images.
[0006] To achieve the above object, the first aspect of the embodiments of this application proposes a method for generating a tongue diagnosis image, including: obtaining reference tongue coating state information; obtaining a randomly generated first noise image, inputting the first noise image and the reference tongue coating state information into a diffusion model, and performing denoising processing on the first noise image with the reference tongue coating state information as a constraint condition to generate a target tongue diagnosis image.
[0007] In some embodiments, the step of inputting the first noise image and the reference tongue coating state information into a diffusion model, performing denoising processing on the first noise image with the reference tongue coating state information as a constraint condition to generate a target tongue diagnosis image includes: inputting the first noise image and the reference tongue coating state information into the diffusion model, performing noise prediction with the reference tongue coating state information as a constraint condition to obtain a first predicted noise; performing denoising processing on the first noise image based on the first predicted noise to obtain a target tongue diagnosis image.
[0008] In some embodiments, the diffusion module includes a Unet network, the Unet network includes multiple levels of downsampling blocks and multiple levels of upsampling blocks, the downsampling blocks include convolutional layers and attention layers, and inputting the first noise image and the reference tongue coating state information into the diffusion model to perform noise prediction with the reference tongue coating state information as a constraint condition to obtain the first predicted noise includes: inputting the first noise image into the diffusion model, performing convolutional processing on the first noise image based on the convolutional layer in the first-level downsampling block to obtain a first intermediate feature; determining a time step, and fusing the time step and the reference tongue coating state information to obtain a conditional parameter; performing attention processing on the first intermediate feature and the conditional parameter based on the attention layer in the first-level downsampling block to obtain a second intermediate feature; and mapping the second intermediate feature based on the remaining downsampling blocks and upsampling blocks to obtain the first predicted noise.
[0009] In some embodiments, performing attention processing on the first intermediate feature and the conditional parameter based on the attention layer in the first-level downsampling block to obtain a second intermediate feature includes: performing channel attention processing on the fused first intermediate feature and conditional parameter based on the attention layer in the first-level downsampling block to obtain a first attention feature, and performing spatial attention processing on the fused first intermediate feature and conditional parameter to obtain a second attention feature; performing self-attention processing on the fused first attention feature and second attention feature to obtain a third attention feature; and performing cross-attention processing on the third attention feature and the conditional parameter to obtain the second intermediate feature. In some embodiments, before obtaining a randomly generated first noise image, inputting the first noise image and the reference tongue coating state information into a diffusion model to perform denoising processing on the first noise image with the reference tongue coating state information as a constraint condition to generate a target tongue diagnosis image, the tongue diagnosis image generation method further includes: obtaining a sample tongue diagnosis image and sample tongue coating state information, adding a first sample noise to the sample tongue diagnosis image to obtain a second noise image; inputting the second noise image and the sample tongue coating state information into the diffusion model to perform noise prediction on the second noise image with the reference tongue coating state information as a constraint condition to obtain a second predicted noise; determining a first loss based on the difference between the first sample noise and the second predicted noise; denoising the sample tongue diagnosis image based on the second predicted noise to obtain a reconstructed tongue diagnosis image; determining a second loss based on the difference between the sample tongue diagnosis image and the reconstructed tongue diagnosis image; and training the diffusion model based on the first loss and the second loss.
[0010] In some embodiments, inputting the first noise image and the reference tongue coating state information into a diffusion model, and performing denoising processing on the first noise image with the reference tongue coating state information as a constraint condition to generate a target tongue diagnosis image includes: obtaining a disease description text corresponding to the reference tongue coating state information; inputting the disease description text into a large language model for text generation to obtain a first tongue coating description text; inputting the first noise image, the reference tongue coating state information, and the first tongue coating description text into the diffusion model, and performing denoising processing on the first noise image with the reference tongue coating state information and the first tongue coating description text as constraint conditions to generate a target tongue diagnosis image.
[0011] In some embodiments, inputting the first noise image, the reference tongue coating state information, and the first tongue coating description text into a diffusion model, and performing denoising processing on the first noise image with the reference tongue coating state information and the first tongue coating description text as constraint conditions to generate a target tongue diagnosis image includes: obtaining a reference tongue diagnosis image corresponding to the reference tongue coating state information; performing edge detection on the reference tongue diagnosis image to obtain an edge image; inputting the first noise image, the reference tongue coating state information, the first tongue coating description text, and the edge image into the diffusion model, and performing denoising processing on the first noise image with the reference tongue coating state information, the first tongue coating description text, and the edge image as constraint conditions to generate a target tongue diagnosis image.
[0012] To achieve the above object, a second aspect of the embodiments of the present application provides a tongue diagnosis image generation device, including: an acquisition module, configured to acquire reference tongue coating state information; a generation module, configured to acquire a randomly generated first noise image, input the first noise image and the reference tongue coating state information into a diffusion model, and perform denoising processing on the first noise image with the reference tongue coating state information as a constraint condition to generate a target tongue diagnosis image.
[0013] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the tongue diagnosis image generation method described in the first aspect above is implemented.
[0014] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, where the storage medium is a computer-readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the tongue diagnosis image generation method described in the first aspect above is implemented.
[0015] The embodiments of the present application at least include the following beneficial effects: obtaining reference tongue coating state information, then obtaining a randomly generated first noise image, inputting the first noise image and the reference tongue coating state information into a diffusion model, and performing denoising processing on the first noise image with the reference tongue coating state information as a constraint condition. The reference tongue coating state information can guide the diffusion model to reconstruct a target tongue diagnosis image related to the reference tongue coating state information, thereby ensuring that the target tongue diagnosis image has sufficient authenticity, and thus obtaining a high-quality target tongue diagnosis image. Moreover, the first noise image is randomly generated, enabling the diffusion model to reconstruct diverse target tongue diagnosis images during the denoising process, realizing the exploration of a wider solution space while maintaining the quality of the target tongue diagnosis image, and being able to improve the diversity of the target tongue diagnosis image while ensuring the sufficient authenticity of the generated target tongue diagnosis image.
[0016] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. They are used together with the embodiments of the present application to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0018] Figure 1 It is an optional flowchart of the tongue diagnosis image generation method provided by the embodiments of the present application;
[0019] Figure 2 It is an optional flowchart of the denoising processing provided by the embodiments of the present application;
[0020] Figure 3 It is an optional flowchart of the noise prediction provided by the embodiments of the present application;
[0021] Figure 4 It is an optional flowchart of the attention processing provided by the embodiments of the present application;
[0022] Figure 5 It is an optional flowchart of the model training provided by the embodiments of the present application;
[0023] Figure 6 It is an optional flowchart of the linked large language model provided by the embodiments of the present application;
[0024] Figure 7 It is an optional flowchart of adding an edge image as a constraint condition provided by the embodiments of the present application;
[0025] Figure 8 An optional architecture diagram of the Unet network provided by the embodiment of the present application;
[0026] Figure 9 An optional architecture diagram of the residual unit provided by the embodiment of the present application;
[0027] Figure 10 An optional process schematic diagram of the processing within the convolutional layer provided by the embodiment of the present application;
[0028] Figure 11 An optional specific process schematic diagram of the attention processing provided by the embodiment of the present application;
[0029] Figure 12 An optional process schematic diagram of the diffusion model training provided by the embodiment of the present application;
[0030] Figure 13 An optional process schematic diagram of the model training method provided by the embodiment of the present application;
[0031] Figure 14 An optional structural schematic diagram of the tongue diagnosis image generation device provided by the embodiment of the present application;
[0032] Figure 15 An optional hardware structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0033] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0034] It should be noted that in each specific implementation manner of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or attribute information set, etc., the permission or consent of the target object will be obtained first. Moreover, the collection, use and processing of these data will comply with relevant laws, regulations and standards. Among them, the target object can be a user. In addition, when the embodiment of the present application needs to obtain the target object attribute information, it will obtain the separate permission or separate consent of the target object through pop-up windows or jumping to the confirmation page, etc. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiment of the present application will be obtained.
[0035] In the description of this application, the meaning of "several" is one or more, the meaning of "multiple" is two or more, "greater than", "less than", "exceeding", etc. are understood as not including the corresponding number, and "above", "below", "within", etc. are understood as including the corresponding number.
[0036] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the device or a different sequence from that in the flowchart. Terms such as "first", "second", etc. in the specification, claims or the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0037] To facilitate the understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application are explained here first:
[0038] Artificial Intelligence (AI) is to use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0039] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0040] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0041] Tongue diagnosis images are image data that record the appearance characteristics of the tongue in traditional Chinese medicine inspection. To improve the prediction accuracy of the tongue image classification model, it is necessary to ensure that the data scale of the tongue diagnosis images used for training the model is large enough.
[0042] Currently, due to the limited number of real collected tongue diagnosis images, in order to increase the number of tongue diagnosis images, related technologies are to scale, rotate or perform affine transformation on the original tongue diagnosis images to generate new tongue diagnosis images. However, the newly generated tongue diagnosis images are too similar to the original tongue diagnosis images, resulting in low diversity of the tongue diagnosis images used for training the model.
[0043] Aiming at the problem of low diversity of the generated images, this application provides a tongue diagnosis image generation method, device, equipment and storage medium. The method includes: obtaining reference tongue coating state information; obtaining a randomly generated first noise image, inputting the first noise image and the reference tongue coating state information into a diffusion model, and performing denoising processing on the first noise image with the reference tongue coating state information as a constraint condition to generate a target tongue diagnosis image. According to the solution provided by the embodiments of this application, the reference tongue coating state information is obtained, then the randomly generated first noise image is obtained, the first noise image and the reference tongue coating state information are input into the diffusion model, and the first noise image is denoised with the reference tongue coating state information as a constraint condition. The reference tongue coating state information can guide the diffusion model to reconstruct a target tongue diagnosis image related to the reference tongue coating state information, thereby ensuring that the target tongue diagnosis image has sufficient authenticity, and thus obtaining a high-quality target tongue diagnosis image. Moreover, the first noise image is randomly generated, so that the diffusion model can reconstruct a diverse target tongue diagnosis image during the denoising process, realizing exploring a wider solution space while maintaining the quality of the target tongue diagnosis image, and being able to improve the diversity of the target tongue diagnosis image while ensuring that the generated target tongue diagnosis image has sufficient authenticity.
[0044] The tongue diagnosis image generation method, device, equipment and storage medium provided by the embodiments of this application will be specifically described through the following embodiments. First, the tongue diagnosis image generation method in the embodiments of this application will be described.
[0045] The tongue diagnosis image generation method provided by the embodiments of this application relates to the field of computer technology. The tongue diagnosis image generation method provided by the embodiments of this application can be applied to a terminal, or can be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the tongue diagnosis image generation method, etc., but is not limited to the above forms.
[0046] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0047] The following further elaborates on the embodiments of this application with reference to the accompanying drawings.
[0048] As Figure 1 shown, Figure 1 is an optional flowchart of the tongue diagnosis image generation method provided by the embodiments of this application. This tongue diagnosis image generation method can be executed by a server, or can also be executed by a terminal, or can also be executed by a server in cooperation with a terminal. This tongue diagnosis image generation method includes, but is not limited to, the following steps S110 to step S120:
[0049] Step S110, obtain reference tongue coating state information.
[0050] Step S120, obtain a randomly generated first noise image, input the first noise image and the reference tongue coating state information into a diffusion model, and perform denoising processing on the first noise image with the reference tongue coating state information as a constraint condition to generate a target tongue diagnosis image.
[0051] It should be noted that the tongue image is an important basis for evaluating the health status of patients by observing and recording the appearance characteristics of the tongue in traditional Chinese medicine inspection. These characteristics, such as the color of the tongue coating, the shape of the tongue body, the distribution of cracks, etc., are usually saved in the form of images or text descriptions for assisting in diagnosis and the formulation of treatment plans.
[0052] Among them, the reference tongue coating state information can be used to characterize the color and texture of the tongue coating. For example, the reference tongue coating state information can be white greasy coating, thin white coating, yellow greasy coating, thin yellow coating, or grayish-black coating, etc.
[0053] Among them, the first noise image can be randomly generated by Gaussian noise.
[0054] Among them, the diffusion model has two main processes: the forward process and the backward process; the forward process is also called the diffusion process. In the forward process, the diffusion model gradually adds Gaussian noise to the image until the image becomes completely random noise. Exemplarily, the overall diffusion process can be a parameterized Markov chain; the backward process is also called the inverse diffusion process. In the backward process, the diffusion model uses a series of Markov chains to gradually remove the predicted noise at each time step, so as to recover the data from the Gaussian noise; therefore, the diffusion model can perform denoising processing on the first noise image based on the reference tongue coating state information in the backward process, so as to generate the target tongue diagnosis image.
[0055] The specific generation process of the target tongue diagnosis image is as follows: First, use the diffusion model to gradually remove the noise in the first noise image through multiple time steps. During the process of removing the noise, at the same time, use the reference tongue coating state information as a constraint condition, and a high-quality and detail-rich target tongue diagnosis image can be recovered. Since the constraint condition includes the reference tongue coating state information, when predicting the noise at each time step, the reference tongue coating state information is injected into the diffusion model, so that the diffusion model considers the reference tongue coating state information in each step of the diffusion process to ensure that the target tongue diagnosis image matches the reference tongue coating state information respectively, and a more natural tongue diagnosis image is generated.
[0056] It should be noted that the target tongue diagnosis image can be various types of images. For example, the target tongue diagnosis image can be an RGB image, a grayscale image, a binary image, etc. The embodiments of the present disclosure do not make limitations here.
[0057] It should be noted that when the diffusion model adopts the implicit diffusion method and the non-implicit diffusion method, different methods are needed to perform denoising processing on the first noise image.
[0058] The following is a detailed description of the implicit diffusion method, which requires the introduction of an encoding network and a decoding network. The encoding network is used to map the input of the diffusion model in the original data space to the target latent space, and the decoding network is used to restore the output of the diffusion model in the target latent space to the original data space. Specifically, the encoding network encodes the input data distribution into a probability distribution of a latent variable, usually a Gaussian distribution, and then samples a latent variable from this probability distribution. The decoding network then remaps this latent variable to the pixel space to reconstruct the original image. For example, after the encoding network encodes the image in the original space, a 64×64 image becomes a 16×16 latent vector. Therefore, after determining the first noise image, first, the encoding network encodes the first noise image to obtain a first encoded image, then denoises the first encoded image with the reference tongue coating state information as a constraint condition to obtain a second encoded image, and the second encoded image can be decoded by the decoding network to be transformed into the target tongue diagnosis image.
[0059] It can be understood that the diffusion model can effectively reduce the computational complexity and improve the training efficiency of the model through the implicit diffusion method. At the same time, the representation in the latent space is more compact and abstract, which helps the model to capture the key features and structures of the data more accurately, thereby improving the quality of the generated images.
[0060] Based on this, obtain the reference tongue coating state information, then obtain a randomly generated first noise image, input the first noise image and the reference tongue coating state information into the diffusion model, and denoise the first noise image with the reference tongue coating state information as a constraint condition. The reference tongue coating state information can guide the diffusion model to reconstruct the target tongue diagnosis image related to the reference tongue coating state information, thereby ensuring that the target tongue diagnosis image has sufficient authenticity, so as to obtain a high-quality target tongue diagnosis image. Moreover, the first noise image is randomly generated, so that the diffusion model can reconstruct a diverse target tongue diagnosis image during the denoising process, realizing the exploration of a wider solution space while maintaining the quality of the target tongue diagnosis image, and being able to improve the diversity of the target tongue diagnosis image while ensuring the sufficient authenticity of the generated target tongue diagnosis image. In addition, referring to Figure 2 , in an embodiment, input the first noise image and the reference tongue coating state information into the diffusion model, and denoise the first noise image with the reference tongue coating state information as a constraint condition to generate the target tongue diagnosis image, including but not limited to the following steps:
[0061] Step S210, input the first noise image and the reference tongue coating state information into the diffusion model, and perform noise prediction with the reference tongue coating state information as a constraint condition to obtain a first predicted noise.
[0062] Step S220: Denoise the first noise image based on the first predicted noise to obtain the target tongue diagnosis image.
[0063] Among them, the dimension of the first predicted noise is the same as that of the first noise image. Therefore, the denoising of the first noise image can be achieved by subtracting the first predicted noise from the first noise image.
[0064] It should be noted that the denoising of the first noise image is carried out step by step multiple times. Therefore, each denoising process corresponds to a time step. That is, after denoising the first noise image at the current time step, it is necessary to denoise the first noise image again at the next time step.
[0065] Among them, since multiple predicted noises are made, it is necessary to determine the current time step each time a predicted noise is made. As the denoising process progresses, the first determined time step will gradually decrease to the initial time step. For example, the first time step is pre-set to 50, the next time step is 49, and the next time step is 48, until the time step is 0, which indicates that the noise has been completely removed.
[0066] Based on this, since the diffusion model gradually predicts the noise of the first noise image to obtain the first predicted noise, and then denoises the first noise image based on the first predicted noise to obtain the target tongue diagnosis image, it can explore a wider solution space while maintaining the quality of the target tongue diagnosis image, thereby improving the diversity of generating the target tongue diagnosis image.
[0067] In addition, referring to Figure 3 , in one embodiment, the diffusion module includes a Unet network. The Unet network includes multiple levels of downsampling blocks and multiple levels of upsampling blocks. The downsampling block includes a convolutional layer and an attention layer. Input the first noise image and the reference tongue coating state information into the diffusion model, and use the reference tongue coating state information as a constraint condition to predict the noise to obtain the first predicted noise, including but not limited to the following steps:
[0068] Step S310: Input the first noise image into the diffusion model, and perform convolutional processing on the first noise image based on the convolutional layer in the first-level downsampling block to obtain the first intermediate feature.
[0069] Step S320: Determine the time step, and fuse the time step and the reference tongue coating state information to obtain the conditional parameter.
[0070] Step S330: Based on the attention layer in the first-level downsampling block, perform attention processing on the first intermediate feature and the conditional parameter to obtain the second intermediate feature.
[0071] Step S340: Map the second intermediate feature based on the remaining downsampling blocks, the remaining attention layers, and the upsampling block to obtain the first predicted noise.
[0072] Among them, the time step needs to be embedded as the first embedding vector corresponding to the time step, and the reference tongue coating state information needs to be embedded as the second embedding vector corresponding to the reference tongue coating state information step. The first embedding vector and the second embedding vector are concatenated to obtain the third embedding vector corresponding to the conditional parameter, and then the third embedding vector is input into the model to provide time information and reference tongue coating state information.
[0073] It should be noted that, referring to Figure 8 , Figure 8 is an optional architecture diagram of the Unet network provided by the embodiments of the present application. The upsampling block may include three sequentially connected residual units, and the number of upsampling blocks is the same as the number of downsampling blocks;
[0074] The convolutional layer of the downsampling layer obtains a residual unit through a residual connection. The downsampling block includes two attention layers and two residual units, and the attention layer and the residual unit are alternately connected. For example, the connection order of the residual unit, the attention layer, the residual unit, and the attention layer. After the output of the last attention layer in each downsampling block is downsampled, it is input into the downsampling block or the residual unit of the next level. The upsampling blocks at each level correspond one-to-one with the downsampling blocks at each level;
[0075] Specifically, a convolutional layer for adjusting the number of channels of the first noise image is provided before the downsampling block of the first level, that is, the downsampling block of the topmost level. Specifically, the number of channels of the first noise image can be compressed;
[0076] Then, since the shared denoising Unet network is used to predict the injection noise at any step, and the injection noise is time-varying, the training difficulty and instability of the Unet network during the training process are relatively large. Therefore, the output of the last attention layer at the same level also needs to be multiplied by one-half of the square root before being input into the first residual unit of the upsampling layer at the same level to alleviate the instability of the diffusion model and reduce the training difficulty;
[0077] The output of the downsampling block of the bottommost level passes through two residual units in sequence and then is input into the upsampling block of the bottommost level;
[0078] A convolutional layer is provided after the upsampling block of the topmost level. This convolutional layer is a one-dimensional convolutional layer, which is used to restore the output of the upsampling block of the topmost level to the original number of channels to obtain the first predicted noise. Therefore, the dimension of the finally output first predicted noise is the same as the dimension of the first noise image.
[0079] Among them, referring to Figure 9, Figure 9 It is an optional architecture diagram of the residual unit provided by the embodiment of the present application. The residual unit includes three convolutional layers. For example, the first input feature is respectively input into the first two convolutional layers. One of the first two convolutional layers is a one-dimensional convolutional layer. The outputs of the first two convolutional layers are added to obtain a third fusion result. The third fusion result is input into the third convolutional layer, and the output of the third convolutional layer is added to the third fusion result to obtain the first output feature.
[0080] Among them, referring to Figure 10 , Figure 10 It is an optional process schematic diagram of the processing within the convolutional layer provided by the embodiment of the present application. For example, after the second input feature is input into the convolutional layer and undergoes convolution, normalization, and activation in sequence, the second output feature is obtained.
[0081] Based on this, the first noise image is input into the diffusion model. The first noise image is downsampled based on the downsampling block of the first level to obtain a first intermediate feature, so as to adjust the number of channels, reduce the computational amount of the model, and improve the training and inference speed of the model. Then, the time step is determined, and the time step and the reference tongue coating state information are fused to obtain a conditional parameter. Based on the attention layer connected to the downsampling block of the first level, attention processing is performed on the first intermediate feature and the conditional parameter to obtain a second intermediate feature. Finally, the second intermediate feature is mapped based on the remaining downsampling blocks, the remaining attention layers, and the upsampling block to obtain a first predicted noise, which can enhance the model's attention to important information during feature extraction, is beneficial to the extraction of fine features, and improves the picture quality of the generated target tongue diagnosis image.
[0082] In addition, referring to Figure 4 , in one embodiment, based on the attention layer in the downsampling block of the first level, performing attention processing on the first intermediate feature and the conditional parameter to obtain a second intermediate feature includes, but is not limited to, the following steps:
[0083] Step S410, based on the attention layer in the downsampling block of the first level, after fusing the first intermediate feature and the conditional parameter, perform channel attention processing to obtain a first attention feature, and after fusing the first intermediate feature and the conditional parameter, perform spatial attention processing to obtain a second attention feature.
[0084] Step S420, based on the fusion of the first attention feature and the second attention feature, perform self-attention processing to obtain a third attention feature.
[0085] Step S430, based on the third attention feature and the conditional parameter, perform cross-attention processing to obtain a second intermediate feature.
[0086] Among them, referring to Figure 11 , Figure 11It is an optional specific process schematic diagram of attention processing provided by an embodiment of the present application. After the first intermediate feature is normalized, it is fused with the conditional parameter to reduce the adverse effects of mini-batch data type training and improve the training efficiency. This step helps to reduce the scale difference of the input features, enabling the subsequent layers to process the data more effectively;
[0087] Specifically, the fusion process is to concatenate the first intermediate feature and the third embedding vector corresponding to the conditional parameter, or to convert the third embedding vector corresponding to the conditional parameter into the same dimension as the first intermediate feature and then add it to the first intermediate feature. The embodiments of the present disclosure do not limit this here;
[0088] Then, channel attention processing is performed on the first fusion result of the above fusion process, that is, after global average pooling and global max pooling on the first fusion result, a shared multi-layer perceptron (MLP) is used to process the results of global average pooling and global max pooling to generate channel weights, and the first fusion result is multiplied by the channel weights to obtain the first attention feature. The channel attention module selectively emphasizes the feature channels with high contributions by learning the importance weights of each channel;
[0089] And spatial attention processing is performed on the first fusion result of the above fusion process, that is, global average pooling and global max pooling are performed on the first fusion result in the channel dimension, and then a convolutional layer is applied to process the results of global average pooling and global max pooling to generate spatial weights, and the first fusion result is multiplied by the spatial weights to obtain the second attention feature. The spatial attention module selectively emphasizes the feature positions with high contributions by learning the importance weights of each spatial position;
[0090] Then, the first attention feature and the second attention feature are added to obtain the second fusion result. The process of self-attention processing of the second fusion result is as follows: The first numerical parameter matrix, the first query parameter matrix, and the first key-value parameter matrix are pre-trained. The first numerical parameter matrix, the first query parameter matrix, and the first key-value parameter matrix are multiplied by the second fusion result respectively to obtain the first numerical matrix V1, the first query matrix Q1, and the first key-value matrix K1. After performing Rotary Position Embedding (ROPE) on the first query matrix Q1 and the first key-value matrix K1, scaled dot-product attention processing is performed with the first numerical matrix V1 to obtain the first intermediate attention feature. After performing convolution on the first intermediate attention feature, it is added to the second fusion result to obtain the third attention feature;
[0091] The process of cross-attention processing is as follows: Randomly generate a second numerical parameter matrix, a second query parameter matrix, and a second key-value parameter matrix. The second numerical parameter matrix and the second key-value parameter matrix are respectively multiplied by the third embedding vector corresponding to the conditional parameter to obtain a second numerical matrix V2 and a second key-value matrix K2. The second query parameter matrix is multiplied by the third attention feature to obtain a second query matrix Q2. The second numerical matrix V2 and the second key-value matrix K2 perform scaled dot-product attention processing with the second query matrix Q2 to obtain a second intermediate attention feature. After performing convolution on the second intermediate attention feature, it is added to the third attention feature to obtain a third intermediate attention feature. After performing convolution on the third intermediate attention feature, it is added to the second fusion result to obtain a second intermediate feature.
[0092] It should be noted that the rotational position encoding uses a rotation matrix to convert the absolute position encoding into relative position encoding. The specific steps are as follows: First, use the sine function and cosine function to generate the absolute position encoding, and the absolute position encoding can be represented by the following formula:
[0093]
[0094] where p x and p y are respectively the abscissa and ordinate of the pixel position, i is the dimension index of the embedding vector, specifically the dimension index of the first query matrix Q1 or the first key-value matrix K1, d model is the dimension of the model, which is equivalent to the number of channel dimensions of the first query matrix Q1 or the first key-value matrix K1. PE(p x , 2i) and PE(p y , 2i) are both absolute position encodings of even dimensions, and PE(p x , 2i + 1) and PE(p y , 2i + 1) are both absolute position encodings of odd dimensions;
[0095] Then, based on the absolute position encoding, determine the rotation matrix. For the position (p x , p y ) of the pixel in the image, the corresponding rotation matrix is represented by the following formula:
[0096]
[0097] where R(θ x ) is the rotation matrix corresponding to the abscissa of the pixel, R(θ y ) is the rotation matrix corresponding to the ordinate of the pixel, θ x and θ y are both frequency parameters, p x and p yare the abscissa and ordinate of the pixel position respectively, i is the dimension index of the embedding vector, specifically the dimension index of the first query matrix Q1 or the first key-value matrix K1, and d model is the dimension of the model, which is equivalent to the number of channel dimensions of the first query matrix Q1 or the first key-value matrix K1.
[0098] Then, based on the rotation matrix and pixel values, the relative position encoding is determined, and the relative position encoding can be expressed by the following formula:
[0099]
[0100] where, x′(p x , 2i) is the relative position encoding of the pixel with abscissa p x and even dimensions, x′(p x , 2i + 1) is the relative position encoding of the pixel with abscissa p x and odd dimensions, x(p x , 2i) is the pixel value of the pixel with abscissa p x and even dimensions, x(p x , 2i + 1) is the pixel value of the pixel with abscissa p x and odd dimensions, R(θ x ) is the rotation matrix corresponding to the abscissa of the pixel; y′(p y , 2i) is the relative position encoding of the pixel with abscissa p y and even dimensions, y′(p y , 2i + 1) is the relative position encoding of the pixel with abscissa p y and odd dimensions, y(p y , 2i) is the pixel value of the pixel with abscissa p y and even dimensions, y(p y , 2i + 1) is the pixel value of the pixel with abscissa p y and odd dimensions, R(θ y ) is the rotation matrix corresponding to the ordinate of the pixel.
[0101] It can be understood that ROPE not only retains the intuitiveness of the absolute position encoding but also introduces the flexibility of the relative position encoding, improves the performance of the model in processing image data, and further enhances the ability of self-attention processing to capture the dependencies between different positions in the image.
[0102] Based on this, based on the attention layer, after fusing the first intermediate feature and the conditional parameter, channel attention processing is performed to obtain the first attention feature, and spatial attention processing is performed after fusing the first intermediate feature and the conditional parameter to obtain the second attention feature, selectively emphasizing the feature channels with high contribution, and then self-attention processing is performed based on the fusion of the first attention feature and the second attention feature to obtain the third attention feature, which can capture the dependencies between positions in the image to improve the network's ability to learn features, and then cross-attention processing is performed based on the third attention feature and the conditional parameter to obtain the second intermediate feature to enhance conditional generation, enabling the model to more effectively learn category conditional information such as the reference tongue coating state information.
[0103] In addition, referring to Figure 5 , in one embodiment, a randomly generated first noise image is obtained, and the first noise image and the reference tongue coating state information are input into the diffusion model. Before generating the target tongue diagnosis image by denoising the first noise image with the reference tongue coating state information as a constraint condition, the tongue diagnosis image generation method further includes but is not limited to the following steps:
[0104] Step S510, obtain a sample tongue diagnosis image and sample tongue coating state information, and add a first sample noise to the sample tongue diagnosis image to obtain a second noise image.
[0105] Step S520, input the second noise image and the sample tongue coating state information into the diffusion model, and perform noise prediction on the second noise image with the reference tongue coating state information as a constraint condition to obtain a second predicted noise.
[0106] Step S530, determine a first loss based on the difference between the first sample noise and the second predicted noise.
[0107] Step S540, denoise the sample tongue diagnosis image based on the second predicted noise to obtain a reconstructed tongue diagnosis image.
[0108] Step S550, determine a second loss based on the difference between the sample tongue diagnosis image and the reconstructed tongue diagnosis image.
[0109] Step S560, train the diffusion model based on the first loss and the second loss.
[0110] It should be noted that in the forward process, the sample tongue diagnosis image is added with the first sample noise multiple times at multiple time steps to obtain the second noise image. In terms of data distribution, the sample tongue diagnosis images at two adjacent time steps follow a Gaussian transition process, as shown in the following formula:
[0111]
[0112] where t represents the time step, x tis the state of the sample tongue diagnosis image at time step t, x t-1 is the state of the second noise image at time step t-1, q(x t |x t-1 ) represents the probability distribution of x t-1 under the condition of x t , N(.) represents the Gaussian distribution, and β t is the noise intensity, and I is the identity matrix.
[0113] Among them, x t can be directly obtained through reparameterization of x t-1 as shown in the following formula:
[0114]
[0115] α t := 1-β t
[0116]
[0117] Among them, t represents the time step, and x t is the state of the sample tongue diagnosis image at time step t, and x 0 is the state of the sample tongue diagnosis image before adding the first sample noise. ε follows a normal distribution with a mean of 0 and a variance of 1, and β t is the noise intensity, and α t is the proportion of not adding noise, that is, the proportion of retaining the original information. is the proportion of retaining the original information accumulated from the beginning to the current time step t. Therefore, through the above formula, the second noise image can finally be obtained.
[0118] It should be noted that in the backward process, that is, in the inverse diffusion process, the second noise image is gradually denoised to the reconstructed tongue diagnosis image at each time step. The inverse diffusion process can be expressed by the following formula:
[0119]
[0120] p(x T ) = N(x T ; 0,1)
[0121]
[0122] Among them, t represents the time step, and p θ (x 0 : T ) is the inverse diffusion process, p(x T ) is the data distribution of the second noise image, x T is the second noise image, and p θ (xt-1 |x t ) is a parameterized Gaussian distribution, where x t is the state of the second noisy image at time step t, and x t-1 is the state of the second noisy image at time step t - 1. μ θ (x t , t) is the mean, and Σ θ (x t , t) is the variance. μ θ (x t , t) and Σ θ (x t , t) can be calculated by the diffusion model. Therefore, through the above formula, the reconstructed tongue diagnosis image can finally be obtained.
[0123] It should be noted that the optimization objective of the diffusion model is to minimize the KL divergence between the true posterior distribution and the posterior distribution estimated by the model:
[0124]
[0125] where Y is the objective function of the diffusion model, q(x t-1 |x t , x 0 ) is the true posterior distribution, p θ (x t-1 |x t ) is the posterior distribution estimated by the model, t represents the time step, t belongs to 1 to T. For q(x t-1 |x t , x 0 ), x t is the state of the sample tongue diagnosis image at time step t, and x t-1 is the state of the sample tongue diagnosis image at time step t - 1. For p θ (x t-1 |x t ), x t is the state of the second noisy image at time step t, and x t-1 is the state of the second noisy image at time step t - 1, and x 0 is the state of the sample tongue diagnosis image before adding the first sample noise.
[0126] It should be noted that the Unet network is used to predict the first sample noise added at each time step, and its loss function, i.e., the first loss, can be expressed by the following formula:
[0127]
[0128] where, λ tis a weighting scheme for DDPM (Denoising Diffusion Probabilistic Models), β t is the noise intensity, α t is the proportion without adding noise, that is, the proportion of preserving the original information. Therefore, L simpls is the first loss, ε is the first sample noise, ε θ is the second predicted noise, t represents the time step, x t is the state of the second noisy image before denoising at time step t, and c is the constraint condition.
[0129] Among them, the goal of the Unet network is to learn the noise of the image damaged by different levels of noise. The closer the time step is to T, the higher the level of noise.
[0130] It should be noted that by resetting the weighting scheme of the objective function, the diffusion model can converge faster and improve performance. The new weighting scheme can be expressed by the following formula:
[0131]
[0132] Among them, t is the time step, λ t is the weighting scheme of the original DDPM, λ′ t is the new weighting scheme of DDPM, SNR(t) is the signal-to-noise ratio corresponding to different time steps t, is the proportion of preserving the original information accumulated from the beginning to the current time step t. γ is a hyperparameter, γ controls the intensity of reducing the weight in the cleaning stage, k is also a hyperparameter, k can prevent the weight explosion under extremely low signal-to-noise ratios, and determines the clarity of the weighting scheme; weight is the weight corresponding to the new weighting scheme.
[0133] It can be understood that SNR(t) is a monotonically decreasing function. In the inverse diffusion process, according to the size of the signal-to-noise ratio, it can be divided into three stages. The early stage is the rough stage, with a relatively low signal-to-noise ratio and more noise in the image. The model learns rough features, such as the global color structure; the middle stage is the content stage, where the model perceives rich content. The last stage is the cleaning stage, with a relatively high signal-to-noise ratio and only a small amount of noise in the image. The model learns imperceptible details in this stage; medical images are content-sensitive and need to learn more accurate features. The new weighting scheme selects to assign the minimum weight to the unnecessary cleaning stage and larger weights to other stages, thereby encouraging the model to learn richer visual concepts.
[0134] It should be noted that the information of the constraint conditions can be randomly discarded during training through Classifier-Free Guidance (CFG), enabling the model to have the ability to learn unconditional generation and conditional generation simultaneously; in the backward process, CFG predicts the second predicted noise at each time step through the following formula:
[0135]
[0136] Among them, is the actually predicted second predicted noise, ε θ (x t ,t,c) is the noise prediction for conditional generation, that is, the second predicted noise predicted when considering the constraint conditions, ε θ (x t ,t) is the noise prediction for unconditional generation, that is, the second predicted noise predicted without considering the constraint conditions, ω is the balance coefficient (guidance scale) between unconditional generation and conditional generation, and the details and fidelity of the second predicted noise can be controlled by adjusting ω. A higher guidance scale will make the generated result more conform to the conditional information, but may sacrifice a certain amount of diversity, while a lower guidance scale will increase the diversity of the generated result, but may reduce the consistency with the conditional information.
[0137] Therefore, combining the new weighting scheme and the prediction scheme based on CFG, the first loss can be expressed by the following formula:
[0138]
[0139] Among them, L′ sinple is the new first loss, λ′ t is the weighting scheme of the new DDPM, β t is the noise intensity, α t is the proportion without adding noise, that is, the proportion of retaining the original information. Therefore, weight is the weight corresponding to the new weighting scheme, ε is the first sample noise, is the actually predicted second predicted noise, t is the time step, x t is the state of the second noise image at time step t, and c is the constraint condition.
[0140] It should be noted that when the diffusion model adopts the implicit diffusion method and the non-implicit diffusion method, different methods are required to add noise to the first noise image, determine the first loss, and determine the second loss.
[0141] Among them, when adopting the non-implicit diffusion method, noise can be directly added to the sample tongue diagnosis image, noise prediction can be directly performed on the second noise image when determining the first loss, and the second loss can be directly determined based on the difference between the reconstructed tongue diagnosis image and the sample tongue diagnosis image; refer to Figure 12 , Figure 12 is an optional process schematic diagram for training the diffusion model provided by the embodiments of the present application. When adopting the implicit diffusion method, it is necessary to encode the sample tongue diagnosis image through an encoding network to obtain a third encoded image. Then, during the diffusion process, starting from the initial time step 0, the third encoded image stack is noise-added step by step based on the first sample noise. When the time step is T, a fourth encoded image is obtained. When determining the first loss, starting from the time step T, the time step and the reference tongue coating state information input to the Unet network are used for noise prediction, and the fourth encoded image is denoised based on the second predicted noise. When the time step is 0, a fifth encoded image is obtained. Finally, the fifth encoded image is decoded based on the decoding network to obtain the reconstructed tongue diagnosis image, and the second loss is determined based on the difference between the reconstructed tongue diagnosis image and the sample tongue diagnosis image.
[0142] It should be noted that when adopting the implicit diffusion method, the formula for the second loss is as follows:
[0143] L loss =A·MSE + B·(1 - SSIM) + C·LPIPS
[0144] Among them, L loss is the second loss, and A, B, and C are weight coefficients respectively, used to adjust the contribution of each part of the loss to the total loss; MSE is the first sub-loss of the second loss, used to measure the mean square difference between the predicted value and the true value; SSIM is the second sub-loss of the second loss, used to quantify the structural similarity between two images; LPIPS is the third sub-loss of the second loss, used to measure the difference between two images;
[0145] Specifically, the first sub-loss of the second loss can be expressed by the following formula:
[0146]
[0147] Among them, m is the number of sample tongue diagnosis images used to train the diffusion model, x i is the i-th sample tongue diagnosis image, and y i is the predicted value of the i-th sample tongue diagnosis image, that is, the reconstructed tongue diagnosis image.
[0148] Specifically, the second sub-loss of the second loss can be expressed by the following formula:
[0149] SSIM = l α ·cβ ·s γ
[0150]
[0151] Among them, SSIM is the second sub-loss of the second loss, l is the luminance, c is the contrast, s is the structure, and α, β, and γ are used to represent the relative importance of the corresponding metrics; μ x represents the mean of the sample tongue diagnosis image, μ y represents the mean of the reconstructed tongue diagnosis image; σ x represents the variance of the sample tongue diagnosis image, σ y represents the variance of the reconstructed tongue diagnosis image; σ xy represents the covariance between the sample tongue diagnosis image and the reconstructed tongue diagnosis image;
[0152] It can be understood that, different from MSE, SSIM measures the structural similarity by imitating the Human Visual System (HSV), is more sensitive to the perception of local result changes of the image, and the value range of SSIM is [-1, 1]. The closer it is to 1, the more similar the two images are, and the closer it is to -1, the less similar the two images are.
[0153] In addition, c 1 、c 2 and c 3 are respectively represented by the following formulas:
[0154] c 1 =(k 1 L) 2
[0155] c 2 =(k 2 L)
[0156]
[0157] Among them, c 1 、c 2 and c 3 respectively represent three constants, k 1 is the constant 0.01, k 2 is the constant 0.03, and L represents the range of image pixel values.
[0158] Specifically, the third sub-loss of the second loss can be represented by the following formula:
[0159]
[0160] Among them, LPIPS is the third sub-loss of the second loss. LPIPS calculates the distance between the reconstructed image and the original image feature map through a pre-trained neural network and weights it with learning parameters to better reflect the differences in human perception. The smaller the value of LPIPS, the more similar the sample tongue diagnosis image is to the reconstructed tongue diagnosis image; x represents the sample tongue diagnosis image, y represents the reconstructed tongue diagnosis image, l represents the network layer in the encoder or decoder, H l represents the height of the feature map of the l-th layer, W l represents the width of the feature map of the l-th layer, represents the normalized feature vector map of two images at the (h, w) position on the l-th layer, w l represents the weight of the l-th layer, ⊙ represents element-wise multiplication, represents the squared Euclidean distance.
[0161] It can be understood that in order to make the reconstructed tongue diagnosis image as similar as possible to the sample tongue diagnosis image and the perceptual features of the latent space similar to the original image, by minimizing the reconstruction loss function L containing three parts: MSE, LPIPS, and SSIM loss to achieve the goal. By calculating the loss generated by forward propagation, using backpropagation to calculate the gradient, and then using the optimizer to adjust the model parameters to gradually reduce the loss, so that the reconstructed tongue diagnosis image is as similar as possible to the sample tongue diagnosis image at the pixel level, perceptual features, and structural features.
[0162] Based on this, the second predicted noise is obtained through diffusion model prediction. Then, taking the first sample noise as the label data, the first loss is determined based on the difference between the first sample noise and the second predicted noise, and the second loss is determined based on the difference between the sample tongue diagnosis image and the reconstructed tongue diagnosis image. Finally, the diffusion model is trained based on the first loss and the second loss, so that the diffusion model can learn the data distribution from the real data distribution, thereby predicting a more appropriate noise, and then performing denoising processing based on the predicted noise to generate a higher-quality reconstructed tongue diagnosis image, which can improve the training effect of other deep learning models.
[0163] In addition, referring to Figure 6 , in one embodiment, the first noise image and the reference tongue coating state information are input into the diffusion model, and the first noise image is denoised with the reference tongue coating state information as the constraint condition to generate the target tongue diagnosis image, including but not limited to the following steps:
[0164] Step S610, obtaining the disease description text corresponding to the reference tongue coating state information.
[0165] Step S620, inputting the disease description text into the large language model for text generation to obtain the first tongue coating description text.
[0166] Step S630: Input the first noise image, the reference tongue coating status information, and the first tongue coating description text into the diffusion model, and use the reference tongue coating status information and the first tongue coating description text as constraint conditions to denoise the first noise image to generate a target tongue diagnosis image.
[0167] Among them, the disease description text corresponding to the reference tongue coating status information can be a description of a specific disease. For example, if the reference tongue coating status information is "yellow and greasy tongue coating", the corresponding disease description text can be "Digestive system problems: such as gastrointestinal dysfunction, indigestion, loss of appetite, abdominal distension and discomfort, etc.", "Hepatobiliary diseases: Dampness and heat in the liver and gallbladder may cause symptoms such as bitter taste in the mouth, pain in the hypochondrium, nausea, etc.", or "Other symptoms: It may also present symptoms such as fatigue, weight gain, and mood irritability".
[0168] Among them, it is necessary to combine the disease description text with the corresponding prompt words and then input them into the large language model for text generation.
[0169] Among them, the first tongue coating description text can be a description other than the tongue coating color and texture. For example, by combining the disease description text with the corresponding prompt words, we get "Having digestive system problems: such as gastrointestinal dysfunction, indigestion, loss of appetite, abdominal distension and discomfort, etc. What are the external manifestations of the tongue?", and then obtain the output from the large language model: "Tongue shape: It may present a swollen tongue, indicating spleen deficiency or dampness; it may also present a thin and emaciated tongue, suggesting qi and blood deficiency".
[0170] Based on this, by obtaining the first tongue coating description text through the large language model and adding the first tongue coating description text to the constraint conditions, it is possible to utilize the knowledge contained within the large language model to guide the training and reasoning of the diffusion model, thereby improving the class accuracy and diversity of the images generated by the diffusion model.
[0171] In addition, referring to Figure 7 , in an embodiment, inputting the first noise image, the reference tongue coating status information, and the first tongue coating description text into the diffusion model, and using the reference tongue coating status information and the first tongue coating description text as constraint conditions to denoise the first noise image to generate a target tongue diagnosis image includes, but is not limited to, the following steps:
[0172] Step S710: Obtain a reference tongue diagnosis image corresponding to the reference tongue coating status information.
[0173] Step S720: Perform edge detection on the reference tongue diagnosis image to obtain an edge image.
[0174] Step S730: Input the first noise image, the reference tongue coating state information, the first tongue coating description text, and the edge image into the diffusion model, and perform denoising processing on the first noise image with the reference tongue coating state information, the first tongue coating description text, and the edge image as constraint conditions to generate a target tongue diagnosis image.
[0175] Among them, the reference tongue diagnosis image can be prepared in advance through network search, and the embodiments of the present disclosure do not limit this here.
[0176] It should be noted that the image shown in the reference tongue diagnosis image may not only correspond to the reference tongue coating state information, but may also correspond to other dimensions of tongue image features. For example, the image shown in the reference tongue diagnosis image corresponds to the yellow and greasy coating of the tongue coating and also corresponds to the enlarged shape of the tongue.
[0177] Among them, the method of edge detection can be a gradient operator, a Laplace operator, a Canny edge detector, or a morphological gradient, etc., and the embodiments of the present disclosure do not limit this here.
[0178] Based on this, by performing edge detection on the reference tongue diagnosis image to obtain an edge image, the attention of the diffusion model to the edge of the reference tongue diagnosis image can be improved, so as to improve the understanding of the shape of the tongue by the diffusion model, which can play a guiding role in the training and inference of the diffusion model, thereby improving the category accuracy and diversity of the images generated by the diffusion model.
[0179] As Figure 13 shown, Figure 13 is an optional process schematic diagram of the model training method provided by the embodiments of the present application. This model training method can be executed by a server, or can also be executed by a terminal, or can also be executed by the server in cooperation with the terminal. This model training method includes, but is not limited to, the following steps S1310 to step S1330:
[0180] Step S1310: Obtain a target tongue diagnosis image and a tongue coating state label corresponding to the target tongue diagnosis image.
[0181] Step S1320: Input the target tongue diagnosis image into the tongue image classification model for classification to determine the tongue coating state result corresponding to the target tongue diagnosis image.
[0182] Step S1330: Determine the model loss based on the tongue coating state result and the tongue coating state label, and train the tongue image classification model based on the model loss.
[0183] Among them, the target tongue diagnosis image is generated by the above-mentioned tongue diagnosis image generation method. Therefore, the tongue coating state label can be used to represent the tongue coating color and tongue coating texture, etc. For example, the tongue coating state label can be white greasy coating, thin white coating, yellow greasy coating, thin yellow coating, or grayish black coating, etc.
[0184] Based on this, through the tongue diagnosis image generation method, a large-scale target tongue diagnosis image with sufficient authenticity and diversity can be generated. Then, the tongue coating state label, which is the expected output of the tongue image classification model, is obtained, and the target tongue diagnosis image is input into the tongue image classification model for classification. The model loss is determined according to the actual output tongue coating state result of the tongue image classification model and the tongue coating state label as the expected output. Training the tongue image classification model based on the model loss can improve the classification accuracy of the tongue image classification model.
[0185] In addition, referring to Figure 14 , this application also provides a tongue diagnosis image generation device 1400, including:
[0186] An acquisition module 1410, configured to acquire reference tongue coating state information;
[0187] A generation module 1420, configured to acquire a randomly generated first noise image, input the first noise image and the reference tongue coating state information into a diffusion model, and perform denoising processing on the first noise image with the reference tongue coating state information as a constraint condition to generate a target tongue diagnosis image.
[0188] It can be understood that the specific implementation manner of the tongue diagnosis image generation device 1400 is basically the same as the specific embodiment of the above tongue diagnosis image generation method, and will not be elaborated here.
[0189] In addition, referring to Figure 15 , Figure 15 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0190] A processor 1501, which can be implemented in a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application;
[0191] A memory 1502, which can be implemented in the form of a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM), etc. The memory 1502 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1502 and are called by the processor 1501 to execute the tongue diagnosis image generation method of the embodiments of this application.
[0192] Input / output interface 1503 is used to implement information input and output;
[0193] Communication interface 1504 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0194] Bus 1505 transmits information between various components of the device (such as processor 1501, memory 1502, input / output interface 1503, and communication interface 1504);
[0195] Among them, processor 1501, memory 1502, input / output interface 1503, and communication interface 1504 achieve communication connections with each other inside the device through bus 1505.
[0196] The embodiment of the present application also provides a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned tongue diagnosis image generation method.
[0197] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0198] The embodiments described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0199] Those skilled in the art can understand that Figures 1 to 15 the technical solutions shown in do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0200] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0201] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0202] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0203] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression thereof refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0204] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0205] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0206] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0207] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.
[0208] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.
Claims
1. A tongue diagnosis image generation method, characterized in that: include: Obtain reference tongue coating status information; A randomly generated first noise image is obtained, the first noise image and the reference tongue coating state information are input into a diffusion model, the first noise image is denoised using the reference tongue coating state information as a constraint condition, and a target tongue diagnosis image is generated.
2. The tongue diagnosis image generation method according to claim 1, characterized in that: The step of inputting the first noise image and the reference tongue coating state information into a diffusion model, performing denoising on the first noise image with the reference tongue coating state information as a constraint condition, and generating a target tongue diagnosis image includes: Inputting the first noise image and the reference tongue coating state information into the diffusion model, performing noise prediction with the reference tongue coating state information as a constraint condition, and obtaining a first predicted noise; The first noise image is denoised based on the first predicted noise to obtain a target tongue diagnosis image.
3. The tongue diagnosis image generation method according to claim 2, characterized in that: The diffusion module includes a Unet network, the Unet network includes multiple levels of downsampling blocks and multiple levels of upsampling blocks, the downsampling blocks include a convolution layer and an attention layer, the first noise image and the reference tongue coating state information are input into the diffusion model, the reference tongue coating state information is used as a constraint condition to perform noise prediction, and the first predicted noise is obtained, including: Inputting the first noise image into the diffusion model, and performing convolution processing on the first noise image based on the convolution layer in the downsampling block of the first level to obtain a first intermediate feature; Determine a time step, and fuse the time step with reference tongue coating state information to obtain a condition parameter; Based on the attention layer in the downsampling block of the first level, performing attention processing on the first intermediate feature and the conditional parameter to obtain a second intermediate feature; The second intermediate feature is mapped based on the remaining down-sampling blocks and the up-sampling block to obtain a first predicted noise.
4. The tongue diagnosis image generation method according to claim 3, characterized in that: The attention layer in the downsampling block based on the first level performs attention processing on the first intermediate feature and the condition parameter to obtain a second intermediate feature, including: Based on the attention layer in the downsampling block of the first level, the first intermediate feature and the conditional parameter are fused and then subjected to channel attention processing to obtain a first attention feature, and the first intermediate feature and the conditional parameter are fused and then subjected to spatial attention processing to obtain a second attention feature; Performing self-attention processing based on the fusion of the first attention feature and the second attention feature to obtain a third attention feature; Cross-attention processing is performed based on the third attention feature and the conditional parameter to obtain a second intermediate feature.
5. The tongue diagnosis image generation method according to claim 1, characterized in that: Before acquiring a randomly generated first noise image, inputting the first noise image and the reference tongue coating state information into a diffusion model, and performing denoising on the first noise image with the reference tongue coating state information as a constraint condition, and generating a target tongue diagnosis image, the tongue diagnosis image generation method further includes: Acquire a sample tongue diagnosis image and sample tongue coating state information, and add a first sample noise to the sample tongue diagnosis image to obtain a second noise image; Inputting the second noise image and the sample tongue coating state information into the diffusion model, and performing noise prediction on the second noise image with the reference tongue coating state information as a constraint condition to obtain a second predicted noise; determining a first loss based on a difference between the first sample noise and the second predicted noise; De-noising the sample tongue diagnosis image based on the second predicted noise to obtain a reconstructed tongue diagnosis image; determining a second loss based on a difference between the sample tongue diagnosis image and the reconstructed tongue diagnosis image; The diffusion model is trained based on the first loss and the second loss.
6. The tongue diagnosis image generation method according to claim 1, characterized in that: The step of inputting the first noise image and the reference tongue coating state information into a diffusion model, performing denoising on the first noise image with the reference tongue coating state information as a constraint condition, and generating a target tongue diagnosis image includes: Obtaining a disease description text corresponding to the reference tongue coating state information; Inputting the disease description text into a large language model for text generation to obtain a first tongue coating description text; The first noise image, the reference tongue coating state information and the first tongue coating description text are input into the diffusion model, and the first noise image is denoised using the reference tongue coating state information and the first tongue coating description text as constraints to generate a target tongue diagnosis image.
7. The tongue diagnosis image generation method according to claim 6, characterized in that: The step of inputting the first noise image, the reference tongue coating state information, and the first tongue coating description text into a diffusion model, performing denoising on the first noise image with the reference tongue coating state information and the first tongue coating description text as constraints, and generating a target tongue diagnosis image includes: Acquire a reference tongue diagnosis image corresponding to the reference tongue coating state information; Performing edge detection on the reference tongue diagnosis image to obtain an edge image; The first noise image, the reference tongue coating state information, the first tongue coating description text and the edge image are input into the diffusion model, and the first noise image is denoised using the reference tongue coating state information, the first tongue coating description text and the edge image as constraints to generate a target tongue diagnosis image.
8. A tongue diagnosis image generating device, characterized in that: include: An acquisition module, used for acquiring reference tongue coating status information; A generation module is used to obtain a randomly generated first noise image, input the first noise image and the reference tongue coating state information into a diffusion model, denoise the first noise image using the reference tongue coating state information as a constraint condition, and generate a target tongue diagnosis image.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the tongue diagnosis image generation method according to any one of claims 1 to 7 when executing the computer program.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the tongue diagnosis image generating method according to any one of claims 1 to 7 is implemented.