Traditional Chinese medicine tongue diagnosis auxiliary method and system based on image reconstruction and tongue picture detection
By adopting the super-resolution reconstruction and tongue image detection method of the dual-task feedback learning framework in traditional Chinese medicine tongue diagnosis, the image blur problem caused by poor imaging conditions in tongue image detection is solved, and the accuracy and objectivity of tongue diagnosis are improved.
Patent Information
- Application Number
- CN202411990685.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In traditional Chinese medicine tongue diagnosis, the image blurring caused by poor imaging conditions, affecting the accuracy and objectivity of the diagnosis.
The dual-task feedback learning framework is adopted to improve the accuracy of tongue image detection through super-resolution reconstruction, and the image reconstruction quality is optimized by tongue image detection to solve the image blur problem.
It improves the objectivity and accuracy of tongue diagnosis in traditional Chinese medicine, solves the image blur problem through super-resolution reconstruction module, and enhances the accuracy and reliability of tongue image detection.
Smart Images

Figure CN119941652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing and artificial intelligence technology, and specifically designs a traditional Chinese medicine tongue diagnosis auxiliary method and system based on image reconstruction and tongue image detection. Background Art
[0002] Tongue diagnosis in traditional Chinese medicine is an important part of the diagnosis system. By observing the color, shape, and coating of the tongue, it reflects the patient's internal organs, qi, blood, and body fluids. Tongue diagnosis is widely used in daily health monitoring and chronic disease management due to its non-invasive and convenient advantages. However, traditional tongue diagnosis relies on the subjective experience of doctors and there are individual differences, which affects the objectivity and accuracy of the diagnosis. Therefore, the development of a tongue image detection system based on image analysis is crucial to improving the scientificity and consistency of tongue diagnosis. Tongue image detection provides a basis for doctors' diagnosis by automatically identifying features such as tongue coating thickness, color changes, and abnormal tongue shape. Thick and greasy tongue coating may indicate moisture in the body, while pale tongue color may indicate insufficient qi and blood. Accurate tongue image detection can help doctors detect health problems early and achieve preventive medical treatment.
[0003] In recent years, the application of artificial intelligence (AI) in tongue image detection has promoted the modernization of tongue diagnosis in traditional Chinese medicine. As an important basis for reflecting the intrinsic health status of the body, tongue image diagnosis traditionally relies on the experience of doctors and lacks standardization and quantitative basis. The introduction of AI has greatly improved the accuracy and efficiency of diagnosis by automatically analyzing the color, shape, texture and other characteristics of the tongue through image processing and deep learning algorithms. Studies have shown that the tongue image detection system based on convolutional neural network (CNN) and transformer model has achieved accurate classification of diseases. Combined with big data and multimodal analysis, AI can also integrate tongue images with other physiological indicators, significantly enhance the accuracy of disease prediction, and promote the standardization and intelligent development of traditional Chinese medicine diagnosis and treatment.
[0004] At the same time, the development and application of image reconstruction technology in tongue diagnosis has greatly improved the accuracy and clarity of tongue analysis. Traditional tongue images may be affected by factors such as lighting and angle, affecting the diagnostic results. In recent years, deep learning-based image reconstruction methods, especially super-resolution reconstruction (SR), have made significant progress in the restoration and detail enhancement of tongue images. These technologies restore the high-frequency details of the tongue, making the texture, color and other features of the tongue clearer, providing a more reliable data basis for subsequent AI model analysis. In addition, combined with cutting-edge methods such as generative adversarial networks (GAN), researchers are able to further improve the quality of tongue images, making them closer to real clinical observations. These developments have not only improved the accuracy of tongue diagnosis, but also promoted the application of AI in TCM tongue diagnosis.
[0005] Although AI technology has shown great potential in tongue detection and image reconstruction, and has promoted the development of intelligent and modern Chinese medicine, the field still faces many challenges. First, tongue detection data is scarce and the annotation quality is not high, which limits the generalization ability of the model. Secondly, the image acquisition quality is unstable and easily affected by factors such as equipment performance and lighting conditions, resulting in inaccurate detection results. In addition, current research usually performs tongue detection and image reconstruction tasks independently, and fails to fully utilize the feature information extracted in the detection, resulting in poor reconstruction effect, which in turn affects the accuracy and effectiveness of diagnosis. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention proposes a TCM tongue diagnosis auxiliary method and system based on image reconstruction and tongue image detection. The tongue image detection accuracy is improved through super-resolution reconstruction, and the image reconstruction quality is optimized by tongue image detection to solve the image blur problem caused by poor imaging conditions during tongue image detection in TCM tongue diagnosis. The system adopts a dual-task feedback learning framework to improve the objectivity and accuracy of TCM tongue diagnosis.
[0007] The present invention discloses a TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection, comprising the following steps:
[0008] Collecting the true value image of the tongue and processing it to obtain a low-resolution image of the tongue;
[0009] Construct a tongue image detection learning model: perform super-resolution reconstruction on the low-resolution tongue image; perform image-level enhancement processing on the reconstructed image; combine the true value image of the tongue image, the reconstructed image and the enhanced image, perform feature-level enhancement processing to obtain a fused high-dimensional tongue image feature map; perform tongue image detection on the high-dimensional tongue image feature map to obtain the type of tongue image;
[0010] The tongue image detection learning model is trained, and feature alignment loss is introduced between super-resolution reconstruction and tongue image detection. The task-specific features extracted by tongue image detection are passed back to the super-resolution reconstruction process, and the super-resolution reconstruction network parameters are adjusted.
[0011] Furthermore, the tongue image is down-sampled by bicubic interpolation to obtain a corresponding low-resolution tongue image, and the tongue image and the corresponding low-resolution tongue image are used as a training data pair.
[0012] Furthermore, the low-resolution tongue image is input into the convolutional neural network and Transformer network architecture to gradually extract the spatial and contextual texture and subtle structure information of the image and represent it in a high-dimensional feature space; the decoder of the multi-layer neural perception network is used to perform upsampling and restoration to obtain the reconstructed super-resolution image.
[0013] Furthermore, the image-level enhancement processing specifically includes taking the super-resolution reconstructed image and the tongue image truth image as input, introducing a binary matrix of the same size as the image as a random mask to control the image fusion method, and then the super-resolution reconstructed image I SR and tongue image truth image I HR The corresponding pixels of are added element by element according to the binary mask to obtain the fused image I processed by image level enhancement. aug .
[0014] Furthermore, the feature-level data enhancement process includes converting the true value image I of the input tongue image HR , the super-resolution reconstructed image I obtained in S2 SR , the corresponding fused image I obtained in S3 aug , respectively using the same feature extraction encoder in the tongue image detection model Extract high-dimensional features to obtain the high-dimensional features F of the true value image HR , the high-dimensional features F of the super-resolution image SR , and the high-dimensional features F of the fused image aug , directly perform element-by-element weighted combination of high-dimensional features to obtain fused high-dimensional tongue image features.
[0015] Preferably, the tongue image detection model is composed of the RPN network structure in FasterR-CNN, the ROI pooling layer, and the fully connected layer of classification regression.
[0016] Furthermore, the loss function in super-resolution reconstruction is:
[0017]
[0018] Where N represents the total number of pixels in the image, that is, H×W, I LR,i , I HR,i Represent the value of the i-th pixel in the low-resolution tongue image and the true value image respectively.
[0019] Furthermore, the loss of the tongue image detection learning model is:
[0020]
[0021] in, Represents network model connection, Represents image stitching. L SR represents the super-resolution reconstruction loss, I LR , I HR Represent the input low-resolution tongue image and the tongue image truth image, and are paired with their corresponding task labels y. AUG Refers to the high-dimensional tongue feature map generated by image-level enhancement. Represents a trainable parameter θ SR The super-resolution reconstruction network is used to obtain Generated super-resolution image I SR Then, through the tongue image detection model Conduct, among which, is the feature extractor, Indicates that tongue image detection is performed to obtain the type of tongue image.
[0022] Based on the same inventive concept, the present invention also provides an electronic device, including:
[0023] one or more processors;
[0024] A storage device for storing one or more programs;
[0025] When one or more programs are executed by the one or more processors, the one or more processors implement a TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection.
[0026] Based on the same inventive concept, the present invention also designs a computer-readable medium on which a computer program is stored. When the program is executed by a processor, a TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection is implemented.
[0027] The advantages of the present invention are:
[0028] In the present invention, a super-resolution reconstruction module is used to reconstruct a low-resolution tongue image into a high-resolution image. This module receives a low-resolution tongue image taken from an input device, such as an ordinary camera, and reconstructs the image through a pre-trained super-resolution neural network. The super-resolution network generates a tongue image with a higher resolution through multiple convolutional layers, upsampling layers, and optimization of the fusion loss function. The super-resolution reconstruction module can solve the image blur problem caused by insufficient illumination and poor performance of the imaging device. For the tongue image detection part, after the super-resolution module generates a high-quality image, the image is input into the tongue image detection module. The tongue image detection module analyzes and identifies the features of the tongue image in combination with a pre-trained classifier. The module identifies the main area of the tongue image through a multi-layer feature extractor, and outputs the detection result through a feature restoration decoder. Different from the existing research, in order to improve the interactivity of the super-resolution reconstruction and the tongue image detection module, the present invention introduces a dual-task feedback learning framework. Through feature alignment loss. Super-resolution reconstruction does not only rely on its own loss function for training, but also obtains task information related to tongue image detection from tongue image detection through a feedback mechanism. The feature alignment loss ensures that super-resolution reconstruction can better adapt to the requirements of the tongue image detection task when generating images, so that super-resolution reconstruction can more accurately optimize tongue image features during training. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flow chart of the auxiliary method of traditional Chinese medicine tongue diagnosis based on image reconstruction and tongue image detection of the present invention.
[0030] Figure 2 It is a network structure diagram of tongue diagnosis-assisted image reconstruction and tongue image detection in the present invention.
[0031] Figure 3 It is a framework diagram of data enhancement and mask generation of random masks in the present invention.
[0032] Figure 4 It is a flow chart of the dual-task feedback learning framework proposed in the present invention.
[0033] Figure 5 It is a schematic diagram of the classification results of tongue coating detection in tongue image detection in tongue diagnosis of the present invention. DETAILED DESCRIPTION
[0034] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical scheme in the present invention is clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the present invention is not limited to the following embodiments, and the specific implementation method can be determined according to the technical scheme of the present invention and the actual situation. In order to avoid confusing the essence of the present invention, the known methods, processes and procedures are not described in detail.
[0035] Embodiment 1
[0036] The present invention proposes a tongue detection method based on a dual-task feedback learning framework, which is used to improve the quality of tongue images and optimize the accuracy of tongue detection. The system includes two main task modules: a super-resolution reconstruction (SR) module and a tongue detection (TD) module. The operation process of the method is as follows: first, the input tongue image is super-reconstructed, and then the reconstructed high-quality image is processed by the tongue detection network, and the training is alternated to achieve accurate tongue detection.
[0037] For the tongue image example of the present invention, the TCMID-Tongue tongue image dataset published on the Traditional Chinese Medicine Comprehensive Database (http: / / www.megabionet.org / tcmid / ) is used. This dataset contains tongue images of different resolutions and five types of tongue coating: mirror-like, thin white coating, white greasy coating, yellow greasy coating, and gray-black coating. Figure 5 As shown in Figure 2. Each image is annotated in XML format, including the tongue bounding box and category label.
[0038] The main steps and flow chart of the TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection disclosed in this embodiment are shown in Figure 1 , specifically including:
[0039] S1: Preprocess the image of the tongue image dataset. Input the tongue image true value image I HR , to I HR Perform bicubic interpolation downsampling to obtain the corresponding low-resolution tongue image I LR , build training data for the super-resolution model.
[0040] The specific embodiments are as follows:
[0041] In a specific embodiment of the present invention, firstly, a high-quality true tongue image I is obtained. HR , and use the bicubic interpolation method, according to the fixed downsampling ratio r s In this embodiment, a fixed magnification such as ×4 or ×8 is used to perform downsampling processing to generate a corresponding low-resolution tongue image I LR , simulating the low-quality images produced by imaging devices. Bicubic interpolation reduces the resolution by weighted average of surrounding pixels, preserving details and reducing distortion. Next, the low-resolution image I LR As input data, the ground-truth high-resolution image I HR As output labels, paired training datasets are constructed for training super-resolution models. The specific definitions are as follows:
[0042] I LR =Bicubic(I HR ,r s )
[0043] Among them, the high-quality true value tongue image I HR The dimension is I HR ∈R H×W×3 , low resolution tongue image I LR The dimension is Downsampling scale r s ∈{4,8}. Bicubic is a bicubic interpolation downsampling method.
[0044] S2: Perform super-resolution reconstruction on the tongue image. LR Input super-resolution reconstruction model In the process, feature extraction is performed through the encoder: the low-resolution tongue image I after downsampling S1 is LRThe input is a network architecture based on convolutional neural network (CNN) and Transformer, which gradually extracts key information such as the spatial and contextual texture and subtle structure of the image. It is represented in a high-dimensional feature space. After feature extraction, the decoder learns the mapping relationship between the input high-order features and the high-resolution image, and uses the decoder of the multi-layer neural perception network to perform upsampling and restoration to obtain the reconstructed super-resolution image I SR .
[0045] The specific embodiments are as follows:
[0046] like Figure 2 As shown, the low-resolution tongue image I generated by S1 LR Input to Reconstructed Model Super-resolution reconstruction is performed in. Reconstruction model The network architecture is composed of a convolutional neural network (CNN) and a Transformer. In this embodiment, the existing EDSR-baseline, SwinIR, SRFomer and other super-resolution reconstruction network architectures are selected as the reconstruction model of the present invention. In the encoding stage, a low-resolution tongue image I is input. LR ,Model Through multi-layer convolution operations, key information such as the spatial features and local detail texture of the tongue image is gradually extracted and embedded in pixels to obtain the high-dimensional feature encoding of the input tongue image. Then, the multi-head attention mechanism of Transformer is used to further model the context information of the encoded high-dimensional features and capture the global dependencies, so that the information of the input low-resolution tongue image can be fully represented in the high-dimensional feature space. In the decoding stage, the model maps the high-dimensional features back to the high-resolution image space, that is, a multi-layer neural perception network is used for upsampling, and the super-resolution tongue image I is gradually restored and reconstructed. SR The specific definitions are as follows:
[0047]
[0048] Among them, the input super-resolution reconstructed low-resolution image I LR The dimension is is the super-resolution reconstruction network, I SR It is the reconstructed super-resolution image mapped back to the high-resolution image space, and its dimension is the same as the tongue image truth image I HR The same, that is, I SR ∈R H×W×3 .
[0049] S3: Implement image-level data enhancement strategy. SRAs the basis of data enhancement, more model reconstruction details are retained. HR A reference image for comparing the reconstruction effect. A binary matrix of the same size as the image is introduced as a random mask to control the image fusion. After applying the random mask, the reconstructed image I SR and the true value image I HR The corresponding pixels of are added element by element according to the binary mask, and the data is fused to obtain the corresponding fused image I aug , and used as input data for subsequent segmentation training to help super-resolution reconstruction models More robust features are learned at different pixel levels.
[0050] The specific implementation is as follows:
[0051] like Figure 3 As shown, the super-resolution reconstructed image I obtained in S2 SR Based on, retain the super-resolution model In the reconstruction process, the learned information is captured and the tongue image I HR That is, the original high-resolution image is used as a comparison standard to enrich and verify the reconstruction effect. In order to achieve effective image-level data fusion, this method uses a binary matrix M with the same size as the image as a random mask to control the reconstructed image I SR and the true value image I HR The fusion method at each pixel position. After applying the random mask, the super-resolution image I nR and the true value image I HR Add pixel by pixel, the specific definition is as follows:
[0052] I AUG (x,y)=M(x,y)·I SR (x,y)+(1-M(x,y))·I HR (x,y)
[0053] The random mask M(x,y) is a binary matrix with the same size as the image, and its dimension is M∈R H×W×3 , I SR (x,y) and I HR (x, y) are the pixel values of the super-resolution image and the tongue image truth image at the corresponding coordinates (x, y). It can be seen that -W≤x≤W, -H≤y≤H.
[0054] Through the above element-by-element addition operation, the data enhanced fusion image I is obtained. AUG The details in the super-resolution reconstructed image are retained while combining the original information of the true image. According to the value of the random mask, the pixel of the true image or the generated super-resolution image is randomly determined, specifically:
[0055]
[0056] The fused image I AUG It will be used as input in the subsequent tongue image detection training. The dimension is I AUG ∈R H×W×3 . It can help the model learn more robust representations at different pixel levels.
[0057] S4: Implement feature-level data enhancement strategy. HR , the super-resolution reconstructed image I obtained in S2 SR , the corresponding fused image I obtained in S3 aug , respectively using the same feature extraction encoder in the tongue image detection model Extract high-dimensional features to obtain the high-dimensional features F of the true value image HR , the high-dimensional features F of the super-resolution image SR , and the high-dimensional features F of the fused image aug . Directly perform element-by-element weighted combination of high-dimensional features to create diversified and complex features for subsequent tongue image detection tasks, thereby improving the analysis capability of subsequent tongue image detection models for different tongue image features.
[0058] The specific implementation is as follows:
[0059] like Figure 2 As shown, by HR , the super-resolution reconstructed image I obtained in S2 SR , the corresponding data enhanced fused image I obtained in S3 aug , using the same tongue image features to propose models To extract high-dimensional features, the present invention adopts the FasterR-CNN network architecture for encoding to obtain the high-dimensional features F of the true value image. HR , the high-dimensional features F of the super-resolution image SR , and the high-dimensional features F of the fused image aug . Specifically expressed as:
[0060]
[0061] Among them, the dimension of the generated features is based on the network. In order to generate more diverse and complex feature representations, we weighted and combined the above high-dimensional features element by element. And manually set the combination coefficients α, β, and γ to correspond to the feature fusion weights of the fused super-resolution image, the real image, and the fused image, respectively, to perform data enhancement. The specific expression is:
[0062] F input =αFSR +β·F HR +γ·Fa aug
[0063] where F input To input the fused high-dimensional features into the subsequent tongue image detection model decoder, it is necessary to satisfy α+β+γ=1 to ensure the rationality of the weighted combination.
[0064] S5: Decode tongue features and implement tongue detection. The high-dimensional tongue features obtained in S4 contain rich detailed information about the tongue. The fused features are restored into visual output to achieve automatic tongue image detection.
[0065] The specific implementation is as follows:
[0066] The fused high-dimensional tongue image feature F obtained in S4 input , which contains rich details about the tongue image. This information includes the global contextual relationship and local content of the tongue image, such as color, texture, and shape. Therefore, this method uses it as the decoder in the tongue image detection model. The decoder is used to gradually restore the features of the tongue image detection results in the image space. The RPN network structure, ROI pooling layer, and fully connected layer of classification regression in FasterR-CNN are retained. The output results include detection box and classification probability. Tongue classification is as follows: Figure 5 As shown. Specifically expressed as:
[0067]
[0068] B is the output of the tongue image detection model, and its dimension is N×4, where N is the number of tongue images detected and 4 is the parameter (x, y, w, h) of each bounding box. The classification output is P, with dimension N×C, where C=5 represents the number of categories, including the category probability of each tongue image, and is restored to a visual output.
[0069] S6: Training based on the dual-task feedback learning framework. The tongue image feature extraction model in S4 And the tongue image detection model in S5 Parameter freezing and super-resolution model Training, at this stage, the super-resolution model Focus on optimizing the tongue image fusion features obtained in tongue image detection. And ensure that the feature encoder is extracted in the detection model And the decoder in the detection model The output features of the tongue image feature extraction model are trained alternately. Decoder in detection model Super-resolution reconstruction model To improve the accuracy and reliability of overall tongue image detection.
[0070] The specific implementation is as follows:
[0071] Different from the previous open-loop framework, the present invention uses a closed-loop framework of dual-task feedback learning, aiming to improve the accuracy of tongue diagnosis while improving the quality of tongue images. Figure 4 As shown, unlike the traditional open-loop structure, the method of the present invention uses a feedback connection feature alignment loss function L FA The Super Resolution (SR) and Tongue Detection (TD) task modules are combined, and the feature alignment loss is the specific implementation of the feedback connection. Figure 4 As shown in the figure, during the training process, the forward connection passes the image generated by SR to the TD module, enabling TD to use higher quality images for detection; the feedback connection passes the task-specific features extracted by TD back to the SR module, and adjusts the SR network parameters through back propagation to generate super-resolution images optimized for detection tasks. This design promotes effective collaboration between the two modules and maximizes the use of feature information by transferring feature priors, thereby improving detection accuracy.
[0072] The feature alignment loss function is specifically defined as:
[0073]
[0074] like Figure 4 As shown, according to the closed-loop framework of dual-task feedback learning in the present invention, in addition to the feature-level loss function L FA , which is used to guide the SR network to obtain the specific task-driven features required in the downstream task TD network. In the SR task, the loss function also retains the original pixel-level loss to ensure that the structure and texture details of the image can be effectively retained during model training. During training, the combined loss function of the SR task is specifically defined as:
[0075]
[0076] Where N represents the total number of pixels in the image, that is, H×W. LR,i , I HR,i They represent the value of the i-th pixel in the low-resolution image and the high-resolution true value image respectively.
[0077] When training the network for the tongue detection (TD) task, the present invention adopts the loss function introduced in Faster R-CNN, which includes the classification loss L cls and the bounding box regression loss L bbox. Classification loss usually uses cross entropy to evaluate the accuracy of category prediction, while bounding box regression loss uses smooth L1 loss to measure the difference between the predicted bounding box and the true bounding box. By optimizing these two loss functions, the performance of object detection is improved. Therefore, during training, the combined loss function of the TD task is specifically defined as:
[0078] L TD =L cls +L bbox
[0079] In summary, the goal of the duet feedback learning framework in the present invention is to obtain the rabbit tongue image I from the low-resolution rabbit tongue image I LR Reconstruct the model through super-resolution Generate super-resolution image I SR , to improve the target detection model The performance is formally expressed as:
[0080]
[0081] Among them, I LR , I HR I represents the input low-resolution tongue image and the ground-truth high-resolution tongue image, and is paired with its corresponding task label y. AUG Refers to the fused image generated by the image-level quality data enhancement strategy. Represents a trainable parameter θ SR The super-resolution reconstruction network is used to obtain Generated super-resolution image I SR After that, the tongue image detection model is Among them, is a feature extractor, Including the subsequent RPN, ROL pooling and fully connected layers in Faster R-CNN, the tongue image classification of tongue diagnosis in the output results is shown in Figure 5.
[0082] In general, the overall architecture of the dual-task feedback learning framework in the present invention mainly consists of two parts: super-resolution reconstruction (SR) network and tongue detection (TD) network. In the traditional open-loop framework dual-task training, the SR module is fed with low-resolution input I LR Generate super-resolution image I SR , and then through the TD module to I SR However, this open-loop design limits the information flow between the two modules, resulting in unsatisfactory detection results due to the lack of task-specific details. To address this problem, the present invention proposes a closed-loop structure by introducing a feedback connection between TD and SR. FA), we implement feature-driven prior information transfer from TD to SR, establishing strong dependencies between tasks. This approach ensures that SR receives high-frequency information relevant to the task, thereby improving image reconstruction, while TD benefits from the improved image quality and achieves more accurate detection. In addition, quality fusion enhancement and alternating training strategies further improve the feature alignment loss L FA effects and overall performance.
[0083] Embodiment 2
[0084] Based on the same inventive concept, the present invention also provides an electronic device, including one or more processors; a storage device for storing one or more programs; when one or more programs are executed by the one or more processors, the one or more processors implement the method described in Example 1.
[0085] Since the device introduced in the second embodiment of the present invention is an electronic device used to implement the Chinese medicine tongue diagnosis auxiliary method based on image reconstruction and tongue image detection in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, the technical personnel in the field can understand the specific structure and deformation of the electronic device, so it is not repeated here. All electronic devices used in a method of the embodiment of the present invention belong to the scope of protection of the present invention.
[0086] Embodiment 3
[0087] Based on the same inventive concept, the present invention further provides a computer-readable medium on which a computer program is stored, and when the program is executed by a processor, the method described in the first embodiment is implemented.
[0088] Since the device introduced in the third embodiment of the present invention is a computer-readable medium used to implement the auxiliary method of traditional Chinese medicine tongue diagnosis based on image reconstruction and tongue image detection in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, the technical personnel in the field can understand the specific structure and deformation of the electronic device, so it is not repeated here. All electronic devices used in a method of the embodiment of the present invention belong to the scope of protection of the present invention.
[0089] The above content is combined with the accompanying drawings to illustrate the preferred embodiments of the present application, but does not limit the scope of the rights of the embodiments of the present application. Without departing from the scope and core ideas of the embodiments of the present application, any modification, equivalent replacement or improvement made by those skilled in the art should be included in the scope of the rights of the embodiments of the present application.
Claims
1. A TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection, characterized in that: The following steps are involved: Collecting the true value image of the tongue and processing it to obtain a low-resolution image of the tongue; Construct a tongue image detection learning model: perform super-resolution reconstruction on the low-resolution tongue image; perform image-level enhancement processing on the reconstructed image; combine the true value image of the tongue image, the reconstructed image and the enhanced image, perform feature-level enhancement processing to obtain a fused high-dimensional tongue image feature map; perform tongue image detection on the high-dimensional tongue image feature map to obtain the type of tongue image; The tongue image detection learning model is trained, and feature alignment loss is introduced between super-resolution reconstruction and tongue image detection. The task-specific features extracted by tongue image detection are passed back to the super-resolution reconstruction process, and the super-resolution reconstruction network parameters are adjusted.
2. The TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection according to claim 1, characterized in that: The tongue image is down-sampled by bicubic interpolation to obtain the corresponding low-resolution tongue image, and the tongue image and the corresponding low-resolution tongue image are used as training data pairs.
3. The TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection according to claim 2, characterized in that: The low-resolution tongue image is input into the convolutional neural network and Transformer network architecture to gradually extract the spatial and contextual texture and subtle structure information of the image and represent it in a high-dimensional feature space; the decoder of the multi-layer neural perception network is used to perform upsampling and restoration to obtain the reconstructed super-resolution image.
4. The TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection according to claim 1, characterized in that: The image-level enhancement process specifically includes taking the super-resolution reconstructed image and the tongue image truth value image as input, introducing a binary matrix of the same size as the image as a random mask to control the image fusion method, and then the super-resolution reconstructed image I SR and tongue image truth image I HR The corresponding pixels of are added element by element according to the binary mask to obtain the fused image I processed by image level enhancement. aug .
5. The TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection according to claim 1, characterized in that: The feature-level data enhancement process includes converting the true value image I of the input tongue image into HR , the super-resolution reconstructed image I obtained in S2 SR , the corresponding fused image I obtained in S3 aug , respectively using the same feature extraction encoder in the tongue image detection model Extract high-dimensional features to obtain the high-dimensional features F of the true value image HR , the high-dimensional features F of the super-resolution image SR , and the high-dimensional features F of the fused image aug , directly perform element-by-element weighted combination of high-dimensional features to obtain fused high-dimensional tongue image features.
6. The TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection according to claim 1, characterized in that: The tongue image detection model consists of the RPN network structure in FasterR-CNN, the ROI pooling layer, and the fully connected layer for classification and regression.
7. The TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection according to claim 1, characterized in that: The loss function in super-resolution reconstruction is: Where N represents the total number of pixels in the image, that is, H×W, I LR,i , I HR,i Represent the value of the i-th pixel in the low-resolution tongue image and the true value image respectively.
8. The TCM tongue diagnosis auxiliary method based on image reconstruction and tongue image detection according to claim 1, characterized in that: The loss of the tongue image detection learning model is: in, Represents network model connection, represents image stitching, L SR represents the super-resolution reconstruction loss, I LR , I HR Represent the input low-resolution tongue image and the tongue image truth image, and are paired with their corresponding task labels y. AUG Refers to the high-dimensional tongue feature map generated by image-level enhancement. Represents a trainable parameter θ SR The super-resolution reconstruction network is used to obtain Generated super-resolution image I SR After that, the tongue image detection model is Conduct, among which, is the feature extractor, Indicates that tongue image detection is performed to obtain the type of tongue image.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the learning index construction method as described in any one of claims 1-8.
10. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the learning index construction method as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Image enhancement-based power transmission line fine hardware defect detection method and system
CN111524135A
Target detection method and device based on medium and low resolution remote sensing images and equipment
CN113705532A
Field wheat scab detection method based on unmanned aerial vehicle image
CN114627385A
Low-resolution image small target detection method based on image super-resolution reconstruction
CN117409186A
Image super-resolution method based on multi-scale Transform and texture enhancement
CN118350999A