Intelligent diagnosis system and method based on multi-mode tongue picture feature fusion

The intelligent diagnosis system based on multimodal tongue feature fusion solves the shortcomings of tongue feature extraction and diagnosis methods in the existing technology, realizes the generation of intuitive and accurate tongue diagnosis results, reduces dependence on physician experience, and improves the adaptability and practicality of the system.

CN120635637APending Publication Date: 2025-09-12NINGBO BEILUN DISTRICT HOSPITAL OF TRADITIONAL CHINESE MEDICINE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510603529.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing tongue feature extraction and diagnosis methods have deficiencies in multimodal information fusion, comprehensive feature extraction, and model generalization capabilities, which affect the diagnostic accuracy and reliability of the system, and rely on physician experience, resulting in less intuitive diagnostic results.

Method used

An intelligent diagnostic system based on multimodal tongue feature fusion is adopted, including a tongue image acquisition module, a multimodal data processing module, a feature fusion analysis module, a diagnosis result generation module and a user interaction module. It uses a high-resolution camera, a light source adjustment device, an ambient light compensation sensor, a deep learning model and a user interaction module to realize the acquisition, processing, feature extraction and generation of diagnosis results of tongue images.

Benefits of technology

It enables intuitive acquisition of diagnostic results based on tongue characteristics, reduces dependence on the experience of traditional Chinese medicine practitioners, improves the accuracy and reliability of diagnosis, and enhances the adaptability and practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635637A_ABST
    Figure CN120635637A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of general image data processing or generation, in particular to an intelligent diagnosis system and method based on multi-modal tongue picture feature fusion, and the system comprises a tongue picture collection module, a multi-modal data processing module, a feature fusion analysis module, a diagnosis result generation module and a user interaction module. Through cooperation of the tongue picture acquisition module, the multi-modal data processing module, the feature fusion analysis module, the diagnosis result generation module and the user interaction module, the diagnosis result can be visually obtained according to the tongue picture features, so that dependence on experience of traditional Chinese medicine doctors is reduced, and the method is effectively different from the method of depending on the experience of the doctors in the prior art; and the diagnosis result is not visual enough and is lack of result readability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the general field of image data processing or generation technology, and in particular to an intelligent diagnosis system based on multimodal tongue image feature fusion, and also to an intelligent diagnosis method. Background Art

[0002] With the rapid development of artificial intelligence and Traditional Chinese Medicine (TCM) diagnostic technologies, tongue-based diagnosis has gained widespread attention in the fields of health monitoring and disease prediction. However, existing tongue feature extraction and diagnosis methods still have shortcomings in multimodal information fusion, comprehensive feature extraction, and model generalization capabilities, which affect the diagnostic accuracy and reliability of the system.

[0003] The prior art has a Chinese invention patent numbered CN117611581B, entitled A tongue image recognition method, device, and electronic device based on multimodal information. It discloses that by integrating tongue images and symptom information, image features, text features, and graph features based on a tongue diagnosis knowledge graph are extracted, and the three are fused to obtain target tongue image features, thereby improving the accuracy and reliability of tongue image recognition. However, in this technical solution, the acquisition of symptom information depends on the user's input of the main complaint and symptom description, which may result in subjective bias or incomplete information, thereby affecting the effect of feature fusion. In addition, this method has high requirements for the construction of a tongue diagnosis knowledge graph. If the knowledge graph has limited coverage or is not updated in a timely manner, it may lead to deviations in the feature fusion results, thereby reducing the accuracy of the diagnosis.

[0004] Based on this, the existing tongue feature extraction and diagnosis methods still have certain shortcomings in terms of multimodal information fusion, comprehensive feature extraction, and adaptability to complex diagnostic scenarios. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent diagnostic system and method based on the fusion of multimodal tongue features to solve the technical problems existing in traditional Chinese medicine tongue diagnosis, such as reliance on physician experience, lack of intuitive diagnostic results, and difficulty in interpretation using digital technology.

[0006] According to one aspect of the present invention, an intelligent diagnosis system based on multimodal tongue image feature fusion is provided, comprising a processor, a tongue image acquisition module, a multimodal data processing module, a feature fusion analysis module, a diagnosis result generation module and a user interaction module; wherein the tongue image acquisition module is data-connected to the multimodal data processing module, the multimodal data processing module is data-connected to the feature fusion analysis module, the feature fusion analysis module is data-connected to the diagnosis result generation module, and the diagnosis result generation module is data-connected to the user interaction module.

[0007] In some embodiments, the tongue image acquisition module includes at least a high-resolution camera, a light source adjustment device, and an ambient light compensation sensor; wherein, the high-resolution camera is used to acquire tongue images, the light source adjustment device includes a ring-shaped LED light group and a light angle adjuster, and the ambient light compensation sensor is used to monitor the ambient light intensity in real time and transmit the data to the light source adjustment device for dynamic adjustment.

[0008] In some embodiments, the multimodal data processing module includes at least an image preprocessing unit, a color correction unit, and a texture extraction unit; wherein, the image preprocessing unit performs denoising, sharpening, and contrast enhancement on the collected tongue image, the color correction unit corrects the color of the tongue image, and the texture extraction unit extracts the texture features of the tongue coating to generate a texture feature vector.

[0009] In some embodiments, the texture extraction unit extracts the texture features of the tongue coating using a gray level co-occurrence matrix algorithm.

[0010] In some embodiments, the feature fusion analysis module includes at least a deep learning model, a feature weight allocation unit, and a classification decision unit; wherein, the deep learning model is a convolutional neural network architecture, the feature weight allocation unit assigns weight values ​​to each type of feature, and the classification decision unit is used to generate the final health status classification result.

[0011] In some embodiments, the diagnosis result generation module includes at least a health status assessment unit and a diagnosis report generation unit; wherein, the health status assessment unit is used to generate a health status assessment report of the user, and the diagnosis report generation unit is used to generate visual charts and voice prompt content.

[0012] In some embodiments, the user interaction module includes at least a touch screen, a voice prompt device and a remote communication interface; wherein, the touch screen is used to display diagnostic results and a user operation interface, the voice prompt device includes at least a speaker and a microphone, and the remote communication interface is connected to an external medical system via a wireless communication protocol.

[0013] According to another aspect of the present invention, an intelligent diagnosis method based on multimodal tongue image feature fusion is provided. The method is executed by a processor and includes: connecting to a tongue image acquisition module to acquire tongue images; connecting to a multimodal data processing module to acquire the texture features of the tongue coating based on the tongue image and generate a texture feature vector; connecting to a feature fusion analysis module to obtain health status classification results; connecting to a diagnosis result generation module to generate visual charts and voice prompt content; and connecting to a user interaction module to perform remote interaction.

[0014] Compared with the existing technology, the present invention has the following beneficial effects: through the cooperation of the tongue image acquisition module, the multimodal data processing module, the feature fusion analysis module, the diagnosis result generation module and the user interaction module, the diagnosis results can be intuitively obtained based on the tongue image characteristics, thereby reducing the dependence on the experience of traditional Chinese medicine practitioners, which is effectively different from the existing technology that relies on the experience of doctors, resulting in the diagnosis results being not intuitive and lacking readability. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 It is a structural diagram of the intelligent diagnosis system of the present invention;

[0017] Figure 2 It is a flow chart of the intelligent diagnosis method of the present invention. DETAILED DESCRIPTION

[0018] The following is a combination of the embodiments of the present invention Figure 1-2 The technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0019] Example 1

[0020] Figure 1The structure diagram of the intelligent diagnostic system provided in an embodiment of the present invention includes a processor. The processor can be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The system also includes a tongue image acquisition module, a multimodal data processing module, a feature fusion analysis module, a diagnosis result generation module, and a user interaction module; wherein the tongue image acquisition module is data-connected to the multimodal data processing module, the multimodal data processing module is data-connected to the feature fusion analysis module, the feature fusion analysis module is data-connected to the diagnosis result generation module, and the diagnosis result generation module is data-connected to the user interaction module.

[0021] Furthermore, the tongue image acquisition module includes at least a high-resolution camera, a light source adjustment device, and an ambient light compensation sensor; wherein, in this embodiment, the high-resolution camera can be installed at the front end of the target device, that is, the position of the device closest to the tongue, for collecting tongue images, with a resolution of 12 million pixels, supporting macro photography. The light source adjustment device includes a ring-shaped LED light group and a light angle adjuster. It should be noted that the ring-shaped LED light group is installed around the camera, and the light angle adjuster controls the angle of the LED light group through a motor to adapt to the tongue position of different users. The ambient light compensation sensor is installed on the top of the target device, for real-time monitoring of the ambient light intensity, and transmitting the data to the light source adjustment device for dynamic adjustment to ensure that the collected tongue image has stable brightness and color. wherein, in this embodiment, the light source adjustment device includes at least a medical examination lamp, a portable oral lighting device, an integrated oral mirror light source, and a fiber optic light guide device.

[0022] Preferably, the tongue image acquisition module also includes a tongue positioning assistance system, which includes an infrared ranging sensor and an autofocus device. The infrared ranging sensor, mounted at the front end of the device, detects the distance between the user's tongue and the camera and transmits this data to the autofocus device. The autofocus device drives the camera lens via a stepper motor, achieving fast and accurate focusing to ensure the clarity of the tongue image. The infrared ranging sensor has a measurement range of 5cm to 30cm with an accuracy of ±0.1cm, and the autofocus device has a response time of 0.3 seconds.

[0023] Furthermore, the multimodal data processing module includes at least an image preprocessing unit, a color correction unit, and a texture extraction unit; wherein the image preprocessing unit performs denoising, sharpening, and contrast enhancement on the collected tongue image. The color correction unit corrects the color of the tongue image based on the data from the ambient light compensation sensor to eliminate ambient light interference. The texture extraction unit extracts the texture features of the tongue coating and generates a texture feature vector. It should be noted that the texture extraction unit extracts the texture features of the tongue coating and generates a texture feature vector using a grayscale co-occurrence matrix algorithm. The grayscale co-occurrence matrix is ​​a classic algorithm for image texture analysis, which extracts texture features by statistically analyzing the grayscale distribution patterns of pixel pairs with specific spatial relationships in the image.

[0024] Furthermore, the feature fusion analysis module includes at least a deep learning model, a feature weight assignment unit, and a classification decision unit. The deep learning model is a convolutional neural network architecture, comprising an input layer, a convolution layer, a pooling layer, and a fully connected layer. The input layer receives tongue image data processed by the multimodal data processing module, the convolution layer and the pooling layer extract high-level features of the image layer by layer, and the fully connected layer outputs preliminary classification results. The feature weight assignment unit assigns weight values ​​to each feature, based on the importance of color, morphology, and fur texture. The weight values ​​range from 0.1 to 0.5, and the specific assignment is determined by experimental data. The classification decision unit is used to generate the final health status classification result. The classification decision unit combines the output of the deep learning model with the results of the feature weight assignment unit to generate the final health status classification result.

[0025] Preferably, the feature fusion analysis module also includes a dynamic feature update mechanism, which includes an online learning unit and a feature library update unit. The online learning unit receives user feedback data to adjust the parameters of the deep learning model in real time and optimize feature extraction capabilities. The feature library update unit regularly updates the weight values ​​of the feature weight allocation unit based on new tongue image data to ensure that the system can adapt to changes in tongue image characteristics of different populations. The learning rate of the online learning unit is set to 0.01, and the update cycle of the feature library update unit is 7 days.

[0026] Furthermore, the diagnosis result generation module includes at least a health status assessment unit and a diagnosis report generation unit. The health status assessment unit, based on the results of the classification decision unit, accesses a preset health status database and generates a health status assessment report for the user. The diagnosis report generation unit converts the health status assessment report into structured text and, based on the needs of the user interaction module, generates visual charts and voice prompts.

[0027] Preferably, the diagnosis result generation module also includes a health trend prediction unit. The health trend prediction unit uses a time series analysis algorithm to predict the user's health status change trend based on the user's multiple diagnosis data. The health trend prediction unit has a prediction period of 30 days and a prediction error range of ±5%.

[0028] Furthermore, the user interaction module includes at least a touch screen display, a voice prompt device, and a remote communication interface; wherein the touch screen display is mounted on the front of the target device and is used to display diagnostic results and a user operation interface. In this embodiment, the voice prompt device includes a speaker and a microphone for providing voice guidance and receiving user instructions. The remote communication interface is connected to an external medical system via a wireless communication protocol, which supports the upload of diagnostic data and remote consultation functions.

[0029] Preferably, the user interaction module also includes a personalized recommendation system, comprising a diet recommendation generation unit and an exercise plan development unit. The diet recommendation generation unit, based on the user's health status assessment results, accesses a preset diet database to generate personalized diet recommendations. The exercise plan development unit, based on the user's health status assessment results, accesses a preset exercise database to generate a personalized exercise plan. The output of the diet recommendation generation unit and exercise plan development unit is presented to the user via a touchscreen display and voice prompts.

[0030] Based on the above, it should be noted that the tongue positioning auxiliary system is connected to the tongue image acquisition module through an infrared ranging sensor and an automatic focusing device, and is connected to the main control unit through a signal transceiver to receive instructions from the central processor in real time and adjust the focus parameters of the camera.

[0031] The dynamic feature update mechanism is connected to the feature fusion analysis module through the online learning unit and the feature library update unit, and is connected to the main control unit through a signal transceiver. It receives instructions from the central processor in real time and updates the deep learning model and feature weights.

[0032] The health trend prediction unit is connected to the diagnosis result generation module through a time series analysis algorithm, and is connected to the main control unit through a signal transceiver, receiving instructions from the central processor in real time to generate health trend prediction results.

[0033] The personalized recommendation system is connected to the user interaction module through the diet recommendation generation unit and the exercise plan formulation unit, and is connected to the main control unit through a signal transceiver, receiving instructions from the central processor in real time to generate personalized recommendation content.

[0034] It is understandable that the tongue positioning assistance system can adjust the focus parameters of the camera in real time through the infrared ranging sensor and automatic focusing device when the user's tongue position changes, ensuring that the tongue image is always clear, significantly improving the stability and accuracy of tongue image acquisition; the dynamic feature update mechanism adjusts the parameters and feature weights of the deep learning model in real time through the online learning unit and the feature library update unit, ensuring that the system can adapt to the changes in tongue image characteristics of different populations, improving the system's adaptability and diagnostic accuracy; the health trend prediction unit is based on the time series analysis algorithm, and can predict the trend of changes in health status based on the user's multiple diagnostic data, providing users with early warning and health management recommendations, enhancing the practicality and foresight of the system.

[0035] Example 2

[0036] Based on the same inventive concept as the intelligent diagnosis system based on multimodal tongue image feature fusion in the aforementioned embodiment, Figure 2 As shown, the present invention also provides an intelligent diagnosis method based on multimodal tongue image feature fusion, which is executed by a processor. The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0037] The method specifically includes: connecting the tongue image acquisition module to execute tongue image acquisition; connecting the multimodal data processing module to execute tongue coating texture feature acquisition based on the tongue image and generate texture feature vectors; connecting the feature fusion analysis module to execute acquisition of health status classification results; connecting the diagnosis result generation module to execute generation of visual charts and voice prompt content; connecting the user interaction module to execute remote interaction.

[0038] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, from all points of view, the embodiments should be regarded as illustrative and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and range of equivalents of the claims are included in the present invention. Any reference signs in the claims should not be construed as limiting the claim to which they relate.

[0039] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. An intelligent diagnosis system based on multimodal tongue image feature fusion, comprising a processor, characterized in that: It also includes a tongue image acquisition module, a multimodal data processing module, a feature fusion analysis module, a diagnosis result generation module and a user interaction module; wherein the tongue image acquisition module is data-connected to the multimodal data processing module, the multimodal data processing module is data-connected to the feature fusion analysis module, the feature fusion analysis module is data-connected to the diagnosis result generation module, and the diagnosis result generation module is data-connected to the user interaction module.

2. The system according to claim 1, wherein: The tongue image acquisition module includes at least a high-resolution camera, a light source adjustment device and an ambient light compensation sensor; wherein, the high-resolution camera is used to collect tongue image, the light source adjustment device includes a ring LED light group and a light angle adjuster, and the ambient light compensation sensor is used to monitor the ambient light intensity in real time and transmit the data to the light source adjustment device for dynamic adjustment.

3. The system according to claim 1, wherein: The multimodal data processing module includes at least an image preprocessing unit, a color correction unit and a texture extraction unit; wherein, the image preprocessing unit performs denoising, sharpening and contrast enhancement on the collected tongue image, the color correction unit corrects the color of the tongue image, and the texture extraction unit extracts the texture features of the tongue coating and generates a texture feature vector.

4. The system according to claim 3, characterized in that The texture extraction unit extracts the texture features of the tongue coating using a gray level co-occurrence matrix algorithm.

5. The system according to claim 1, wherein: The feature fusion analysis module includes at least a deep learning model, a feature weight allocation unit and a classification decision unit; wherein the deep learning model is a convolutional neural network architecture, the feature weight allocation unit assigns weight values ​​to various features, and the classification decision unit is used to generate the final health status classification result.

6. The system according to claim 1, wherein: The diagnosis result generation module includes at least a health status assessment unit and a diagnosis report generation unit; wherein, the health status assessment unit is used to generate a health status assessment report of the user, and the diagnosis report generation unit is used to generate a visual chart and voice prompt content.

7. The system according to claim 1, wherein: The user interaction module includes at least a touch screen, a voice prompt device and a remote communication interface; wherein, the touch screen is used to display diagnostic results and a user operation interface, the voice prompt device includes at least a speaker and a microphone, and the remote communication interface is connected to an external medical system via a wireless communication protocol.

8. An intelligent diagnosis method based on multimodal tongue image feature fusion, the method being executed by a processor and characterized in that: include: Connect to the tongue image acquisition module to acquire tongue images; Connecting to the multimodal data processing module, performing tongue coating texture feature acquisition based on the tongue image, and generating a texture feature vector; Connect the feature fusion analysis module to obtain the health status classification results; Connect to the diagnosis result generation module to generate visual charts and voice prompts; Connect to the user interaction module to perform remote interaction.

Citation Information

Patent Citations

  • Tongue image recognition method, device and electronic device based on multimodal information

    CN117611581B