A method and system for tongue diagnosis image recognition

CN121661058BActive Publication Date: 2026-04-10TIANJIN ACAD OF TRADITIONAL CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-06
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional tongue diagnosis relies on the physician's personal experience and subjective visual judgment, which makes it difficult to guarantee the accuracy and consistency of the diagnostic results. Furthermore, it is difficult to conduct precise quantification and longitudinal comparison, limiting its application in scenarios that require a high degree of consistency and repeatability.

Method used

By generating color correction and pose compensation matrices, the color and pose of the tongue image are corrected, pixel-level multispectral feature vectors are constructed, and the tongue image features are quantized and segmented using the K-means clustering algorithm to output standardized analysis results.

Benefits of technology

It achieves objectivity and standardization in tongue diagnosis image recognition, improves the accuracy and consistency of diagnostic results, and can deeply explore the subtle features of the tongue image to output standardized analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661058B_ABST
    Figure CN121661058B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer-aided diagnosis, in particular to a tongue diagnosis image recognition-based method and system, comprising the following steps: generating a color correction matrix and an initial spatial coordinate system through a standard color chart, capturing a tongue image and correcting its color and posture, constructing a pixel-level multispectral feature vector, dividing pixels through a K-means clustering algorithm, extracting area, color and shape features and fusing them into a comprehensive vector, matching a discrimination rule and outputting a tongue appearance classification analysis result.In the present application, the image color is standardized through a calibration color chart to ensure accurate tongue color analysis, position offset and physical shaking are compensated in combination with a posture reading, a pixel-level multispectral feature vector is constructed to mine tongue appearance details, a clustering algorithm is used for quantitative segmentation, objective data such as area and color are calculated, morphological and color features are fused for comprehensive discrimination, a standardized analysis result is output, and a change from subjective experience to objective data driving is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer-aided diagnosis, in particular to a tongue diagnosis image recognition method and system. BACKGROUND

[0002] The technical field of computer-aided diagnosis is an important branch of the cross-application of medical informatization and artificial intelligence, mainly involving the analysis and judgment of clinical symptoms or medical images by means of computer image processing, pattern recognition and medical knowledge base, etc. Its core matters include medical image data acquisition and preprocessing, feature extraction and recognition, diagnosis model establishment and reasoning, and output of auxiliary decision-making. This technical field aims to realize the objectivity and standardization of diagnosis through digitalization and intelligent analysis of physiological signals and image information, and widely covers medical image recognition, clinical data mining, pathological image auxiliary analysis and other aspects. Among them, traditional tongue diagnosis refers to observing the external manifestations of the patient's tongue shape, tongue color, tongue fur thickness, etc. in the process of traditional Chinese medicine diagnosis and treatment to infer the internal pathological state of the body. The technical matters targeted are diagnosis requirements based on tongue feature acquisition and analysis. Traditional tongue diagnosis usually relies on the naked eye observation and experience judgment of doctors to complete information extraction and analysis. Common methods include identifying symptom characteristics by manually comparing tongue image atlases or manually judging according to the theory of traditional Chinese medicine tongue diagnosis, thereby completing the auxiliary diagnosis of the patient's health status.

[0003] The diagnosis process of the prior art relies too much on the personal experience and subjective visual judgment of doctors, and the accuracy and consistency of the diagnosis results are easily affected. For example, observing the same tongue image under different lighting conditions will show significant differences in color and luster, leading to different judgments by doctors. At the same time, slight changes in angle or position during collection will also cause visual distortion of the tongue shape features. Moreover, the descriptive conclusions of traditional tongue diagnosis are difficult to accurately quantify and longitudinally compare, making it lack objective basis for dynamic tracking of subtle changes in the disease. This non-standardized diagnosis method limits its application in scenarios requiring high consistency and repeatability. SUMMARY

[0004] The purpose of the present application is to solve the problems existing in the prior art and to provide a tongue diagnosis image recognition method and system.

[0005] In order to achieve the above purpose, the present application adopts the following technical scheme: a tongue diagnosis image recognition method, comprising the following steps:

[0006] S1: generating a mapping relationship between the red, green and blue three-channel values of the standard color chart image generated by illuminating the standard color chart with a monochromatic light source and a white light source and the pre-stored standard color values, generating a color correction matrix, and constructing an initial spatial coordinate system based on the initial posture reading collected;

[0007] S2: capture the tongue image and collect real-time posture readings, call the color correction matrix to transform the color values of the tongue image, generate a multispectral image sequence and a real-time posture reading set;

[0008] S3: generate a perspective compensation matrix and a posture compensation transformation parameter according to the deviation of the real-time posture reading set and the initial spatial coordinate system, perform global perspective correction on the multispectral image sequence, then select an initial tongue region from the white light tongue image, simultaneously call the posture compensation transformation parameter to correct the initial tongue region, calculate the second-order central moment of the initial tongue region, derive the tongue main axis tilt angle, and construct a rotation matrix to perform main axis alignment transformation on the initial tongue region, and combine pixel brightness values to construct a pixel-level multispectral feature vector set;

[0009] S4: input the pixel-level multispectral feature vector set into the K-means clustering algorithm to divide the pixel categories, calculate the area ratio and average color value of the pixel categories, extract the shape value of the initial tongue region, and fuse it into a comprehensive feature vector, match the comprehensive feature vector with the discrimination rule, and output the tongue image classification analysis result.

[0010] As a further scheme of the present application, the color correction matrix includes red channel correction coefficients, green channel correction coefficients and blue channel correction coefficients, the initial spatial coordinate system specifically refers to three-axis rotation components and three-axis translation components, the multispectral image sequence includes white light tongue image frames and monochromatic light tongue image frames, the real-time posture reading set specifically refers to real-time readings of a gyroscope and real-time readings of an accelerometer, the posture compensation transformation parameter includes a rotation matrix and a translation vector, the pixel-level multispectral feature vector set includes a pixel position index and a multispectral response vector, and the comprehensive feature vector includes an area ratio value, an average color value and a tongue shape value. The tongue image classification analysis result specifically refers to a tongue image category identifier and a confidence score.

[0011] As a further scheme of the present application, the initial spatial coordinate system acquisition step specifically includes:

[0012] S101: acquire a standard color plate, and irradiate it with a monochromatic light source and a white light source respectively, simultaneously collect standard color plate images, analyze the standard color plate images, extract the pixel values of the red channel, green channel and blue channel in the image frames, and combine the extracted pixel values into a multidimensional array to generate a standard color plate image pixel value matrix;

[0013] S102: Call the standard color plate image pixel value matrix, and obtain the pre-stored standard color value. A linear transformation equation is constructed between the channel values in the standard color plate image pixel value matrix and the color components corresponding to the standard color value. The linear transformation equation is solved to obtain a transformation coefficient. The transformation coefficient is organized into a three-order matrix form to generate a color correction matrix.

[0014] S103: Collect initial attitude readings of roll angle, pitch angle, and yaw angle data from the attitude sensor. According to the initial attitude readings, the roll angle, the pitch angle, and the yaw angle data are defined as the first, second, and third rotation axes of the spatial coordinate system, respectively, and are used as a reference to construct the initial spatial coordinate system.

[0015] As a further scheme of the present application, the step of obtaining the multi-spectral image sequence and the real-time attitude reading set is specifically:

[0016] S201: Monitor the image sensor to capture a single frame of tongue image, and synchronously monitor the attitude sensor to collect real-time attitude readings. Color information of a plurality of pixel points in the tongue image is deconstructed into red, green, and blue three-channel intensity values to generate an original tongue RGB image. Gyroscope and accelerometer original data in the real-time attitude readings are integrated to establish an initial attitude parameter set.

[0017] S202: Obtain a preset color correction matrix. A pixel color vector is constructed according to the red, green, and blue three-channel intensity values of each pixel position in the original tongue RGB image. The color correction matrix and the pixel color vector are called to perform matrix multiplication operation. The transformation vector generated by the operation is used to replace the red, green, and blue three-channel intensity values to obtain a corrected spectral data frame.

[0018] S203: For a plurality of corrected spectral data frames and the initial attitude parameter set obtained in a continuous sampling period, a multi-spectral image sequence is constructed by stacking the plurality of corrected spectral data frames along the time axis dimension. The acquisition time stamp corresponding to the initial attitude parameter set is paired with the image frame index of the multi-spectral image sequence to generate the multi-spectral image sequence and the real-time attitude reading set.

[0019] As a further scheme of the present application, the step of obtaining the pixel-level multi-spectral feature vector set is specifically:

[0020] S301: Calculate homography matrix based on the deviation of the real-time pose reading set and the initial spatial coordinate system, eliminate the perspective distortion caused by the shooting angle tilt, calculate the zero-order moment M00 and the first-order moment M10, M01 of the initial tongue region to determine the centroid coordinates, and combine the second-order central moments μ20, μ11, μ02 to calculate the deflection angle α of the tongue main axis relative to the vertical direction, and generate the pose compensation transformation parameters including rotation and translation components;

[0021] S302: Call the white light tongue image of the multispectral image sequence, set the pixel chrominance range and the brightness lower threshold, traverse all pixels of the white light tongue image, remove background pixels that do not meet the pixel chrominance range and the brightness lower threshold, and screen to obtain an initial tongue region. Call the pose compensation transformation parameters to perform inverse geometric transformation on the pixel coordinates of the initial tongue region to obtain a corrected tongue region pixel set.

[0022] S303: Traverse the pixel position recorded in the corrected tongue region pixel set, extract pixel brightness values from N spectral channel images of the multispectral image sequence according to the pixel position one by one, arrange the extracted N pixel brightness values in order of wavelength from short to long to form an N-dimensional feature vector, aggregate all N-dimensional feature vectors of the pixels to construct a pixel-level multispectral feature vector set.

[0023] As a further scheme of the present application, the acquisition step of the tongue image classification analysis result is specifically:

[0024] S401: For the pixel-level multispectral feature vector set, iteratively divide according to the Euclidean distance between vectors until the cluster center is stable, form a plurality of pixel categories, and calculate the area ratio by counting the proportion of the number of pixels in a plurality of pixel categories to the total number of pixels in the initial tongue region. At the same time, perform arithmetic average operation on the color values of all pixels in a plurality of pixel categories to obtain average color values. Finally, combine the area ratio and the average color value of all pixel categories to generate a pixel category attribute set.

[0025] S402: According to the contour coordinates of the initial tongue region, calculate the ratio of its circumference to area as a shape value, call the pixel category attribute set, and serialize the shape value and the area ratio and the average color value in the pixel category attribute set to establish a comprehensive feature vector.

[0026] S403: Compare the area ratio, the average color value, and the shape value of the comprehensive feature vector with the numerical range threshold preset for different tongue image categories in the discrimination rule one by one. When all the values of the comprehensive feature vector fall within the preset threshold range of a tongue image category, it is determined that the matching is successful, and the tongue image classification analysis result is obtained.

[0027] As a further scheme of the present application, in the screening step of the initial tongue body region, the setting method of the pixel chrominance range and the luminance lower limit threshold is specifically:

[0028] Obtaining the white light tongue image, calculating the three-channel pixel value distribution histogram of the white light tongue image in the hue, saturation and luminance color space, generating a hue distribution map, a saturation distribution map and a luminance distribution map;

[0029] Analyzing the luminance distribution map, using a peak detection algorithm to identify the main peak and the secondary peak in the luminance distribution map, defining the luminance value corresponding to the main peak as the background luminance center, and defining the luminance value corresponding to the secondary peak as the target luminance center;

[0030] According to the target luminance center and the preset luminance standard deviation multiple, the luminance lower limit threshold is calculated, and the pixels with luminance values in the neighborhood of the target luminance center are analyzed, and the distribution interval of the pixels in the hue distribution map and the saturation distribution map is determined, and the pixel distribution interval is taken as the pixel chrominance range;

[0031] Based on the calculated luminance lower limit threshold and the pixel chrominance range, each pixel point in the white light tongue image is traversed and judged, and the pixel points meeting the conditions are screened to generate the initial tongue body region;

[0032] The preset luminance standard deviation multiple is an empirical value determined by statistical analysis of multiple sample tongue images, and the value range is 2 to 4.

[0033] As a further scheme of the present application, in the step of dividing the pixel classes by the K-means clustering algorithm, the judgment condition for the stability of the clustering center is specifically:

[0034] Randomly selecting K multi-spectral response vectors from the set of pixel-level multi-spectral feature vectors as initial clustering centers;

[0035] For each multi-spectral response vector in the set of pixel-level multi-spectral feature vectors, calculate the Euclidean distance from the multi-spectral response vector to multiple clustering centers, and assign the multi-spectral response vector to the pixel class corresponding to the nearest clustering center;

[0036] Recalculating the mean of all multi-spectral response vectors in each pixel class, and updating the mean as the current clustering center;

[0037] Repeat the assignment and update steps, and judge whether the clustering process converges by the following formula:

[0038]

[0039] wherein, AC represents the average variation of all cluster centers between two consecutive iterations, K represents the preset number of pixel categories, k is the pixel category index from 1 to K, represents the cluster center vector of the kth pixel category at the tth iteration, represents the cluster center vector of the pixel category at the previous iteration, i.e., the (t-1)th iteration, represents the Euclidean norm of the vector, represents the preset cluster center convergence threshold;

[0040] The preset number of pixel categories K is determined according to the common classification number of the quality and color of the coating in the tongue diagnosis theory, and the value range is 3 to 7;

[0041] When the AC is less than or equal to the preset cluster center convergence threshold, it is determined that the clustering process converges, and a plurality of pixel categories are generated.

[0042] As a further scheme of the present application, the calculation step of the confidence score in the tongue classification analysis result is specifically:

[0043] When the comprehensive feature vector matches the discriminant rule of a tongue category successfully, the preset numerical range threshold of the tongue category is obtained, including the upper limit and the lower limit of each feature;

[0044] The area ratio value, the average color value and the tongue body shape value in the comprehensive feature vector are called, and the confidence score is calculated by the following formula:

[0045]

[0046] wherein, S conf represents the confidence score, M represents the total number of dimensions of the comprehensive feature vector, j is the feature index from 1 to M, w j represents the preset weight coefficient of the jth feature and the sum of all weight coefficients is 1, v j represents the actual value of the jth dimension in the comprehensive feature vector, U j represents the upper limit of the preset numerical range of the jth feature matched with the tongue category, L j represents the lower limit of the preset numerical range of the feature, and δ represents a smoothing factor set to avoid zero denominator;

[0047] The preset weight coefficient w j of the jth feature is determined according to the contribution of the differentiated feature to the tongue classification result, by expert scoring method or machine learning model training;

[0048] ​The calculated confidence score is combined with tongue image category identifiers of the tongue image categories to jointly generate the tongue image classification analysis result.

[0049] A tongue diagnosis image recognition-based system for implementing the tongue diagnosis image recognition-based method described above, the system comprising:

[0050] A reference construction module for generating a mapping relationship between red, green, and blue three-channel values of a standard color chart image generated by illuminating a standard color chart with a monochromatic light source and a white light source and pre-stored standard color values, generating a color correction matrix, and constructing an initial spatial coordinate system based on an initial posture reading collected, and transmitting the color correction matrix and the initial spatial coordinate system to a tongue image collection and color correction module;

[0051] A tongue image collection and color correction module for capturing a tongue image and collecting a real-time posture reading, calling the color correction matrix to transform color values of the tongue image, generating a multispectral image sequence and a real-time posture reading set, and transmitting the multispectral image sequence, the real-time posture reading set, and the initial spatial coordinate system to a posture correction and feature construction module;

[0052] A posture correction and feature construction module for generating posture compensation transformation parameters according to deviations of the real-time posture reading set and the initial spatial coordinate system, screening an initial tongue region from a white light tongue image of the multispectral image sequence, simultaneously correcting the initial tongue region by calling the posture compensation transformation parameters, combining pixel brightness values to construct a pixel-level multispectral feature vector set, and transmitting the pixel-level multispectral feature vector set to a comprehensive discriminant analysis module;

[0053] A comprehensive discriminant analysis module for inputting the pixel-level multispectral feature vector set to a K-means clustering algorithm to divide pixel categories, calculating area proportions and average color values of the pixel categories, extracting shape values of the initial tongue region, fusing the shape values into a comprehensive feature vector, matching the comprehensive feature vector with a discriminant rule, and outputting a tongue image classification analysis result.

[0054] Compared with the prior art, the present application has the following advantages and positive effects:

[0055] In the present application, the standard color chart under monochromatic and white light sources is calibrated to regulate the image color, ensuring the accuracy and consistency of tongue color analysis, combining posture reading to compensate for positional deviation in the collection process, eliminating the interference of physical shaking on tongue shape recognition, constructing pixel-level multispectral feature vectors to deeply mine subtle feature information of tongue images, and using a clustering algorithm to quantitatively segment tongue image features, calculate objective data such as area proportions and color means of each region, fuse morphological and color features for comprehensive discrimination, and output standardized analysis results, realizing the transition from subjective experience dependence to objective data driving. Attached Figure Description

[0056] Figure 1 This is an overall flowchart of the tongue diagnosis image recognition method of the present invention;

[0057] Figure 2 This is a flowchart illustrating the construction process of the color correction matrix and initial spatial coordinate system of this invention.

[0058] Figure 3 This is a flowchart illustrating the process of acquiring multispectral image sequences and real-time attitude reading sets according to the present invention.

[0059] Figure 4 This is a flowchart of the tongue region correction and feature vector construction process of the present invention;

[0060] Figure 5 This is a flowchart of the output of tongue image classification and analysis results in this invention. Detailed Implementation

[0061] In the description of this invention, the system architecture relationships or data processing flows indicated by terms such as "layer," "module," "interface," "data flow," "client," and "server" are all defined based on the architecture diagram or flowchart corresponding to the embodiments. This way of describing is only used to clearly illustrate the logical relationships between the elements in the technical solution, and not to limit the physical deployment form. The term "multiple" includes two or more technical units, including but not limited to multiple data nodes, processing threads, service instances, or functional components and other scalable elements. The specific number is determined according to the actual business scenario and needs to be specifically specified.

[0062] Please see Figure 1 and Figure 2 This invention provides a method for tongue diagnosis image recognition, comprising the following steps:

[0063] S1: The mapping relationship between the red, green, and blue channel values ​​of the standard color chart image generated by illuminating the standard color chart with a monochromatic light source and a white light source and the pre-stored standard color values ​​is used to generate a color correction matrix, and an initial spatial coordinate system is constructed based on the collected initial attitude readings.

[0064] The color correction matrix includes red channel correction coefficients, green channel correction coefficients, and blue channel correction coefficients;

[0065] The initial spatial coordinate system consists of three rotational components and three translational components.

[0066] The specific steps for obtaining the initial spatial coordinate system are as follows:

[0067] S101: Obtain a standard color chart, and respectively use a single-color light source and a white light source to irradiate it, synchronously collect standard color chart images, analyze the standard color chart images, extract pixel values in red, green and blue channels in the image frames, and combine the extracted pixel values into a multi-dimensional array to generate a standard color chart image pixel value matrix;

[0068] S102: Call the standard color chart image pixel value matrix, and obtain the pre-stored standard color values, construct a linear transformation equation between the multiple channel values in the standard color chart image pixel value matrix and the corresponding color components of the standard color values, solve the linear transformation equation to obtain the transformation coefficients, organize the transformation coefficients into a three-order matrix form to generate a color correction matrix;

[0069] S103: Collect initial attitude readings of roll angle, pitch angle and yaw angle data from the attitude sensor, and define the roll angle, pitch angle and yaw angle data as the first, second and third rotation axes of the spatial coordinate system according to the initial attitude readings, and take them as the reference to construct an initial spatial coordinate system.

[0070] In the process of generating the color correction matrix, a standard X-Rite ColorCheckerClassic color chart is prepared first, which contains 24 standard color blocks. The CIELAB standard color values of each color block are pre-stored in the local database. In a closed environment with constant illumination of 500 Lux and color temperature of 6500K, the standard color chart is placed in front of the image acquisition device at a distance of 30 centimeters. The image acquisition device uses a CMOS sensor with a resolution of 1920x1080 pixels. First, turn on the single-color red LED light source with a wavelength of 630 nm to irradiate the standard color chart, and collect and store 10 frames of standard color chart images. Then, turn off the red light source, turn on the single-color green LED light source with a wavelength of 525 nm and the single-color blue LED light source with a wavelength of 465 nm in turn, respectively irradiate the color chart and collect 10 frames of images for each. Finally, turn off all single-color light sources, turn on the white LED light source, and collect 10 frames of images again.

[0071] For the collected images, taking one frame under white light as an example for analysis. The system first locates the 24 standard color block regions in the image, and calculates the arithmetic mean of the pixel values of all pixel points in each color block region in the red (R), green (G), and blue (B) channels; for example, for the color block 'dark skin' (Darkskin), the pre-stored standard RGB color value and the actual RGB value obtained by image acquisition and analysis are as shown in the following table; similarly, for the color block 'blue sky' (Bluesky), the standard value and the collected value are as shown in the following table; select three non-linearly related color blocks, for example, 'dark skin', 'blue sky', and 'bluish green' (Bluishgreen), the standard value and the actual collected value are as follows:

[0072]

[0073] According to the three groups of data, a linear transformation equation set is constructed: , By analogy, by solving the linear equation set, a 3x3 transformation coefficient matrix is obtained. In this embodiment, by least square fitting calculation on all 24 color block data, a more universal color correction matrix is finally obtained, and the specific value is:

[0074]

[0075] The matrix is the color correction matrix called in the subsequent steps, where the first row [0.92, -0.11, 0.05] is the red channel correction coefficient, the second row [0.08, 0.95, -0.03] is the green channel correction coefficient, and the third row [-0.04, 0.07, 0.91] is the blue channel correction coefficient.

[0076] In the step of constructing the initial spatial coordinate system, a six-axis attitude sensor (MPU-6050) is integrated inside the image acquisition device. In the state that the device is stationary on the horizontal detection platform, 100 data points are continuously collected from the attitude sensor. Each data point contains the readings of roll angle, pitch angle and yaw angle. The arithmetic mean of the 100 data points is calculated to eliminate random noise. For example, the calculated average values are: roll angle 0.12°, pitch angle -0.08°, and yaw angle 25.3°. Taking this set of initial attitude readings as the reference, a right-handed coordinate system is defined: the initial rotation state of the roll angle 0.12° as the X axis (first rotation axis), the initial rotation state of the pitch angle -0.08° as the Y axis (second rotation axis), and the initial rotation state of the yaw angle 25.3° as the Z axis (third rotation axis). At the same time, the translation position of the device in the three-dimensional space at this time (0, 0, 0) is recorded as the reference of the translation component. Thus, an initial spatial coordinate system containing three-axis rotation components (0.12°, -0.08°, 25.3°) and three-axis translation components (0, 0, 0) is constructed.

[0077] Please refer to Figure 1 and Figure 3 , S2: capturing the tongue image and collecting real-time attitude readings, calling the color correction matrix to transform the color values of the tongue image, generating a multispectral image sequence and a real-time attitude reading set;

[0078] The multispectral image sequence includes white light tongue image frames and monochromatic light tongue image frames.

[0079] The real-time attitude reading set specifically refers to the real-time readings of the gyroscope and the real-time readings of the accelerometer.

[0080] The steps of obtaining the multispectral image sequence and the real-time attitude reading set are specifically: S201: monitoring the image sensor to capture a single-frame tongue image, synchronously monitoring the attitude sensor to collect real-time attitude readings, deconstructing the color information of multiple pixel points in the tongue image into red, green and blue three-channel intensity values, generating an original tongue RGB image, and integrating the original data of the gyroscope and the accelerometer in the real-time attitude readings to establish an initial attitude parameter set;

[0081] S202: obtaining a preset color correction matrix, constructing a pixel color vector according to the red, green and blue three-channel intensity values of each pixel position in the original tongue RGB image, calling the color correction matrix and the pixel color vector to perform matrix multiplication operation, replacing the red, green and blue three-channel intensity values with the generated transformation vector to obtain a corrected spectral data frame;

[0082] S203: For the plurality of corrected spectral data frames obtained in the continuous sampling period and the initial set of pose parameters, stack the plurality of corrected spectral data frames along the time axis dimension to construct a multispectral image sequence, and pair the acquisition time stamp corresponding to the initial set of pose parameters with the image frame index of the multispectral image sequence to generate the multispectral image sequence and the set of real-time pose readings.

[0083] When the user aligns the device to the tongue and starts the acquisition procedure, the image sensor captures tongue images at a rate of 30 frames per second, while the pose sensor acquires real-time pose readings at a frequency of 100 times per second. The system processes one frame of image and the synchronized pose data. Assume that at time stamp T = 0.033 seconds, the image sensor captures one frame of tongue image, and synchronously monitors a set of real-time pose readings of the pose sensor. For this frame of tongue image, select one pixel point P in the tongue tip region for processing. The original color information of this pixel point is decomposed into intensity values of three channels of red, green, and blue, with specific values of (210, 125, 110), which is a pixel color vector in the original tongue RGB image. At the same time, integrate the raw data of the gyroscope (e.g., angular velocity of (0.5 rad / s, -0.2 rad / s, 0.1 rad / s)) and the raw data of the accelerometer (e.g., acceleration of (0.01 g, 0.02 g, -0.99 g)) acquired by the pose sensor at T = 0.033 seconds to establish the initial set of pose parameters at this moment.

[0084] Next, obtain the preset color correction matrix generated in step S1. Construct the red, green, and blue channel intensity values of pixel point P (210, 125, 110) into a 1x3 pixel color vector V orig. Call the color correction matrix M calib generated in step S1:

[0085]

[0086] Perform matrix multiplication operation of the color correction matrix and the pixel color vector V orig, and the calculation process is: R corr=0.92*210+(-0.11)*125+0.05*110=193.2-13.75+5.5=184.95;

[0087] G corr=0.08*210+0.95*125+(-0.03)*110=16.8+118.75-3.3=132.25;

[0088] B corr=(-0.04)*210+0.07*125+0.91*110=-8.4+8.75+100.1=100.45;

[0089] The generated transform vector V_corr = [184.95, 132.25, 100.45]. The result is rounded to obtain new RGB three-channel intensity values ​​(185, 132, 100), which replace the original intensity values. This process is repeated for each pixel in the image to generate a frame of corrected spectral data. This process is also applicable to tongue image frames acquired under monochromatic light sources (e.g., red or green light) to generate corresponding monochromatic tongue image frames.

[0090] Within a continuous sampling period, e.g., 1 second, the system acquires 30 frames of calibrated spectral data (including white light tongue image frames and monochromatic light tongue image frames) and 100 sets of initial attitude parameters. The system stacks these 30 frames of calibrated spectral data along the time axis in chronological order of acquisition, constructing a multispectral image sequence. Subsequently, the system retrieves the acquisition timestamp of each frame and pairs it with the attitude parameter set whose acquisition time is closest to the first frame in the 100 sets of initial attitude parameters. For example, if the timestamp of the first frame is T=0.033s, it is paired with the attitude parameter set whose acquisition time is closest to 0.033s; if the timestamp of the second frame is T=0.066s, it is paired with the attitude parameter set whose acquisition time is closest to 0.066s. In this way, each frame index of the multispectral image sequence is precisely paired with a real-time attitude reading, ultimately generating a multispectral image sequence and a set of real-time attitude readings containing 30 pairs of "calibrated spectral data frame - real-time attitude reading" data.

[0091] Please see Figure 1 and Figure 4 S3: Based on the deviation between the real-time attitude reading set and the initial spatial coordinate system, a perspective compensation matrix and attitude compensation transformation parameters are generated, and global perspective correction is performed on the multispectral image sequence. Then, the initial tongue region is selected from the white light tongue image, and the attitude compensation transformation parameters are called to correct the initial tongue region. The second central moment of the initial tongue region is calculated, the principal axis tilt angle of the tongue is derived, and a rotation matrix is ​​constructed to align the principal axis of the tongue, and a set of pixel-level multispectral feature vectors is constructed.

[0092] The attitude compensation transformation parameters include the rotation matrix and the translation vector;

[0093] The pixel-level multispectral feature vector set includes pixel location indices and multispectral response vectors;

[0094] The specific steps for obtaining the pixel-level multispectral feature vector set are as follows:

[0095] S301: Perform global perspective correction, based on the real-time posture reading set, the system extracts the posture data of a frame (for example, the 15th frame) from the real-time posture reading set; assume that the inclination of the collection device relative to the initial spatial coordinate system is monitored (for example, the pitch angle deviation is -10°, that is, the lens is tilted downward); at this time, physical alignment alone cannot eliminate the perspective distortion of the image (that is, the root of the tongue appears wider than the actual, and the tip of the tongue appears narrower than the actual); the system uses the angle deviation to construct a perspective compensation matrix (PerspectiveMatrix) to perform inverse perspective projection transformation on the entire image to restore the image view to the standard orthoview plane;

[0096] S302: Perform tongue region segmentation and principal axis calculation, call the white light tongue image after perspective correction, determine the brightness threshold using the "double peak method" (as described above, the main peak background is 45, the secondary peak tongue is 175, and the threshold is 130), and combine the color range to remove the background to obtain the initial tongue region mask after binaryzation. At this time, although the perspective has been corrected, the patient may have a "tilted head" or "tilted tongue" phenomenon. In order to obtain accurate shape values, principal axis alignment must be performed. The system performs image moment (ImageMoments) analysis on the binary region:

[0097] Calculate the centroid: count the coordinates (x, y) of the white pixels in the region, calculate the zero-order distance M00 (area) and the first-order moment M10, M01;

[0098] The centroid coordinates are Cx=M10 / M00 and Cy=M01 / M00;

[0099] Calculate the deflection angle: calculate the second-order central moments μ20=∑(x-Cx)2, μ02=∑(y-Cy)2, μ11=∑(x-Cx)(y-Cy); use the formula

[0100] α=0.5Xarctan(2μ11 / (μ20-μ02)) to solve the deflection angle α of the tongue principal axis;

[0101] Perform alignment: assume that α=8° (8 degrees to the right) is calculated, the system constructs a rotation matrix to rotate the tongue region counterclockwise by 8 degrees with the centroid as the origin;

[0102] Through the above steps, the corrected tongue region pixel set not only eliminates the shooting angle error of the device, but also eliminates the posture error of the patient, achieving standardization;

[0103] S303: Traverse the pixel positions recorded in the corrected tongue body region pixel set, extract the pixel brightness values from the N spectral channel images of the multispectral image sequence according to the pixel positions one by one, arrange the extracted N pixel brightness values in order from short to long according to the spectral channel wavelength, form an N-dimensional feature vector, aggregate the N-dimensional feature vectors of all pixels, and construct a pixel-level multispectral feature vector set;

[0104] In the screening step of the initial tongue body region, the setting method of the pixel chrominance range and the brightness lower limit threshold is specifically:

[0105] Obtain the white light tongue image, calculate the three-channel pixel value distribution histogram of the white light tongue image in the hue, saturation and brightness color space, and generate the hue distribution map, the saturation distribution map and the brightness distribution map;

[0106] Analyze the brightness distribution map, identify the main peak and the secondary peak in the brightness distribution map using a peak detection algorithm, define the brightness value corresponding to the main peak as the background brightness center, and define the brightness value corresponding to the secondary peak as the target brightness center;

[0107] According to the target brightness center and the preset brightness standard deviation multiple, the brightness lower limit threshold is calculated, at the same time, the pixels whose brightness values are in the neighborhood of the target brightness center are analyzed, and the distribution interval of the pixels in the hue distribution map and the saturation distribution map is determined, and the pixel distribution interval is taken as the pixel chrominance range;

[0108] Based on the calculated brightness lower limit threshold and the pixel chrominance range, each pixel point in the white light tongue image is traversed and judged, the pixel points meeting the conditions are screened, and the initial tongue body region is generated;

[0109] The preset brightness standard deviation multiple is an empirical value determined according to statistical analysis of multiple sample tongue images, and the value range is 2 to 4.

[0110] The system extracts the pose data of a certain frame (e.g., frame 15, timestamp T=0.5s) from the real-time pose reading set, which, after processing, obtains the following pose angles: roll angle 0.32°, pitch angle -0.15°, and yaw angle 25.8°. At the same time, the initial spatial coordinate system constructed in step S1 is monitored, and its reference pose is (0.12°, -0.08°, 25.3°). The real-time pose data is spatially aligned with the initial spatial coordinate system, and the deviation values of the two in the three rotational degrees of freedom are calculated: roll angle deviation = 0.32°-0.12°=0.20° pitch angle deviation =-0.15°-(-0.08°)=-0.07° yaw angle deviation = 25.8°-25.3°=0.50° Based on these deviation values, the system constructs a perspective compensation matrix (HomographyMatrix) for inverse perspective projection transformation of the entire image to eliminate the perspective distortion of the image caused by the tilting of the collection device (e.g., shooting from above or from the side), ensuring the true geometric proportion of the tongue projection on the two-dimensional plane.

[0111] Next, the white light tongue image frame (frame 15) in the multispectral image sequence is called to screen the initial tongue region. To do this, it is necessary to set the pixel chroma range and the lower limit threshold of brightness. The setting process first converts the white light tongue image from RGB space to HSV (hue, saturation, brightness) space. Then, the pixel value distribution histogram of the entire image in the H, S, and V channels is calculated to generate the hue distribution map, saturation distribution map, and brightness distribution map. Analyzing the brightness distribution map, it is found that it presents a bimodal pattern, with a main peak at a brightness value of 45 and a secondary peak at a brightness value of 175. The system defines the brightness value 45 corresponding to the main peak as the background brightness center and the brightness value 175 corresponding to the secondary peak as the target brightness center.

[0112] The calculation of the lower limit threshold of brightness is based on the target brightness center and the preset brightness standard deviation multiple. The determination of the preset brightness standard deviation multiple is based on the statistical experiment on the database of 500 tongue samples under different light conditions and individual differences. In the experiment, the standard deviation multiple is taken as 2.0, 2.5, 3.0, 3.5, and 4.0, respectively, to calculate the accuracy, recall rate, and F1 score of the segmentation result. The experimental data show that when the multiple value is 3.0, the F1 score reaches the peak, indicating that the segmentation effect is best. Therefore, the preset brightness standard deviation multiple is determined as 3.0 in the embodiment. After calculation, the standard deviation of the pixel brightness near the target brightness center 175 is 15. Therefore, the lower limit threshold of brightness = 175-3.0*15 = 130. At the same time, the pixels with the brightness value in the (neighborhood of the target brightness center 175) range are analyzed, and the distribution intervals of these pixels in the hue distribution graph and the saturation distribution graph are counted. The statistical results show that the hue values of these pixels are mainly distributed in the intervals [0, 0.1] and [0.9, 1.0], and the saturation values are distributed in the interval [0.3, 0.8]. This interval is set as the pixel chroma range. Based on the calculated lower limit threshold of brightness 130 and the above pixel chroma range, the system traverses each pixel point of the white light tongue image of the frame; if the V value of a pixel is greater than 130, the H value is in [0, 0.1] or [0.9, 1.0], and the S value is in [0.3, 0.8], the pixel is retained, otherwise it is excluded. All the retained pixels together constitute the initial tongue body region. Subsequently, the previously generated posture compensation transformation parameters (rotation matrix R) are called to perform inverse geometric transformation on the coordinates of each pixel in the initial tongue body region to obtain the corrected tongue body region pixel set.

[0113] Finally, a pixel-level multispectral feature vector set is constructed. Assuming that the multispectral image sequence in the embodiment contains 5 spectral channels (N = 5), arranged in order from short to long wavelengths: blue channel (465 nm), green channel (525 nm), white channel, red channel (630 nm), and near-infrared channel (850 nm). The system traverses each pixel position recorded in the corrected tongue body region pixel set. For example, for the pixel position (x, y), the system extracts the pixel brightness values of this position from the 5 spectral channel images respectively. The extracted values may be: blue channel brightness 95, green channel brightness 150, white channel brightness 178, red channel brightness 205, and near-infrared channel brightness 130. These 5 brightness values are arranged in order to form a 5-dimensional feature vector. This vector is a multispectral response vector. Aggregating the 5-dimensional feature vectors of all pixels in the corrected tongue body region, a pixel-level multispectral feature vector set containing all pixel position indexes and their corresponding multispectral response vectors is finally constructed.

[0114] See Figure 1 and Figure 5S4: input the pixel-level multi-spectral feature vector set into the K-means clustering algorithm to divide the pixel categories, calculate the area proportion and average color value of the pixel categories, extract the shape value of the initial tongue region, fuse the comprehensive feature vector, match the comprehensive feature vector with the discrimination rule, and output the tongue image classification analysis result.

[0115] The comprehensive feature vector includes the area proportion value, the average color value, and the tongue shape value.

[0116] The tongue image classification analysis result specifically refers to a tongue image category identifier and a confidence score.

[0117] The acquisition steps of the tongue image classification analysis result are specifically:

[0118] S401: for the pixel-level multi-spectral feature vector set, iteratively divide according to the Euclidean distance between vectors until the clustering center is stable, form a plurality of pixel categories, and calculate the area proportion by calculating the proportion of the number of pixels in the plurality of pixel categories to the total number of pixels in the initial tongue region. At the same time, the average color value is obtained by performing arithmetic average operation on the color values of all pixels in the plurality of pixel categories. Finally, the area proportion and the average color value of all pixel categories are combined to generate a pixel category attribute set;

[0119] S402: calculate the ratio of the perimeter to the area of the initial tongue region as the shape value according to the contour coordinates, call the pixel category attribute set, and serialize and splice the shape value, the area proportion, and the average color value in the pixel category attribute set to establish a comprehensive feature vector;

[0120] S403: compare the area proportion, the average color value, and the shape value of the comprehensive feature vector with the numerical range threshold preset for the differentiated tongue image category in the discrimination rule one by one. When all the values of the comprehensive feature vector fall within the preset threshold range of a tongue image category, it is determined that the matching is successful, and the tongue image classification analysis result is obtained.

[0121] In the step of dividing the pixel categories by the K-means clustering algorithm, the judgment condition for the stable clustering center is specifically:

[0122] Randomly select K multi-spectral response vectors from the pixel-level multi-spectral feature vector set as the initial clustering center;

[0123] For each multi-spectral response vector in the pixel-level multi-spectral feature vector set, calculate the Euclidean distance from the plurality of clustering centers, and assign the multi-spectral response vector to the pixel category corresponding to the nearest clustering center;

[0124] Recalculate the mean of all multi-spectral response vectors in each pixel category, and update the mean as the current clustering center;

[0125] The assigning and updating steps are repeatedly performed, and whether the clustering process converges is determined by the following formula:

[0126]

[0127] wherein, ΔC represents the average change of all cluster centers between two consecutive iterations, K represents the preset number of pixel categories, k is the pixel category index from 1 to K, represents the cluster center vector of the kth pixel category at the tth iteration, represents the cluster center vector of the pixel category at the previous iteration, i.e., the (t-1)th iteration, represents the Euclidean norm of the vector, represents the preset cluster center convergence threshold;

[0128] The preset number of pixel categories K is determined according to the common classification number of the quality and color of the coating in the tongue diagnosis theory, and the value range is 3 to 7;

[0129] When ΔC is less than or equal to , it is determined that the clustering process converges, and a plurality of pixel categories are generated;

[0130] The calculation steps of the confidence score in the tongue classification analysis result are as follows:

[0131] When the comprehensive feature vector matches a tongue category successfully, the preset numerical range threshold of the tongue category is obtained, including the upper limit and the lower limit of each feature;

[0132] The area proportion value, the average color value and the tongue shape value in the comprehensive feature vector are called, and the confidence score is calculated by the following formula:

[0133]

[0134] wherein, S conf represents the confidence score, M represents the total number of dimensions of the comprehensive feature vector, j is the feature index from 1 to M, w j represents the preset weight coefficient of the jth feature and the sum of all weight coefficients is 1, v j represents the actual value of the jth dimension in the comprehensive feature vector, U j represents the upper limit of the preset numerical range of the jth feature matched with the tongue category, L j represents the lower limit of the preset numerical range of the feature, and δ represents a smoothing factor set to avoid zero denominator;

[0135] The preset weight coefficient w j of the jth feature is determined according to the contribution of the differentiated feature to the tongue classification result, by expert scoring method or machine learning model training;

[0136] The calculated confidence score is combined with the tongue image category identifier of the tongue image category to jointly generate a tongue image classification analysis result.

[0137] The system inputs the pixel-level multi-spectral feature vector set to the K-means clustering algorithm. The preset number of pixel categories K is determined according to the common combination classification number of the tongue fur texture (such as thick, thin, moist, dry, greasy) and fur color (such as white, yellow, gray-black) in the theory of tongue diagnosis of traditional Chinese medicine. In order to cover common cases, 1000 tongue image samples labeled by experts need to be pre-clustered. In the experiment, K values are 3, 4, 5, 6, and 7, and the silhouette coefficient and Calinski-Harabasz index are used as evaluation indexes. The experimental results show that when K = 5, the comprehensive score of the two indexes is the highest, indicating that the clustering effect at this time is the most consistent with the classification of traditional Chinese medicine theory. Therefore, K = 5 is set in this embodiment. The algorithm first randomly selects 5 multi-spectral response vectors from the feature vector set as the initial cluster centers. Then, for each vector in the set, the Euclidean distance between it and the 5 cluster centers is calculated, and it is assigned to the pixel category corresponding to the cluster center with the closest distance. After all the vectors are assigned, the mean of all vectors in each pixel category is recalculated, and the mean is updated as the new cluster center. This assignment and update step is repeated.

[0138] After each iteration, it is judged whether the clustering process converges by the following formula:

[0139]

[0140] This formula is used to calculate the average change of all cluster centers between two consecutive iterations.

[0141] Where ΔC represents the average change; K is the preset number of pixel categories, which is 5 in this example; k is the pixel category index, from 1 to 5; represents the cluster center vector of the kth pixel category at the current tth iteration, which is a 5-dimensional vector; represents the cluster center vector of the category at the previous iteration (t-1th); represents the Euclidean norm (i.e. Euclidean distance) between the two vectors; is a preset cluster center convergence threshold. The setting of this threshold is based on the trade-off experiment between convergence speed and accuracy. By monitoring the number of iterations and the stability of the final clustering result under different values (such as 0.1, 0.01, 0.001), it is determined that is a value that can stably converge within 100 iterations and has minimal result fluctuations.

[0142] Example: Suppose in the 10th and 11th iterations, the changes of the 5 cluster centers are as follows (for simplicity, take two-dimensional vectors as an example): … (other 3 centers) Calculate AC:

[0143] ;

[0144] Suppose the distances of the other three centers are 0.15, 0.18, and 0.12, respectively.

[0145] .

[0146] Since 0.1582 > 0.01, it is determined that the clustering process has not converged, and the next iteration continues. When the AC value calculated in a certain iteration is 0.008, since 0.008 ≤ 0.01, it is determined that the clustering process has converged, forming 5 pixel categories. Subsequently, the proportion of the number of pixels in each pixel category to the total number of pixels in the initial tongue body region is calculated to obtain the area proportion. At the same time, the arithmetic mean operation is performed on the color values (in the CIELAB space) of all pixels in each category to obtain the average color value.

[0147] Next, according to the contour coordinates of the initial tongue body region, the ratio of its perimeter (e.g. 800 pixels) to area (e.g. 30000 square pixels) is calculated to obtain a shape value of 800 / 30000 ≈ 0.0267; this shape value is serialized and spliced with the area proportions and average color values of the 5 pixel categories to establish a comprehensive feature vector, for example: [shape value, class 1 area, class 1 L, class 1 a, class 1 b, class 2 area, …, class 5 b].

[0148] Finally, the comprehensive feature vector is compared with the discrimination rule; the discrimination rule is pre-set in the database, for example, the numerical value range threshold corresponding to 'pale white tongue thin white fur' is: shape value [0.02, 0.03], main fur color category (assuming class 1) area proportion [0.7, 0.9], main fur color L value, a value [-5, 5], b value, etc.; when all the values of the comprehensive feature vector fall within the pre-set threshold range of a certain tongue appearance category, it is determined that the matching is successful.

[0149] After matching is successful, the confidence score is calculated by the formula:

[0150]

[0151] The formula is used to quantify the fitting degree of the feature value v j in its pre-set valid interval The closer the feature value is to the midpoint of the interval, the higher its contribution value to the total confidence score, making the score result more discriminative. Example: Suppose the weight of the shape value w1 = 0.2, and its range , actual value v1 = 0.0267; class 1 area ratio weight w2 = 0.3, the range , actual value v2 = 0.75.

[0152] Calculate the score item of the first feature:

[0153] Center point = (0.03 + 0.02) / 2 = 0.025;

[0154] Halfway = (0.03 - 0.02) / 2 = 0.005;

[0155] Score item ;

[0156] Calculate the score item of the second feature:

[0157] Center point = (0.9 + 0.7) / 2 = 0.8;

[0158] Halfway = (0.9 - 0.7) / 2 = 0.1;

[0159] Score item ;

[0160] Add the score items of all M dimensions to get the total confidence score S conf . For example, the final calculation score is 0.89. The result shows that the confidence of the current tongue appearance being identified as "pale white tongue thin white fur" is 89%. Finally, the system outputs the tongue appearance category identification "pale white tongue thin white fur" and the confidence score 0.89, together generating the tongue classification analysis result.

[0161] A system based on tongue diagnosis image recognition, the system based on tongue diagnosis image recognition is used to execute the above-mentioned method based on tongue diagnosis image recognition, the system comprises:

[0162] The reference construction module is used to generate the mapping relationship between the red, green and blue three-channel numerical values of the standard color chart image generated by the standard color chart irradiated by the monochromatic light source and the white light source and the pre-stored standard color numerical values, generate a color correction matrix, and based on the initial posture reading collected, construct an initial spatial coordinate system, and pass the color correction matrix and the initial spatial coordinate system to the tongue appearance collection and color correction module;

[0163] The tongue appearance collection and color correction module is used to capture tongue images and collect real-time posture readings, call the color correction matrix to transform the color values of the tongue images, generate a multi-spectral image sequence and a real-time posture reading set, and pass the multi-spectral image sequence, the real-time posture reading set and the initial spatial coordinate system to the posture correction and feature construction module;

[0164] The posture correction and feature construction module is configured to generate posture compensation transformation parameters according to deviations of a real-time posture reading set and an initial spatial coordinate system, screen an initial tongue body region from a white light tongue image of the multispectral image sequence, correct the initial tongue body region by calling the posture compensation transformation parameters, combine pixel brightness values to construct a pixel-level multispectral feature vector set, and deliver the pixel-level multispectral feature vector set to the comprehensive discriminant analysis module.

[0165] The comprehensive discriminant analysis module is configured to divide pixel categories by inputting the pixel-level multispectral feature vector set to a K-means clustering algorithm, calculate area proportions and average color values of the pixel categories, extract shape values of the initial tongue body region, fuse the shape values into a comprehensive feature vector, match the comprehensive feature vector with a discriminant rule, and output a tongue image classification analysis result.

[0166] The above embodiments demonstrate preferred embodiments of the present application, and any equivalent adjustment of the technical solutions based on a software engineering method falls within the protection scope, including but not limited to: technical improvements such as implementation of algorithm logic in different programming languages, service reconstruction of functional modules, adjustment of data interaction protocols, and optimization of resource scheduling strategies. Any implementation scheme derived by reasonable modification of a data processing flow, a service calling link, or a system architecture level without departing from the technical core of the present application should be considered within the protection scope defined by the claims of the present application.

Claims

1. A method of image recognition based on a tongue diagnosis, the method comprising: obtaining a tongue image; and identifying a tongue diagnosis based on the tongue image. The method comprises the following steps: S1: generating a mapping relationship between red, green and blue three-channel values of a standard color chart image generated by illuminating a standard color chart by a monochromatic light source and a white light source and pre-stored standard color values, generating a color correction matrix, and constructing an initial spatial coordinate system based on an initial attitude reading collected; S2: capturing a tongue image and collecting real-time attitude readings, calling the color correction matrix to transform color values of the tongue image, generating a multispectral image sequence and a real-time attitude reading set; S3: generating a perspective compensation matrix and attitude compensation transformation parameters according to deviations of the real-time attitude reading set and the initial spatial coordinate system, performing global perspective correction on the multispectral image sequence, then screening an initial tongue region from a white light tongue image, simultaneously calling the attitude compensation transformation parameters to correct the initial tongue region, calculating a second-order central moment of the initial tongue region, deriving a tongue main axis tilt angle, and constructing a rotation matrix to perform main axis alignment transformation on the initial tongue region, and combining pixel brightness values to construct a pixel-level multispectral feature vector set; S4: inputting the pixel-level multispectral feature vector set into a K-means clustering algorithm to divide pixel categories, calculating area proportions and average color values of the pixel categories, extracting shape values of the initial tongue region, and fusing them into a comprehensive feature vector, matching the comprehensive feature vector with a discrimination rule, and outputting a tongue image classification analysis result.

2. The method of identifying based on a tongue diagnosis image according to claim 1, wherein, The color correction matrix comprises red channel correction coefficients, green channel correction coefficients and blue channel correction coefficients, the initial spatial coordinate system specifically comprises three-axis rotation components and three-axis translation components, the multispectral image sequence comprises white light tongue image frames and monochromatic light tongue image frames, the real-time attitude reading set specifically refers to real-time readings of a gyroscope and real-time readings of an accelerometer, the attitude compensation transformation parameters comprise a rotation matrix and a translation vector, the pixel-level multispectral feature vector set comprises a pixel position index and a multispectral response vector, the comprehensive feature vector comprises an area proportion value, an average color value and a tongue shape value, and the tongue image classification analysis result specifically refers to a tongue image category identifier and a confidence score. 3.The tongue diagnosis image recognition-based method according to claim 2, characterized in that, The obtaining step of the initial spatial coordinate system specifically comprises: S101: obtaining a standard color chart, illuminating it by a monochromatic light source and a white light source respectively, synchronously collecting a standard color chart image, analyzing the standard color chart image, extracting pixel values of red, green and blue channels in the image frame, combining the extracted pixel values into a multidimensional array, and generating a standard color chart image pixel value matrix; S102: calling the standard color chart image pixel value matrix, obtaining pre-stored standard color values, constructing a linear transformation equation between multiple channel values in the standard color chart image pixel value matrix and color components of the standard color values, solving the linear transformation equation to obtain transformation coefficients, organizing the transformation coefficients into a three-order matrix form, and generating a color correction matrix; S103: Collect initial attitude reading of roll angle, pitch angle and yaw angle data from the attitude sensor, and define the roll angle, the pitch angle and the yaw angle data as the first, second and third rotation axes of the initial spatial coordinate system respectively according to the initial attitude reading, and take it as the reference to construct the initial spatial coordinate system. 4.The tongue diagnosis image recognition-based method according to claim 3, characterized in that, The acquisition step of the multispectral image sequence and the real-time attitude reading set is specifically: S201: Monitor the image sensor to capture a single frame of tongue image, and synchronously monitor the attitude sensor to collect real-time attitude reading. The color information of a plurality of pixel points in the tongue image is decomposed into red, green and blue three-channel intensity values to generate an original tongue RGB image. The original data of the gyroscope and the accelerometer in the real-time attitude reading are integrated to establish an initial attitude parameter set; S202: Obtain a preset color correction matrix, construct a pixel color vector according to the red, green and blue three-channel intensity values of each pixel position in the original tongue RGB image, perform matrix multiplication operation on the color correction matrix and the pixel color vector, replace the red, green and blue three-channel intensity values with the generated transformation vector to obtain a corrected spectral data frame; S203: For a plurality of corrected spectral data frames and the initial attitude parameter set acquired in a continuous sampling period, stack a plurality of corrected spectral data frames along the time axis dimension to construct a multispectral image sequence, and pair the acquisition time stamp corresponding to the initial attitude parameter set with the image frame index of the multispectral image sequence to generate the multispectral image sequence and the real-time attitude reading set.

5. The method of identifying based on a tongue diagnosis image according to claim 4, wherein, The acquisition step of the pixel-level multispectral feature vector set is specifically: S301: Calculate the homography matrix based on the deviation of the real-time attitude reading set and the initial spatial coordinate system, eliminate the perspective distortion caused by the shooting angle tilt, calculate the zeroth moment M00 and the first moment M10, M01 of the initial tongue region to determine the centroid coordinates, and calculate the deflection angle a of the tongue main axis relative to the vertical direction combined with the second central moments μ20, μ11, μ02 to generate the attitude compensation transformation parameters including rotation and translation components; S302: Call the white light tongue image of the multispectral image sequence, set the pixel chroma range and the brightness lower threshold, traverse all pixels of the white light tongue image, remove background pixels that do not meet the pixel chroma range and the brightness lower threshold, and screen to obtain an initial tongue region. Call the attitude compensation transformation parameters to perform inverse geometric transformation on the pixel coordinates of the initial tongue region to obtain a corrected tongue region pixel set; S303: Traverse the pixel positions recorded in the corrected tongue region pixel set, extract pixel brightness values from N spectral channel images of the multispectral image sequence one by one according to the pixel positions, arrange the extracted N pixel brightness values in the order of wavelength from short to long to form an N-dimensional feature vector, aggregate all the N-dimensional feature vectors of the pixels to construct a pixel-level multispectral feature vector set.

6. The method of identifying based on a tongue diagnosis image according to claim 5, wherein, The acquisition step of the tongue image classification analysis result is specifically: S401: For the pixel-level multi-spectral feature vector set, iteratively divide according to the Euclidean distance between vectors until the cluster center is stable, form a plurality of pixel categories, and count the proportion of the number of pixels in a plurality of the pixel categories to the total number of pixels in the initial tongue body region, calculate the area proportion, and simultaneously perform an arithmetic average operation on the color values of all pixels in a plurality of the pixel categories to obtain the average color value, and finally combine the area proportion and the average color value of all the pixel categories to generate a pixel category attribute set; S402: According to the contour coordinates of the initial tongue body region, calculate the ratio of its circumference to area as a shape value, call the pixel category attribute set, and serialize the shape value, the area proportion and the average color value in the pixel category attribute set to establish a comprehensive feature vector; S403: Compare the area proportion, the average color value and the shape value of the comprehensive feature vector with the numerical range threshold preset in the discrimination rule for the differentiated tongue image category, and when all the values of the comprehensive feature vector fall within the preset threshold range of a tongue image category, it is determined that the matching is successful, and the tongue image classification analysis result is obtained. 7.The tongue diagnosis image recognition-based method according to claim 5, characterized in that, In the screening step of the initial tongue body region, the setting method of the pixel chroma range and the brightness lower limit threshold is specifically: Obtain the white light tongue image, calculate the three-channel pixel value distribution histogram of the white light tongue image in the hue, saturation and brightness color space, generate a hue distribution graph, a saturation distribution graph and a brightness distribution graph; Analyze the brightness distribution graph, identify the main peak and the secondary peak in the brightness distribution graph using a peak detection algorithm, define the brightness value corresponding to the main peak as the background brightness center, and define the brightness value corresponding to the secondary peak as the target brightness center; According to the target brightness center and the preset brightness standard deviation multiple, the brightness lower limit threshold is calculated, and the pixels with brightness values in the neighborhood of the target brightness center are analyzed, and the distribution interval of the pixels in the hue distribution graph and the saturation distribution graph is determined, and the pixel distribution interval is taken as the pixel chroma range; Based on the calculated brightness lower limit threshold and the pixel chroma range, each pixel point in the white light tongue image is traversed and judged, and the pixel points meeting the conditions are screened to generate the initial tongue body region; The preset brightness standard deviation multiple is an empirical value determined by statistical analysis of a plurality of sample tongue images, and the value range is 2 to 4. 8.The tongue diagnosis image recognition-based method according to claim 6, characterized in that, In the step of dividing the pixel categories by the K-means clustering algorithm, the judgment condition for the stable cluster center is specifically: Randomly select K multi-spectral response vectors from the pixel-level multi-spectral feature vector set as initial cluster centers; For each multi-spectral response vector in the pixel-level multi-spectral feature vector set, calculate the Euclidean distance from the plurality of cluster centers, and assign the multi-spectral response vector to the pixel category corresponding to the nearest cluster center; recalculating the mean of all the multi-spectral response vectors in each of the pixel classes and updating the mean as the current cluster center; repeating the assigning and updating steps and determining whether the clustering process converges by the following formula: , wherein, AC represents the average variation of all cluster centers between two successive iterations, K represents the preset number of pixel categories, k is the pixel category index from 1 to K, represents the cluster center vector of the kth pixel category at the tth iteration, represents the cluster center vector of the pixel category at the previous iteration, i.e., the (t-1)th iteration, represents the Euclidean norm of the vector, represents the preset cluster center convergence threshold; The preset number of pixel classes K is determined according to the common classification number of the quality and color of the tongue coating in the tongue diagnosis theory, and the value range is 3 to 7. When the AC is less than or equal to the convergence of the clustering process is determined, generating a number of the pixel classes. 9.The tongue diagnosis image recognition-based method according to claim 6, characterized in that, The calculation step of the confidence score in the tongue classification analysis result is specifically: After the comprehensive feature vector matches a tongue class successfully, a preset numerical range threshold of the tongue class is obtained, including the upper and lower limits of each feature; The area ratio value, the average color value and the tongue shape value in the comprehensive feature vector are called, and the confidence score is calculated by the following formula: , wherein S conf represents the confidence score, M represents the total number of dimensions of the comprehensive feature vector, j is a feature index from 1 to M, w j represents a preset weight coefficient of the jth feature and the sum of all weight coefficients is 1, v j represents an actual value of the jth dimension in the comprehensive feature vector, U j represents an upper limit of a preset value range of the jth feature matched with the tongue appearance category, L j represents a lower limit of a preset value range of the feature, and δ represents a smoothing factor set to avoid a zero denominator. The preset weight coefficient w of the jth feature j is determined by expert scoring method or machine learning model training according to the contribution degree of the differentiated feature to the tongue appearance classification result. The calculated confidence score and the tongue class identifier of the tongue class are combined to generate the tongue classification analysis result together.

10. A system for image recognition based on a tongue diagnosis, characterized by, The system is used to implement the tongue diagnosis image recognition method of any one of claims 1-9, and the system comprises: A reference construction module is configured to generate a color correction matrix by mapping the red, green and blue three-channel values of a standard color chart image generated by illuminating a standard color chart with a monochromatic light source and a white light source with a pre-stored standard color value, and to construct an initial spatial coordinate system based on an initial posture reading, and to transfer the color correction matrix and the initial spatial coordinate system to a tongue image acquisition and color correction module; A tongue image acquisition and color correction module is configured to capture a tongue image and acquire a real-time posture reading, call the color correction matrix to transform the color value of the tongue image, generate a multi-spectral image sequence and a real-time posture reading set, and transfer the multi-spectral image sequence, the real-time posture reading set and the initial spatial coordinate system to a posture correction and feature construction module; A posture correction and feature construction module is configured to generate a posture compensation transformation parameter according to the deviation of the real-time posture reading set and the initial spatial coordinate system, select an initial tongue region from the white light tongue image of the multi-spectral image sequence, correct the initial tongue region by calling the posture compensation transformation parameter, combine the pixel brightness value to construct a pixel-level multi-spectral feature vector set, and transfer it to a comprehensive discriminant analysis module; A comprehensive discriminant analysis module is configured to input the pixel-level multi-spectral feature vector set into a K-means clustering algorithm to divide pixel classes, calculate the area ratio and average color value of the pixel classes, extract the shape value of the initial tongue region, and fuse it into a comprehensive feature vector, match the comprehensive feature vector with a discriminant rule, and output a tongue classification analysis result.

Citation Information

Patent Citations

  • Traditional Chinese medicine tongue picture feature extraction method and system based on depth model

    CN119068486A

  • Visual data processing system applied to pediatric department of traditional Chinese medicine

    CN120260960A