Tongue image feature extraction and health assessment method based on deep learning
By setting the preset standards for mouthwash and presetting light and support, and pre-processing and analyzing the tongue image data in combination with Gaussian pyramid and optical flow method, the problem of inaccurate data in traditional tongue image acquisition technology is solved, and more accurate tongue image feature extraction and health assessment are achieved.
Patent Information
- Application Number
- CN202510065404.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Traditional tongue image collection technology has problems such as insufficient oral cleanliness, uneven light, and unstable head, resulting in inaccurate tongue image data, and the analysis of tongue movement characteristics is not accurate and comprehensive enough.
By setting the preset standards for mouthwash, presetting light and support, dynamic data of tongue images are collected. The Gaussian pyramid, DoG pyramid and detection of local extreme points are used to preprocess the dynamic video, extract the tongue image features, and calculate the tongue motion vector by optical flow method.
The accuracy and stability of tongue image data are improved, the tongue image features and tongue movement parameters are accurately extracted, and a more comprehensive and detailed health assessment system is established, which can more accurately reflect the functional status and pathological changes of the tongue.
Smart Images

Figure CN119453950B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to health assessment technology, and specifically to a tongue image feature extraction and health assessment method based on deep learning. Background Art
[0002] Tongue image refers to a simple and effective method to assist in the diagnosis and identification of diseases by observing the changes in the color, shape, and tongue coating of the tongue. In traditional Chinese medicine, tongue diagnosis is an important part of visual diagnosis, and by observing the tongue image, one can understand the pathological changes in the body.
[0003] Traditional tongue image collection focuses more on static tongue images, and the collection of dynamic changes is not comprehensive and standardized. Dynamic tongue images can capture the dynamic information of tongue changes over time and can capture more comprehensive information. Dynamic tongue images can more accurately reflect the functional status and pathological changes of the tongue by observing the dynamic changes of the tongue, which helps to improve the accuracy of TCM diagnosis.
[0004] The lack of effective methods to judge the degree of oral hygiene of patients leads to unclean oral cavity, which affects the correct judgment of tongue coating color and shape; uneven lighting makes some areas in the image too bright or too dark, thus affecting the accurate recognition and analysis of tongue features; head instability can cause the captured tongue image to be blurred, interfering with the recognition and analysis of tongue features;
[0005] The analysis of tongue movement characteristics is not accurate and comprehensive enough. There is a lack of systematic methods to accurately track the tongue movement trajectory, and it is difficult to obtain detailed movement parameters. The health assessment method based on tongue image is relatively simple and subjective, and lacks in-depth exploration of the complex relationship between tongue movement characteristics and health status.
[0006] In view of the above technical problems, this application proposes a solution. Summary of the invention
[0007] The purpose of the present invention is to effectively solve the problem of inaccurate tongue image data caused by incomplete oral cleaning and unstable collection environment during the collection process by setting preset standards for mouthwash, lighting and support; constructing Gaussian pyramids, DoG pyramids and detecting local extreme points to pre-process dynamic videos, improve image stability and feature point accuracy, and compare with demonstration videos to quantitatively analyze the tongue motion vector in terms of quantity, direction, tremor amplitude and offset direction, calculate health factors and disease possibilities, and solve the problems of incomplete capture of traditional static tongue image information, difficulty in normal judgment of tongue image due to oral cleanliness, lighting influence and unstable support, and inability to accurately track the tongue movement trajectory and obtain detailed parameters, and propose a tongue image feature extraction and health assessment method based on deep learning.
[0008] The purpose of the present invention can be achieved through the following technical solutions:
[0009] The tongue image feature extraction and health assessment method based on deep learning includes the following steps:
[0010] PG1. Tongue dynamic data collection: Before collection, the patient rinses his mouth until the mouthwash meets the preset standard; the head fixation bracket and lighting equipment are preset according to the patient's head size and oral light intensity data, and then the dynamic changes of the tongue in a natural state are collected using high-speed camera equipment; dynamic changes include extension, movement and tremor;
[0011] PG2. Tongue feature extraction and analysis: The patient's tongue image is collected by a high-speed camera, and the collected dynamic video data is preprocessed and cropped to a uniform size. The dynamic video is then divided into image data according to the number of frames, and the image data is denoised. The tongue features are then extracted and analyzed.
[0012] PG3. Health assessment based on tongue feature analysis: The analysis results include tremor amplitude and deviation direction, and the patient's health status is assessed based on the tremor amplitude and deviation direction.
[0013] As a preferred embodiment of the present invention, the preset standards of the mouthwash in PG1 are as follows:
[0014] Step 1: Before the patient performs the dynamic tongue image collection, the patient is instructed to rinse the mouth with a set amount of mild water. The initial volume and initial weight of the water are and During the mouthwash process, the number of cheek movements on both sides is recorded, and the detection distance between the cheek and the distance detector increases from small to large and then decreases as one movement. After the number of movements reaches the preset number, the patient is prompted to spit out the water in the mouth, and the volume of the spitted water is and weight Take measurements;
[0015] Step 2: Compare the volume and weight of the spitted water with the initial volume and weight of the water. , it is judged that the oral cavity is clean and the tongue image collection operation can be performed.
[0016] As a preferred implementation of the present invention, the dynamic video preprocessing steps in PG2 are as follows:
[0017] Step 1: Divide the dynamic video into image data according to each frame, and use the original image of one frame of image data as the first layer of the Gaussian pyramid. The layer image is processed by Gaussian filtering using different smoothing coefficients, and then the Gaussian filtered image is double-sampled to generate the first layer of the Gaussian pyramid. Layer image;
[0018] Step 2: Repeat the operation of step 1 until the number of layers of the Gaussian pyramid reaches the preset number of layers, and then subtract the first layer from the second layer of the Gaussian pyramid to get the first layer of the DoG pyramid. The layer image is composed of the Gaussian pyramid Layer minus the The local extreme points are detected in each layer of the DoG pyramid, that is, for each pixel point in the DoG pyramid, it is compared with the eight adjacent points of the same scale and the 9×2 points corresponding to the upper and lower adjacent scales, a total of 26 points. If the point is the maximum or minimum value among these 26 points, these extreme points are recorded as feature points.
[0019] Step 3: For two adjacent frames in the dynamic video, match the feature points detected in the previous frame with the feature points in the next frame, calculate the Euclidean distance, and select the minimum Euclidean distance. As matching point pairs, perform affine transformation on the matching point pairs to obtain the transformed pixel coordinates , calculate the weighted average of the coordinate and the four surrounding integer coordinate points to get the coordinate The pixel value of
[0020] Step 4: Preset video Frame, then The translation of the frame relative to the first frame is , is the horizontal translation amount, and the default window size is , then the trajectory of the horizontal translation Smoothed trajectory , is the dimension of the space, Indicates that in calculating The frame number involved in the smoothing of the trajectory of the frame horizontal translation amount is used to complete the smoothing of the changing trajectory.
[0021] As a preferred implementation manner of the present invention, the noise processing steps in PG2 are as follows:
[0022] Step 1: Image Perform a 2D Fourier transform , calculate the proportion of high-frequency area energy to total energy , preset frequency threshold , the frequency area above the threshold is regarded as the high-frequency area, , if the preset ratio threshold , it is determined that there is noise, and the image data is denoised; otherwise, the tongue image data is feature extracted;
[0023] Step 2: Tongue Image ,in M-1, N-1; calculate the mean of the image , and then calculate the image standard deviation based on the calculated image mean , preset two noise standard deviation thresholds and , ;like , set the filter window size to ;like , set the filter window size to ;like , set the filter window size to ;
[0024] Step 3: For each pixel in the tongue image , starting from the upper left corner of the tongue image, check each pixel in the image one by one in the order from left to right and from top to bottom. For each pixel in the tongue image, a filtering window of the corresponding window size is established with itself as the center, and the grayscale values of all pixels in the window are obtained. The grayscale values of all pixels in the window are added and then divided by the total number of pixels in the window. The average value is taken as the new grayscale value of the pixel in the center of the filtering window; the same processing is performed on each pixel in the tongue image to achieve the denoising effect.
[0025] As a preferred embodiment of the present invention, the steps of extracting tongue features in PG2 are as follows:
[0026] Step 1: Select four feature points on the tongue surface. At the moment The pixel brightness is recorded as , then at time Time, point The pixel brightness is and For point exist Displacement data within time; two adjacent frames of images, the current frame is preset as , the next frame is , then the time gradient , in actual calculation, Usually 1, time gradient Simplified to the pixel value difference between two adjacent frames;
[0027] Step 2: Calculate the image in the horizontal direction and vertical direction The spatial gradient of and ,right exist Perform Taylor expansion at and ignore higher-order terms, and Set to 1, , , we get the optical flow equation , and The required motion vector is and Directional weight;
[0028] Step 3: Under the spatial consistency assumption, all pixels in the window share the same motion vector , for a window, we can get equation, and the motion vector is obtained by solving it using the least squares method ; For each feature point in the tongue image, repeat the above steps to calculate the motion vector corresponding to each feature point, and combine the motion vectors of all pixels to obtain the optical flow field of the entire image;
[0029] Step 4: For each pixel in the optical flow field, Frame and The motion vector of the frame is interpolated to obtain its rate of change over time. and ; for pixel points Its neighboring pixels The motion vector difference is compared to calculate the pixel point Spatial rate of change in horizontal and vertical directions and ; Substitute the calculated change rate into the weight ratio formula to obtain the factor for determining the tongue movement status , , , and is the weight coefficient; if the preset condition threshold , then the rate of change of the surface motion vector is too large, and the tongue movement is judged to be irregular.
[0030] As a preferred embodiment of the present invention, the tongue image characteristics are analyzed, and the analysis steps are as follows:
[0031] Step 1: Divide the tongue motion vectors in the demonstration video into the same intervals, and count the number and direction to obtain the number of healthy standards in each interval. and the average direction ; For the motion vectors that are determined to be moving irregularly, the direction of the motion vector is divided into several intervals of equal size according to the set angle size, and the number of motion vectors in each interval is counted and the average direction , is the interval number;
[0032] Step 2: Calculate the number difference and average direction difference of the motion vectors in each interval gate. , the average direction difference , and compare the calculated quantity difference value and average direction difference value with the preset quantity threshold and preset direction threshold, and record the number of corresponding items exceeding the threshold ;
[0033] Step 3: If , it indicates that the number and direction deviation of motion vectors in the corresponding interval are small. is the preset ratio; substitute the quantity difference value and the average direction difference value into the weight formula to obtain the health factor , and is the weight coefficient. If the health threshold is preset , the patient is judged to have potential health problems and the number and movement direction of the tongue are analyzed; otherwise, the patient is judged to be in good health and has no problems; if , it indicates that the number and direction deviation of the motion vectors in the corresponding interval are large, and it is determined that the patient has obvious health problems, and the number and movement direction of the tongue are analyzed;
[0034] Step 4: For each frame of the patient and demonstration video in the corresponding time period, count the number of motion vectors in each direction interval and spatial position interval. Suppose the number of motion vectors in a certain direction interval and spatial position interval in the patient video is , the number of corresponding intervals in the demonstration video is , is the interval number, Number the spatial position interval; perform difference calculation on the two obtained quantity data to obtain the difference in the number of motion vectors in each interval , the difference in the number of motion vectors calculated for each interval Compare, if the corresponding interval , it indicates that the patient's tongue tremors in that direction and position.
[0035] As a preferred embodiment of the present invention, the steps of assessing the health status of the patient are as follows:
[0036] Step 1: If the corresponding interval and , it means that the patient's tongue is trembling and deviating in direction, and it is judged that the patient may have a disease related to the nervous system. It is recommended that the patient go to the hospital for examination. The possibility of disease , and is the weight coefficient;
[0037] Step 2: If the preset disease determination threshold is 1 , it means that the patient's tongue tremor amplitude and direction deviation are small, and it is judged that there may be early Parkinson's disease or mild neurological dysfunction. The patient is recommended to go to the neurology department for further examination; if the preset disease judgment threshold is two There are two possibilities. One is that the patient's tremor amplitude is large but the number of directional deviations is small. It is judged that the patient may have a local muscle disease and a change in the excitability of the nervous system. The patient is recommended to go to the stomatology department for further examination of the tongue muscles. The other is that the patient's tremor amplitude is small but the number of directional deviations is large. It is judged that the patient may have a neurodegenerative disease of the multiple system atrophy. The patient is recommended to go to the neurology department of the hospital for further examination. If the preset disease judgment threshold is two , indicating that the patient's tremor has a large amplitude and many directional deviations, indicating that the patient may have early manifestations of neurological dysfunction. It is recommended that the patient go to the neurology department for further examination.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] 1. By setting the preset standard of mouthwash, the volume and weight changes before and after rinsing with water are accurately measured to judge the degree of oral cleanliness, ensuring that the collected tongue image can truly reflect the physical condition; at the same time, the head fixing bracket and lighting equipment are accurately preset using distance sensors, pressure sensors and light intensity sampling points to ensure that the patient's head is stable and the tongue is evenly illuminated during the collection process, effectively solving the problem of inaccurate tongue image data caused by incomplete oral cleaning and unstable collection environment during the collection process;
[0040] 2. In the feature extraction process, Gaussian pyramids and DoG pyramids are constructed and local extreme points are detected to pre-process dynamic videos to improve image stability and feature point accuracy; the tongue motion vector is calculated by combining the optical flow method with time and space gradients, and its rate of change is further analyzed. At the same time, image noise is systematically detected, the noise situation is judged according to the energy proportion of the high-frequency area, and a suitable filter window is selected for denoising, so as to more accurately extract tongue image features, overcoming the problem of inaccurate analysis of tongue motion features and susceptibility to noise interference in the prior art;
[0041] 3. A more comprehensive and detailed health assessment system has been established. By comparing with the demonstration video, the tongue motion vector is quantitatively analyzed in terms of quantity, direction, tremor amplitude and offset direction, and the health factor and disease possibility are calculated. The health status is judged according to different assessment index thresholds, and detailed department examination recommendations are given for different situations. This has achieved a transformation from qualitative to quantitative, and from single feature to multi-feature comprehensive health assessment, solving the problems of inaccuracy and lack of specificity of traditional health assessment methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0043] Figure 1 It is a structural diagram of the process method of the present invention;
[0044] Figure 2 It is a structural diagram of tongue image collection of the present invention. DETAILED DESCRIPTION
[0045] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] Example:
[0047] See also Figure 1-2 As shown in the figure, the tongue image feature extraction and health assessment method based on deep learning includes:
[0048] Mouthwash preset standard: Before the dynamic collection of tongue images, the patient is instructed to rinse the mouth with a set amount of mild water. The initial volume and initial weight of the water are and During the mouthwash process, the number of cheek movements on both sides is recorded, and the detection distance between the cheek and the distance detector increases from small to large and then decreases as one movement. After the number of movements reaches the preset number, the patient is prompted to spit out the water in the mouth, and the volume of the spitted water is and weight Measure and compare the volume and weight of the spitted water with the initial volume and weight of the water. , then the oral cavity is judged to be clean and the tongue image collection operation can be performed;
[0049] Head fixation bracket preset: the patient places his head on the inside of the head fixation bracket, and the distance sensor on the inside of the head fixation bracket detects the distance between the head and the patient's head, and then clamps the patient's head according to the detected distance data. During the clamping process, the pressure value between the patient's head and the clamping component is detected by the pressure sensor on the clamping component. When the pressure value reaches the preset pressure value, the clamping operation of the clamping component is stopped, and the distance data between the clamping components on both sides and the head fixation bracket are compared. If the distance data between the clamping components on both sides and the fixed bracket are equal, it is determined that the patient's head is in the center of the head fixation frame. Otherwise, it is determined that the side with the larger distance data between the clamping component and the fixed bracket is the left or right side, and the voice prompts the patient to adjust the position of the head to the left or right, and the distance data between the clamping components on both sides and the fixed bracket are equal;
[0050] Lighting device presets: Figure 2 As shown in A, after the head portrait is centered, the camera is aimed at the tongue in the patient's mouth, and the two ends of the tongue root that are farthest apart are connected once, and the midpoint of the connection line is connected twice with the tip of the tongue. A square figure is drawn with the midpoint of the secondary connection line as the center of the square, and a third connection line is made at the midpoint of the opposite side of the square. The four corner points of the square, the midpoints of the four sides and the midpoints of the third connection line are used as light intensity sampling points. The light intensity data of the light intensity sampling points are collected by a light meter, and the average light intensity data of the nine light intensity sampling points is calculated. , and the light intensity data corresponding to the light intensity sampling point The mean value of light intensity data For comparison, if , it is determined that the illumination on the tongue surface is uneven, and the illumination intensity of the illumination device is adjusted;
[0051] Tongue image acquisition: Fix the high-speed camera in the patient's oral cavity, make the center of the high-speed camera lens and the center of the patient's tongue root on the same horizontal line, and adjust the shooting ratio of the camera lens so that the tongue occupies the set proportion in the picture. , and then a dynamic video of a preset time length is shot by a high-speed camera. During the dynamic video shooting process, the display screen connected to the high-speed camera plays the set screen according to the set time after the dynamic video starts recording, prompting the patient to perform tongue extension, lip licking and tremor movements according to the demonstration video playback screen during the dynamic video recording process. After completing the dynamic video shooting operation, a folder is created to store the video file compressed in a lossless manner, and an explanation document is created in the folder corresponding to the video file, and the explanation document contains the patient's basic information and the collected basic information;
[0052] Dynamic video preprocessing: Reduce the impact of slight head movement of the patient or device shaking during the acquisition process, stabilize the dynamic video, and perform the following processing on each frame of the dynamic video: use the original image as the first layer of the Gaussian pyramid, The layer image is processed by Gaussian filtering using different smoothing coefficients, and then the Gaussian filtered image is double-sampled to generate the first layer of the Gaussian pyramid. The Gaussian filtering and double sampling operations are repeated until the number of layers of the Gaussian pyramid reaches the preset number of layers; then the first layer of the DoG pyramid is obtained by subtracting the first layer from the second layer of the Gaussian pyramid. The layer image is composed of the first Layer minus the The local extreme points are detected in each layer of the DoG pyramid, that is, for each pixel in the DoG pyramid, it is compared with eight adjacent points of the same scale and 9×2 points corresponding to the upper and lower adjacent scales, a total of 26 points. If the point is the maximum or minimum value among these 26 points, these extreme points are recorded as feature points.
[0053] For two adjacent frames in a dynamic video, the feature points detected in the previous frame are matched with the feature points in the next frame. The feature point descriptor vectors of the two adjacent frames are preset as follows: , , the Euclidean distance between two feature points corresponding to two adjacent frames of images , For the number of dimensions of the space, select the minimum Euclidean distance As matching point pairs, perform affine transformation on the matching point pairs to obtain the transformed pixel coordinates , calculate the weighted average of the coordinate and the four surrounding integer coordinate points to get the coordinate The pixel value of the preset video is Frame, then The translation of the frame relative to the first frame is , is the horizontal translation amount, and the default window size is , then the trajectory of the horizontal translation Smoothed trajectory , Indicates that in calculating The frame number involved in the smoothing trajectory of the frame horizontal translation. The smoothed transformation trajectory can make the video movement more natural and reduce the instability caused by local jitter.
[0054] Each frame of the dynamic video is cropped according to the tongue shape, and then the width data at both ends of the cropped image data are measured, and the mean of the width data is calculated. Taking the mean of the width data as the standard, the image width data is divided by the mean of the width data to obtain the scaling ratio, and the other image data is scaled to the same size as the mean width data image according to the scaling ratio of the corresponding image;
[0055] Image noise detection: Perform a 2D Fourier transform , , , , and are the number of rows and columns of the image, respectively. and is the frequency domain coordinate; under normal circumstances, the spectrum energy of the tongue image is mainly concentrated in the low-frequency area. This is because most of the information in the tongue image changes relatively slowly and belongs to the low-frequency component. If there is noise in the image, especially high-frequency noise, it will cause the spectrum energy to be abnormally distributed in the high-frequency area; calculate the proportion of high-frequency area energy to total energy , preset frequency threshold , the frequency area above the threshold is regarded as the high-frequency area, , if the preset ratio threshold , it is determined that there is noise, and the image data is denoised; otherwise, the tongue image data is feature extracted;
[0056] Tongue denoising: tongue image ,in M-1, N-1; calculate the mean of the image , and then calculate the image standard deviation based on the calculated image mean In noisy tongue images, noise causes large fluctuations in pixel values, resulting in increased variance and standard deviation. Two noise standard deviation thresholds are preset. and , ;like , set the filter window size to ;like , set the filter window size to ;like , set the filter window size to ; For each pixel in the tongue image , starting from the upper left corner of the tongue image, in the order from left to right and from top to bottom, check each pixel in the image one by one. For each pixel in the tongue image, a filter window of corresponding window size is established with itself as the center, and the grayscale values of all pixels in the window are obtained. The grayscale values of all pixels in the window are added, and then divided by the total number of pixels in the window. The average value is taken as the new grayscale value of the center pixel of the filter window; each pixel in the tongue image is processed in the same way to achieve the effect of denoising;
[0057] Tongue feature extraction: Figure 2 As shown in B, the tip of the tongue is selected as a feature point in the tongue image, and a line is drawn between the tip of the tongue feature point and the end point of the tongue root. A vertical line is drawn at the midpoint of the line. The point where the vertical line intersects the tongue boundary is marked as a feature point. The intersection point of the vertical lines on both sides at the center of the tongue is also marked as a feature point. At the moment The pixel brightness is recorded as , then at time Time, point The pixel brightness is and For point exist Displacement data within time; two adjacent frames of images, the current frame is preset as , the next frame is , then the time gradient , in actual calculation, Usually 1, time gradient Simplified to the pixel value difference between two adjacent frames, calculate the image in the horizontal direction and vertical direction The spatial gradient of and , according to the brightness constant assumption, exist Perform Taylor expansion at and ignore higher-order terms, and Set to 1, , , we get the optical flow equation , and The required motion vector is and The optical flow equation can be listed for each pixel in the selected neighborhood window. Under the spatial consistency assumption, all pixels in the window share the same motion vector , for a window, we can get equation, and the motion vector is obtained by solving it using the least squares method ; For each feature point in the tongue image, repeat the above steps to calculate the motion vector corresponding to each feature point, and combine the motion vectors of all pixels to obtain the optical flow field of the entire image;
[0058] For each pixel in the optical flow field, Frame and The motion vector of the frame is interpolated to obtain its rate of change over time. and ; for pixel points Its neighboring pixels The motion vector difference is compared to calculate the pixel point Spatial rate of change in horizontal and vertical directions and ; Substitute the calculated change rate into the weight ratio formula to obtain the factor for determining the tongue movement status , , , and is the weight coefficient; if the preset condition threshold , then the rate of change of the surface motion vector is too large, and the tongue movement is judged to be irregular;
[0059] The tongue motion vectors in the demonstration video are divided into the same intervals, and the number and direction are counted to obtain the number of healthy standards in each interval. and the average direction ; For the motion vectors that are determined to be moving irregularly, the direction of the motion vector is divided into several intervals of equal size according to the set angle size, and the number of motion vectors in each interval is counted and the average direction , is the interval number; then calculate the number difference and average direction difference of motion vectors in each interval, the number difference value , the average direction difference , and compare the calculated quantity difference value and average direction difference value with the preset quantity threshold and preset direction threshold, and record the number of corresponding items exceeding the threshold ,like , it indicates that the number and direction deviation of motion vectors in the corresponding interval are small. is the preset ratio; substitute the quantity difference value and the average direction difference value into the weight formula to obtain the health factor , and is the weight coefficient. If the health threshold is preset , the patient is judged to have potential health problems and the number and movement direction of the tongue are analyzed; otherwise, the patient is judged to be in good health and has no problems; if , it indicates that the number and direction deviation of the motion vectors in the corresponding interval are large, and it is determined that the patient has obvious health problems, and the number and movement direction of the tongue are analyzed;
[0060] In the demonstration video, the corresponding stretching and lip licking actions are two time periods. For each frame of the patient and the demonstration video in the corresponding time period, the number of motion vectors in each direction interval and spatial position interval is counted. Suppose the number of motion vectors in a certain direction interval and spatial position interval in the patient video is , the number of corresponding intervals in the demonstration video is , is the interval number, Number the spatial position interval; perform difference calculation on the two obtained quantity data to obtain the difference in the number of motion vectors in each interval , the difference in the number of motion vectors calculated for each interval Compare, if the corresponding interval , it means that the patient's tongue trembles in this direction and position, and the number of tremors is The tremor velocity data is calculated by the tremor times and the time corresponding to each frame of the image; the average direction of the motion vector in each interval is calculated, and the average direction of the motion vector in the demonstration video is calculated, and the difference between the two average direction data is calculated to obtain the motion vector direction difference value , if the corresponding interval , it means that the patient's tongue movement in this direction deviates from the normal direction, and the number of times it deviates from the normal direction is ;
[0061] If the corresponding interval and , it means that the patient's tongue is trembling and deviating in direction, and it is judged that the patient may have a disease related to the nervous system. It is recommended that the patient go to the hospital for examination. The possibility of disease , and is the weight coefficient. If the disease judgment threshold is set to , it means that the patient's tongue tremor amplitude and direction deviation are small, and it is judged that there may be early Parkinson's disease or mild neurological dysfunction. The patient is recommended to go to the neurology department for further examination; if the preset disease judgment threshold is two There are two possibilities. One is that the patient's tremor amplitude is large but the number of directional deviations is small. It is judged that the patient may have a local muscle disease and a change in the excitability of the nervous system. The patient is recommended to go to the stomatology department for further examination of the tongue muscles. The other is that the patient's tremor amplitude is small but the number of directional deviations is large. It is judged that the patient may have a neurodegenerative disease of the multiple system atrophy. The patient is recommended to go to the neurology department of the hospital for further examination. If the preset disease judgment threshold is two , indicating that the patient's tremor amplitude is large and the number of directional deviations is high, indicating that the patient may be an early manifestation of neurological dysfunction. It is recommended that the patient go to the neurology department for further examination; if the corresponding interval is but There are two possibilities. One side indicates that the direction deviation is less frequent. The patient may have early nerve compression or mild nerve dysfunction. It is recommended that the patient go to the neurology department for a simple neurological function assessment. The other side indicates that the direction deviation is more frequent. The patient may have severe nerve damage or neurological disease. It is recommended that the patient go to the neurosurgery department for imaging examination to check whether there are structural lesions in the brain or nerves. If the corresponding interval is but There are two possibilities. One is that the tremor occurs less frequently and the patient may have a mild electrolyte disorder. It is recommended that the patient go to the laboratory for a blood biochemical test, including blood calcium, magnesium and other electrolyte level tests. The other is that the direction deviation occurs more frequently and the patient may have certain autoimmune diseases involving the nervous system. It is recommended that the patient go to the neurology department for a comprehensive neurological function test and screening for related diseases.
[0062] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A tongue image feature extraction and health assessment method based on deep learning, characterized in that: The following steps are involved: PG1. Tongue dynamic data collection: Before collection, the patient rinses his mouth until the mouthwash meets the preset standard; the head fixation bracket and lighting equipment are preset according to the patient's head size and oral light intensity data, and then the dynamic changes of the tongue in a natural state are collected using high-speed camera equipment; dynamic changes include extension, movement and tremor; PG2. Tongue feature extraction and analysis: The patient's tongue image is collected by a high-speed camera, and the collected dynamic video data is preprocessed and cropped to a uniform size. The dynamic video is then divided into image data according to the number of frames, and the image data is denoised. The tongue features are then extracted and analyzed. The steps of tongue image feature extraction are as follows; Step 1: Select four feature points on the tongue surface. At the moment The pixel brightness is recorded as , then at time Time, point The pixel brightness is and For point exist Displacement data within time; two adjacent frames of images, the current frame is preset as , the next frame is , then the time gradient , in actual calculation, Usually 1, time gradient Simplified to the pixel value difference between two adjacent frames; Step 2: Calculate the image in the horizontal direction and vertical direction The spatial gradient of and ,right exist Perform Taylor expansion at and ignore higher-order terms, and Set to 1, , , we get the optical flow equation , and The required motion vector is and Directional weight; Step 3: Under the spatial consistency assumption, all pixels in the window share the same motion vector , for a window, we can get equation, and the motion vector is obtained by solving it using the least squares method ; For each feature point in the tongue image, repeat the above steps to calculate the motion vector corresponding to each feature point, and combine the motion vectors of all pixels to obtain the optical flow field of the entire image; Step 4: For each pixel in the optical flow field, Frame and The motion vector of the frame is interpolated to obtain its rate of change over time. and ; for pixel points Its neighboring pixels The motion vector difference is compared to calculate the pixel point Spatial rate of change in horizontal and vertical directions and ; Substitute the calculated change rate into the weight ratio formula to obtain the factor for determining the tongue movement status , , , and is the weight coefficient; if the preset condition threshold , then the rate of change of the surface motion vector is too large, and the tongue movement is judged to be irregular; PG3. Health assessment based on tongue feature analysis: The analysis results include tremor amplitude and deviation direction, and the patient's health status is assessed based on the tremor amplitude and deviation direction.
2. The tongue image feature extraction and health assessment method based on deep learning according to claim 1 is characterized in that: The preset standards for mouthwash in PG1 are as follows: Step 1: Before the patient performs the dynamic tongue image collection, the patient is instructed to rinse the mouth with a set amount of mild water. The initial volume and initial weight of the water are and During the mouthwash process, the number of cheek movements on both sides is recorded, and the detection distance between the cheek and the distance detector increases from small to large and then decreases as one movement. After the number of movements reaches the preset number, the patient is prompted to spit out the water in the mouth, and the volume of the spitted water is and weight Take measurements; Step 2: Compare the volume and weight of the spitted water with the initial volume and weight of the water. , it is judged that the oral cavity is clean and the tongue image collection operation can be performed.
3. The tongue image feature extraction and health assessment method based on deep learning according to claim 1 is characterized in that: The steps of dynamic video preprocessing in PG2 are as follows: Step 1: Divide the dynamic video into image data according to each frame, and use the original image of one frame of image data as the first layer of the Gaussian pyramid. The layer image is processed by Gaussian filtering using different smoothing coefficients, and then the Gaussian filtered image is double-sampled to generate the first layer of the Gaussian pyramid. Layer image; Step 2: Repeat the operation of step 1 until the number of layers of the Gaussian pyramid reaches the preset number of layers, and then subtract the first layer from the second layer of the Gaussian pyramid to get the first layer of the DoG pyramid. The layer image is composed of the Gaussian pyramid Layer minus the The local extreme points are detected in each layer of the DoG pyramid, that is, for each pixel point in the DoG pyramid, it is compared with the eight adjacent points of the same scale and the 9×2 points corresponding to the upper and lower adjacent scales, a total of 26 points. If the point is the maximum or minimum value among these 26 points, these extreme points are recorded as feature points. Step 3: For two adjacent frames in the dynamic video, match the feature points detected in the previous frame with the feature points in the next frame, calculate the Euclidean distance, and select the minimum Euclidean distance. As matching point pairs, perform affine transformation on the matching point pairs to obtain the transformed pixel coordinates , calculate the weighted average of the coordinate and the four surrounding integer coordinate points to get the coordinate The pixel value of Step 4: Preset video Frame, then The translation of the frame relative to the first frame is , is the horizontal translation amount, and the default window size is , then the trajectory of the horizontal translation Smoothed trajectory , is the dimension of the space, Indicates that in calculating The frame number involved in the smoothing of the trajectory of the frame horizontal translation amount is used to complete the smoothing of the changing trajectory.
4. The tongue image feature extraction and health assessment method based on deep learning according to claim 1 is characterized in that: The noise processing steps in PG2 are as follows: Step 1: Image Perform a 2D Fourier transform , calculate the proportion of high-frequency area energy to total energy , preset frequency threshold , the frequency area above the threshold is regarded as the high-frequency area, , if the preset ratio threshold , it is determined that there is noise, and the image data is denoised; otherwise, the tongue image data is feature extracted; Step 2: Tongue Image ,in M-1, N-1 ; Calculate the mean of the image , and then calculate the image standard deviation based on the calculated image mean , preset two noise standard deviation thresholds and , ;like , set the filter window size to ;like , set the filter window size to ;like , set the filter window size to ; Step 3: For each pixel in the tongue image , starting from the upper left corner of the tongue image, check each pixel in the image one by one in the order from left to right and from top to bottom. For each pixel in the tongue image, a filtering window of the corresponding window size is established with itself as the center, and the grayscale values of all pixels in the window are obtained. The grayscale values of all pixels in the window are added and then divided by the total number of pixels in the window. The average value is taken as the new grayscale value of the pixel in the center of the filtering window; the same processing is performed on each pixel in the tongue image to achieve the denoising effect.
5. The tongue image feature extraction and health assessment method based on deep learning according to claim 1 is characterized in that: The tongue characteristics are analyzed as follows: Step 1: Divide the tongue motion vectors in the demonstration video into the same intervals, and count the number and direction to obtain the number of healthy standards in each interval. and the average direction ; For the motion vectors that are determined to be moving irregularly, the direction of the motion vector is divided into several intervals of equal size according to the set angle size, and the number of motion vectors in each interval is counted and the average direction , is the interval number; Step 2: Calculate the number difference and average direction difference of the motion vectors in each interval gate. , the average direction difference , and compare the calculated quantity difference value and average direction difference value with the preset quantity threshold and preset direction threshold, and record the number of corresponding items exceeding the threshold ; Step 3: If , it indicates that the number and direction deviation of motion vectors in the corresponding interval are small. is the preset ratio; Substitute the quantity difference value and the average direction difference value into the weight formula to obtain the health factor , and is the weight coefficient. If the health threshold is preset , the patient is judged to have potential health problems and the number and movement direction of the tongue are analyzed; otherwise, the patient is judged to be in good health and has no problems; if , it indicates that the number and direction deviation of the motion vectors in the corresponding interval are large, and it is determined that the patient has obvious health problems, and the number and movement direction of the tongue are analyzed; Step 4: For each frame of the patient and demonstration video in the corresponding time period, count the number of motion vectors in each direction interval and spatial position interval. Suppose the number of motion vectors in a certain direction interval and spatial position interval in the patient video is , the number of corresponding intervals in the demonstration video is , is the interval number, Number the spatial position interval; perform difference calculation on the two obtained quantity data to obtain the difference in the number of motion vectors in each interval , the difference in the number of motion vectors calculated for each interval Compare, if the corresponding interval , it indicates that the patient's tongue tremors in that direction and position.
6. The tongue image feature extraction and health assessment method based on deep learning according to claim 5 is characterized in that: The steps for assessing the patient's health status are as follows: Step 1: If the corresponding interval and , it means that the patient's tongue is trembling and deviating in direction, and it is judged that the patient may have a disease related to the nervous system. It is recommended that the patient go to the hospital for examination. The possibility of disease , and is the weight coefficient; Step 2: If the preset disease determination threshold is 1 , it means that the patient's tongue tremor amplitude and direction deviation are small, and it is judged that there may be early Parkinson's disease or mild neurological dysfunction. The patient is recommended to go to the neurology department for further examination; if the preset disease judgment threshold is two There are two possibilities. One is that the patient's tremor amplitude is large but the number of directional deviations is small. It is judged that the patient may have a local muscle disease and a change in the excitability of the nervous system. The patient is recommended to go to the stomatology department for further examination of the tongue muscles. The other is that the patient's tremor amplitude is small but the number of directional deviations is large. It is judged that the patient may have a neurodegenerative disease of the multiple system atrophy. The patient is recommended to go to the neurology department of the hospital for further examination. If the preset disease judgment threshold is two , indicating that the patient's tremor has a large amplitude and many directional deviations, indicating that the patient may have early manifestations of neurological dysfunction. It is recommended that the patient go to the neurology department for further examination.
Citation Information
Patent Citations
Device and method for acquiring and recognizing pulsation information and tongue inspection information
CN101129261A
Tongue picture characteristic change rule analysis method based on rehabilitation process of chronic disease patient
CN119314690A