Tongue diagnosis data acquisition method and device, computer equipment and readable storage medium

By combining high-resolution and high-frame-rate cameras with a pre-trained face and tongue joint detection model, the problem of poor equipment coordination in tongue diagnosis data collection is solved, high-precision tongue diagnosis data collection and calibration are achieved, data quality is improved, and reliable data support is provided for Traditional Chinese Medicine diagnosis.

CN120643186APending Publication Date: 2025-09-16CHINESE ACAD OF PREVENTIVE MEDICINE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510704453.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing tongue diagnosis data collection method cannot accurately collect tongue diagnosis data, and the equipment coordination is poor, resulting in spatial deviation in the data and affecting data quality.

Method used

A high-resolution camera and a high-frame rate camera are combined with a pre-trained face and tongue joint detection model. High-resolution tongue diagnosis images and high-frame rate tongue diagnosis dynamic videos are collected through motion and time voice prompts. They are calibrated through dual-camera geometric projection calibration to generate spatially aligned target images and videos.

Benefits of technology

The accuracy and quality of tongue diagnosis data collection are improved, providing a reliable data basis for subsequent tongue diagnosis analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120643186A_ABST
    Figure CN120643186A_ABST
Patent Text Reader

Abstract

The invention discloses a tongue diagnosis data acquisition method and device, computer equipment and a readable storage medium, and the method comprises the steps: firstly responding to a tongue diagnosis data acquisition instruction, starting a high-resolution camera and a high-frame-rate camera, and calling a pre-trained face and tongue joint detection model to carry out face detection; after a human face is detected, an action and time voice prompt aiming at the tongue is sent out, high-resolution tongue diagnosis images and high-frame-rate tongue diagnosis dynamic videos are collected through double cameras, finally collected data are calibrated, and target images and videos which are spatially aligned are obtained, so that the accuracy and the data quality of tongue diagnosis data collection are improved. And a reliable data basis is provided for subsequent tongue diagnosis analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and in particular to a tongue diagnosis data acquisition method, device, computer equipment and readable storage medium. Background Art

[0002] Traditional tongue diagnosis relies primarily on visual observation by Traditional Chinese Medicine practitioners, which is subject to significant subjectivity and a lack of quantitative standards. With the advancement of computer vision technology, image and video-based tongue diagnosis data collection methods are gaining popularity. However, existing methods are unable to accurately capture tongue diagnosis data. For example, poor interoperability between image and video capture devices can lead to spatial deviations in the data, compromising data quality. Summary of the Invention

[0003] The object of the present invention is to provide a tongue diagnosis data collection method, device, computer equipment and readable storage medium.

[0004] In a first aspect, an embodiment of the present invention provides a tongue diagnosis data collection method, comprising:

[0005] In response to the tongue diagnosis data collection instruction, the high-resolution camera and the high-frame rate camera are started, and the pre-trained face and tongue joint detection model is called to perform face detection in the preset area;

[0006] When a face detection frame is detected in the preset area, a movement voice prompt and a time voice prompt for the user's tongue are initiated, and a high-resolution tongue diagnosis image and a high-frame-rate tongue diagnosis dynamic video determined based on the movement voice prompt and the time voice prompt are captured by the high-resolution camera and the high-frame-rate camera;

[0007] The high-resolution tongue diagnosis image and the high-frame-rate tongue diagnosis dynamic video are calibrated to obtain spatially aligned target high-resolution tongue diagnosis image and target high-frame-rate tongue diagnosis dynamic video.

[0008] In a possible implementation, the face and tongue joint detection model is trained by the following method, including:

[0009] Get the basic model built based on yolov5;

[0010] The prediction head included in the basic model adds regression prediction values ​​for the mouth key points and tongue key points, and sets the corresponding joint loss function to obtain the improved basic model. The joint loss function is: L = λ f L face +λ t L tongue +λ k L keypoints , where L face is the face detection loss, is the tongue detection loss, L keypoints The loss function for the mouth key point regression is, f ,λ t ,λ k Respectively represent the weights of the corresponding losses;

[0011] The improved basic model is trained based on the joint loss function, and when the training termination condition is met, the face and tongue joint detection model that has completed training is obtained.

[0012] In a possible implementation, the action voice prompt includes a tongue surface action voice prompt and a tongue base action voice prompt, and the time voice prompt includes an extension time voice prompt;

[0013] The initiating of a motion voice prompt and a time voice prompt for the user's tongue, and collecting a high-resolution tongue diagnosis image and a high-frame-rate tongue diagnosis dynamic video determined based on the motion voice prompt and the time voice prompt by the high-resolution camera and the high-frame-rate camera, includes:

[0014] Initiate voice prompts for tongue movements on the user's tongue;

[0015] Initiating a voice prompt for the tongue extension time when it is determined based on the face and tongue joint detection model that the user's tongue is correctly extended according to the tongue action voice prompt;

[0016] When the confidence level of the tongue surface detection of the user's tongue surface in consecutive n frames based on the voice prompt of the tongue extension time is greater than a preset confidence threshold, a high-resolution camera shooting function is called to control the high-resolution camera to obtain a high-resolution tongue surface image, and a dynamic video of the tongue surface within a preset time range is obtained through the high-frame rate camera; n is a positive integer;

[0017] Initiating a voice prompt of the tongue bottom movement directed to the user's tongue bottom;

[0018] Initiating a voice prompt for the extension time when it is determined based on the face and tongue joint detection model that the user's tongue base is correctly extended according to the tongue base action voice prompt;

[0019] When the confidence level of the tongue bottom detection of the user's tongue bottom based on the voice prompt of the extension time in consecutive n frames is greater than a preset confidence threshold, a high-resolution camera shooting function is called to control the high-resolution camera to obtain a high-resolution tongue bottom image, and a dynamic video of the tongue bottom within a preset time range is obtained through the high-frame rate camera;

[0020] The high-resolution tongue surface image and the high-resolution tongue base image are used as the high-resolution tongue diagnosis image, and the tongue surface dynamic video and the tongue base dynamic video are used as the high-frame rate tongue diagnosis dynamic video.

[0021] In one possible implementation, the method further includes:

[0022] Performing clarity detection on the high-resolution tongue surface image and the high-resolution tongue base image based on grayscale variance;

[0023] In the case that the high-resolution tongue surface image and / or the high-resolution tongue base image fails the clarity detection, the corresponding step of returning to execute the step of initiating the voice prompt of the tongue surface movement for the user's tongue surface or the step of initiating the voice prompt of the tongue base movement for the user's tongue base is performed.

[0024] In a possible implementation, calibrating the high-resolution tongue diagnosis image and the high-frame-rate tongue diagnosis dynamic video to obtain a spatially aligned target high-resolution tongue diagnosis image and a target high-frame-rate tongue diagnosis dynamic video includes:

[0025] The high-resolution tongue diagnosis image and high-frame-rate tongue diagnosis dynamic video are calibrated using the dual-camera geometric projection formula: Performing calibration to obtain the target high-resolution tongue diagnosis image and the target high-frame-rate tongue diagnosis dynamic video that are spatially aligned;

[0026] Wherein, K1 is the camera intrinsic parameter matrix of the high-resolution camera, K2 is the camera intrinsic parameter matrix of the high-frame rate camera, p1 is the pixel coordinate corresponding to the spatial point p in the image of the high-resolution camera, p2 is the pixel coordinate corresponding to the spatial point p in the image of the high-frame rate camera, and [R|t] is the rotation and translation relationship between the high-resolution camera and the high-frame rate camera.

[0027] In a possible implementation, after the spatial alignment, the method further includes:

[0028] Performing U-Net-based tongue semantic segmentation on the target high-resolution tongue diagnosis image to generate a segmented image containing a tongue region mask;

[0029] Performing tongue edge detection frame by frame on the target high-frame-rate tongue diagnosis dynamic video, and eliminating tongue edge breaks in the dynamic video through adaptive morphological closing operation;

[0030] Among them, the tongue semantic segmentation adopts a shadow-aware attention mechanism, which dynamically adjusts the weight distribution of the segmentation network to the tongue reflective area by calculating the gradient histogram of the V channel of the HSV color space.

[0031] In one possible implementation, the method further includes:

[0032] The motion features of the segmented tongue dynamic video are extracted, including:

[0033] The displacement vector field of the tongue area in adjacent frames is calculated using the optical flow method;

[0034] A time-domain feature encoder is constructed based on the LSTM network. The displacement vector field of each frame and the corresponding tongue mask are channel-concatenated and then input into the encoder.

[0035] The dynamic feature vector output by the encoding is fused with the texture feature vector of the high-resolution tongue diagnosis image at multiple scales to generate a joint feature representation for TCM syndrome differentiation.

[0036] In a second aspect, an embodiment of the present invention provides a tongue diagnosis data collection device, comprising:

[0037] A startup module is used to start the high-resolution camera and the high-frame rate camera in response to the tongue diagnosis data collection instruction, and call the pre-trained face and tongue joint detection model to perform face detection in a preset area;

[0038] The acquisition module initiates action voice prompts and time voice prompts for the user's tongue when a face detection frame is detected in the preset area, and acquires high-resolution tongue diagnosis images and high-frame-rate tongue diagnosis dynamic videos determined based on the action voice prompts and the time voice prompts through the high-resolution camera and the high-frame-rate camera; calibrates the high-resolution tongue diagnosis images and high-frame-rate tongue diagnosis dynamic videos to obtain spatially aligned target high-resolution tongue diagnosis images and target high-frame-rate tongue diagnosis dynamic videos.

[0039] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, wherein when the computer instructions are executed by the processor, the computer device executes the method described in the first aspect.

[0040] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, wherein the readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method described in the first aspect.

[0041] Compared to existing technologies, the present invention provides the following advantages: Using the disclosed tongue diagnosis data collection method, apparatus, computer device, and readable storage medium, the system activates a high-resolution camera and a high-frame-rate camera in response to a tongue diagnosis data collection instruction, and calls a pre-trained face and tongue joint detection model to perform face detection. Upon detecting a face, a voice prompt is issued regarding the tongue's movement and timing. Based on this, dual cameras are used to capture high-resolution tongue diagnosis images and high-frame-rate tongue diagnosis dynamic videos. Finally, the captured data is calibrated to obtain spatially aligned target images and videos. This improves the accuracy and quality of tongue diagnosis data collection, providing a reliable data foundation for subsequent tongue diagnosis analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly describes the drawings required for use in the embodiments. It should be understood that the following drawings illustrate only certain embodiments of the present invention and should not be construed as limiting the scope of the present invention. Those skilled in the art can, without inventive effort, derive other relevant drawings from these drawings.

[0043] Figure 1 A schematic diagram of the steps of the tongue diagnosis data collection method provided by an embodiment of the present invention;

[0044] Figure 2 Schematic diagram of the face and tongue joint detection model framework based on the improved yolov5 provided in an embodiment of the present invention;

[0045] Figure 3 A schematic block diagram of the structure of a tongue diagnosis data acquisition device provided in an embodiment of the present invention;

[0046] Figure 4 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.

[0048] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0049] In order to solve the technical problems in the above background technology, Figure 1 This is a flow chart of the tongue diagnosis data collection method provided in an embodiment of the present disclosure. The tongue diagnosis data collection method is introduced in detail below.

[0050] Step S201: In response to the tongue diagnosis data acquisition instruction, a high-resolution camera and a high-frame rate camera are activated, and a pre-trained face and tongue joint detection model is called to perform face detection in a preset area;

[0051] Step S202: When a face detection frame is detected in the preset area, a movement voice prompt and a time voice prompt for the user's tongue are initiated, and a high-resolution tongue diagnosis image and a high-frame-rate tongue diagnosis dynamic video determined based on the movement voice prompt and the time voice prompt are captured by the high-resolution camera and the high-frame-rate camera;

[0052] Step S203 , calibrating the high-resolution tongue diagnosis image and the high-frame-rate tongue diagnosis dynamic video to obtain spatially aligned target high-resolution tongue diagnosis image and target high-frame-rate tongue diagnosis dynamic video.

[0053] In this embodiment of the present invention, for example, assume that a patient comes to a modern traditional Chinese medicine clinic for treatment. The doctor uses a terminal device in the clinic to send a tongue diagnosis data collection instruction to the server. After receiving the instruction, the server immediately starts to execute the relevant operations.

[0054] The server establishes a connection with the high-resolution camera and high-frame-rate camera installed in the clinic and activates both cameras. The high-resolution camera captures high-definition images of the tongue, capturing subtle features such as texture and color. The high-frame-rate camera records dynamic tongue video, capturing detailed changes in tongue movement.

[0055] At the same time, the server calls a pre-trained joint face and tongue detection model. This model is obtained through a complex training process. For example, the server obtains a base model built based on Yolov5, then adds regression predictions for mouth and tongue key points to the prediction head included in the base model and sets the corresponding joint loss function. The improved base model is trained based on this joint loss function. When the training termination condition is met, the trained joint face and tongue detection model is obtained.

[0056] At this point, the model begins face detection in a preset area. This can be a specific area in front of the camera lens, such as the area where the patient is sitting in front of the camera with their face roughly in the center of the frame.

[0057] For example, a patient sits in a specific acquisition position, facing the camera. The server captures the real-time image through the camera, and the face-tongue joint detection model begins searching for faces in the image. It analyzes various features in the image, such as facial contours and the relative positions of facial features, to determine whether a face is present. If other people or objects are present in the image, the model automatically filters out these distracting features and focuses on finding areas that match the facial features.

[0058] Once the face and tongue joint detection model successfully detects a face in the preset area, it generates a face detection frame.

[0059] The server issues voice prompts for the user's tongue movements, including those for the tongue surface and tongue base. For example, the server uses the clinic's voice playback device to clearly signal the patient: "Please stick out your tongue as far as possible, exposing the tongue surface." This guides the patient through the correct tongue movements.

[0060] Based on a joint face and tongue detection model, the server analyzes the patient's movements in real time to see if they meet the requirements. The model detects the position and state of key points on the mouth and tongue to determine whether the tongue is correctly extended as prompted. For example, the model observes features such as the length and angle of tongue extension, as well as whether the tongue surface is fully exposed. If the patient's tongue is not extended long enough or not fully extended, the model will detect this.

[0061] When the model determines that the patient's tongue is correctly extended according to the voice prompt of the tongue movement, the server immediately initiates a voice prompt for the extension time: "Please maintain the current action for 3 seconds." At this time, the high-resolution camera and the high-frame rate camera begin to prepare to collect data.

[0062] While the patient's tongue is maintaining movement based on the voice prompts for the extension time, the server analyzes the video images transmitted in real time by the high-frame rate camera to detect whether the confidence of the tongue detection in n consecutive frames is greater than the preset confidence threshold. For example, assume that n is 5 and the preset confidence threshold is 0.8. When the confidence of the tongue detection is greater than 0.8 in 5 consecutive frames, the server believes that the current tongue state is stable and meets the requirements, so it calls the high-resolution camera photo function to control the high-resolution camera to quickly take a high-resolution tongue image. At the same time, the high-frame rate camera is like a "faithful recorder", obtaining a dynamic video of the tongue within a preset time range (such as 5 seconds) to record the subtle changes in the tongue's movements during this process.

[0063] After collecting the tongue image, the server further processes it. The high-resolution tongue image is tested for clarity based on grayscale variance. This is like using a "clarity ruler" to measure the image's quality. If the grayscale variance is within an appropriate range, the image has high clarity; conversely, if the grayscale variance is too small, the image may be blurry.

[0064] If the high-resolution tongue image fails the clarity test, the server, like a rigorous quality inspector, will deem the image unacceptable and then return to initiate voice prompts for the user's tongue movements. For example, if the patient slightly shakes their head while maintaining the movement, causing the image to blur, the server will remind the patient again: "Please stick out your tongue as far as possible, exposing the tongue surface, and remain steady."

[0065] After completing the tongue surface data collection, the server continues to issue a voice prompt for the user's tongue base movement: "Now please roll your tongue upward to expose the tongue base." Similarly, after the face and tongue joint detection model determines that the patient's tongue base has been correctly extended according to the tongue base movement voice prompt, it issues another extension time voice prompt: "Please maintain the current position for 3 seconds."

[0066] Similar to collecting tongue surface data, when the confidence level of the tongue base detection in consecutive n frames based on the voice prompt of the patient's tongue extension time is greater than the preset confidence threshold, the server calls the high-resolution camera photo function to control the high-resolution camera, obtains a high-resolution tongue base image, and obtains a dynamic video of the tongue base within a preset time range through the high-frame rate camera.

[0067] Finally, the server uses the high-resolution tongue surface image and the high-resolution tongue bottom image as high-resolution tongue diagnosis images, and uses the tongue surface dynamic video and the tongue bottom dynamic video as high-frame rate tongue diagnosis dynamic video.

[0068] Since high-resolution cameras and high-frame rate cameras may have differences in physical position and parameter settings, the tongue position and angle in the collected high-resolution tongue diagnosis images and high-frame rate tongue diagnosis dynamic videos may be inconsistent, which requires calibration.

[0069] The server uses the dual-camera geometric projection calibration formula to calibrate the collected data, and adjusts the tongue position in the tongue diagnosis images taken by the high-resolution camera and the tongue diagnosis dynamic video recorded by the high-frame rate camera so that they are spatially aligned.

[0070] Specifically, the server "adjusts" the tongue in images and videos based on the respective characteristics of the high-resolution and high-frame-rate cameras (similar to their "personal characteristics"), as well as their relative positional relationship (rotation and translation). After calibration, the target high-resolution tongue diagnosis image and target high-frame-rate tongue diagnosis dynamic video are spatially aligned. This ensures spatial consistency of the data for subsequent analysis and processing, laying the foundation for accurate tongue diagnosis.

[0071] In an embodiment of the present invention, the face and tongue joint detection model is trained in the following manner and can be implemented through the following examples.

[0072] Get the basic model built based on yolov5;

[0073] The prediction head included in the basic model adds regression prediction values ​​for the mouth key points and tongue key points, and sets the corresponding joint loss function to obtain the improved basic model. The joint loss function is: L = λ f L face +λ t L tongue +λ k L keypoints , where L face is the face detection loss, is the tongue detection loss, L keypoints The loss function for the mouth key point regression is, f ,λ t ,λk Respectively represent the weights of the corresponding losses;

[0074] The improved basic model is trained based on the joint loss function, and when the training termination condition is met, the face and tongue joint detection model that has completed training is obtained.

[0075] In the embodiment of the present invention, for example, first, the server obtains a basic model built based on yolov5. This is like having a rough frame before building a house. The yolov5 basic model has certain image detection capabilities, but it needs to be optimized for tongue diagnosis scenarios.

[0076] Next, the server adds regression predictions for the mouth and tongue keypoints to the prediction head of this base model and sets the corresponding joint loss function. The prediction head of the base model is a "detector" that may originally only respond to common targets. Now, the server adds the ability to detect the mouth and tongue keypoints.

[0077] The server then trains the improved basic model based on this joint loss function. The server continuously feeds the model with large amounts of image data containing annotations of key points of the face, tongue, and mouth. The model evaluates the quality of its detection results based on the joint loss function and adjusts its parameters accordingly. As training progresses, the model's detection accuracy increases. When the training termination conditions are met, such as the model's detection accuracy on the validation set improving minimally for multiple consecutive times, or when it reaches the accuracy standard preset by researchers, the server obtains a fully trained joint face and tongue detection model. This model can then accurately detect face and tongue-related information in tongue diagnosis data collection.

[0078] In an embodiment of the present invention, the action voice prompt includes a tongue surface action voice prompt and a tongue base action voice prompt, and the time voice prompt includes an extension time voice prompt;

[0079] The initiation of action voice prompts and time voice prompts for the user's tongue, and the acquisition of high-resolution tongue diagnosis images and high-frame rate tongue diagnosis dynamic videos determined based on the action voice prompts and the time voice prompts by the high-resolution camera and the high-frame rate camera can be implemented through the following examples.

[0080] Initiate voice prompts for tongue movements on the user's tongue;

[0081] Initiating a voice prompt for the tongue extension time when it is determined based on the face and tongue joint detection model that the user's tongue is correctly extended according to the tongue action voice prompt;

[0082] When the confidence level of the tongue surface detection of the user's tongue surface in consecutive n frames based on the voice prompt of the tongue extension time is greater than a preset confidence threshold, a high-resolution camera shooting function is called to control the high-resolution camera to obtain a high-resolution tongue surface image, and a dynamic video of the tongue surface within a preset time range is obtained through the high-frame rate camera; n is a positive integer;

[0083] Initiating a voice prompt of the tongue bottom movement directed to the user's tongue bottom;

[0084] Initiating a voice prompt for the extension time when it is determined based on the face and tongue joint detection model that the user's tongue base is correctly extended according to the tongue base action voice prompt;

[0085] When the confidence level of the tongue bottom detection of the user's tongue bottom based on the voice prompt of the extension time in consecutive n frames is greater than a preset confidence threshold, a high-resolution camera shooting function is called to control the high-resolution camera to obtain a high-resolution tongue bottom image, and a dynamic video of the tongue bottom within a preset time range is obtained through the high-frame rate camera;

[0086] The high-resolution tongue surface image and the high-resolution tongue base image are used as the high-resolution tongue diagnosis image, and the tongue surface dynamic video and the tongue base dynamic video are used as the high-frame rate tongue diagnosis dynamic video.

[0087] In an embodiment of the present invention, for example, when the server receives a tongue diagnosis data collection instruction and the face and tongue joint detection model detects a face detection frame in a preset area, it starts the movement voice prompt and time voice prompt process for the user's tongue, and coordinates the high-resolution camera and the high-frame rate camera for data collection.

[0088] First, the server uses the voice playback device in the clinic to clearly initiate a voice prompt for the user's tongue movements: "Please stick out your tongue naturally, stretch it as far as possible, and fully expose the tongue." This prompt is like a professional guide, clearly informing the patient of the movements they need to make. At this time, the face and tongue joint detection model is like a rigorous "supervisor", paying close attention to the patient's movements. Based on the knowledge acquired in previous training, it analyzes the real-time images captured by the camera to determine whether the user's tongue is correctly extended according to the voice prompt for the tongue movements. For example, the model will observe whether the tongue is extended long enough, whether the tongue surface is flat and spread out, and whether there is any curling or obstruction.

[0089] Once the model determines that the user's tongue has been correctly extended, the server immediately issues a voice prompt for the extension time: "Please maintain the current tongue extension state for 3 seconds." During these 3 seconds, the high-resolution camera and high-frame-rate camera remain in full alert. The server continuously analyzes the real-time video streamed by the high-frame-rate camera to determine the tongue detection confidence level. This confidence level can be understood as the model's assessment of the accuracy of the currently detected tongue state. When the tongue detection confidence level exceeds a preset confidence threshold (e.g., 0.8) for n consecutive frames (assuming n is 5), the server deems the tongue state stable and meeting the requirements. Much like a referee determining that an athlete has met the required standards, the server immediately calls the high-resolution camera's capture function, precisely controlling the high-resolution camera to capture a high-resolution image of the tongue, capturing every subtle feature. Simultaneously, the high-frame-rate camera begins capturing dynamic tongue video for a preset timeframe (e.g., 5 seconds), recording any subtle tremors or changes that may occur during this timeframe.

[0090] After completing the tongue surface data collection, the server then issues a voice prompt for the user's tongue base through the voice device: "Next, please roll your tongue upwards to fully expose the tongue base." Similarly, the face and tongue joint detection model once again takes on the responsibility of judging, analyzing the image to see if the user's tongue base is correctly extended according to the prompt. If it is determined that the user's tongue base is correctly extended, the server issues another voice prompt for the extension time: "Please keep the tongue base extended for 3 seconds."

[0091] The server continuously monitors the images transmitted by the high-frame-rate camera, just as it previously monitored the tongue surface. When the tongue base detection confidence exceeds a preset confidence threshold for n consecutive frames based on the voice prompt for the tongue extension time, the server quickly calls the high-resolution camera's capture function, controlling the high-resolution camera to obtain a high-resolution image of the tongue base. Simultaneously, the high-frame-rate camera captures a dynamic video of the tongue base within a preset time range, comprehensively recording the tongue base's condition.

[0092] Finally, the server integrates the carefully collected high-resolution tongue surface and tongue base images into a high-resolution tongue diagnosis image. It also aggregates the acquired dynamic tongue surface and tongue base videos into a high-frame-rate dynamic tongue diagnosis video. This data will provide comprehensive and accurate information for subsequent tongue diagnosis analysis in Traditional Chinese Medicine (TCM), helping doctors more accurately assess patients' physical conditions.

[0093] In the embodiments of the present invention, the following implementation modes are also provided.

[0094] Performing clarity detection on the high-resolution tongue surface image and the high-resolution tongue base image based on grayscale variance;

[0095] In the case that the high-resolution tongue surface image and / or the high-resolution tongue base image fails the clarity detection, the corresponding step of returning to execute the step of initiating the voice prompt of the tongue surface movement for the user's tongue surface or the step of initiating the voice prompt of the tongue base movement for the user's tongue base is performed.

[0096] In an embodiment of the present invention, for example, in a traditional Chinese medicine clinic, after the server completes the acquisition of high-resolution tongue surface and tongue base images, it will immediately perform a clarity test. The image clarity is measured by calculating the change in grayscale in the image, that is, the grayscale variance. For example, in a clear tongue surface image, the texture of the tongue coating and the color change of the tongue will cause obvious and reasonable fluctuations in the grayscale, which is reflected in the grayscale variance as a suitable numerical range. The server compares the calculated grayscale variance with a pre-set standard to determine the image clarity.

[0097] Similarly, for high-resolution tongue base images, the server uses the same method to detect clarity based on grayscale variance. Details such as the vascular distribution and mucosal condition of the tongue base exhibit specific grayscale variation patterns in clear images, and the server uses the grayscale variance criteria corresponding to these patterns to determine clarity.

[0098] If a high-resolution tongue image fails the clarity test, the server responds quickly, much like a product failing quality inspection. It immediately responds by initiating voice prompts for the user's tongue movements. For example, the server clearly informs the patient again through the voice device: "Please extend your tongue naturally, stretch it as far as possible, and fully expose the tongue." This guides the patient to perform the correct tongue movements again to recapture a clear tongue image.

[0099] If the high-resolution tongue base image fails the clarity test, the server will return to initiating a voice prompt for the user's tongue base. The server will prompt, "Next, please roll your tongue upward to fully expose the tongue base," reminding the patient to re-expose the tongue base as required, allowing for another clear tongue base image to be captured. Only when both the high-resolution tongue surface image and the high-resolution tongue base image pass the clarity test will the server deem this round of tongue diagnosis image acquisition qualified and proceed with subsequent tongue diagnosis data processing, providing reliable image data for Traditional Chinese Medicine diagnosis.

[0100] In the embodiment of the present invention, the high-resolution tongue diagnosis image and the high-frame-rate tongue diagnosis dynamic video are calibrated to obtain a spatially aligned target high-resolution tongue diagnosis image and a target high-frame-rate tongue diagnosis dynamic video, which can be implemented through the following examples.

[0101] The high-resolution tongue diagnosis image and high-frame-rate tongue diagnosis dynamic video are calibrated using the dual-camera geometric projection formula: Performing calibration to obtain the target high-resolution tongue diagnosis image and the target high-frame-rate tongue diagnosis dynamic video that are spatially aligned;

[0102] Wherein, K1 is the camera intrinsic parameter matrix of the high-resolution camera, K2 is the camera intrinsic parameter matrix of the high-frame rate camera, p1 is the pixel coordinate corresponding to the spatial point p in the image of the high-resolution camera, p2 is the pixel coordinate corresponding to the spatial point p in the image of the high-frame rate camera, and [R|t] is the rotation and translation relationship between the high-resolution camera and the high-frame rate camera.

[0103] In this embodiment of the present invention, a high-resolution camera and a high-frame-rate camera serve as two "observers" with different perspectives, simultaneously recording the patient's tongue. However, due to their different positions, angles, and inherent characteristics, the spatial position of the tongue in the captured images and videos may differ.

[0104] The server uses a dual-camera geometric projection calibration formula to calibrate these differences. The server first obtains the camera intrinsic parameter matrix K1 of the high-resolution camera, which records its own characteristics such as focal length and imaging plane, and also obtains the camera intrinsic parameter matrix K2 of the high-frame rate camera.

[0105] Then, for any point p in space, the server finds its corresponding pixel coordinates p1 in the high-resolution camera image and p2 in the high-frame-rate camera image. Point p here can be a feature point on the tongue, such as a point on the tip of the tongue.

[0106] At the same time, the server also holds the rotation and translation relationship [R|t] between the high-resolution camera and the high-frame rate camera, which describes the relative position and angle relationship between the two cameras in space.

[0107] Based on this information, the server processes the high-resolution tongue diagnosis images and high-frame-rate tongue diagnosis videos using a specific calibration method. The tongue position in the high-resolution tongue diagnosis images and high-frame-rate tongue diagnosis videos is precisely adjusted based on the relationship between the cameras and their individual characteristics.

[0108] After this calibration operation, the images and videos that originally had different spatial positions are spatially aligned, resulting in the spatially aligned target high-resolution tongue diagnosis images and target high-frame-rate tongue diagnosis dynamic videos. This makes subsequent analysis of tongue diagnosis data more accurate, because both the static characteristics of the tongue (via the target high-resolution tongue diagnosis images) and dynamic changes (via the target high-frame-rate tongue diagnosis dynamic videos) are observed based on a unified spatial coordinate system, providing a more reliable data foundation for TCM practitioners to accurately diagnose conditions.

[0109] In an embodiment of the present invention, after the spatial alignment, the following implementation manner is further provided.

[0110] Performing U-Net-based tongue semantic segmentation on the target high-resolution tongue diagnosis image to generate a segmented image containing a tongue region mask;

[0111] Performing tongue edge detection frame by frame on the target high-frame-rate tongue diagnosis dynamic video, and eliminating tongue edge breaks in the dynamic video through adaptive morphological closing operation;

[0112] Among them, the tongue semantic segmentation adopts a shadow-aware attention mechanism, which dynamically adjusts the weight distribution of the segmentation network to the tongue reflective area by calculating the gradient histogram of the V channel of the HSV color space.

[0113] In an embodiment of the present invention, illustratively, first, the server performs U-Net-based tongue semantic segmentation on the target high-resolution tongue diagnosis image.

[0114] During the segmentation process, tongue semantic segmentation utilizes a shadow-aware attention mechanism. For example, the tongue can sometimes have reflective areas due to lighting issues, which can interfere with segmentation accuracy. The server addresses this issue by calculating the gradient histogram of the V channel in the HSV color space. This acts as a "special guide" for the U-Net. By analyzing the V channel gradient histogram, the server understands the brightness variations in different regions of the tongue and dynamically adjusts the segmentation network's weighting of reflective areas. If a particular area is highly reflective, the server increases the segmentation network's attention to that area, allowing the U-Net to more accurately identify this part as belonging to the tongue, ultimately generating a segmented image containing a mask of the tongue area.

[0115] At the same time, the server processes the target high-frame-rate tongue diagnosis dynamic video. However, in dynamic videos, the edges of the tongue may be broken due to various factors.

[0116] To address this issue, the server uses adaptive morphological closing operations to eliminate broken tongue edges. For example, if a small break in the tongue edge is detected in a frame, the "brush" automatically "draws" appropriate lines based on the morphological characteristics of the surrounding tongue edges to connect the broken area, ensuring that the tongue edge is complete and continuous in every frame.

[0117] Through this series of operations, the server processes the spatially aligned tongue diagnosis data more finely, providing a higher-quality data basis for subsequent TCM syndrome differentiation analysis based on tongue characteristics, helping doctors to more accurately judge the patient's health status.

[0118] In the embodiments of the present invention, the following implementation modes are also provided.

[0119] The motion features of the segmented tongue dynamic video are extracted, including:

[0120] The displacement vector field of the tongue area in adjacent frames is calculated using the optical flow method;

[0121] A time-domain feature encoder is constructed based on the LSTM network. The displacement vector field of each frame and the corresponding tongue mask are channel-concatenated and then input into the encoder.

[0122] The dynamic feature vector output by the encoding is fused with the texture feature vector of the high-resolution tongue diagnosis image at multiple scales to generate a joint feature representation for TCM syndrome differentiation.

[0123] In an embodiment of the present invention, for example, the server first calculates the displacement vector field of the tongue region between adjacent frames using the optical flow method. The server uses the optical flow method to track the movement trajectory of the tongue. It carefully compares the changes in the tongue region between two adjacent frames and calculates the direction and distance of movement of each point on the tongue from one frame to the next. This movement information is aggregated to form a displacement vector field. For example, if the tongue trembles slightly over a period of time, the optical flow method can accurately capture the direction and amplitude of the tremor of each part of the tongue, presenting it in the form of a vector.

[0124] Next, the server constructs a time-domain feature encoder based on an LSTM network. The LSTM network integrates information describing tongue movement (the displacement vector field) and tongue contour information (the tongue mask) to form a richer information set. This set is then input into the time-domain feature encoder built on the LSTM network. The encoder begins to work, deeply analyzing and processing the input information to extract the tongue's motion characteristics in the time dimension.

[0125] Finally, the server performs multi-scale feature fusion on the dynamic feature vectors output by the encoding with the texture feature vectors of the high-resolution tongue diagnosis image. The high-resolution tongue diagnosis image records the texture information of the tongue, such as the thickness of the tongue coating and the texture of the tongue. These texture feature vectors represent the static characteristics of the tongue. The server fuses the feature vectors representing the dynamic movement of the tongue and the feature vectors representing the static texture of the tongue at different scales. It comprehensively considers these features from multiple angles, allowing the dynamic and static information to complement each other. For example, the combination of the tremor of the tongue (dynamic feature) and the texture of the tongue coating (static feature) generates a more comprehensive and representative joint feature representation. This joint feature representation contains rich information about the tongue and can provide strong support for TCM doctors in their diagnosis and diagnosis, helping them to more accurately judge the patient's physical condition.

[0126] In order to more clearly describe the method provided by the embodiment of the present invention, a relatively complete implementation method is provided below.

[0127] Step 1: Start the device and the high frame rate camera starts working continuously to detect whether there is a face entering the device;

[0128] Step 2: The system detects whether a face appears in the device through the improved algorithm of face and tongue joint detection based on yolov5 proposed by the present invention.

[0129] The algorithm improved in this paper does not change the Backbone and Neck fusion part. The improvement occurs in the prediction head and loss function at the end of the model. The regression prediction values ​​of the mouth and tongue key points are added and optimized by the proposed joint loss function. Please refer to Figure 2 , Figure 2 Schematic diagram of the face and tongue joint detection model framework based on the improvement of yolov5 provided in an embodiment of the present invention.

[0130] The improved algorithm reuses features by sharing feature maps, key point detection, and tongue detection, thereby improving detection accuracy and efficiency. Feature maps are multi-scale features extracted from the model's backbone and neck, containing both spatial and semantic information about the image. Facial key points (such as the corners or center of the mouth) provide clues to where the tongue may appear, limiting the prediction range. Global features (such as the face or mouth contour) used in tongue detection help to achieve more accurate key point detection.

[0131] L face =λ cls L cls +λ obj L obj +λ box L box

[0132] Where, Face detection loss λ cls ,λ obj ,λ box These are the weight coefficients for classification loss, confidence loss, and bounding box loss, respectively, which are used to adjust the importance of different loss items in the total loss. These weight coefficients can be adjusted and optimized based on specific tasks and datasets.

[0133] L tongue =λ cls L cls +λ obj L obj +λ box L box

[0134] Where, Tongue detection loss, other parameters are the same as above.

[0135] Add the mouth key point detection module in YOLOv5 as an auxiliary task to improve the accuracy of tongue positioning. The Euclidean distance between the predicted key points and the real key points is minimized through model training. Assume that there are K key points in the detection box, and their coordinates are (x i ,y i ),i=1,2,…,K, then The regression loss of the mouth keypoint is defined as:

[0136]

[0137] Where, is the key point coordinate predicted by the model, (x i ,y i ) is the real key point coordinate, L key It is the key point regression loss function in the form of mean square error (MSE). The posture angle is calculated based on the key point coordinates obtained by regression, and the extension angle of the tongue can be calculated.

[0138]

[0139] Where (x1, y1) and (x2, y2) are the coordinates of two key points respectively. The joint loss function of multi-task learning is:

[0140] L=λ f L face +λ t L tongue +λ k L keypoints

[0141] Among them, λ f ,λ t ,λ k They represent the weights of the corresponding losses.

[0142] Step 3: After the face detection frame is obtained through the above target detection model, the voice prompt "Please stick out your tongue" is broadcast.

[0143] Step 4: Continue to call the above model to make the model detect whether the tongue is correctly extended. After detecting that it is correctly extended, the voice prompt "Please hold for three seconds"

[0144] Step 5: When the confidence level of tongue detection for n consecutive frames reaches the threshold m1, the high-resolution camera photo function is called to obtain a high-resolution tongue image and a dynamic video of the high-frame rate camera for 3 seconds before the photo is taken.

[0145] Step 6: Perform a clarity test on the tongue surface image obtained in step 5 to avoid blurring caused by tongue shaking at the moment of taking the photo. The specific method is to calculate the SMD (grayscale variance):

[0146] D(f)=∑ y ∑ x (|f(x,y)-f(x,y-1)|+|f(x,y)-f(x+1,y)|)

[0147] When the grayscale variance D is lower than the threshold, it means the image clarity is low, and jump back to step 3. When it is higher than the threshold, it means the image clarity is normal. Go to step 7

[0148] In step 7, a voice prompt "Please stick out the bottom of your tongue" is broadcast, and the above-mentioned target detection model is called to detect whether the bottom of the tongue is correctly extended.

[0149] Step 8: Continue to call the above model to make the model detect whether the tongue base is correctly extended. After detecting that it is correctly extended, the voice prompt "Please hold for three seconds"

[0150] Step 9: When the confidence level of tongue base detection for n consecutive frames reaches the threshold m1, the high-resolution camera photo function is called to obtain a high-resolution tongue base image and a dynamic video of the high-frame rate camera 3 seconds before the photo is taken.

[0151] Step 10: Perform a clarity check on the tongue base image obtained in step 9 to avoid blurring caused by tongue shaking at the moment of taking the photo. The specific method used is to calculate the grayscale variance (SMD). If the grayscale variance D is below the threshold, the image clarity is low, and the process returns to step 7. If it is above the threshold, the image clarity is normal, and the process proceeds to step 11.

[0152] Step 11: The high-resolution images and high-frame-rate videos collected by the above operation have different camera positions and angles, resulting in mismatched tongue coordinates, which affects the subsequent analysis process. Therefore, a dual-camera geometric projection calibration algorithm is used for image calibration.

[0153] In order to solve the problem of target position offset caused by camera position and angle differences in the dual-camera acquisition system, this section designs a dual-camera geometric projection calibration algorithm based on feature point matching. This paper uses a 9×6 corner point checkerboard as the calibration plate, with a spacing of 25mm between each corner point. The two cameras simultaneously capture images of the calibration plate at different angles (0°, ±15°, ±30°) to establish a mapping relationship between the camera coordinate systems. According to the pinhole camera imaging model, the projection relationship between camera 1 and camera 2 on the spatial point P can be expressed as: According to the camera imaging model, the following relationship can be obtained:

[0154] p1=K1[I|0]P;

[0155] p2=K2[R|t]P;

[0156] Where K1 and K2 are the camera intrinsic parameter matrices, including the focal length and principal point coordinates; [R|t] represents the rotation and translation relationship between the two cameras. The corresponding pixel coordinates in the image of camera 1 are p1, and the corresponding pixel coordinates in the image of camera 2 are p2.

[0157] The simultaneous equations yield:

[0158]

[0159] By performing the above calculations on multiple pairs of feature points, we can obtain a set of linear equations related to R and t. By solving the R and t parameters in Equation (3) using the Levenberg-Marquardt optimization algorithm, we can obtain the relative position between camera 1 and camera 2.

[0160] In step 12, we finally obtain spatially aligned high-resolution images of the tongue surface and tongue base, along with a high-frame-rate dynamic video, laying a solid foundation for future analysis. Returning to step 1, we control the device to continue performing subsequent user detection tasks.

[0161] Please refer to Figure 3 , Figure 3 A tongue diagnosis data collection device 110 provided in an embodiment of the present invention includes:

[0162] The starting module 1101 is configured to start the high-resolution camera and the high-frame rate camera in response to the tongue diagnosis data collection instruction, and call a pre-trained face and tongue joint detection model to perform face detection in a preset area;

[0163] The acquisition module 1102, when a face detection frame is detected in the preset area, initiates a movement voice prompt and a time voice prompt for the user's tongue, and acquires a high-resolution tongue diagnosis image and a high-frame-rate tongue diagnosis dynamic video determined based on the movement voice prompt and the time voice prompt through the high-resolution camera and the high-frame-rate camera; calibrates the high-resolution tongue diagnosis image and the high-frame-rate tongue diagnosis dynamic video to obtain a spatially aligned target high-resolution tongue diagnosis image and a target high-frame-rate tongue diagnosis dynamic video.

[0164] It should be noted that the implementation principles of the tongue diagnosis data acquisition device 110 can be referenced from the implementation principles of the tongue diagnosis data acquisition method described above and will not be elaborated upon here. It should be understood that the division of the various modules of the above device is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into a single physical entity or physically separated. Furthermore, these modules can be implemented entirely in software invoked by a processing element; or entirely in hardware; or some modules can be implemented in software invoked by a processing element, while others can be implemented in hardware. For example, the tongue diagnosis data acquisition device 110 can be a separate processing element, or integrated into a chip of the above device. Furthermore, it can be stored in the form of program code in the memory of the above device, invoked by a processing element of the above device and perform the functions of the tongue diagnosis data acquisition device 110. The implementation of the other modules is similar. Furthermore, these modules can be fully or partially integrated together or implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. During implementation, the steps of the above method or the above modules can be performed by hardware integrated logic circuits in the processor element or by software instructions.

[0165] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented by scheduling program code on a processing element, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0166] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the tongue diagnosis data collection device 110. Figure 4 As shown, Figure 4 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a tongue diagnosis data acquisition device 110 , a memory 111 , a processor 112 , and a communication unit 113 .

[0167] In order to realize data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, the electrical connection between these components can be achieved through one or more communication buses or signal lines. The tongue diagnosis data acquisition device 110 includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the tongue diagnosis data acquisition device 110 stored in the memory 111, such as the software function modules and computer programs included in the tongue diagnosis data acquisition device 110.

[0168] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, it controls the computer device where the readable storage medium is located to execute the aforementioned tongue diagnosis data acquisition device 110.

[0169] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.

Claims

1. A tongue diagnosis data collection method, characterized in that: include: In response to the tongue diagnosis data collection instruction, the high-resolution camera and the high-frame rate camera are started, and the pre-trained face and tongue joint detection model is called to perform face detection in the preset area; When a face detection frame is detected in the preset area, a movement voice prompt and a time voice prompt for the user's tongue are initiated, and a high-resolution tongue diagnosis image and a high-frame-rate tongue diagnosis dynamic video determined based on the movement voice prompt and the time voice prompt are captured by the high-resolution camera and the high-frame-rate camera; The high-resolution tongue diagnosis image and the high-frame-rate tongue diagnosis dynamic video are calibrated to obtain spatially aligned target high-resolution tongue diagnosis image and target high-frame-rate tongue diagnosis dynamic video.

2. The method according to claim 1, characterized in that The face and tongue joint detection model is trained by the following methods, including: Get the basic model built based on yolov5; The prediction head included in the basic model adds regression prediction values ​​for the mouth key points and tongue key points, and sets the corresponding joint loss function to obtain the improved basic model. The joint loss function is: K = λ f L face +λ t L tongue +λ k L keypo ints , where L face is the face detection loss, is the tongue detection loss, L keypo ints The loss function for the mouth key point regression is, f ,λ t ,λ k Respectively represent the weights of the corresponding losses; The improved basic model is trained based on the joint loss function, and when the training termination condition is met, the face and tongue joint detection model that has completed training is obtained.

3. The method according to claim 1, characterized in that The action voice prompt includes a tongue surface action voice prompt and a tongue base action voice prompt, and the time voice prompt includes an extension time voice prompt; The initiating of a motion voice prompt and a time voice prompt for the user's tongue, and collecting a high-resolution tongue diagnosis image and a high-frame-rate tongue diagnosis dynamic video determined based on the motion voice prompt and the time voice prompt by the high-resolution camera and the high-frame-rate camera, includes: Initiate voice prompts for tongue movements on the user's tongue; Initiating a voice prompt for the tongue extension time when it is determined based on the face and tongue joint detection model that the user's tongue is correctly extended according to the tongue action voice prompt; When the confidence level of the tongue surface detection of the user's tongue surface in consecutive n frames based on the voice prompt of the tongue extension time is greater than a preset confidence threshold, a high-resolution camera shooting function is called to control the high-resolution camera to obtain a high-resolution tongue surface image, and a dynamic video of the tongue surface within a preset time range is obtained through the high-frame rate camera; n is a positive integer; Initiating a voice prompt of the tongue bottom movement directed to the user's tongue bottom; Initiating a voice prompt for the extension time when it is determined based on the face and tongue joint detection model that the user's tongue base is correctly extended according to the tongue base action voice prompt; When the confidence level of the tongue bottom detection of the user's tongue bottom based on the voice prompt of the extension time in consecutive n frames is greater than a preset confidence threshold, a high-resolution camera shooting function is called to control the high-resolution camera to obtain a high-resolution tongue bottom image, and a dynamic video of the tongue bottom within a preset time range is obtained through the high-frame rate camera; The high-resolution tongue surface image and the high-resolution tongue base image are used as the high-resolution tongue diagnosis image, and the tongue surface dynamic video and the tongue base dynamic video are used as the high-frame rate tongue diagnosis dynamic video.

4. The method according to claim 3, characterized in that The method further comprises: Performing clarity detection on the high-resolution tongue surface image and the high-resolution tongue base image based on grayscale variance; In the case that the high-resolution tongue surface image and / or the high-resolution tongue base image fails the clarity detection, the corresponding step of returning to execute the step of initiating the voice prompt of the tongue surface movement for the user's tongue surface or the step of initiating the voice prompt of the tongue base movement for the user's tongue base is performed.

5. The method according to claim 1, wherein The calibrating the high-resolution tongue diagnosis image and the high-frame-rate tongue diagnosis dynamic video to obtain a spatially aligned target high-resolution tongue diagnosis image and a target high-frame-rate tongue diagnosis dynamic video includes: The high-resolution tongue diagnosis image and high-frame-rate tongue diagnosis dynamic video are calibrated using the dual-camera geometric projection formula: Performing calibration to obtain the target high-resolution tongue diagnosis image and the target high-frame-rate tongue diagnosis dynamic video that are spatially aligned; Wherein, K1 is the camera intrinsic parameter matrix of the high-resolution camera, K2 is the camera intrinsic parameter matrix of the high-frame rate camera, p1 is the pixel coordinate corresponding to the spatial point p in the image of the high-resolution camera, p2 is the pixel coordinate corresponding to the spatial point p in the image of the high-frame rate camera, and [R|t] is the rotation and translation relationship between the high-resolution camera and the high-frame rate camera.

6. The method according to claim 1, characterized in that After the spatial alignment, the method further includes: Performing U-Net-based tongue semantic segmentation on the target high-resolution tongue diagnosis image to generate a segmented image containing a tongue region mask; Performing tongue edge detection frame by frame on the target high-frame-rate tongue diagnosis dynamic video, and eliminating tongue edge breaks in the dynamic video through adaptive morphological closing operation; Among them, the tongue semantic segmentation adopts a shadow-aware attention mechanism, which dynamically adjusts the weight distribution of the segmentation network to the tongue reflective area by calculating the gradient histogram of the V channel of the HSV color space.

7. The method according to claim 6, characterized in that The method further comprises: The motion features of the segmented tongue dynamic video are extracted, including: The displacement vector field of the tongue area in adjacent frames is calculated using the optical flow method; A time-domain feature encoder is constructed based on the LSTM network. The displacement vector field of each frame and the corresponding tongue mask are channel-concatenated and then input into the encoder. The dynamic feature vector output by the encoding is fused with the texture feature vector of the high-resolution tongue diagnosis image at multiple scales to generate a joint feature representation for TCM syndrome differentiation.

8. A tongue diagnosis data collection device, characterized in that: include: A startup module is used to start the high-resolution camera and the high-frame rate camera in response to the tongue diagnosis data collection instruction, and call the pre-trained face and tongue joint detection model to perform face detection in a preset area; The acquisition module initiates action voice prompts and time voice prompts for the user's tongue when a face detection frame is detected in the preset area, and acquires high-resolution tongue diagnosis images and high-frame-rate tongue diagnosis dynamic videos determined based on the action voice prompts and the time voice prompts through the high-resolution camera and the high-frame-rate camera; calibrates the high-resolution tongue diagnosis images and high-frame-rate tongue diagnosis dynamic videos to obtain spatially aligned target high-resolution tongue diagnosis images and target high-frame-rate tongue diagnosis dynamic videos.

9. A computer device, characterized in that: The computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A tongue image segmentation method based on single target region segmentation

    CN109584251A

  • Tongue body automatic segmentation method based on U-net model

    CN111260619A

  • Tongue diagnosis detection method based on AI semantic segmentation and image recognition

    CN117541574A

  • Method for accurately segmenting and registering multispectral tongue body image

    CN119417848A

  • Tongue picture feature extraction and health assessment method based on deep learning

    CN119453950A