Tongue surface image recognition method and system based on three-dimensional face and tongue body reconstruction technology
By using 3D face and tongue reconstruction technology and employing a multimodal acquisition device and RGB-D array channels, a tongue simulation model was constructed. This solved the problem that static acquisition could not capture dynamic changes, and achieved the accuracy of tongue point cloud data and the comprehensiveness of tongue image information.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, static acquisition cannot effectively capture the dynamic changes of the tongue, resulting in inaccurate tongue point cloud data and affecting the accuracy of tongue image information.
By using 3D face and tongue reconstruction technology, a multimodal acquisition device is used to input 3D face point cloud data, capture face motion anchor points, identify tongue point cloud data, and connect them through RGB-D array channels to construct a tongue simulation model, extract tongue surface features in different areas, and form a detailed set of tongue surface features in different areas.
This ensured the accuracy of point cloud data, obtained more comprehensive tongue image information, and improved the accuracy and completeness of tongue feature analysis.
Smart Images

Figure CN121147993B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a tongue surface image recognition method and system based on three-dimensional face and tongue body reconstruction technology. BACKGROUND
[0002] Tongue surface information collection mainly relies on static collection methods, which obtain tongue surface features through single-frame images or two-dimensional images with limited angles. Static images can only reflect the state of the tongue body at a certain moment, while the tongue body has slight dynamic changes in the natural state. These dynamic information is crucial for a comprehensive and accurate description of the real state of the tongue surface. Static collection cannot effectively capture dynamic changes, resulting in inaccurate tongue point cloud data, which affects the analysis of tongue shape and color changes.
[0003] In summary, the prior art has the technical problem of inaccurate tongue point cloud data due to the inability of static collection to effectively capture dynamic changes, which affects the accuracy of tongue information. SUMMARY
[0004] The purpose of the present application is to provide a tongue surface image recognition method and system based on three-dimensional face and tongue body reconstruction technology to solve the technical problem of inaccurate tongue point cloud data due to the inability of static collection to effectively capture dynamic changes, which affects the accuracy of tongue information in the prior art.
[0005] In view of the above problems, the present application provides a tongue surface image recognition method and system based on three-dimensional face and tongue body reconstruction technology.
[0006] In a first aspect, the present application provides a tongue surface image recognition method based on three-dimensional face and tongue body reconstruction technology, which is realized by a tongue surface image recognition system based on three-dimensional face and tongue body reconstruction technology. The tongue surface image recognition method based on three-dimensional face and tongue body reconstruction technology comprises: inputting three-dimensional face point cloud data of a current user according to a multi-modal collection device; capturing a face action anchor point of the user according to the three-dimensional face point cloud data, and recording the face action anchor point for archiving; when the user performs an action according to a set tongue body collection action, identifying a tongue body point cloud data based on the face action anchor point; connecting the tongue body point cloud data with an RGB-D array channel, outputting tongue body point cloud data processed based on RGB parameters, and constructing a tongue body simulation model; extracting partition tongue surface features based on the tongue body simulation model, outputting a partition tongue surface feature set, and transmitting the partition tongue surface feature set to a management end of an electronic medical record corresponding to the user.
[0007] Optionally, the facial action anchor point is a key action anchor point of the face accompanying tongue action, the facial action anchor point is obtained by capturing changes of the three-dimensional face point cloud data about the anchor point region through a deep learning model, wherein the anchor point region at least includes a perioral region and a mandibular region; time sequence information and spatial parameters of each facial action anchor point in the perioral region and the mandibular region are recorded, the time sequence information includes action start time and time sequence time, and the spatial parameters include displacement vector and angle change; a joint motion trajectory of the perioral region and the mandibular region is obtained through time sequence modeling, and tongue point cloud data is identified according to the joint motion trajectory.
[0008] Optionally, when the user performs an action according to a set tongue collection action, it is judged whether the joint motion trajectory is in a stable state at the current time sequence; if the spatial parameter change rate of the joint motion trajectory at the current time sequence is greater than or equal to a preset threshold, a tongue collection instruction is triggered, and tongue point cloud data is obtained according to the tongue collection instruction; if the spatial parameter change rate of the joint motion trajectory at the current time sequence is less than the preset threshold, the tongue collection instruction is not triggered.
[0009] Optionally, the change correlation of the facial action anchor point is calculated, and the weight attribute of the facial action anchor point is configured according to the change correlation; the weight attribute is sent to the multi-modal collection device, and the multi-modal collection device performs collection view angle adjustment according to the weight attribute of the facial action anchor point.
[0010] Optionally, after the multi-modal collection device performs collection view angle adjustment according to the weight attribute of the facial action anchor point, the multi-modal collection device starts a magnification collection instruction; the focal length of an imaging device in the multi-modal collection device is adjusted according to the magnification collection instruction, fine-grained tongue point cloud data is obtained, and the tongue simulation model is updated according to the fine-grained tongue point cloud data.
[0011] Optionally, a face model is constructed according to the three-dimensional face point cloud data; an alignment constraint condition is constructed, the alignment constraint condition includes key point matching and normal vector consistency; the face model and the tongue simulation model are aligned in the same three-dimensional coordinate system under the alignment constraint condition, a face-tongue model is output, and the facial action anchor point is recalculated and updated according to the face-tongue model.
[0012] Optionally, facial features are identified according to the face model, a partitioned face feature set is obtained, the partitioned face feature set is added to a partitioned tongue surface feature set as auxiliary feature data, and the partitioned face feature set and the partitioned tongue surface feature set are transmitted to a management end of an electronic medical record corresponding to the user.
[0013] Optionally, tongue surface partition feature extraction is performed based on the tongue body simulation model, and a partitioned tongue surface feature set is output, the partitioned tongue surface feature set at least including a tongue surface feature set of a first region, a tongue surface feature set of a second region, a tongue surface feature set of a third region, and a tongue surface feature set of a fourth region; wherein the first region is a tongue tip region, the second region is a tongue middle region, the third region is a tongue root region, and the fourth region is a tongue lateral margin region.
[0014] Optionally, multi-modal features are constructed, including shape features, texture features, color features, and three-dimensional features; and tongue surface partition feature extraction is performed based on the multi-modal features according to the tongue body simulation model, and a partitioned tongue surface feature set is output.
[0015] In a second aspect, the present application also provides a tongue surface image recognition system based on three-dimensional face and tongue body reconstruction technology, which is used to execute the tongue surface image recognition method based on three-dimensional face and tongue body reconstruction technology as described in the first aspect, wherein the tongue surface image recognition system based on three-dimensional face and tongue body reconstruction technology includes: a face data input module, which is used to input three-dimensional face point cloud data of a current user according to a multi-modal acquisition device; a tongue body data recognition module, which is used to capture face action anchor points of the user according to the three-dimensional face point cloud data, and record and archive the face action anchor points, and when the user performs an action according to a set tongue body acquisition action, tongue body point cloud data is recognized based on the face action anchor points; a simulation model construction module, which is used to connect the tongue body point cloud data with an RGB-D array channel, output tongue body point cloud data processed based on RGB parameters, and construct a tongue body simulation model; and a tongue surface feature extraction module, which is used to perform tongue surface partition feature extraction based on the tongue body simulation model, output a partitioned tongue surface feature set, and transmit the partitioned tongue surface feature set to a management terminal of an electronic medical record corresponding to the user.
[0016] One or more technical solutions provided in the present application have at least the following beneficial effects:
[0017] The process involves: inputting the user's 3D facial point cloud data using a multimodal acquisition device; capturing and archiving the user's facial motion anchor points based on the 3D facial point cloud data; obtaining tongue point cloud data based on the facial motion anchor points when the user performs a set tongue action; connecting the tongue point cloud data to an RGB-D array channel to output tongue point cloud data processed with RGB parameters, thus constructing a tongue simulation model; extracting tongue surface features based on the tongue simulation model, outputting a set of tongue surface features for each region, and transmitting this set of tongue surface features to the management terminal of the user's corresponding electronic medical record. In other words, the process involves inputting the user's 3D facial point cloud data through image acquisition and a 3D point cloud model, recording the user's facial motion anchor points through facial motion capture, dynamically recognizing the tongue point cloud data based on the facial motion images, processing the RGB parameters to construct a tongue simulation model, dividing the tongue surface into regions and extracting features to form a detailed set of tongue surface features for each region, ensuring the accuracy of the point cloud data, and thus obtaining more comprehensive tongue image information.
[0018] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the tongue image recognition method based on 3D face and tongue reconstruction technology proposed in this application.
[0021] Figure 2 This is a schematic diagram of the tongue surface image recognition system based on three-dimensional face and tongue reconstruction technology in this application.
[0022] Figure labeling: 11. Face data entry module, 12. Tongue data recognition module, 13. Simulation model construction module, 14. Tongue surface feature extraction module. Detailed Implementation
[0023] The tongue surface image recognition method and system based on the three-dimensional face and tongue body reconstruction technology are provided, and the technical problem that the tongue body point cloud data is inaccurate due to the fact that the static acquisition cannot effectively capture the dynamic changes, thereby affecting the accuracy of the tongue information in the prior art is solved. The three-dimensional face point cloud data of the user is input through image acquisition and three-dimensional point cloud model, and the action anchor point of the user's face is recorded through facial motion capture. The tongue body point cloud data is dynamically identified according to the facial motion image, the tongue body simulation model is constructed after RGB parameter processing, the tongue surface is partitioned and the features are extracted, the detailed partitioned tongue surface feature set is formed, the accuracy of the point cloud data is ensured, and more comprehensive tongue information is obtained.
[0024] Hereinafter, the technical solutions in the present application will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application. In addition, it should be noted that, for the convenience of description, only the parts related to the present application are shown in the drawings, not all.
[0025] Embodiment one, please refer to the accompanying Figure 1 The tongue surface image recognition method based on the three-dimensional face and tongue body reconstruction technology is provided, wherein the tongue surface image recognition method based on the three-dimensional face and tongue body reconstruction technology is executed by the tongue surface image recognition system based on the three-dimensional face and tongue body reconstruction technology. The tongue surface image recognition method based on the three-dimensional face and tongue body reconstruction technology specifically includes the following steps:
[0026] S100: Input the three-dimensional face point cloud data of the current user according to the multi-modal acquisition device.
[0027] Specifically, before scanning the device, the user keeps the head stable, the face information collection device starts, and the face information is collected through the built-in multi-modal collection device, so as to construct the three-dimensional face point cloud data of the current user, and accurately represent the geometric features of the face, such as the shape of the nose bridge, eye socket, and mouth. The multi-modal collection device refers to a device that can simultaneously collect multiple data types, including depth perception cameras, infrared sensors, etc., which can simultaneously obtain visible light images and three-dimensional point cloud data, such as three-dimensional geometric data, color information, depth information, etc., to improve the comprehensiveness and accuracy of data collection. The scanning device emits structured light or laser and receives the reflected signal through the camera sensor in the multi-modal collection device, calculates the depth information of each pixel point, and generates three-dimensional point cloud data of the face. During the collection process, the scanning device usually needs to be scanned from multiple angles (front, side, top, etc.) to ensure complete recording of the geometric information of the face and reduce the occlusion error. Alternatively, the device adopts a multi-angle cooperative collection mechanism: the front main camera captures the core features of the nose bridge, eye socket, mouth, and tongue, the two side auxiliary sensors capture the facial contour curve, the top view module records the forehead and chin shape, and the bottom view module records the chin and lower part of the tongue. The multi-source data fusion method can effectively eliminate the face occlusion error under single view, and the geometric error of the reconstructed three-dimensional model is controlled within ±0.1 mm.
[0028] Due to the possibility of noise (such as light interference or device jitter) during the collection process, point cloud filtering needs to be used for optimization, such as through statistical outlier filtering to remove noise points or outliers from the data set. By calculating the mean and standard deviation of the distance between each point and its neighborhood points, it is determined whether the point deviates from the statistical distribution of its neighborhood, so as to mark it as an outlier and remove it, ensuring the accuracy and consistency of the face information.
[0029] Three-dimensional face point cloud data refers to the three-dimensional coordinate data of the face captured by the multi-modal collection device, a data set composed of a large number of three-dimensional coordinate points (x, y, z), used to describe the shape and structure of the face. By entering the three-dimensional face point cloud data through the multi-modal collection device, a high-precision and high-resolution three-dimensional face model is obtained, which provides real and detailed three-dimensional face data for subsequent face motion capture, facial feature analysis, etc., thereby improving the accuracy and usability of the overall data.
[0030] S200: Capture the face motion anchor point of the user according to the three-dimensional face point cloud data, and record the face motion anchor point, when the user performs the action according to the set tongue collection action, identify the tongue point cloud data based on the face motion anchor point.
[0031] Further, the S200 of the present application comprises:
[0032] The face action anchor point is a key action anchor point of the face accompanying tongue movement, and is obtained by capturing changes of the three-dimensional face point cloud data with respect to an anchor point region through a deep learning model, wherein the anchor point region at least includes a perioral region and a mandibular region; time sequence information and spatial parameters of each face action anchor point in the perioral region and the mandibular region are recorded, the time sequence information includes action start time and time sequence time, and the spatial parameters include a displacement vector and an angle change; a joint motion trajectory of the perioral region and the mandibular region is obtained through time sequence modeling, and tongue point cloud data is identified according to the joint motion trajectory.
[0033] Further, the present application further includes the following steps:
[0034] When the user performs an action according to the set tongue collection action, it is judged whether the joint motion trajectory is in a stable state at the current time sequence; if the spatial parameter change rate of the joint motion trajectory at the current time sequence is greater than or equal to a preset threshold, a tongue collection instruction is triggered, and tongue point cloud data is obtained according to the tongue collection instruction; if the spatial parameter change rate of the joint motion trajectory at the current time sequence is less than the preset threshold, the tongue collection instruction is not triggered.
[0035] Specifically, three-dimensional face point cloud data of a user is collected by a multi-modal collection device, and when the perioral region and the mandibular region change, the time sequence information and the spatial parameters of each face action anchor point of the perioral region and the mandibular region are recorded in real time. In other words, when the user performs tongue movement, such as tongue extension, left and right swinging, etc., the key action anchor points of the face accompanying tongue movement are obtained, ensuring that the anchor points of the perioral region and the mandibular region can be recorded completely, and the recorded data is archived. The face action anchor point is a specific point that can reflect the key motion characteristics of the face when performing a specific action (such as speaking, chewing, tongue extension, etc.), that is, a key point in a specific region of the face that changes significantly when there is an action. For example, when the tongue is extended or moved, the lips, mandible, cheeks and other parts will deform and displace accordingly, and the key points of these positions are called face action anchor points.
[0036] A deep learning model is applied to the three-dimensional face point cloud data of the user, focusing on the key anchor points in the perioral region and the mandibular region, automatically identifying the geometric features of these regions, and outputting the spatial coordinates of the anchor points. The anchor point region refers to the representative key parts on the face, which are used to track facial movements, including the perioral region, the mandibular region, etc. The perioral region includes the lips, the corners of the mouth and other parts, which will produce obvious displacement when the tongue moves or the facial expression changes; the mandibular region mainly refers to the chin and its surrounding area, and its movement is usually accompanied by dynamic changes of the whole face, such as upward or downward movement of the mandible.
[0037] In the perioral region, the deep learning model extracts the corners of the mouth and the center of the lips as anchor points; in the mandibular region, the center of the chin and the key points of the mandibular edge are extracted. The position of each anchor point can be represented by (x, y, z) coordinates, and the displacement and rotation angle of these points at different time points are the focus of dynamic capture. A labeled dataset is obtained, which includes a large amount of three-dimensional face point cloud data and labeled key anchor points in the perioral and mandibular regions, including the corners of the mouth, the centers of the upper and lower lips, the center of the chin, and the edge of the mandibular edge. The labeled dataset is input into a multi-layer neural network structure, which extracts features from each input point cloud point by point, and aggregates global information using symmetric functions such as max pooling. The network learns how to distinguish which points belong to the perioral or mandibular region by continuously optimizing parameters. During training, the model's output (such as predicted key point coordinates or region segmentation map) is compared with the true value labeled by humans, and the network parameters are adjusted through error feedback. Training is stopped until the model's accuracy reaches 98% or above, and the deep learning model is obtained.
[0038] The input three-dimensional face point cloud data is input into the trained deep learning model, which automatically extracts geometric features from the point cloud through its multi-layer neural network structure, and internally divides the region to identify which parts belong to the perioral, mandibular, and other anchor point regions. The user performs actions, and the face moves with the tongue, continuously collecting three-dimensional face point cloud data, and real-time capturing and tracking the anchor points in the continuous frame data, outputting the displacement and angle change data of the anchor points over time. Specifically, multi-modal acquisition devices are used to record continuous frame data in a dynamic environment. In each frame, the pre-trained deep learning model automatically locates the key anchor points in the perioral and mandibular regions. For example, in the mouth region, the model identifies the corners of the mouth, the boundary points of the upper and lower lips; in the mandibular region, the center of the chin and the edge of the mandibular edge are identified. Through the changes in the positions of these key points, the motion state of the face can be described.
[0039] For each face action anchor point in the perioral and mandibular regions, record the timing information and spatial parameters, the timing information refers to the time sequence information of the action, including the start time and duration of the action, the spatial parameters refer to the position and direction information of each face action anchor point in three-dimensional space, such as displacement vector and angle change. For example, when the face starts to smile or open the mouth, mark this moment; if the face action lasts for 0.5 seconds, record the time sequence data from 0 seconds to 0.5 seconds. At each time node, record the translation of the key action anchor point in the X, Y, Z directions relative to the previous time point, as well as the angle change, to determine the direction and amount of movement of each point, and determine the attitude change of the action.
[0040] According to the time sequence information and spatial parameters of each facial action anchor point, the motion data of the perioral area and the mandibular area are first determined, the time sequence modeling technology is used to model the changes of each anchor point in the continuous frame data, and the motion trajectory is constructed. By the joint modeling method, the motion correlation between the two areas is considered. For example, when the corners of the mouth are displaced, the mandible often adjusts the position accordingly, and the motion of the two forms a linkage mechanism. In the specific execution process, the displacement vector and the angle change at the continuous time points are taken as the input, the dynamic relationship between each time point is learned, and an overall joint motion trajectory is generated. The trajectory can intuitively show the motion path of each facial action anchor point in the whole action process.
[0041] After time sequence modeling, a trajectory curve describing the joint motion of the perioral area and the mandibular area is generated. For example, in an experiment, when the subject performs a smiling action, the key anchor points of the corners of the mouth and the mandible are displaced and angularly changed from the static state within 0.5 seconds. Experimental data show that the average displacement of the corners of the mouth increases by about 2.5 mm, and the angular change of the mandible reaches about 4°. The entire joint trajectory shows a highly consistent motion trend in dynamic capture.
[0042] Exemplarily, when the user 11:15:31 is in the action of opening the mouth and stretching the tongue, the time sequence information and spatial parameters of the key action anchor points of the perioral area and the mandibular area are captured from the continuous three-dimensional face point cloud data, and the whole process lasts 0.8s. The perioral area anchor point record (taking the upper lip center as an example) is shown in Table 1, the mandibular area anchor point record (taking the lower jaw center as an example) is shown in Table 2, and the joint motion trajectory of the perioral area and the mandibular area is shown in Table 3.
[0043] Table 1 Perioral area anchor point record
[0044]
[0045]
[0046] The starting time 11:15:31 is when the action just starts, the upper lip center position is taken as the reference, the displacement vector is zero, and the angle has no change; at 11:15:35, the upper lip center moves 1.2 mm in the right and upward direction during the action process, and the angle changes slightly by 2°, indicating that the upper lip starts to move upward due to the opening of the mouth; at 11:15:39, the action is close to completion, the upper lip center further moves to the (2.5, -1.0, 0.7) mm position, and the angle change accumulates to 4°, reflecting a clear motion trend.
[0047] Table 2 Mandibular area anchor point record
[0048] Time Timing time (s) Displacement vector (x, y, z) (mm) Angle change (°) 11:15:31 0.0 (0.0,0.0,0.0) 0 11:15:35 0.4 (1.8,1.0,-0.4) 1 11:15:39 0.8 (1.6,2.0,-0.8) 2
[0049] The initial time 11:15:31 is when the action just starts, the center of the chin is in the initial state without displacement and angle change; at 11:15:35, the center of the chin starts to move slightly to the lower right direction (for example, x increases by 0.8 mm, y increases by 1.0 mm, and z direction changes by-0.4 mm, indicating fine adjustment of rear or forward movement), and the angle is slightly rotated by 1°; at 11:15:39, after the action, the center of the chin accumulates to (1.6, 2.0, -0.8) mm, and the angle changes to 3°, indicating that the whole mandibular region presents a downward and outward movement trend.
[0050] Table 3 Joint motion trajectory of perioral region and mandibular region
[0051]
[0052]
[0053] At the start of the action, both the perioral region and the mandibular region are in the initial static state; as the action proceeds, the perioral region (such as the center of the upper lip) gradually moves upward, while the mandibular region (such as the center of the chin) exhibits downward movement, which is a typical feature of the action of opening the mouth and sticking out the tongue.
[0054] According to the joint motion trajectory of the user when performing the action, the current state of the user is identified, whether it is closed or open and stick out the tongue, and whether to start collecting tongue point cloud data is determined accordingly. When the user performs the action according to the pre-set tongue collection action, the multi-modal collection device monitors the time sequence information and spatial parameters of each facial action anchor point of the user in real time, and comprehensively obtains the joint motion trajectory of each anchor point of the perioral region and the mandibular region, reflecting the overall motion state of the whole face during the action. The set tongue collection action is a specified action, that is, the user is required to perform during data collection, so as to capture the specific state of the tongue, which usually refers to opening the mouth and sticking out the tongue.
[0055] It is judged whether the joint motion trajectory is in a stable state at the current time sequence, that is, whether the spatial parameter change rate of the joint motion trajectory at the current time sequence exceeds the pre-set threshold. If the spatial parameter change rate of the joint motion trajectory at the current time sequence is greater than or equal to the pre-set threshold, the tongue collection instruction is triggered, and the point cloud data of the tongue is collected. If the spatial parameter change rate of the joint motion trajectory at the current time sequence is less than the pre-set threshold, the joint motion trajectory of the perioral region and the mandibular region may hardly change, indicating that the facial motion is in a stable state, and the current state of the user is closed, so the tongue collection instruction is not triggered, thereby avoiding collecting invalid or incorrect tongue data.
[0056] The preset threshold is a pre-set judgment standard to determine whether it is a dynamic state and whether it can start collecting tongue information of the user. The spatial parameter change rate is the speed of displacement and angle change of the key area of the face in unit time, reflecting the dynamic degree of the action, which is calculated according to the time sequence information and the spatial parameter. For example, when the displacement of any anchor point is greater than or equal to 5 mm or the angle change is greater than or equal to 15°, it is considered that the current user state is mouth opening and tongue stretching, the anchor point changes, and therefore the tongue collection instruction is triggered. By setting a reasonable preset threshold, dynamic actions and static states can be effectively distinguished, and false data collection when the mouth is closed or the action is stable can be avoided, thereby improving the overall robustness and data quality.
[0057] By using real-time monitoring of the joint motion trajectory of the face and automatic judgment of the spatial parameter change rate, accurate collection decisions are realized when the user performs the tongue collection action, which significantly improves the collection accuracy and real-time performance of the tongue point cloud data, reduces unnecessary collection, and avoids data errors caused by unstable tongue action.
[0058] Further, the application also includes the following steps:
[0059] The change correlation of the face action anchor points is calculated, and the weight attribute of the face action anchor points is configured according to the change correlation; the weight attribute is sent to the multi-modal collection device, and the multi-modal collection device adjusts the collection angle according to the weight attribute of the face action anchor points.
[0060] Specifically, according to the time sequence information and the spatial parameter of each face action anchor point, the change correlation between different anchor points in the action process is evaluated by statistical or mathematical methods (such as correlation coefficient calculation). The change correlation refers to the degree of correlation between the changes of different face action anchor points in the action process, that is, how the change of one anchor point affects or correlates to the change of other anchor points. For example, when a person opens his mouth, the movement of the corners of the mouth and the chin may be highly correlated. The Pearson correlation coefficient is an index for measuring the degree of linear correlation between two variables, and its value ranges from -1 to 1.
[0061] First, according to the change data of each action anchor point at multiple time points, including the displacement vector or angle change of each anchor point at each time point. For each pair of anchor points, the Pearson correlation coefficient formula is used to calculate the correlation between them. The closer the value of the calculated correlation coefficient is to 1, the more positively correlated the changes between the two anchor points are; the closer the value is to -1, the more negatively correlated they are; and the closer the value is to 0, the less linearly correlated they are. According to the calculated correlation coefficient, a weight is assigned to each anchor point.
[0062] The higher the weight of the anchor point, the higher the importance of the anchor point in the whole facial motion, which may play a key role in motion recognition and tongue acquisition. The weight configuration is usually configured according to a preset rule by changing the correlation, for example, the weight of the anchor point with a correlation greater than 0.8 is set to 1.0, the weight of the anchor point with a correlation between 0.5 and 0.8 is set to 0.7, and the weight of the anchor point with a correlation less than 0.5 is set to 0.5. The weight attribute is a numerical value determined according to the change correlation, which reflects the importance of each anchor point in subsequent motion capture and data processing. For example, anchor points with higher correlation may be assigned higher weights because they play a key role in motion recognition.
[0063] The configured weight attribute is sent to the multi-modal acquisition device through a data interface. After receiving the weight information, the multi-modal acquisition device adjusts the acquisition view angle according to the position and change of the anchor point with high weight. The algorithm built-in the device calculates the position distribution of the current anchor point in the acquisition picture in real time. If it is detected that the anchor point with high weight is at the edge of the picture or may be blocked, the shooting angle or focal length of the camera is adjusted to ensure that all key anchor points are within the optimal acquisition view angle.
[0064] The multi-modal acquisition device adjusts the acquisition view angle according to the weight attribute of the face motion anchor point to ensure that all anchor points are within the acquisition view angle, avoiding anchor point loss or data deviation caused by poor view angle. For example, when it is detected that a certain anchor point has a high weight, the camera may fine-tune the direction or focal length to better capture the details of the anchor point, ensuring data integrity.
[0065] For example, when the user performs the action of opening the mouth and sticking out the tongue, the correlation coefficient of the center of the upper lip is 0.88, the weight is assigned as 1.0, the upper lip is at the edge of the picture, and the camera angle is recommended to be adjusted upward; the correlation coefficient of the left corner of the mouth is 0.82, the weight is assigned as 1.0, the left edge deviates from the center, and the view angle is recommended to be moved to the left; the correlation coefficient of the right corner of the mouth is 0.80, the weight is assigned as 1.0, the right side is slightly deviated, and the view angle is recommended to be moved to the right; the correlation coefficient of the center of the chin is 0.65, the weight is assigned as 0.7, the position is basically in the center, and no adjustment is recommended; the correlation coefficient of the edge of the lower jaw is 0.50, the weight is assigned as 0.5, and the edge is located, but the weight is low, and no adjustment is required.
[0066] By calculating the change correlation of the face motion anchor point and sending the weight attribute to the multi-modal acquisition device according to the configuration, the most representative areas are focused on during dynamic acquisition, and the accuracy of facial motion capture is improved. The camera view angle is dynamically adjusted to keep all key anchor points within the optimal acquisition area, reducing data loss or noise caused by view angle deviation, and ensuring that the acquired point cloud data is complete and clear.
[0067] S300: Connect the tongue point cloud data with the RGB-D array channel, output the tongue point cloud data processed based on RGB parameters, and construct a tongue simulation model.
[0068] Specifically, when the user performs an action, the tongue point cloud data, which is a set of three-dimensional coordinate points on the tongue surface, is collected by a multi-modal acquisition device. This data includes X, Y, Z three-dimensional coordinate information and may also include color or depth values, which are used to reconstruct the three-dimensional shape of the tongue. The collection angle should comprehensively cover the tongue, including tongue color, tongue shape, moss color, moss quality, fluid, point stimulation, blood stasis points and patches, peeling, tooth marks, crack quantity and location, and sublingual plexus conditions, etc.
[0069] RGB-D refers to the combination of red, green, and blue (RGB) color information and depth (Depth) information acquisition method, usually provided by RGB-D sensors, which can capture color and depth information simultaneously, making three-dimensional reconstruction more accurate. RGB-D array channel refers to the collection of data channels collected by multiple RGB-D sensors. These channels can provide multi-angle, multi-view tongue data, which helps to improve the completeness and accuracy of the point cloud data.
[0070] By using RGB-D array channel, color (RGB) and depth (D) data of the tongue are simultaneously acquired, and three-dimensional coordinates and corresponding color information of the tongue are accessed. Multiple angle RGB-D sensors (such as 0°, 30°, 60°, 90°) are set to avoid data loss caused by a single angle. RGB parameters are used to process tongue point cloud data, including adjusting color balance, contrast, and brightness, etc., to enhance the visualization effect of the tongue. RGB color values are adjusted to a standard range to eliminate deviations under different lighting conditions; tongue color contrast is adjusted to enhance the tongue surface boundary, making it clearer; RGB information is mapped to three-dimensional point cloud data to ensure color and shape matching.
[0071] Using the processed tongue point cloud data, a tongue simulation model is constructed. This model is a three-dimensional representation that shows the shape, size, and color of the tongue. The tongue simulation model is a three-dimensional digital model of the tongue constructed using computer graphics and modeling algorithms such as surface reconstruction, texture mapping, and mesh optimization. The specific process is as follows: Since the data collected at different angles may have overlapping areas, redundant points need to be removed. A downsampling algorithm in three-dimensional modeling software is used to remove redundant points. Color information on the tongue surface is mapped to the three-dimensional model to make it more realistic. Combined with the motion characteristics of the tongue, dynamic behavior of the tongue is simulated through animation technology. The optimized RGB color information is projected onto the three-dimensional model to achieve realistic texture mapping.
[0072] Through multi-angle acquisition by the RGB-D array channel, the problem of single-angle data loss is avoided, the integrity of the tongue point cloud data is ensured, and the tongue simulation model is closer to the real shape in combination with the RGB parameter processing. A high-precision tongue simulation model is constructed, the three-dimensional structure of the tongue is displayed, and accurate surface color information is included.
[0073] Further, the S300 of the present application comprises:
[0074] According to the three-dimensional face point cloud data, a face model is constructed; an alignment constraint condition is constructed, which includes key point matching and normal vector consistency; the face model and the tongue simulation model are aligned in the same three-dimensional coordinate system under the alignment constraint condition, and a face-tongue model is outputted. According to the face-tongue model, the face action anchor point is recalculated and updated.
[0075] Specifically, similar to the aforementioned tongue simulation model construction, according to the input three-dimensional face point cloud data, a face model is constructed. The three-dimensional face point cloud data must include front and side views. After processing (denoising, downsampling, surface fitting, etc.) of the three-dimensional face point cloud data, the cloud processing software locates the key points such as the eye corner, the nose tip, and the mandibular angle. A complete 3D grid model is constructed from the three-dimensional face point cloud data, and the construction of the face model is completed.
[0076] The alignment constraint condition is a series of rules for ensuring the correct alignment of the face model and the tongue simulation model, including key point matching and normal vector consistency. Key point matching is to align the model by matching the feature points (such as the nose tip, the eye corner, and the mandibular angle). For example, the nose tip, the mandibular angle, and other landmark points of the face model are matched with the tongue tip, the tongue root, and other landmark points of the tongue simulation model. Normal vector consistency ensures that the surface normal vectors of the two models are in the same direction, eliminating local deformation errors. The normal vector is a vector perpendicular to the model surface. Maintaining normal vector consistency can ensure the smoothness and continuity of the model surface.
[0077] According to the alignment constraint condition, the face model and the tongue simulation model are accurately aligned in the same coordinate system. Using the alignment constraint condition, the relative position and direction of the two models in space are ensured to be correct, thereby creating an accurate face-tongue model. Through landmark point matching, the two models are accurately aligned, ensuring that the nose tip and the tongue tip, the mandibular angle and the tongue root, and other landmark points are correctly matched, and the surface normal vectors of the two models are in the same direction. If there is a normal vector deviation, the angle deviation is calculated, and the model normal vector direction is adjusted to make it consistent.
[0078] After aligning the face model and the tongue simulation model, the face action anchor points are recalculated and updated. Since the pose and position of the face model may have changed during the alignment process, the anchor points need to be updated to reflect these changes. The changes in the face key points after alignment are calculated, and the anchor point weights are adjusted to reflect the relative position and dynamic relationship between the face and the tongue. By aligning the face model and the tongue simulation model with the alignment constraints, a face-tongue model is output, accurately analyzing and understanding the interaction and dynamic relationship between the face and the tongue, and improving the restoration degree of the model to the real biological structure.
[0079] Further, the present application also includes the following steps:
[0080] According to the face model, the facial features are identified, and a partitioned face feature set is obtained. The partitioned face feature set is added as auxiliary feature data to the partitioned tongue surface feature set. The partitioned face feature set and the partitioned tongue surface feature set are transmitted to the management end of the user's electronic medical record.
[0081] Specifically, according to the face model, the facial features are identified, and the face model is divided into specific regions (such as eyes, lip color, ears, nose, eyebrows, cheeks, etc.), and the feature data of these regions is extracted, including skin color, surface texture, deformation information, etc. For example, the eyes include eye socket depth, eye slit size, eyelid state, etc., the lip color includes the color and state of the lips, the ears include the ear wheel shape, the earlobe state, etc., the nose includes the nose bridge height, the nose wing width, etc., the eyebrows include the inter-brow wrinkles, the skin state, etc., and the cheeks include the shape, color and skin condition of the cheeks, etc.
[0082] According to the tongue simulation model, the tongue surface is partitioned into the tongue tip area, the tongue middle area, the tongue root area, and the tongue lateral margin area, and the corresponding features are extracted to obtain a partitioned tongue surface feature set. Because the tongue surface features and the face features are correlated in some cases, the partitioned face feature set is used as auxiliary feature data and fused with the partitioned tongue surface feature set. The correlation between different regional features (such as the similarity of lip color and tongue tip color) is calculated, and a weighted average model is used for data fusion, giving different weights to different regional features. The partitioned face feature set and the partitioned tongue surface feature set are transmitted to the user's electronic medical record management end through a secure network connection for recording and analyzing the user's face and tongue surface feature information for subsequent analysis.
[0083] The management end of the electronic medical record stores the partitioned face feature set and the partitioned tongue surface feature set of the user, so that the doctor can obtain a more comprehensive patient biological feature view. By adding facial features as auxiliary information, the recognition effect of tongue surface features is optimized, the error of single modal data is reduced, and the accuracy and reliability of tongue surface feature recognition are improved.
[0084] Further, the present application further comprises the following steps:
[0085] After the multi-modal acquisition device adjusts the collection angle according to the weight attribute of the facial action anchor point, the multi-modal acquisition device starts the zoom-in collection instruction; adjusts the focal length of the imaging device in the multi-modal acquisition device according to the zoom-in collection instruction, acquires fine-grained tongue point cloud data, and updates the tongue simulation model according to the fine-grained tongue point cloud data.
[0086] Specifically, according to the weight attribute of the facial action anchor point, the multi-modal acquisition device adjusts the working angle and direction of its camera or sensor, ensures that all important anchor points are within the collection angle, and pays more attention to the key area (such as the tongue). Once the collection angle adjustment is completed, the multi-modal acquisition device will receive the instruction to start zoom-in collection for acquiring more detailed tongue data. The zoom-in collection instruction means that after the preliminary collection is completed, the device further narrows the field of view, increases the resolution, and improves the collection accuracy of the target area (such as the tongue surface details).
[0087] According to the zoom-in collection instruction, the focal length of the imaging device (such as the camera) in the multi-modal acquisition device is adjusted to collect the tongue area with higher resolution. For example, the overall tongue data is acquired at a wide angle, and the high-precision tongue surface point cloud data is acquired at a narrow angle. Increasing the focal length can make the imaging device focus more on the tongue, so as to acquire fine-grained tongue point cloud data, which contains more details of the tongue, such as texture, color change, and slight concave-convex. Compared with the coarse-grained point cloud data acquired in the preliminary collection, the fine-grained point cloud data has higher point cloud density and more accurate surface information. For example, the preliminary collection may only be able to distinguish the outline of the tongue, while the fine-grained data can capture microscopic features such as tongue coating thickness and surface cracks.
[0088] According to the newly acquired fine-grained tongue point cloud data, the tongue simulation model is optimized, and the updated tongue simulation model will contain more details of the tongue. The fine-grained tongue point cloud data is aligned with the original tongue simulation model, and the corresponding data is adjusted to more accurately reflect the actual shape and characteristics of the tongue. At the same time, through the normal vector consistency correction, the local collection error is eliminated, and the tongue surface is made smoother. Through zoom-in collection and zoom control, high-resolution point cloud data collection is realized, so that the tongue simulation model is more accurate. Dynamic angle adjustment is adopted to ensure that all key areas are collected, and the integrity of the face and tongue data fusion is improved.
[0089] S400: Extracting tongue surface partition features based on the tongue simulation model, outputting a partition tongue surface feature set, and transmitting the partition tongue surface feature set to the management end of the electronic medical record of the user.
[0090] Further, the application S400 comprises:
[0091] Based on the tongue simulation model, tongue surface partition feature extraction is performed, and a partitioned tongue surface feature set is output, which at least includes a tongue surface feature set of a first area, a tongue surface feature set of a second area, a tongue surface feature set of a third area, and a tongue surface feature set of a fourth area; wherein the first area is a tongue tip area, the second area is a tongue middle area, the third area is a tongue root area, and the fourth area is a tongue lateral margin area.
[0092] Specifically, according to the constructed tongue simulation model, the tongue surface is divided into different areas, such as the tongue tip area, the tongue middle area, the tongue root area, and the tongue lateral margin area. For each area, corresponding features are extracted, such as the shape, size, color, etc. of the area, forming a partitioned tongue surface feature set. K-Means clustering + RGB color space conversion is used to extract tongue color; local binary model is used to analyze tongue fur texture, and thin fur, thick fur, and greasy fur are distinguished. Through PCA analysis of morphological principal components, tongue thickness, width, crack, etc. parameters are calculated to determine tongue shape; glossiness detection algorithm is used to analyze tongue surface reflectivity to determine whether the body fluid is sufficient.
[0093] The feature set of each area is output to form a partitioned tongue surface feature set, which includes four tongue surface feature sets, i.e. tongue tip area feature set, tongue middle area feature set, tongue root area feature set, and tongue lateral margin area feature set.
[0094] Tongue appearance is like a mirror of health, showing the health status of the body constitution and viscera at all times. According to the recognized tongue surface feature information, the user state is evaluated through the database. The tongue appearance feature content recognition project table is shown in Table 4:
[0095] Table 4 Tongue appearance feature content recognition project table
[0096]
[0097]
[0098] The partitioned tongue surface feature set is transmitted to the user's electronic medical record management terminal through a secure communication protocol (such as HTTPS, VPN, etc.), ensuring that the data transmission process complies with data protection regulations and user privacy protection requirements. The electronic medical record management terminal is a digital platform for storing, managing, and analyzing user health data, including storing patient tongue surface feature data, enabling long-term health trend analysis; providing tongue surface feature automatic comparison functions, combined with AI-assisted diagnosis functions. After each data entry by the user, the facial feature data and tongue surface feature data are automatically updated and compared with historical data to generate an analysis report. Doctors can view the trend of tongue surface feature changes through the electronic medical record management terminal to assist in diagnosis. By extracting tongue surface features based on a tongue simulation model, the partitioned tongue surface feature set is output, improving the accuracy and completeness of tongue information.
[0099] Further, the present application also includes the following steps:
[0100] Constructing multi-modal features, including shape features, texture features, color features, and three-dimensional features; the tongue simulation model extracts tongue surface features according to the multi-modal features, and outputs a partitioned tongue surface feature set.
[0101] Specifically, multi-modal features refer to features extracted from different sources or different types of data, used to extract shape features, texture features, color features, and three-dimensional features from the tongue simulation model. Shape feature analysis of the tongue's geometric shape, including but not limited to volume, surface area, curvature, convexity, perimeter, etc., helps identify changes in the shape of the tongue. Texture features include analyzing the surface texture of the tongue, such as roughness, spots, lines, etc., to extract features and reveal changes in the fine structure of the tongue surface. Color features extract color information from the tongue, including hue, saturation, brightness, etc., through the RGB-Lab distribution of tongue coating. Color changes may be related to certain conditions of the body. Three-dimensional features use three-dimensional data to extract features such as surface normal vector, volume density depth information, surface, three-dimensional shape description, etc., which help understand the three-dimensional structure of the tongue.
[0102] According to the anatomical structure of the tongue, the tip area, middle area, root area, and lateral margin area are defined. For each area, the corresponding shape features, texture features, color features, and three-dimensional features are extracted according to the multi-modal features, forming a partitioned tongue surface feature set. For shape features, curvature, area, perimeter, etc. are extracted, curvature is used to calculate local and global curvature of the tongue surface to judge the smoothness of the tongue surface; area is used to measure the projected area of the tongue to judge the thickness of the tongue; perimeter is used to calculate the boundary length of the tongue to assist in judging the shape of the tongue.
[0103] For texture features, the contrast, entropy value, homogeneity, etc. are calculated by a gray level co-occurrence matrix to extract the texture information of tongue fur, and to determine thin fur, thick fur, greasy fur, etc. High value of contrast represents rough texture, and low value represents smoothness; entropy value measures the complexity of texture information; homogeneity describes the uniformity of tongue fur distribution. For color features, the overall color distribution of tongue body is analyzed by tongue fur RGB-Lab color distribution, such as red tongue, pale white tongue. L is the brightness, with a value of 0 to 100; a represents the red-green channel (negative value is green, positive value is red); b represents the yellow-blue channel (negative value is blue, positive value is yellow). For three-dimensional features, the local normal vector variation of the tongue surface is calculated to detect the tongue surface relief; the structure tightness of the tongue body is evaluated to determine the tongue fur adhesion.
[0104] By constructing multi-modal features and performing tongue surface partition feature extraction, a partitioned tongue surface feature set is output, high-precision tongue surface partition feature extraction is realized, the accuracy of tongue body information is improved, and more comprehensive tongue information is obtained.
[0105] In summary, the tongue surface image recognition method based on three-dimensional face and tongue body reconstruction technology provided in the present application has the following beneficial effects:
[0106] By inputting the three-dimensional face point cloud data of the current user according to the multi-modal acquisition device; capturing the face action anchor point of the user according to the three-dimensional face point cloud data, and recording the face action anchor point for archiving, when the user performs an action according to the set tongue body acquisition action, tongue body point cloud data is recognized based on the face action anchor point; the tongue body point cloud data is connected with the RGB-D array channel, the tongue body point cloud data processed based on RGB parameters is output, and a tongue body simulation model is constructed; tongue surface partition feature extraction is performed based on the tongue body simulation model, a partitioned tongue surface feature set is output, and the partitioned tongue surface feature set is transmitted to the management end of the electronic medical record corresponding to the user. That is, the three-dimensional face point cloud data of the user is input by image acquisition and three-dimensional point cloud model, the action anchor point of the user's face is recorded by face action capture, tongue body point cloud data is dynamically recognized according to the face action image, a tongue body simulation model is constructed after RGB parameter processing, the tongue surface is partitioned and features are extracted, a detailed partitioned tongue surface feature set is formed, the accuracy of the point cloud data is ensured, and more comprehensive tongue information is obtained.
[0107] Embodiment two, based on the same inventive concept as the tongue surface image recognition method based on three-dimensional face and tongue body reconstruction technology in the foregoing embodiment one, the present application also provides a tongue surface image recognition system based on three-dimensional face and tongue body reconstruction technology, please refer to the accompanying Figure 2 , the tongue surface image recognition system based on three-dimensional face and tongue body reconstruction technology comprises:
[0108] The face data input module 11 is configured to input three-dimensional face point cloud data of a current user according to a multi-modal acquisition device; the tongue data recognition module 12 is configured to capture a face action anchor point of the user according to the three-dimensional face point cloud data, and record and archive the face action anchor point, and when the user performs an action according to a set tongue acquisition action, tongue point cloud data is recognized based on the face action anchor point; the simulation model construction module 13 is configured to connect the tongue point cloud data with an RGB-D array channel, output tongue point cloud data processed based on an RGB parameter, and construct a tongue simulation model; and the tongue surface feature extraction module 14 is configured to extract a partitioned tongue surface feature set based on the tongue simulation model, and transmit the partitioned tongue surface feature set to a management terminal of an electronic medical record corresponding to the user.
[0109] Further, the tongue data recognition module 12 in the tongue surface image recognition system based on the three-dimensional face and tongue reconstruction technology is further configured to:
[0110] The face action anchor point is a key action anchor point of a face accompanying tongue action, and the face action anchor point is captured by a deep learning model to obtain changes of the three-dimensional face point cloud data about an anchor point region, wherein the anchor point region at least includes a perioral region and a mandibular region; time sequence information and spatial parameters of each face action anchor point in the perioral region and the mandibular region are recorded, the time sequence information includes an action start time and a time sequence time, and the spatial parameters include a displacement vector and an angle change; a joint motion trajectory of the perioral region and the mandibular region is obtained through time sequence modeling, and tongue point cloud data is recognized according to the joint motion trajectory.
[0111] Further, the tongue data recognition module 12 in the tongue surface image recognition system based on the three-dimensional face and tongue reconstruction technology is further configured to:
[0112] When the user performs an action according to a set tongue acquisition action, it is judged whether the joint motion trajectory is in a stable state at a current time sequence; if a spatial parameter change rate of the joint motion trajectory at the current time sequence is greater than or equal to a preset threshold, a tongue acquisition instruction is triggered, and tongue point cloud data is obtained according to the tongue acquisition instruction; and if the spatial parameter change rate of the joint motion trajectory at the current time sequence is less than the preset threshold, the tongue acquisition instruction is not triggered.
[0113] Further, the tongue data recognition module 12 in the tongue surface image recognition system based on the three-dimensional face and tongue reconstruction technology is further configured to:
[0114] The change correlation of the face action anchor point is calculated, a weight attribute of the face action anchor point is configured according to the change correlation, and the weight attribute is sent to the multi-modal acquisition device, so that the multi-modal acquisition device performs acquisition angle adjustment according to the weight attribute of the face action anchor point.
[0115] Further, the tongue data recognition module 12 in the tongue surface image recognition system based on the three-dimensional face and tongue reconstruction technology is further used for:
[0116] After the multi-modal acquisition device performs acquisition angle adjustment according to the weight attribute of the face action anchor point, the multi-modal acquisition device starts a magnification acquisition instruction, adjusts the focal length of an imaging device in the multi-modal acquisition device according to the magnification acquisition instruction, acquires fine-grained tongue point cloud data, and updates the tongue simulation model according to the fine-grained tongue point cloud data.
[0117] Further, the simulation model construction module 13 in the tongue surface image recognition system based on the three-dimensional face and tongue reconstruction technology is further used for:
[0118] A face part model is constructed according to the three-dimensional face point cloud data, an alignment constraint condition is constructed, the alignment constraint condition includes key point matching and normal vector consistency, the face part model and the tongue simulation model are aligned in the same three-dimensional coordinate system under the alignment constraint condition, a face-tongue model is output, and face action anchor points are recalculated and updated according to the face-tongue model.
[0119] Further, the simulation model construction module 13 in the tongue surface image recognition system based on the three-dimensional face and tongue reconstruction technology is further used for:
[0120] Face features are identified according to the face part model, a partitioned face feature set is acquired, the partitioned face feature set is added to a partitioned tongue surface feature set as auxiliary feature data, and the partitioned face feature set and the partitioned tongue surface feature set are transmitted to a management end of an electronic medical record corresponding to the user.
[0121] Further, the tongue surface feature extraction module 14 in the tongue surface image recognition system based on the three-dimensional face and tongue reconstruction technology is further used for:
[0122] Tongue surface partitioned feature extraction is performed based on the tongue simulation model, and a partitioned tongue surface feature set is output, the partitioned tongue surface feature set at least includes a tongue tip area tongue surface feature set, a tongue middle area tongue surface feature set, a tongue root area tongue surface feature set, and a tongue side edge area tongue surface feature set; wherein the first area is a tongue tip area, the second area is a tongue middle area, the third area is a tongue root area, and the fourth area is a tongue side edge area.
[0123] Further, the tongue surface feature extraction module 14 in the tongue surface image recognition system based on the three-dimensional face and tongue body reconstruction technology is further used for:
[0124] Constructing multi-modal features including shape features, texture features, color features and three-dimensional features; the tongue body simulation model extracts features of tongue surface partitions according to the multi-modal features, and outputs a set of partitioned tongue surface features.
[0125] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The foregoing Figure 1 The tongue surface image recognition method based on the three-dimensional face and tongue body reconstruction technology in Embodiment One and the specific examples are also applicable to the tongue surface image recognition system based on the three-dimensional face and tongue body reconstruction technology in the present embodiment. Through the foregoing detailed description of the tongue surface image recognition method based on the three-dimensional face and tongue body reconstruction technology, those skilled in the art can clearly know the tongue surface image recognition system based on the three-dimensional face and tongue body reconstruction technology in the present embodiment. Therefore, for the sake of brevity of the specification, it will not be described in detail here. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant part can be referred to the method part description.
[0126] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
[0127] Obviously, for those skilled in the art, without departing from the principles of the present application, the present application can be improved and modified in several ways, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A tongue surface image recognition method based on three-dimensional face and tongue body reconstruction technology, characterized in that, The method comprises the following steps: According to the multi-modal acquisition device, the three-dimensional face point cloud data of the current user is input; According to the three-dimensional face point cloud data, the user's face action anchor point is captured, and the face action anchor point is recorded and archived. When the user performs an action according to the set tongue acquisition action, the tongue point cloud data is recognized based on the face action anchor point; Connect the tongue point cloud data with the RGB-D array channel, output the tongue point cloud data processed based on the RGB parameter, and construct a tongue simulation model; Based on the tongue simulation model, the tongue surface feature set is output by extracting the tongue surface feature set, and the tongue surface feature set is transmitted to the management end of the electronic medical record corresponding to the user; The tongue point cloud data recognized based on the face action anchor point comprises: The face action anchor point is the key action anchor point of the face accompanying the tongue action, and the face action anchor point is captured by a deep learning model to obtain the change of the three-dimensional face point cloud data about the anchor point area, wherein the anchor point area at least includes the perioral area and the mandibular area; Record the time sequence information and spatial parameters of each face action anchor point in the perioral area and the mandibular area. The time sequence information includes the action start time and the time sequence time, and the spatial parameters include the displacement vector and the angle change; Obtain the joint motion trajectory of the perioral area and the mandibular area through time sequence modeling, and recognize the tongue point cloud data according to the joint motion trajectory.
2. The tongue surface image recognition method based on three-dimensional face and tongue reconstruction technology according to claim 1, wherein, Based on the tongue simulation model, the tongue surface feature set is output by extracting the tongue surface feature set, and the tongue surface feature set at least includes the tongue surface feature set of the first area, the tongue surface feature set of the second area, the tongue surface feature set of the third area and the tongue surface feature set of the fourth area; Wherein, the first area is the tongue tip area, the second area is the middle tongue area, the third area is the tongue root area, and the fourth area is the lateral tongue edge area.
3. The tongue surface image recognition method based on three-dimensional face and tongue reconstruction technology according to claim 1, wherein, According to the joint motion trajectory, the tongue point cloud data is recognized, comprising: When the user performs an action according to the set tongue acquisition action, it is judged whether the joint motion trajectory is in a stable state at the current time sequence; If the spatial parameter change rate of the joint motion trajectory at the current time sequence is greater than or equal to a preset threshold, the tongue acquisition instruction is triggered, and the tongue point cloud data is obtained according to the tongue acquisition instruction; If the spatial parameter change rate of the joint motion trajectory at the current time sequence is less than the preset threshold, the tongue acquisition instruction is not triggered.
4. The tongue surface image recognition method based on three-dimensional face and tongue reconstruction technology according to claim 1, wherein, After constructing the tongue simulation model, it further comprises: According to the three-dimensional face point cloud data, a face model is constructed; Construct an alignment constraint condition, which includes key point matching and normal vector consistency; Under the alignment constraint condition, the face model and the tongue simulation model are aligned in the same three-dimensional coordinate system, and a face-tongue model is output. The face action anchor point is recalculated and updated according to the face-tongue model.
5. The tongue surface image recognition method based on three-dimensional face and tongue reconstruction technology according to claim 4, wherein, After constructing the face model according to the three-dimensional face point cloud data, it further comprises: According to the face model, facial features are recognized, and a partitioned face feature set is obtained. The partitioned face feature set is added to the partitioned tongue surface feature set as auxiliary feature data. The partitioned face feature set and the partitioned tongue surface feature set are transmitted to the management end of the electronic medical record corresponding to the user.
6. The tongue surface image recognition method based on three-dimensional face and tongue reconstruction technology according to claim 1, wherein, According to the three-dimensional face point cloud data, the face action anchor point of the user is captured, and the method further comprises: Calculating the change correlation of the face action anchor point, and configuring the weight attribute of the face action anchor point according to the change correlation; The weight attribute is sent to the multi-modal acquisition device, and the multi-modal acquisition device adjusts the collection angle according to the weight attribute of the face action anchor point.
7. The tongue surface image recognition method based on three-dimensional face and tongue reconstruction technology according to claim 6, wherein, After the multi-modal acquisition device adjusts the collection angle according to the weight attribute of the face action anchor point, the multi-modal acquisition device starts the magnification collection instruction; According to the magnification collection instruction, the focal length of the imaging device in the multi-modal acquisition device is adjusted, the fine-grained tongue point cloud data is obtained, and the tongue simulation model is updated according to the fine-grained tongue point cloud data.
8. The tongue surface image recognition method based on three-dimensional face and tongue reconstruction technology according to claim 1, wherein, Based on the tongue simulation model, tongue surface partition feature extraction is performed, and a partitioned tongue surface feature set is output, comprising: Constructing multi-modal features, including shape features, texture features, color features, and three-dimensional features; The tongue simulation model extracts tongue surface partition features according to the multi-modal features, and outputs a partitioned tongue surface feature set.
9. A tongue surface image recognition system based on three-dimensional face and tongue reconstruction technology, characterized by, The steps of the tongue surface image recognition method based on three-dimensional face and tongue reconstruction technology according to any one of claims 1-8, the tongue surface image recognition system based on three-dimensional face and tongue reconstruction technology comprises: A face data input module for inputting three-dimensional face point cloud data of a current user according to a multi-modal acquisition device; A tongue data recognition module for capturing a face action anchor point of the user according to the three-dimensional face point cloud data, and recording the face action anchor point, and when the user performs an action according to a set tongue collection action, tongue point cloud data is recognized based on the face action anchor point; A simulation model construction module for connecting the tongue point cloud data with an RGB-D array channel, outputting tongue point cloud data processed based on RGB parameters, and constructing a tongue simulation model; A tongue surface feature extraction module for extracting tongue surface partition features based on the tongue simulation model, outputting a partitioned tongue surface feature set, and transmitting the partitioned tongue surface feature set to the management end of the electronic medical record corresponding to the user.
Citation Information
Patent Citations
Three-dimensional image reconstruction method
CN107221029A
Face dense feature point detection and expression parameter capture method and device
CN117975536A