Machine learning-based traditional chinese medicine constitution identification analysis method and system
By collecting and analyzing health description texts, tongue images, facial images, and whole-body optical images, a posture structure feature vector is generated, which solves the problem of lacking body structure information collection in existing technologies and realizes the standardization and accuracy improvement of TCM constitution identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI BAOFANG TECHNOLOGY CO LTD
- Filing Date
- 2026-04-03
- Publication Date
- 2026-07-03
Smart Images

Figure CN122337508A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and in particular to a method and system for identifying and analyzing traditional Chinese medicine constitution based on machine learning. Background Technology
[0002] With the development of artificial intelligence technology, some studies have attempted to introduce machine learning methods into the TCM constitution identification process to improve the objectivity and standardization of the identification process. Existing machine learning-based TCM constitution identification schemes typically collect multidimensional user data and complete feature extraction and classification. Data collection methods mainly include: obtaining health description text through a human-computer interaction interface and extracting semantic features of the text using natural language processing technology; obtaining tongue and facial images through image acquisition devices and extracting tongue texture features and facial color features using image recognition algorithms. Some related systems can guide users to fill out electronic consultation forms and upload tongue images, and complete a comprehensive judgment of constitution type based on text analysis and image feature extraction.
[0003] However, in practical applications, the physical characteristics of a person may not only be reflected in the description of internal symptoms and the appearance of the tongue, but may also be related to the individual's body shape and posture structure. Current technical solutions mostly focus on text semantic features and tongue image features in terms of data analysis dimensions, while less attention is paid to posture data that can reflect body shape and structure characteristics.
[0004] For example, in the analysis of a user, the system can process the health description text entered by the user and the uploaded tongue image well, but it lacks a dedicated collection and feature extraction mechanism for posture information that reflects the body structure, such as the user's standing posture and body shape. This difference in data collection dimensions may lead to the feature information used for body constitution classification being unable to fully cover the multifaceted characteristics related to body constitution. Summary of the Invention
[0005] This invention provides a method and system for TCM constitution identification and analysis based on machine learning, with a two-stage judgment process of model classification and standard matching, which reduces the probability of misjudgment by a single model.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a method for identifying and analyzing TCM constitution based on machine learning, the method comprising: Acquire health description text data input by the user through the human-computer interaction interface, and extract the text semantic feature vector; collect the user's tongue image, face image and whole body optical image, extract the tongue image texture feature set and face image color feature set, and locate the bilateral acromion point, bilateral anterior superior iliac spine point, seventh cervical vertebra spinous process point and bilateral patella point in the whole body optical image to obtain the initial posture point set; Construct a posture space coordinate system and map the initial posture point set to the posture space coordinate system; determine the configuration origin based on the ratio of the user's shoulder width to pelvic width; construct a spatial configuration domain with the configuration origin as the center and the user's torso length as the sampling radius; The spatial modeling domain is divided into multiple spatial units at equal intervals along three directions, and the number of initial attitude points contained in each spatial unit is counted. The gradient change sequence of the number of initial attitude points between adjacent spatial units along each axis is calculated, and the gradient change sequences of the three axes are cross-fused in pairs to generate attitude structure feature vectors. The text semantic feature vector, tongue texture feature set, facial color feature set, and posture structure feature vector are concatenated in a feature layer to obtain a comprehensive constitution feature tensor. The comprehensive constitution feature tensor is then input into a pre-trained classification model to obtain the user's preliminary TCM constitution type. The preliminary TCM constitution type is then compared with the standard constitution types in the preset standard constitution database to calculate the similarity and determine the final TCM constitution type based on the similarity.
[0007] Secondly, a machine learning-based TCM constitution identification and analysis system includes: The feature extraction module is used to acquire health description text data input by the user through the human-computer interaction interface, extract the semantic feature vector of the text; collect the user's tongue image, face image and whole body optical image, extract the tongue image texture feature set and face image color feature set, and locate the bilateral acromion point, bilateral anterior superior iliac spine point, seventh cervical vertebra spinous process point and bilateral patellar point in the whole body optical image to obtain the initial posture point set; The configuration domain generation module is used to construct a posture space coordinate system and map the initial posture point set to the posture space coordinate system; determine the configuration origin based on the ratio of the user's shoulder width to pelvic width; construct a spatial configuration domain with the configuration origin as the center and the user's torso length as the sampling radius; The vector generation module is used to divide the spatial modeling domain into multiple spatial units at equal intervals along three directions, and count the number of initial attitude points contained in each spatial unit; calculate the gradient change sequence of the number of initial attitude points between adjacent spatial units along each axis, and cross-merge the gradient change sequences of the three axes to generate attitude structure feature vectors. The type determination module is used to concatenate the text semantic feature vector, tongue texture feature set, facial color feature set and posture structure feature vector into a feature layer to obtain a comprehensive constitution feature tensor; input the comprehensive constitution feature tensor into a pre-trained classification model to obtain the user's preliminary TCM constitution type; calculate the similarity between the preliminary TCM constitution type and the standard constitution type in the preset standard constitution database, and determine the final TCM constitution type based on the similarity.
[0008] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0009] The above-described solution of the present invention has at least the following beneficial effects: This method simultaneously collects four types of data: health description text, tongue image, facial image, and whole-body posture. It overcomes the limitations of relying solely on text and tongue images, incorporating body posture structure into the identification dimension for more comprehensive feature information and avoiding the bias caused by a single dimension. Posture feature quantification extraction enhances the objectivity and standardization of identification. By constructing a posture space coordinate system, spatial morphology domain, and three-dimensional spatial unit segmentation, it digitizes, quantifies, and repeatably extracts features from key points of the human body, reducing judgment bias caused by differences in human experience and achieving standardization of constitution identification. Multi-feature tensor fusion strengthens feature expression and classification accuracy. It performs feature layer splicing and tensor reconstruction of four types of features: text semantics, tongue texture, facial color, and posture structure, fully integrating the correlation information of different modalities and improving the representational ability of comprehensive features. A two-stage judgment mechanism improves the stability and reliability of the final result. First, a pre-trained classification model outputs a preliminary constitution type, then a similarity check is performed with a standard constitution database, forming a two-stage judgment process of model classification and standard matching. This reduces the probability of misjudgment by a single model, making the final output TCM constitution type more consistent with the standard constitution definition. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating the TCM constitution identification and analysis method based on machine learning provided in an embodiment of the present invention.
[0011] Figure 2 This is a schematic diagram of a machine learning-based TCM constitution identification and analysis system provided in an embodiment of the present invention. Detailed Implementation
[0012] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0013] like Figure 1 As shown, embodiments of the present invention propose a method for identifying and analyzing TCM constitution based on machine learning, the method comprising the following steps: Step 1: Obtain health description text data input by the user through the human-computer interaction interface and extract the text semantic feature vector; collect the user's tongue image, face image and whole body optical image, extract the tongue image texture feature set and face image color feature set, and locate the bilateral acromion point, bilateral anterior superior iliac spine point, seventh cervical vertebra spinous process point and bilateral patellar point in the whole body optical image to obtain the initial posture point set; Step 2: Construct a posture space coordinate system and map the initial posture point set to the posture space coordinate system; determine the configuration origin based on the ratio of the user's shoulder width to pelvic width; construct a spatial configuration domain with the configuration origin as the center and the user's torso length as the sampling radius; Step 3: Divide the spatial configuration domain into multiple spatial units at equal intervals along the three directions, and count the number of initial attitude points contained in each spatial unit; calculate the gradient change sequence of the number of initial attitude points between adjacent spatial units along each axis, and cross-merge the gradient change sequences of the three axes to generate attitude structure feature vectors. Step 4: Concatenate the text semantic feature vector, tongue texture feature set, facial color feature set, and posture structure feature vector into a feature layer to obtain a comprehensive constitution feature tensor; input the comprehensive constitution feature tensor into a pre-trained classification model to obtain the user's preliminary TCM constitution type; calculate the similarity between the preliminary TCM constitution type and the standard constitution type in the preset standard constitution database, and determine the final TCM constitution type based on the similarity.
[0014] In this embodiment of the invention, the method simultaneously collects four types of data: health description text, tongue image, facial image, and full-body posture data. This overcomes the limitations of relying solely on text and tongue images, incorporating body posture structure into the identification dimension, resulting in more comprehensive feature information and avoiding the bias caused by a single dimension. Quantitative extraction of posture features enhances the objectivity and standardization of identification. By constructing a posture spatial coordinate system, a spatial morphological domain, and three-dimensional spatial unit segmentation, key points of the human body are digitally, quantitatively, and reproducibly extracted, reducing judgment bias caused by differences in human experience and achieving standardization of body constitution identification. Multi-feature tensor fusion enhances feature expression and classification accuracy. It performs feature layer concatenation and tensor reconstruction on four types of features: text semantics, tongue texture, facial color, and posture structure. It fully integrates the correlation information of different modalities and improves the representation ability of comprehensive features. A two-stage judgment mechanism improves the stability and reliability of the final result. First, a preliminary constitution type is output through a pre-trained classification model. Then, a similarity check is performed with a standard constitution database. This forms a two-stage judgment process of model classification and standard matching, which reduces the probability of misjudgment by a single model and makes the final output TCM constitution type more in line with the standard constitution definition.
[0015] In a preferred embodiment of the present invention, step 1 involves acquiring health description text data input by the user through a human-computer interaction interface and extracting text semantic feature vectors; collecting the user's tongue image, facial image, and full-body optical images, extracting the tongue image texture feature set and facial image color feature set, and locating the bilateral acromion points, bilateral anterior superior iliac spine points, the spinous process point of the seventh cervical vertebra, and bilateral patellar points in the full-body optical images to obtain an initial posture point set; the extraction of text semantic feature vectors includes: Step 100: Segment the health description text data to obtain a word sequence; perform part-of-speech tagging on the word sequence to obtain a tagged word sequence. Specifically, this includes: first, receiving health-related description text data actively input by the user through a built human-computer interaction interface. This text data covers various health-related information such as physical symptoms, discomfort, lifestyle, and dietary preferences. Simultaneously, the image acquisition module is activated, guiding the user to complete standardized acquisition of tongue, facial, and full-body optical images according to preset image acquisition specifications. Among these, the tongue image acquisition must ensure that the tongue surface is complete and clear; the facial image is a frontal view without makeup; and the full-body optical image is a complete frontal view of the user standing naturally. From a body perspective, after acquisition, the three types of images are preprocessed separately. Through operations such as image denoising, grayscale conversion, and feature enhancement, tongue texture feature set that can reflect the details of tongue surface texture is extracted from the tongue image preprocessing image, and facial color feature set that can reflect the depth and distribution of facial color is extracted from the facial image preprocessing image. Human key point detection is performed on the whole body optical preprocessing image. Through feature point matching and localization, key body feature points such as bilateral acromion points, bilateral anterior superior iliac spine points, seventh cervical vertebra spinous process points, and bilateral patellar points are determined. The coordinate information of all the located key feature points is integrated to form an initial pose point set containing the spatial position information of each point.
[0016] After completing the collection and preliminary feature processing of the aforementioned multi-dimensional data, the health description text data underwent word segmentation and format normalization, removing invalid characters such as spaces, line breaks, modal particles, and punctuation marks. Then, based on Chinese semantic expression rules and vocabulary segmentation criteria, combined with a professional thesaurus in the field of Traditional Chinese Medicine, a word segmentation algorithm was used to segment the normalized continuous text, breaking it down into independent words with complete semantic meanings. This included both common everyday vocabulary and TCM professional terms. Following the order of appearance of the words in the original text, all segmented words were arranged in an ordered manner to form a structured word sequence. After obtaining the word sequence, part-of-speech tagging was performed. Based on the grammatical attributes, semantic features, and grammatical function of each word in the sentence, a corresponding standard part-of-speech tag was matched for each word in the word sequence. The part-of-speech tags covered all common grammatical part-of-speech categories, including nouns, verbs, adjectives, adverbs, and prepositions. After tagging, following the order of the words in the original sequence, the words with matched part-of-speech tags were integrated to form a tagged word sequence, achieving a dual association between the grammatical attributes and semantic information of the words.
[0017] Step 101: Perform semantic role analysis on the labeled word sequence to obtain a semantic role sequence. Input the semantic role sequence into a pre-trained language model. After context encoding by the encoder of the pre-trained language model, obtain the text semantic feature vector. Specifically, this includes: conducting semantic role analysis on the labeled word sequence. First, complete the construction and full-process training of the semantic role labeling model. This model adopts a network structure that combines a bidirectional long short-term memory network and a conditional random field. The bidirectional long short-term memory network has 3 hidden layers, with 256 neurons in each layer. The forward and backward networks capture the forward and backward contextual semantic relationships of the labeled word sequence, respectively. Features are used to address the problem that unidirectional networks cannot capture all contextual information. A Conditional Random Field (CRF) is used as the output layer, based on the output features of a bidirectional Long Short-Term Memory (LSTM) network, to perform global probability optimization on the sequence labeling results, constraining the rationality of the labeling results and avoiding logical contradictions in semantic role labeling. The initial weights of the input layer are set to 0.02, the hidden layer weights to 0.92, the output layer weights to 0.06, the CRF transition matrix weights to 0.85, and the regularization constraint weights to 0.01. After the model is built, a training dataset is constructed, with general Chinese text and health description texts in the field of Traditional Chinese Medicine in a 7:3 ratio. The dataset was integrated in a specific ratio. The health description texts in the field of Traditional Chinese Medicine (TCM) included texts related to TCM symptoms, constitution manifestations, physical signs, and treatment records. The integrated original dataset underwent automated data preprocessing. Text cleaning algorithms were used to remove invalid data such as garbled characters, meaningless characters, and duplicate text. A unified word segmentation and part-of-speech tagging algorithm was then applied to create a standardized preprocessed dataset. This dataset was randomly divided into training, validation, and test sets in an 8:1:1 ratio. The training set was input into the constructed semantic role labeling model in batches of 64. An adaptive moment estimation optimization algorithm was used to optimize the model. The network parameters are iteratively updated with an initial learning rate of 0.001 and a weight decay coefficient of 0.0001. The cross-entropy loss function is used as the core loss criterion for model training. The effect is evaluated on the validation set every 10 iterations, and the training loss value and semantic role recognition accuracy of the model are monitored in real time. When the validation set accuracy does not improve for 3 consecutive iterations, the learning rate is decayed to 0.5. During the iteration process, the input layer weights are iteratively optimized with a step size of 0.001, the hidden layer weights with a step size of 0.005, the output layer weights with a step size of 0.002, and the conditional random field transition matrix weights with a step size of 0.001.The 003 step size iterative optimization continuously improves the model's accuracy in recognizing semantic roles and its classification performance. Simultaneously, during model training, the generalization ability of the model is automatically verified using a test set until the model's loss value on the test set stabilizes and the semantic role recognition accuracy reaches a preset standard (no less than 90%). At this point, the model has fully converged to a stable state, resulting in the trained semantic role annotation model.
[0018] The trained semantic role labeling model is directly applied to the labeled word sequence to be analyzed. The model first performs feature vectorization on the labeled word sequence using a built-in algorithm. Then, combined with the contextual semantic logic and sentence grammatical structure of the sequence itself, it performs word-by-word semantic role identification and category determination for each word in the sequence through forward computation. This clarifies the specific semantic role that each word plays in the health description statement. Semantic roles include core categories such as subject, action, object, modifier, causality, time, and location. At the same time, through the feature extraction and association analysis capabilities trained by the model, the semantic relationships and logical levels between words are deeply sorted out and mined, and the inherent connections between words such as semantic collocation, causal relationship, modification limitation, and time limitation are determined. According to the original order of words in the labeled word sequence, the words labeled with corresponding semantic roles are arranged in an orderly manner to form a semantic role sequence that can fully reflect the semantic logic of the text and the relationship between words.
[0019] After generating the semantic role sequence, a customized pre-trained language model for semantic encoding was constructed and trained. The model uses a 12-layer Transformer encoder as its core network structure. Each Transformer encoder layer includes a multi-head self-attention mechanism layer (with 12 attention heads), a feedforward neural network layer (with 2048 hidden neurons), and a layer normalization module. The initial weights for the query matrix, key matrix, and value matrix in the multi-head self-attention mechanism layer were set to 0.88, 0.88, and 0.90, respectively. The initial weights for the hidden layers in the feedforward neural network layer were set to 0.94, the initial weights for the layer normalization module were set to 0.96, and the dropout regularization weight was set to 0.1. The self-attention mechanism layer enables parallel attention to features at different positions in the sequence, the feedforward neural network layer completes the nonlinear transformation of features, and the layer normalization module solves the gradient vanishing problem in model training. The synergistic effect of the three-layer structure can efficiently capture long-distance semantic associations and local semantic features of the sequence. After the model is built, a large-scale dedicated training corpus is constructed, configured in a ratio of 4:4:2 for TCM professional texts, health description texts, and general semantic texts. The TCM professional texts include original texts of TCM classics, literature on TCM constitution identification, TCM-related knowledge texts, and TCM clinical sign observation records. The health description texts include texts describing human signs, body constitution, and information on human activity, diet, posture, complexion, and tongue appearance. Relevant objective descriptive texts were used; general semantic texts were selected from mainstream Chinese corpora, including daily communication, descriptions of life scenes, and descriptions of objects and behaviors. Based on this corpus, the model was pre-trained using an unsupervised training task of a masked language model (MLM). At a ratio of 15%, words in the corpus were randomly masked using an algorithm, with 80% replaced by mask symbols, 10% replaced by random words, and 10% remaining unchanged. The model was then allowed to predict the original word at the masked position based on contextual information. During training, the batch size was set to 128, the initial learning rate to 0.0001, and the total number of training rounds to 100. The AdamW optimization algorithm was used for iterative optimization of model parameters, and the multi-head self-attention mechanism was layer-wise optimized. The weights of the query, key, and value matrices are iteratively optimized with a step size of 0.0005, the weights of the feedforward neural network layers are iteratively optimized with a step size of 0.0003, and the weights of the layer normalization module are iteratively optimized with a step size of 0.0004. During the iteration process, the deviation of the values of the weight parameters of each layer is continuously corrected, so that the model can fully learn and deeply master the semantic features, language expression rules, and semantic connotations and collocations of TCM professional terms in the field of TCM health. At the same time, through multiple iterations of training, the model's ability to fuse and represent contextual semantic information is continuously improved until the model's masked word prediction accuracy on the validation corpus tends to stabilize (not less than 85%). The model completes training and reaches a stable convergence state, resulting in a pre-trained language model specifically adapted to the TCM constitution identification scenario.
[0020] The semantic role sequence is formatted according to the model input requirements using an algorithm. After binding words with their corresponding semantic role identifiers, it is converted into a vector sequence that the model can recognize. This vector sequence is then input into a pre-trained language model, where deep context encoding is performed using the model's encoder module. This encoder consists of a 12-layer cascaded Transformer network structure, with each layer containing an attention mechanism layer, a feature extraction layer, and an information fusion layer. It can perform multi-level, multi-dimensional feature extraction and context information fusion on the semantic role sequence. The encoding process follows a fixed algorithm flow. First, the feature extraction layer extracts basic semantic features from each word in the semantic role sequence using linear transformations and a non-linear activation function (GELU), resulting in a 768-dimensional basic semantic feature representation of the words. Then, the attention mechanism layer calculates self-attention weights... This method captures the contextual association weights between words in a sequence, strengthening the weights of word features corresponding to physical characteristics, physical state, and body posture, while weakening the weights of word features corresponding to irrelevant and redundant information (such as modal particles and conjunctions). This emphasizes the representation of core semantic information. Furthermore, the semantic features of each word are standardized and vectorized, converting discrete word semantic features into a continuous 768-dimensional numerical vector. Simultaneously, an information fusion layer deeply integrates semantic association information and semantic role attribute information from the context into the feature vector of each word. This ensures that the final feature vector of each word not only contains its own core semantic information and corresponding semantic role attributes but also fully covers its semantic logical associations and collocations with the surrounding words. During the encoding process, the vector transformation of semantic features is completed using a feature mapping formula, which is: ,in The first in the semantic role sequence The 768-dimensional semantic feature vector of each word Representing the Vectorized results of semantic role identifiers corresponding to each word Representing the The feature vector is a fusion of the contextual semantic information of each word.
[0021] After the encoder completes the encoding of all words in the semantic role sequence, the 768-dimensional semantic feature vectors of all individual words output by the encoder are uniformly standardized and integrated. First, the vector dimension unification adjustment operation is performed through the algorithm. For the very few feature vector dimension deviations caused by high semantic specificity, a linear mapping feature dimension expansion or compression method is adopted. Based on the preset 768-dimensional fixed dimension, the semantic feature vectors of all individual words are dimension-mapped and numerically adjusted. The dimension mapping weight is set to 0.92 to ensure that the semantic feature vector of each word is in the same 768-dimensional feature space. Then, according to the original order of words in the semantic role sequence, the vector concatenation algorithm performs continuous feature concatenation operation on all the 768-dimensional semantic feature vectors of words after dimension unification, and merges the feature vectors of individual words into a whole in sequence. Finally, a whole feature vector with fixed dimension and complete and non-redundant feature information is obtained. The dimension of this whole feature vector is the number of words in the sequence × 768. It fully integrates the core semantic information, semantic role attributes and overall semantic logic of all words in the health description text, and can comprehensively represent the overall semantic information of the health description text data.
[0022] This embodiment performs step-by-step processing of health description text data, including word segmentation and part-of-speech tagging, achieving refined decomposition of the text data, effectively eliminating invalid information in the text, and clarifying the grammatical attributes of words. It deeply explores the semantic relationships and logical levels between words in the text, identifying the semantic role of each word in the health description, making the extracted semantic features more consistent with the actual semantic expression of the user's health description, and improving the accuracy of the text semantic features. It uses a pre-trained language model encoder for context encoding, fully integrating the contextual semantic information of the text, so that the generated text semantic feature vector not only contains the semantic information of individual words, but also covers the overall logic and semantic relationships of the sentences, making the representation of text semantic features more comprehensive and three-dimensional, and able to fully reflect the overall semantic connotation of the user's health description. This extraction method is compatible with the extraction logic of features such as tongue texture, facial color, and posture structure.
[0023] In a preferred embodiment of the present invention, step 2 involves constructing a posture space coordinate system and mapping the initial posture point set to the posture space coordinate system; determining the configuration origin based on the ratio of the user's shoulder width to pelvic width; and constructing a spatial configuration domain centered on the configuration origin and with the user's torso length as the sampling radius, including: Step 200: The ratio of the distance between the bilateral acromion points to the distance between the bilateral anterior superior iliac spine points is used as the morphological ratio value. If the morphological ratio value is greater than a preset threshold, the association mode is identified as the first mode, and the midpoint of the line connecting the bilateral acromion points is used as the configuration origin. If the morphological ratio value is less than or equal to the preset threshold, the association mode is identified as the second mode, and the midpoint of the line connecting the bilateral anterior superior iliac spine points is used as the configuration origin. Specifically, this includes: firstly, establishing a three-dimensional rectangular posture space coordinate system, with the geometric center of the whole-body optical image as the coordinate system origin reference, and the vertical direction of the human body as the Z-axis (upward is the positive direction), and the left side of the human body as the Z-axis. With the right horizontal direction as the X-axis (positive direction to the right) and the front-to-back depth direction of the human body as the Y-axis (positive direction forward), a standardized three-dimensional rectangular coordinate system conforming to the human body shape characteristics is established to complete the construction of the posture space coordinate system. This coordinate system provides a unified benchmark for the spatial calculation of all posture points. The spatial coordinate information of the initial posture point set obtained by positioning, including the bilateral acromion points, bilateral anterior superior iliac spine points, the spinous process of the seventh cervical vertebra, and bilateral patellar points, is all mapped to this posture space coordinate system. The spatial coordinate calibration of each initial posture point is completed through a coordinate matching algorithm, realizing the spatial positioning of all posture feature points. Within the attitude space coordinate system after coordinate mapping, the three-dimensional coordinates of the bilateral acromion points (left acromion point and right acromion point) and the bilateral anterior superior iliac spine points (left anterior superior iliac spine point and right anterior superior iliac spine point) are extracted from the initial attitude point set. The distance between the bilateral acromion points and the distance between the bilateral anterior superior iliac spine points are calculated using the distance calculation formula between two points in three-dimensional space. Based on the calculated distance between the bilateral acromion points and the distance between the bilateral anterior superior iliac spine points, the morphological ratio value is calculated. The morphological ratio value is the ratio of the distance between the bilateral acromion points to the distance between the bilateral anterior superior iliac spine points.
[0024] After calculating the morphological proportion values, they are compared with a preset body shape judgment threshold of 1.5. This threshold is determined through standardized three-dimensional human body measurements and statistics of adults from different regions and age groups. The measured samples cover multiple demographic dimensions. During the measurement process, optical three-dimensional scanning equipment is used to collect the spatial coordinates of the bilateral acromion points and anterior superior iliac spine points and calculate the distance ratio. All measured distance ratio data are first subjected to outlier removal using the 3σ criterion, and then the filtered valid data are subjected to normal distribution fitting analysis to calculate the skewness of the data. With kurtosis The value verifies the normal distribution property of the data, and the skewness calculation formula is as follows: The formula for calculating kurtosis is: ,in For effective data sample size, For the first Spacing ratio data, The mean of the data. To determine the standard deviation of the data, the normal distribution characteristics of the data are verified based on the skewness and kurtosis calculation results, ensuring that the data distribution meets the requirements of statistical analysis. For the effective ratio data that conforms to a normal distribution, quartiles are determined, and all effective data are arranged in ascending order to form the lower quartiles. The calculation formula is , median The calculation formula is Upper quartile The calculation formula is The lower quartile, median, and upper quartile of the data were calculated using formulas to clarify the overall distribution range and characteristic quantile values of the ratio data. Finally, based on the ratio quantile values corresponding to the upper quartile as the core reference for the quantile distribution pattern, a critical characteristic value for the ratio distribution within the upper quartile range was selected. This value effectively distinguishes between body types where the distance between the bilateral acromion points exceeds 0.5 times the distance between the bilateral anterior superior iliac spines and body types with balanced shoulder-pelvic width where the difference between the bilateral acromion point and bilateral anterior superior iliac spine distances is within 0.5 times. Assuming this threshold is set at 1.5, this threshold effectively distinguishes between body types where shoulder width is greater than pelvic width and body types with balanced shoulder-pelvic width. A balanced pelvic width has clear statistical significance and actual body shape recognition in body shape differentiation. If the shape ratio value is greater than the preset threshold of 1.5, the current human posture is identified as the first mode. In the posture space coordinate system, the midpoint coordinates of the line connecting the two acromion points are calculated, and this midpoint is determined as the configuration origin of the spatial configuration analysis. If the shape ratio value is less than or equal to the preset threshold of 1.5, the current human posture is identified as the second mode. In the posture space coordinate system, the midpoint coordinates of the line connecting the two anterior superior iliac spines are calculated, and this midpoint is determined as the configuration origin of the spatial configuration analysis.
[0025] Step 201: Calculate the user's trunk length based on the vertical distance between the bilateral acromion points and the bilateral anterior superior iliac spine points. Construct a spherical space with the configuration origin as the center and the trunk length as the radius. The internal region of this spherical space is the spatial configuration domain. Specifically, after determining the configuration origin in step 200, calculate the vertical distance between the bilateral acromion points and the bilateral anterior superior iliac spine points based on the 3D coordinates of the previously calibrated bilateral acromion points and bilateral anterior superior iliac spine points in the posture space coordinate system. This vertical distance represents the length of the human trunk in the vertical direction, i.e., the user's trunk length. First, calculate the vertical coordinates of the center points of the bilateral acromion points and the center points of the bilateral anterior superior iliac spine points, then calculate the trunk length, i.e., trunk length = After completing the trunk length calculation, the configuration origin determined in step 200 (midpoint of the line connecting the bilateral acromion points in the first mode, and midpoint of the line connecting the bilateral anterior superior iliac spine points in the second mode) is used as the center of the three-dimensional spherical space. The calculated trunk length is used as the radius of the spherical space. In the constructed posture space coordinate system, a standard three-dimensional spherical space is drawn with the center and radius as parameters. The spatial range of this spherical space is centered on the configuration origin and extends to the trunk length, fully covering the posture point distribution range of the core area of the human trunk. The entire internal area covered by this three-dimensional spherical space is the spatial configuration domain for human posture feature analysis.
[0026] This embodiment can unify the spatial reference of human posture features by constructing a dedicated posture space coordinate system and mapping posture point sets; it can divide the association pattern based on the ratio of the distance between the bilateral acromion points and the anterior superior iliac spine points and dynamically determine the configuration origin, which can adapt to the body characteristics of users with different body types; and it can delineate the effective posture range of the human torso by constructing a spherical spatial configuration domain with the torso length as a parameter.
[0027] In a preferred embodiment of the present invention, step 3 involves dividing the spatial configuration domain into multiple spatial units at equal intervals along three directions, and counting the number of initial attitude points contained in each spatial unit; calculating the gradient change sequence of the number of initial attitude points between adjacent spatial units along each axis, and cross-merging the gradient change sequences of the three axes pairwise to generate an attitude structure feature vector, including: Step 300: Establish three orthogonal coordinate axes—sagittal, coronal, and vertical—based on the configuration origin. Divide the spatial configuration domain along the sagittal axis at a first preset interval to obtain multiple sagittal slices. Divide each sagittal slice along the coronal axis at a second preset interval to obtain multiple coronal strips. Divide each coronal strip along the vertical axis at a third preset interval to obtain multiple spatial units. Specifically, this includes: establishing pairwise orthogonal sagittal, coronal, and vertical axes based on the configuration origin determined in step 200. The sagittal axis corresponds to the anterior-posterior direction of the human body, the coronal axis corresponds to the lateral direction of the human body, and the vertical axis corresponds to the vertical direction of the human body. The three coordinate axes are pairwise perpendicular and intersect at the configuration origin, together forming a three-dimensional orthogonal basis for spatial segmentation. The constructed spherical spatial domain is divided evenly along the sagittal axis at a first preset interval of 20 mm to obtain multiple parallel and equally spaced sagittal slices. Each sagittal slice is then divided evenly along the coronal axis at a second preset interval of 20 mm to obtain multiple parallel and equally spaced coronal strips. Finally, each coronal strip is divided evenly along the vertical axis at a third preset interval of 20 mm to obtain multiple cubic spatial units with a side length of 20 mm. All spatial units are arranged in a close manner to achieve full coverage of the entire spatial domain. The spatial range of each spatial unit is a cubic region extending 10 mm forward, backward, left, right, up, and down from its own center.
[0028] Step 301: Traverse all spatial units, count the number of initial attitude points falling into each spatial unit, and obtain the preliminary count value of the initial attitude points for each spatial unit; for each spatial unit, construct a point set based on all initial attitude points within the corresponding spatial unit; calculate the geometric moments of each order of the corresponding point set, construct multiple central moments based on the geometric moments of each order, and normalize the central moments to obtain the normalized central moments. Specifically, this includes: traversing all spatial units, performing spatial coordinate inclusion determination for each spatial unit, counting the total number of initial attitude points falling into the spatial unit, and obtaining the preliminary count value of the initial attitude points; integrating all initial attitude points within each spatial unit to construct a local 3D point set, and calculating the zeroth-order geometric moment, first-order geometric moment, second-order geometric moment, and higher-order geometric moments sequentially based on this point set. The calculations are as follows: zeroth-order geometric moment... In the formula, The zeroth order geometric moment, This represents the total number of initial attitude points within the current spatial cell. The attitude point number; first-order geometric moments , , In the formula, , , These are the first-order geometric moments in the coronal, sagittal, and vertical axes, respectively. The first Three-dimensional coordinates of each attitude point; second-order geometric moments , , , , , In the formula It is the second-order axial geometric moment. The second-order cross geometric moments are used; the coordinates of the center of the point set are calculated based on the first-order geometric moments. In the formula These are the three-dimensional coordinates of the geometric center of the point set.
[0029] Construct central moments of each order based on the coordinates of the center of the point set, eliminating the effects of translation and offset, and the second-order central moments. , , , , , In the formula The central moments are second-order; the central moments are normalized by scaling to obtain normalized central moments.
[0030] Step 302 involves performing combination operations on the normalized central moments to obtain a set of moment invariants that remain unchanged under translation, rotation, and scaling. The moment invariants are then compared with a preset standard moment invariant template to calculate the similarity coefficient of the corresponding spatial unit. Finally, the initial count of initial attitude points is fused with the morphological similarity coefficient to obtain the number of initial attitude points for each spatial unit. Specifically, this includes: performing operations on the normalized central moments using preset fixed combination rules, first executing linear operations, then nonlinear operations. Through the synergistic fusion of linear and nonlinear operations, a set of moment invariants that remain numerically unchanged under spatial translation, rotation, and scaling is constructed. Specifically, the linear operation uses the formula (second-order normalized central moment × 0.3) ± (third-order normalized central moment × 0.2) ± (fourth-order and above normalized central moments × 0.1) to achieve weighted integration of normalized central moments of each order; the nonlinear operation uses three sets of formulas executed synergistically: a product combination formula, where normalized central moments of each order are multiplied to strengthen the correlation between features of each order; a power combination formula, where... , Highlighting the differentiated expression of characteristics at different orders; the natural index mapping formula, i.e. This method achieves nonlinear enhancement and normalization calibration of features. The weighted integration result obtained from linear operations is combined with the product combination result, power combination result, and natural exponent mapping result obtained from nonlinear operations, and then summed with equal weights (all weights are 0.25). This deeply combines linear and nonlinear features, and finally constructs a set of moment invariants that remain numerically unchanged under spatial translation, spatial rotation, and scaling transformations. These moment invariants are standard features for 3D morphological recognition, which can ensure that human posture features remain stable when the position moves, the limbs rotate, and the shooting distance changes, avoiding feature deviations caused by external transformations.
[0031] The morphological similarity coefficient = 1 ÷ (1 + Euclidean distance between moment invariants and the standard template). The standard template is constructed by selecting multiple sets of standard human 3D pose data of different body types and postures, dividing each set of standard human 3D pose data into spatial units according to the spatial segmentation rules in step 300, and then obtaining the normalized central moments of each spatial unit of each set of standard human bodies according to the calculation method in step 301. Through linear and nonlinear operations and fusion integration methods in the steps, the moment invariants of each spatial unit of each set of standard human bodies are constructed. The mean value of the moment invariants of corresponding spatial units of all sets of standard human bodies is then obtained. The standard moment invariant template corresponding to the spatial unit is used to complete the construction of the standard template; the value range of the morphological similarity coefficient is between 0 and 1. The closer the value is to 1, the more similar the current point set distribution is to the standard human body shape distribution; the closer the value is to 0, the greater the difference is; the initial count value of the initial pose points and the morphological similarity coefficient are weighted and fused to calculate so that the final count value contains both point density information and morphological matching degree, avoiding simply counting and ignoring the distribution structure. The final number of initial pose points = the initial count value of the initial pose points × the morphological similarity coefficient, thus obtaining the final number of initial pose points for each spatial unit.
[0032] Step 303: Along the sagittal axis, calculate the difference in the number of initial pose points of corresponding spatial units between adjacent sagittal slices to obtain the first gradient sequence; along the coronal axis, calculate the difference in the number of initial pose points of corresponding spatial units between adjacent coronal strips to obtain the second gradient sequence; along the vertical axis, calculate the difference in the number of initial pose points between adjacent spatial units in the vertical direction to obtain the third gradient sequence. Specifically, this includes: along the sagittal axis, sequentially calculating the absolute difference in the number of initial pose points of spatial units corresponding to each other between adjacent sagittal slices, the purpose of which is to extract the density variation of pose points in the anterior-posterior direction of the human body. The gradient characteristic is defined as the gradient difference between adjacent units = |number of current spatial unit attitude points -number of adjacent corresponding spatial unit attitude points|. All calculated differences are arranged sequentially from front to back along the sagittal axis to obtain the first gradient sequence. This first gradient sequence represents the magnitude of change in the initial attitude point distribution along the sagittal direction through the numerical value of the difference. The larger the difference, the greater the magnitude of change; the smaller the difference, the smaller the magnitude of change. The order of the differences represents the trend of change, that is, increasing adjacent differences indicate that the magnitude of change is gradually increasing, decreasing adjacent differences indicate that the magnitude of change is gradually decreasing, and the difference tends to be stable, indicating that the attitude points are evenly distributed.
[0033] Along the coronal axis, the absolute difference in the number of initial pose points of corresponding spatial units between adjacent coronal bands is calculated sequentially, using the same method as above. The purpose is to extract the density variation characteristics of pose points in the left-right direction of the human body. All differences are arranged sequentially from left to right along the coronal axis to obtain the second gradient sequence, whose variation amplitude and trend are represented in the same way as the first gradient sequence. Along the vertical axis, the absolute difference in the number of initial pose points of adjacent spatial units within the same coronal band is calculated sequentially, using the same method as above. The purpose is to extract the density variation characteristics of pose points in the vertical direction of the human body. All differences are arranged sequentially from bottom to top along the vertical axis to obtain the third gradient sequence, whose variation amplitude and trend are represented in the same way as the first gradient sequence.
[0034] Step 304 involves performing an outer product operation on the first gradient sequence and the second gradient sequence to obtain a first fusion matrix; performing an outer product operation on the first gradient sequence and the third gradient sequence to obtain a second fusion matrix; and performing an outer product operation on the second gradient sequence and the third gradient sequence to obtain a third fusion matrix. Specifically, this includes: performing a vector outer product operation on the first gradient sequence and the second gradient sequence to couple the gradient features in the sagittal and coronal directions, obtaining two-dimensional cross features that reflect the joint change law of the attitude point in the plane; the outer product calculation formula is: the value in the i-th row and j-th column of the first fusion matrix = the i-th value of the first gradient sequence × the j-th value of the second gradient sequence. This calculation yields the first fusion matrix representing the mutual coupling of the gradient features in the sagittal and coronal directions; performing a vector outer product operation on the first gradient sequence and the third gradient sequence in the same way to couple the gradient features in the sagittal and vertical directions to obtain a second fusion matrix; and performing a vector outer product operation on the second gradient sequence and the third gradient sequence in the same way to couple the gradient features in the coronal and vertical directions to obtain a third fusion matrix.
[0035] The three sets of fusion matrices characterize the cross-correlation and synergistic change features of the initial attitude point distribution gradients in different spatial directions by the magnitude and distribution pattern of the matrix elements. The larger the matrix element value, the more significant the gradient change and the stronger the cross-correlation in the corresponding two directions; the smaller the matrix element value, the gentler the gradient change and the weaker the cross-correlation in the corresponding two directions. The overall numerical distribution of the matrix can show the synergy of gradient changes in the three spatial directions. If the element values in a certain continuous region of the matrix are all higher than the average value of the matrix as a whole, it indicates that the attitude point corresponding to that region has significant changes in the corresponding two spatial directions, and the trend of change is consistent, thus achieving a comprehensive description of the three-dimensional attitude structure.
[0036] Step 305: Flatten the first fusion matrix, the second fusion matrix, and the third fusion matrix into one-dimensional vectors respectively to obtain the first flattened vector, the second flattened vector, and the third flattened vector; concatenate the first flattened vector, the second flattened vector, and the third flattened vector and input them into a nonlinear activation function for mapping to obtain the pose structure feature vector. Specifically, this includes: flattening the first fusion matrix, the second fusion matrix, and the third fusion matrix into one-dimensional vectors according to a fixed row and column order. The purpose is to convert the two-dimensional matrix features into one-dimensional vectors to obtain the first flattened vector, the second flattened vector, and the third flattened vector with uniform dimensions and an ordered arrangement. The three flattened vectors are concatenated end-to-end in the order of the first flattened vector first, the second flattened vector in the middle, and the third flattened vector last, forming a high-dimensional combined feature vector. The purpose is to integrate the gradient fusion features from the three directions into a unified feature. This combined feature vector is then input into a preset nonlinear activation function for numerical mapping and feature compression. The nonlinear activation function uses a monotonic nonlinear mapping method to map the features to a specific numerical range of [0, 1] and enhance the nonlinear expressive power of the features, i.e., the mapped feature value = Through nonlinear transformation and dimension regularization, a posture structure feature vector with fixed dimensions, complete information, and comprehensive representation of the distribution characteristics of human posture structure is finally obtained. This posture structure feature vector represents the distribution characteristics of human posture structure through the specific numerical magnitude, numerical range, and fixed arrangement order of each element in the vector. That is, the first 1 / 3 of the elements of the vector correspond to the features after flattening the first fusion matrix (sagittal-coronal direction), representing the distribution changes and location correlations of posture points within the horizontal cross-section of the human body (parallel to the ground, divided along the vertical axis direction), specifically corresponding to each layer of the horizontal cross-section from head to feet, with each element corresponding to the sagittal-coronal gradient coupling feature of a spatial unit within that cross-section; the middle 1 / 3 of the vector... The elements correspond to the flattened features of the second fusion matrix (sagittal-vertical direction), representing the distribution changes and location associations of posture points within the human sagittal section (parallel to the anterior-posterior direction of the human body, divided along the coronal axis). Specifically, each element corresponds to a sagittal section of the human body from left to right, and each element corresponds to the sagittal-vertical gradient coupling feature of a spatial unit within that section. The last 1 / 3 elements of the vector correspond to the flattened features of the third fusion matrix (coronal-vertical direction), representing the distribution changes and location associations of posture points within the human coronal section (parallel to the lateral direction of the human body, divided along the sagittal axis). Specifically, each element corresponds to a coronal section of the human body from front to back, and each element corresponds to the coronal-vertical gradient coupling feature of a spatial unit within that section.
[0037] After being mapped by a nonlinear activation function, each element's value falls within the specific range of [0, 1]. A larger value indicates a stronger gradient change in the distribution of posture points within the corresponding spatial unit, meaning a more significant difference in the density of posture points at that location. For example, higher values for elements at joints indicate drastic changes in the distribution of posture points at the joints. Conversely, lower values indicate a smoother gradient change in the distribution of posture points within the corresponding spatial unit, meaning a more uniform distribution of posture points at that location. For example, lower values for elements in the middle of the torso indicate a relatively uniform distribution of posture points in the torso. The element arrangement order corresponds one-to-one with the distribution order of the three-dimensional spatial units. Horizontal sections are arranged from head to toe, sagittal sections from left to right, and coronal sections from front to back. The spatial units within each section are arranged sequentially according to a preset grid order, ensuring that the arrangement of vector elements perfectly matches the three-dimensional spatial structure of the human body. This posture structure feature vector can completely and accurately represent the spatial structure of the overall human posture, the distribution positions of each part (head, torso, limbs), and the density variation of posture points in different parts.
[0038] This embodiment forms standardized spatial units by dividing the spatial domain into three-dimensional equidistant segments, providing a unified and quantifiable spatial basis for posture point distribution statistics and improving the standardization and consistency of posture feature extraction. By combining geometric moments, central moments, and moment invariants, it can stably characterize the morphological features of posture points within the spatial unit, unaffected by translation, rotation, or scaling. Weighted fusion of point quantity and morphological similarity ensures that the final count reflects both the density and distribution of points, enhancing the feature's ability to represent real human postures. Calculating gradient sequences from the sagittal, coronal, and vertical directions captures the changing patterns of posture points in different dimensions, comprehensively reflecting differences in human body structure. Constructing a fusion matrix through outer product operations achieves cross-fusion of multi-directional gradient features.
[0039] In a preferred embodiment of the present invention, step 4 involves concatenating the text semantic feature vector, tongue texture feature set, facial color feature set, and posture structure feature vector at a feature layer to obtain a comprehensive constitution feature tensor; inputting the comprehensive constitution feature tensor into a pre-trained classification model to obtain the user's preliminary TCM constitution type; calculating the similarity between the preliminary TCM constitution type and the standard constitution types in a preset standard constitution database, and determining the final TCM constitution type based on the similarity, including: Step 400: Normalize the tongue texture feature set and facial color feature set to obtain the tongue normalized vector and facial normalized vector. Specifically, this includes: performing independent linear normalization on the tongue texture feature set and facial color feature set to eliminate dimensional differences and numerical range deviations between the two types of features; first, traversing the tongue texture feature set and facial color feature set respectively, and selecting the minimum and maximum values of all original feature values in each feature set, where the minimum and maximum values of tongue features are obtained from the tongue texture feature set, and the minimum and maximum values of facial features are obtained from the facial color feature set; for each original feature value in the two feature sets, substitute it into the linear normalization calculation formula for calculation, i.e., normalized feature value = (original feature value - corresponding value). The formula is: (Minimum value of the feature set) ÷ (Maximum value of the corresponding feature set - Minimum value of the corresponding feature set). For each original tongue feature value in the tongue texture feature set, the extreme value of the tongue feature is substituted into the formula to obtain the tongue normalized value. For each original face feature value in the face color feature set, the extreme value of the face feature is substituted into the formula to obtain the face normalized value. All tongue normalized values are arranged in the order of the original tongue texture feature set to form a tongue normalized vector. All face normalized values are arranged in the order of the original face color feature set to form a face normalized vector, ensuring that the normalized vector corresponds to the original feature set. Finally, tongue and face normalized vectors with uniform numerical ranges, standardized distribution, and no dimensional bias are obtained.
[0040] Step 401: Align the text semantic feature vector, tongue image normalized vector, face image normalized vector, and posture structure feature vector along the feature dimension. Concatenate the four aligned vectors end-to-end according to a preset order to obtain a concatenated feature vector. Reshape the concatenated feature vector into a three-dimensional tensor to obtain a comprehensive physical feature tensor. Specifically, this includes performing a unified feature dimension alignment process on the text semantic feature vector, the tongue image normalized vector obtained in step 400, the face image normalized vector, and the extracted posture structure feature vector. This involves retrieving a preset feature dimension standard, specifying that the feature vector has 1024 dimensions and that the feature arrangement rule is sorted from highest to lowest feature variance. The technical consideration of setting the dimension to 1024 dimensions... To balance the integrity of multi-source feature data with the computational efficiency of the algorithm model, this approach fully preserves the core data dimensions of the four feature classes while avoiding computational redundancy caused by excessively high feature dimensions. Dimensional adaptation is performed on each of the four feature vectors. For feature vectors with fewer than 1024 dimensions, zero-value padding is performed at the end of the vector, filling with zero values without altering the original vector's feature data arrangement or numerical information. For feature vectors with more than 1024 dimensions, a variance sorting method is used for feature selection, retaining the 1024 feature dimensions with the largest variance and eliminating redundant feature data. Through these dimension adjustment operations, the four feature vectors achieve complete uniformity in terms of dimension quantity, numerical precision, and arrangement rules.
[0041] After unifying and aligning the feature dimensions, a head-to-tail concatenation operation is performed on the four feature vectors. The preset order is: text semantic feature vector first, tongue image normalized vector in the middle, facial image normalized vector second, and posture structure feature vector last. This order is set to allow the model to prioritize extracting feature data with high correlation to physical constitution determination during feature learning, thereby improving the effectiveness of model computation. During the concatenation operation, the arrangement order of feature data within each feature vector remains unchanged. The beginning of the tongue image normalized vector is seamlessly connected to the end of the text semantic feature vector, the beginning of the facial image normalized vector is connected to the end of the tongue image normalized vector, and the beginning of the posture structure feature vector is seamlessly connected to the end of the facial image normalized vector. Finally, a one-dimensional concatenated feature vector is generated, which integrates four types of feature data: text semantics, tongue texture, facial color, and human posture structure. The total dimension of this vector is 4096 dimensions (1024 dimensions × 4), realizing the complete integration of multi-source physical constitution-related feature data. After generating the one-dimensional concatenated feature vector, a dimension reshaping operation is performed. First, the total dimension of the one-dimensional concatenated feature vector is confirmed to be 4096. Then, according to the preset three-dimensional spatial dimension rules, the total dimension is split into length, width, and height. The preset three-dimensional spatial dimension rules specify that the length, width, and height dimensions of the three-dimensional tensor are 32, 32, and 4, respectively. The 32×32×4 three-dimensional structure can not only adapt to the convolution operation logic of the model, but also realize the spatial division of the four types of feature data. That is, each type of 1024-dimensional feature data is split into a 32×32 two-dimensional matrix, with the height dimension of 4 corresponding to the four types of feature data, so that each type of feature data is spatially divided in three dimensions. Independent feature data layers are formed intermittently. The specific dimensional transformation operation is as follows: the one-dimensional concatenated feature vector is split in the original order of text semantic features, tongue image normalization features, face image normalization features, and posture structure features. First, the first 1024 dimensions of data are split out and arranged into a two-dimensional feature matrix in a 32×32 dimension. Then, the next 1024 dimensions of data are split out and arranged into a second two-dimensional feature matrix in a 32×32 dimension. Next, the next 1024 dimensions of data are split out and arranged into a third two-dimensional feature matrix in a 32×32 dimension. Finally, the remaining 1024 dimensions of data are split out and arranged into a fourth two-dimensional feature matrix in a 32×32 dimension.
[0042] Four 32×32 two-dimensional feature matrices are superimposed along their height dimension to generate a 32 (length) × 32 (width) × 4 (height) three-dimensional tensor structure. This three-dimensional tensor serves as a spatialized data carrier for multi-source feature data, with its spatial dimensions forming a fixed correspondence with the feature data. Specifically, the four levels of the height dimension correspond one-to-one with the four categories of split feature data, and the superposition order is as follows: text semantic feature matrix, tongue image normalized feature matrix, face image normalized feature matrix, and posture structure feature matrix. Each level is an independent 32×32 two-dimensional feature matrix, enabling independent storage and non-interference of the four categories of feature data in the three-dimensional tensor. The 32×32 two-dimensional plane formed by the length and width dimensions serves as the numerical distribution carrier for single-class feature data. Each coordinate point in the plane corresponds to a feature value, which is then normalized. The values are then in the range of 0 to 1, and the magnitude of the values is used to quantify the corresponding feature data. Among them, the first layer of the height dimension is a 32×32 feature matrix, which stores the quantified data of the text semantic features related to TCM consultation. The values of each coordinate point in the matrix are the quantified results of the text semantic features. The second layer of the height dimension is a 32×32 feature matrix, which stores the normalized quantified data of the tongue texture features. The values of each coordinate point in the matrix are the quantified results of the tongue texture-related features. The third layer of the height dimension is a 32×32 feature matrix, which stores the normalized quantified data of the facial color features. The values of each coordinate point in the matrix are the quantified results of the facial color-related features. The fourth layer of the height dimension is a 32×32 feature matrix, which stores the quantified data of the human body posture structure features. The values of each coordinate point in the matrix are the quantified results of the posture structure-related features.
[0043] By using a fixed four-level hierarchical division in the height dimension, four types of feature data—text semantics, tongue image normalization, facial image normalization, and posture structure—are stored in independent levels, achieving physically isolated classification and storage of different types of feature data. Through a fixed coordinate division of a 32×32 two-dimensional plane composed of length and width dimensions, each feature value of each category of 1024-dimensional feature data is mapped to a unique coordinate point on the two-dimensional plane, achieving precise numerical distribution and positioning of single-class feature data. Each coordinate point is a dedicated storage location for the corresponding feature value. The specific numerical value between 0 and 1 at the coordinate point directly corresponds to the quantification index value of the feature data. The closer the value is to 1, the higher the strength of the quantification index of the corresponding feature data; the closer the value is to 0, the lower the strength of the quantification index of the corresponding feature data. This provides an intuitive numerical representation of the quantification index of the feature data, ultimately forming a three-dimensional structured data carrying system of feature type, data location, and quantification index, generating a standardized comprehensive physical feature tensor.
[0044] Step 402: Input the comprehensive constitution feature tensor into the pre-trained classification model. Through multiple fully connected layers of the classification model, perform nonlinear transformations layer by layer to obtain the probability distribution corresponding to each TCM constitution type. Select the constitution type corresponding to the maximum probability as the initial TCM constitution type. Specifically, this includes: inputting the comprehensive constitution feature tensor into a pre-constructed, trained, and parameter-fixed TCM constitution classification model. This TCM constitution classification model is a deep learning classification model designed for multi-dimensional constitution feature recognition. The model construction process involves building the backbone network based on a deep learning network architecture using a stacked fully connected layer approach. Based on the classification requirements of the nine TCM constitution types, determine the three-layer fully connected network structure, the number of neurons in each layer, and the 9-dimensional... The output dimension completes the model structure definition. The model training process involves collecting massive amounts of multi-source data, including text semantics, tongue images, facial images, and posture structures from users. After unified processing in steps 400 to 401, a comprehensive constitution feature sample set with constitution type labels is generated. The sample set is divided into a training set and a validation set in an 8:2 ratio. Using multi-class cross-entropy as the loss function, the Adam adaptive optimization algorithm is used for forward and backward propagation iterative training. The weight parameters and bias parameters are updated round by round based on the gradient descent rule to continuously reduce the classification loss value until the model classification accuracy improves by less than 0.001 for 10 consecutive rounds and reaches the preset convergence threshold of 92%. After training is completed, all parameters are fixed to obtain the TCM constitution classification model.
[0045] The model contains three fully connected layers connected in series with completely fixed parameters. Each fully connected layer consists of a weight matrix and a bias vector. First, the input features are linearly weighted and transformed, then a ReLU activation function is applied to complete the nonlinear mapping. The combination of linear transformation and activation function achieves nonlinear feature extraction, providing the model with deep mining of high-dimensional features, nonlinear transformation, and multi-class recognition capabilities. Specifically, the preset fixed parameter values for each fully connected layer are as follows: the weight parameters of the first fully connected layer range from -0.05 to 0.05, and the bias parameter is fixed at 0.01; the weight parameters of the second fully connected layer range from -0.03 to 0.03, and the bias parameter is fixed at 0.005; the third… The weight parameters of the three fully connected layers range from -0.01 to 0.01, and the bias parameter is fixed at 0.001. Combining the nonlinear distribution and complexity of the constitution feature data, the shallow fully connected layers (the first two layers) use a relatively wide range of weight and bias values to perform preliminary nonlinear transformation and feature decoupling on the input high-dimensional tensor, and amplify the feature differences between different constitution types to improve feature discrimination. The deep fully connected layers (the third layer) use a progressively narrower range of weight and bias values to map the deep abstract features extracted in the previous layer to the output dimensions of the nine TCM constitution types, while suppressing the overfitting of the model to the training samples, ensuring that the model still has stable generalization ability and judgment accuracy on unknown new samples.
[0046] After the comprehensive physical characteristic tensor is input, it first enters the first fully connected layer of the model. This layer first flattens the 32×32×4 three-dimensional tensor, converting it into a 4096-dimensional one-dimensional feature sequence. Then, a linear transformation is performed, which involves multiplying this one-dimensional feature sequence with the preset weight matrix of this layer. The result is then superimposed with the fixed bias parameter of this layer (with a value of 0.01) to obtain the linearly transformed feature value sequence. Subsequently, a ReLU activation function is applied to complete the non-linear transformation. The ReLU activation function sets the values less than 0 in the linear transformation result to 0, and retains the values greater than 0. The first layer performs nonlinear mapping of features and suppresses invalid features. Simultaneously, this layer performs feature compression through feature dimension filtering and integration, ultimately converting the feature sequence into a 2048-dimensional one-dimensional feature vector, completing the extraction of shallow physical characteristics. This 2048-dimensional shallow feature vector is then input into a second fully connected layer. This layer uses the same linear transformation combined with a ReLU activation function for nonlinear transformation, followed by a ReLU activation function for nonlinear mapping to uncover more discriminative deep physical characteristics. Furthermore, it performs feature compression, shrinking the feature vector... The depth is reduced to 1024 dimensions. Finally, the 1024-dimensional deep feature vector is input to the third fully connected layer. This layer first performs a linear transformation on the feature vector to obtain linear feature results, and then connects it to the Softmax activation function to complete the final nonlinear feature mapping. The Softmax activation function performs exponential normalization on the linear feature results, transforming them into output probability values corresponding to various TCM constitution types. The constitution types output by the model cover the full range of standard TCM constitution types, specifically including balanced constitution, qi deficiency constitution, yang deficiency constitution, yin deficiency constitution, phlegm-dampness constitution, damp-heat constitution, and blood stasis constitution. The model classifies nine constitution types, including Qi stagnation constitution and special constitution. For each of these nine constitution types, it outputs an independent floating-point probability value, resulting in nine floating-point values. Each probability value retains six significant decimal places and is limited to a range of 0 to 1, forming a complete and standardized probability distribution of constitution types. The model then iterates through all nine floating-point probability values in the distribution and compares them one by one. The probability value with the highest value is selected, and the TCM constitution type corresponding to this highest probability value in a fixed output order is determined as the preliminary TCM constitution type, thus completing the initial TCM constitution determination.
[0047] Step 403: Obtain the comprehensive constitution feature tensor corresponding to the preliminary TCM constitution type as the feature vector to be matched. Calculate the cosine similarity between the feature vector to be matched and the standard feature vector corresponding to each standard constitution type pre-stored in the standard constitution database to obtain a similarity set. Specifically, this includes: extracting the comprehensive constitution feature tensor corresponding to the preliminary TCM constitution type from step 402 and using it as the feature vector to be matched. During the extraction process, the dimension of this vector is maintained at 32×32×4, completely consistent with the comprehensive constitution feature tensor output in step 401, avoiding dimensional deviations in the feature data; retrieving the system's preset standard constitution database. The construction process of this standard constitution database involves collecting a large sample size of original constitution feature data for each of the nine TCM standard constitution types: balanced constitution, qi deficiency constitution, yang deficiency constitution, yin deficiency constitution, phlegm-dampness constitution, damp-heat constitution, blood stasis constitution, qi stagnation constitution, and special constitution. The samples cover populations from different regions to ensure the quality of the samples. Representativeness and diversity: The original data of each type of constitution collected are standardized according to the entire process of steps 400 to 401, which includes tongue and facial feature normalization, multi-feature vector dimension alignment, splicing, and dimension reshaping to generate a 32×32×4 comprehensive constitution feature tensor for each sample. The element-wise mean of all comprehensive constitution feature tensors after preprocessing for each type of constitution is calculated to generate a unique mean feature vector for each type of constitution, which serves as the standard feature vector for that type of constitution. The nine types of TCM standard constitutions are associated and stored with their corresponding standard feature vectors to establish a one-to-one mapping relationship between feature vectors and constitution types. At the same time, an index and retrieval interface are configured for the database to complete the construction of the standard constitution database. Each type of TCM standard constitution in the database corresponds to a unique standard feature vector, and the standard feature vector and the feature vector to be matched are completely consistent in terms of the number of dimensions, arrangement rules, and numerical precision.
[0048] The cosine similarity between the feature vector to be matched and the standard feature vector corresponding to each standard body type in the standard body type database is calculated sequentially. Cosine similarity is an index that quantifies the similarity between two feature vectors. The closer the similarity value is to 1, the more similar the feature distributions of the two feature vectors are; the closer the similarity value is to 0, the greater the difference in feature distributions between the two feature vectors. That is, the cosine similarity value = (the dot product of the feature vector to be matched and the standard feature vector) ÷ (the magnitude of the feature vector to be matched × the magnitude of the standard feature vector). After the cosine similarity calculation between the feature vector to be matched and all standard feature vectors in the standard body type database is completed, all the calculated similarity values are sorted and summarized to form a similarity set containing all similarity calculation results.
[0049] Step 404: Select the maximum cosine similarity value and its corresponding target standard constitution type from the similarity set, and determine whether the maximum cosine similarity value is greater than a preset similarity threshold. If it is greater, the target standard constitution type is determined as the final TCM constitution type; if it is not greater, the preliminary TCM constitution type is determined as the final TCM constitution type. Specifically, this includes: performing numerical filtering and matching on the similarity set obtained in step 403, extracting the maximum cosine similarity value in the set, and determining the standard constitution type corresponding to the maximum cosine similarity value, which is then used as the target standard constitution type. During the filtering process, the entire similarity set is traversed, matching the maximum similarity value with the corresponding target standard constitution type, and retrieving the system's preset similarity threshold. This similarity threshold ranges from 0.7 to 0.8, i.e., the similarity threshold value is 0.75. When the threshold is lower than... At a threshold of 0.7, the feature matching criteria are too broad, requiring only a low degree of similarity between the feature vector to be matched and the standard feature vector to be considered a match. At this threshold, the distinction between similar constitution types with highly overlapping features is less than 80%, directly leading to low feature matching accuracy and a false constitution type misclassification rate exceeding 25%, increasing the error rate of the final constitution determination. When the threshold is higher than 0.8, the feature matching criteria are too narrow, requiring an extremely high degree of similarity between the feature vector to be matched and the standard feature vector to be considered a match. Even reasonable samples that conform to the constitution determination logic will be excluded due to individual differences, data collection biases, and other objective factors, resulting in more than 30% of reasonable matching results being excluded because the similarity does not reach the threshold. This directly leads to overly stringent feature matching standards, and the model's initial determination results cannot be effectively verified and corrected through the standard constitution database, reducing the practicality of the overall constitution determination process.
[0050] The maximum cosine similarity value obtained from the screening is compared with the preset similarity threshold (0.75). The final constitution type determination is performed according to the preset logic. If the maximum cosine similarity value is greater than the preset similarity threshold, it indicates that the feature vector to be matched has a high similarity with the standard feature vector corresponding to the target standard constitution type. The preliminary TCM constitution type determination result matches the matching result of the standard constitution database. At this time, the target standard constitution type is determined as the final TCM constitution type. If the maximum cosine similarity value is less than or equal to the preset similarity threshold, it indicates that the feature vector to be matched has a low similarity with the standard feature vector corresponding to the target standard constitution type. The preliminary TCM constitution type determination result deviates from the matching result of the standard constitution database. At this time, the matching result of the standard constitution database is abandoned, and the preliminary TCM constitution type obtained in step 402 is directly determined as the final TCM constitution type, completing the entire process of TCM constitution determination.
[0051] This embodiment normalizes the tongue texture and facial color features, unifying the numerical range of multi-source features and eliminating the judgment interference caused by dimensional differences. It integrates and fused four types of features—textual semantics, tongue appearance, facial appearance, and posture structure—to achieve multi-dimensional information integration for TCM consultation, tongue diagnosis, facial diagnosis, and body posture diagnosis, comprehensively covering the core information required for TCM constitution determination. The spliced features are reshaped into a three-dimensional tensor and combined with a pre-trained classification model for preliminary judgment, fully utilizing the nonlinear feature extraction capabilities of deep learning to ensure the efficiency and basic accuracy of the judgment process. A standard constitution database cosine similarity matching and threshold judgment mechanism performs secondary verification and correction of the preliminary judgment results, effectively reducing the probability of model misjudgment.
[0052] like Figure 2 As shown, embodiments of the present invention also provide a machine learning-based TCM constitution identification and analysis system, including: The feature extraction module is used to acquire health description text data input by the user through the human-computer interaction interface, extract the semantic feature vector of the text; collect the user's tongue image, face image and whole body optical image, extract the tongue image texture feature set and face image color feature set, and locate the bilateral acromion point, bilateral anterior superior iliac spine point, seventh cervical vertebra spinous process point and bilateral patellar point in the whole body optical image to obtain the initial posture point set; The configuration domain generation module is used to construct a posture space coordinate system and map the initial posture point set to the posture space coordinate system; determine the configuration origin based on the ratio of the user's shoulder width to pelvic width; construct a spatial configuration domain with the configuration origin as the center and the user's torso length as the sampling radius; The vector generation module is used to divide the spatial modeling domain into multiple spatial units at equal intervals along three directions, and count the number of initial attitude points contained in each spatial unit; calculate the gradient change sequence of the number of initial attitude points between adjacent spatial units along each axis, and cross-merge the gradient change sequences of the three axes to generate attitude structure feature vectors. The type determination module is used to concatenate the text semantic feature vector, tongue texture feature set, facial color feature set and posture structure feature vector into a feature layer to obtain a comprehensive constitution feature tensor; input the comprehensive constitution feature tensor into a pre-trained classification model to obtain the user's preliminary TCM constitution type; calculate the similarity between the preliminary TCM constitution type and the standard constitution type in the preset standard constitution database, and determine the final TCM constitution type based on the similarity.
[0053] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0054] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0055] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for identifying and analyzing TCM constitution based on machine learning, characterized in that, The method includes: Acquire health description text data input by the user through the human-computer interaction interface, and extract the text semantic feature vector; collect the user's tongue image, face image and whole body optical image, extract the tongue image texture feature set and face image color feature set, and locate the bilateral acromion point, bilateral anterior superior iliac spine point, seventh cervical vertebra spinous process point and bilateral patella point in the whole body optical image to obtain the initial posture point set; Construct a posture space coordinate system and map the initial posture point set to the posture space coordinate system; determine the configuration origin based on the ratio of the user's shoulder width to pelvic width; construct a spatial configuration domain with the configuration origin as the center and the user's torso length as the sampling radius; The spatial modeling domain is divided into multiple spatial units at equal intervals along three directions, and the number of initial attitude points contained in each spatial unit is counted. The gradient change sequence of the number of initial attitude points between adjacent spatial units along each axis is calculated, and the gradient change sequences of the three axes are cross-fused in pairs to generate attitude structure feature vectors. The text semantic feature vector, tongue texture feature set, facial color feature set, and posture structure feature vector are concatenated in a feature layer to obtain a comprehensive constitution feature tensor. The comprehensive constitution feature tensor is then input into a pre-trained classification model to obtain the user's preliminary TCM constitution type. The preliminary TCM constitution type is then compared with the standard constitution types in the preset standard constitution database to calculate the similarity and determine the final TCM constitution type based on the similarity.
2. The method for TCM constitution identification and analysis based on machine learning according to claim 1, characterized in that, The extraction of text semantic feature vectors includes: The health description text data is segmented into words to obtain a word sequence; the word sequence is then tagged with parts of speech to obtain a tagged word sequence. Semantic role analysis is performed on the labeled word sequence to obtain a semantic role sequence; the semantic role sequence is then input into a pre-trained language model, and after context encoding by the encoder of the pre-trained language model, the text semantic feature vector is obtained.
3. The method for TCM constitution identification and analysis based on machine learning according to claim 2, characterized in that, The configuration origin is determined based on the ratio of the user's shoulder width to pelvic width; with the configuration origin as the center and the user's torso length as the sampling radius, a spatial configuration domain is constructed, including: The ratio of the distance between the bilateral acromion points to the distance between the bilateral anterior superior iliac spine points is used as the morphological ratio value. If the morphological ratio value is greater than the preset threshold, the association mode is identified as the first mode, and the midpoint of the line connecting the bilateral acromion points is used as the origin of the configuration. If the morphological ratio value is less than or equal to the preset threshold, the association mode is identified as the second mode, and the midpoint of the line connecting the bilateral anterior superior iliac spine points is used as the origin of the configuration. The user's trunk length is calculated based on the vertical distance between the bilateral acromion points and the bilateral anterior superior iliac spine points; a spherical space is constructed with the configuration origin as the center and the trunk length as the radius, and the internal region of the spherical space is the spatial configuration domain.
4. The method for TCM constitution identification and analysis based on machine learning according to claim 3, characterized in that, The spatial configuration domain is divided into multiple spatial units at equal intervals along three directions, and the number of initial attitude points contained in each spatial unit is counted, including: Three orthogonal coordinate axes—sagittal, coronal, and vertical—are established with the configuration origin as a reference. The spatial configuration domain is divided along the sagittal axis at a first preset interval to obtain multiple sagittal slices. Each sagittal slice is divided along the coronal axis at a second preset interval to obtain multiple coronal strips. Each coronal strip is divided along the vertical axis at a third preset interval to obtain multiple spatial units. Traverse all spatial units, count the number of initial attitude points falling into each spatial unit, and obtain the preliminary count value of the initial attitude points of each spatial unit; for each spatial unit, construct a point set based on all initial attitude points in the corresponding spatial unit; calculate the geometric moments of each order of the corresponding point set, construct multiple central moments based on the geometric moments of each order, and normalize the central moments to obtain the normalized central moments. The normalized central moments are combined to obtain a set of moment invariants that are invariant to translation, rotation, and scaling. The moment invariants are compared with the preset standard moment invariant template to obtain the morphological similarity coefficient of the corresponding spatial unit. The initial count of the initial attitude points is fused with the morphological similarity coefficient to obtain the number of initial attitude points of each spatial unit.
5. The method for TCM constitution identification and analysis based on machine learning according to claim 4, characterized in that, Calculate the gradient change sequence of the number of initial attitude points between adjacent spatial cells along each axis, and perform pairwise cross-fusion of the gradient change sequences of the three axes to generate an attitude structure feature vector, including: Along the sagittal axis, the difference in the number of initial attitude points of corresponding spatial units between adjacent sagittal slices is calculated to obtain the first gradient sequence; along the coronal axis, the difference in the number of initial attitude points of corresponding spatial units between adjacent coronal strips is calculated to obtain the second gradient sequence; along the vertical axis, the difference in the number of initial attitude points between adjacent spatial units in the vertical direction is calculated to obtain the third gradient sequence. The first gradient sequence and the second gradient sequence are outer products to obtain the first fusion matrix; the first gradient sequence and the third gradient sequence are outer products to obtain the second fusion matrix; the second gradient sequence and the third gradient sequence are outer products to obtain the third fusion matrix. The first, second, and third fusion matrices are flattened into one-dimensional vectors to obtain the first flattened vector, the second flattened vector, and the third flattened vector. The first, second, and third flattened vectors are concatenated end to end and then input into a nonlinear activation function for mapping to obtain the pose structure feature vector.
6. The method for TCM constitution identification and analysis based on machine learning according to claim 5, characterized in that, The text semantic feature vector, tongue texture feature set, facial color feature set, and posture structure feature vector are concatenated in a feature layer to obtain a comprehensive physical feature tensor, including: The tongue image texture feature set and the face image color feature set are normalized to obtain the tongue image normalized vector and the face image normalized vector. The text semantic feature vector, tongue image normalized vector, face image normalized vector, and posture structure feature vector are aligned in the feature dimension. The four aligned vectors are then concatenated end to end in a preset order to obtain the concatenated feature vector. The concatenated feature vector is then reshaped into a three-dimensional tensor to obtain the comprehensive physical feature tensor.
7. The method for TCM constitution identification and analysis based on machine learning according to claim 6, characterized in that, The comprehensive constitution feature tensor is input into a pre-trained classification model to obtain the user's preliminary TCM constitution type, including: The comprehensive constitution feature tensor is input into the pre-trained classification model. The model undergoes nonlinear transformation layer by layer through multiple fully connected layers to obtain the probability distribution corresponding to each TCM constitution type. The constitution type corresponding to the maximum probability is selected as the preliminary TCM constitution type.
8. The method for TCM constitution identification and analysis based on machine learning according to claim 7, characterized in that, The preliminary TCM constitution type is compared with the standard constitution types in the pre-set standard constitution database. Based on the similarity, the final TCM constitution type is determined, including: Obtain the comprehensive constitution feature tensor corresponding to the preliminary TCM constitution type as the feature vector to be matched, calculate the cosine similarity between the feature vector to be matched and the standard feature vector corresponding to each standard constitution type pre-stored in the standard constitution database, and obtain the similarity set. Select the maximum cosine similarity value and the corresponding target standard constitution type from the similarity set, and determine whether the maximum cosine similarity value is greater than the preset similarity threshold. If it is greater, the target standard constitution type is determined as the final TCM constitution type. If it is not greater, the preliminary TCM constitution type is determined as the final TCM constitution type.
9. A machine learning-based TCM constitution identification and analysis system, wherein the system implements the method as described in any one of claims 1 to 8, characterized in that, include: The feature extraction module is used to acquire health description text data input by the user through the human-computer interaction interface, extract the semantic feature vector of the text; collect the user's tongue image, face image and whole body optical image, extract the tongue image texture feature set and face image color feature set, and locate the bilateral acromion point, bilateral anterior superior iliac spine point, seventh cervical vertebra spinous process point and bilateral patellar point in the whole body optical image to obtain the initial posture point set; The configuration domain generation module is used to construct a posture space coordinate system and map the initial posture point set to the posture space coordinate system; determine the configuration origin based on the ratio of the user's shoulder width to pelvic width; construct a spatial configuration domain with the configuration origin as the center and the user's torso length as the sampling radius; The vector generation module is used to divide the spatial modeling domain into multiple spatial units at equal intervals along three directions, and count the number of initial attitude points contained in each spatial unit; calculate the gradient change sequence of the number of initial attitude points between adjacent spatial units along each axis, and cross-merge the gradient change sequences of the three axes to generate attitude structure feature vectors. The type determination module is used to concatenate the text semantic feature vector, tongue texture feature set, facial color feature set and posture structure feature vector into a feature layer to obtain a comprehensive constitution feature tensor; input the comprehensive constitution feature tensor into a pre-trained classification model to obtain the user's preliminary TCM constitution type; calculate the similarity between the preliminary TCM constitution type and the standard constitution type in the preset standard constitution database, and determine the final TCM constitution type based on the similarity.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 8.