Traditional Chinese medicine disease differentiation treatment method and system based on machine learning
Through machine learning-based methods, the whole process of traditional Chinese medicine diagnosis and treatment has been standardized and intelligentized, which has solved the problems of inconsistent data quality and difficult to integrate multi-source heterogeneous data in the existing technology, and has achieved objectification of evidence classification and dynamic update of knowledge, improving the accuracy and efficiency of traditional Chinese medicine diagnosis.
Patent Information
- Application Number
- CN202510078771.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
The existing traditional Chinese medicine treatment methods for disease diagnosis and treatment have problems such as inconsistent data quality, difficulty in fusion of multi-source heterogeneous data, lack of quantitative evaluation standards for proof classification and treatment method selection, and lack of systematic mechanisms for knowledge inheritance.
Using a machine learning-based method, the four diagnosis data is transformed through hierarchical semantic standardization processing, multi-dimensional feature transformation is performed using a fusion feature mapping algorithm, deep evidence classification network is used for pattern recognition, treatment serialization process generates prescription compatibility sequences, coordinated correlation calculation and hierarchical clustering analysis extract core features, and target dialectical model library is constructed through progressive knowledge mining and confidence evaluation.
It has realized the standardization and intelligence of the entire process from the collection of four diagnosis data to the accumulation of dialectical treatment knowledge, solved the problems of multi-source heterogeneous data fusion analysis, objectification of proof classification, and dynamic update of knowledge, and improved the accuracy and efficiency of traditional Chinese medicine diagnosis.
Smart Images

Figure CN119993534A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a method and system for disease identification and treatment in Traditional Chinese Medicine based on machine learning. Background Art
[0002] In the existing TCM disease diagnosis and treatment methods, the diagnosis of syndromes and the formulation of treatment plans mainly rely on the doctor's personal experience and subjective judgment. The traditional syndrome differentiation process collects patient symptom information through the four examinations of observation, auscultation, questioning, and palpation. The doctor conducts a comprehensive analysis of this information based on TCM theory and clinical experience to determine the syndrome and formulate a treatment plan. In recent years, with the development of artificial intelligence technology, some studies have begun to try to apply machine learning methods to the field of TCM diagnosis, such as analyzing tongue features through image recognition technology, analyzing sound features through speech processing technology, and analyzing symptom descriptions through natural language processing technology.
[0003] However, the existing methods of TCM disease diagnosis and treatment have the following shortcomings: first, the collection and processing of the four diagnostic data lack a unified standardized process, resulting in uneven data quality and difficulty in effective analysis and utilization; second, the existing machine learning methods mostly analyze a single diagnostic dimension, which makes it difficult to achieve effective integration of multi-source heterogeneous data and fully reflect the holistic characteristics of TCM syndrome differentiation; third, the syndrome classification and treatment selection process lack quantitative evaluation standards, making it difficult to ensure the objectivity and consistency of the diagnostic results; finally, the inheritance and accumulation of TCM knowledge mainly relies on oral transmission of personal experience, and lacks a systematic and standardized knowledge precipitation mechanism. Summary of the invention
[0004] The present application provides a method and system for TCM disease differentiation and treatment based on machine learning, which is used to achieve standardization and intelligence of the entire process from four-diagnosis data collection to syndrome differentiation and treatment knowledge accumulation by constructing a hierarchical data processing flow, and solves technical problems such as multi-source heterogeneous data fusion analysis, objectivity of syndrome type classification, and dynamic updating of knowledge.
[0005] In a first aspect, the present application provides a method for disease differentiation and treatment in traditional Chinese medicine based on machine learning, and the method for disease differentiation and treatment in traditional Chinese medicine based on machine learning comprises: performing clinical data conversion through hierarchical semantic standardization processing according to the patient's four diagnostic information to obtain a traditional Chinese medicine semantic data set, wherein the patient's four diagnostic information includes tongue image data, complexion data, sound data, symptom description data, pulse waveform data and palpation data; performing multi-dimensional feature transformation on the traditional Chinese medicine semantic data set through a fusion feature mapping algorithm to obtain a traditional Chinese medicine feature association matrix; performing pattern recognition through a deep syndrome classification network according to the traditional Chinese medicine feature association matrix to obtain a syndrome feature mapping vector; generating prescription rules based on the syndrome feature mapping vector through treatment method serialization processing to obtain a prescription compatibility sequence; performing multi-dimensional data coupling through collaborative correlation calculation according to the prescription compatibility sequence to obtain a syndrome differentiation feature weight matrix, and performing data dimensionality reduction on the syndrome differentiation feature weight matrix through hierarchical clustering analysis to obtain core syndrome differentiation data indicators; extracting association rules from the core syndrome differentiation data indicators through progressive knowledge mining, screening and verifying the extracted rules through confidence assessment, and obtaining a target syndrome differentiation pattern library.
[0006] In a second aspect, the present application provides a TCM disease differentiation and treatment system based on machine learning, and the TCM disease differentiation and treatment system based on machine learning includes:
[0007] The acquisition module is used to convert clinical data through hierarchical semantic standardization processing according to the patient's four diagnostic information to obtain a TCM semantic data set, wherein the patient's four diagnostic information includes tongue image data, complexion data, voice data, symptom description data, pulse waveform data and palpation data;
[0008] A transformation module, used for performing multi-dimensional feature transformation on the TCM semantic data set through a fusion feature mapping algorithm to obtain a TCM feature association matrix;
[0009] A recognition module, used for performing pattern recognition through a deep syndrome classification network according to the TCM feature association matrix to obtain a syndrome feature mapping vector;
[0010] A generation module, used for generating prescription rules based on the syndrome characteristic mapping vector by serializing the treatment method to obtain a prescription compatibility sequence;
[0011] A coupling module is used to perform multi-dimensional data coupling according to the prescription compatibility sequence through collaborative correlation calculation to obtain a syndrome differentiation feature weight matrix, and perform data dimension reduction on the syndrome differentiation feature weight matrix through hierarchical clustering analysis to obtain core syndrome differentiation data indicators;
[0012] The extraction module is used to extract association rules from the core dialectical data indicators through progressive knowledge mining, screen and verify the extracted rules through confidence evaluation, and obtain a target dialectical pattern library.
[0013] The third aspect of the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the above-mentioned method of disease differentiation and treatment in Traditional Chinese Medicine based on machine learning.
[0014] In the technical solution provided by the present application, clinical data conversion is performed on the patient's four diagnostic information through hierarchical semantic standardization processing, which realizes the standardized collection and preprocessing of multi-source heterogeneous data including tongue image data, complexion data, sound data, symptom description data, pulse waveform data and palpation data, laying a reliable data foundation for subsequent data analysis; a fusion feature mapping algorithm is used for multi-dimensional feature transformation, which effectively solves the problem of feature extraction and fusion of different types of data, so that features from different diagnostic dimensions can be expressed and analyzed in a unified feature space; pattern recognition is performed through a deep syndrome classification network It improves the accuracy and interpretability of syndrome classification and makes the syndrome differentiation process more objective and standardized; it generates prescription rules based on the serialization of treatment methods and realizes the automated reasoning from syndrome to treatment methods and then to prescriptions, thus ensuring the rationality and standardization of prescriptions; it couples and reduces dimensions of multidimensional data through collaborative correlation calculation and hierarchical clustering analysis, effectively extracts the core features of the syndrome differentiation process, reduces data redundancy and improves computational efficiency; finally, through progressive knowledge mining and confidence assessment, it realizes the dynamic accumulation and optimization of syndrome differentiation and treatment knowledge, promotes the inheritance and development of TCM knowledge and improves the accuracy and efficiency of TCM diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0016] Figure 1 This is a schematic diagram of an embodiment of a method for disease identification and treatment in traditional Chinese medicine based on machine learning in an embodiment of the present application;
[0017] Figure 2 This is the RGB color space distribution diagram of the tongue image in the embodiment of the present application;
[0018] Figure 3 is a schematic diagram of the drug combination map in the examples of this application;
[0019] Figure 4 This is a schematic diagram of an embodiment of a TCM disease diagnosis and treatment system based on machine learning in the embodiments of the present application. DETAILED DESCRIPTION
[0020] The present application embodiment provides a kind of TCM disease differentiation and treatment method and system based on machine learning. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable in appropriate circumstances, so that the embodiments described here can be implemented in an order other than the content illustrated or described here. In addition, the terms "including" or "having" and any variation thereof are intended to cover non-exclusive inclusions, for example, the process, method, system, product or equipment comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0021] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of the method for disease differentiation and treatment based on TCM based on machine learning includes:
[0022] Step S101, converting clinical data through hierarchical semantic standardization processing according to the patient's four diagnostic information to obtain a TCM semantic data set, wherein the patient's four diagnostic information includes tongue image data, complexion data, voice data, symptom description data, pulse waveform data and palpation data;
[0023] Step S102, performing multi-dimensional feature transformation on the TCM semantic data set by using a fusion feature mapping algorithm to obtain a TCM feature association matrix;
[0024] Step S103, performing pattern recognition through a deep syndrome classification network according to the TCM feature association matrix to obtain a syndrome feature mapping vector;
[0025] Step S104, generating prescription rules through treatment method serialization processing based on syndrome type feature mapping vectors to obtain a prescription compatibility sequence;
[0026] Step S105, performing multidimensional data coupling by synergistic correlation calculation according to the prescription compatibility sequence to obtain a syndrome differentiation feature weight matrix, and performing data dimension reduction on the syndrome differentiation feature weight matrix by hierarchical clustering analysis to obtain core syndrome differentiation data indicators;
[0027] Step S106: extract association rules from the core syndrome differentiation data indicators through progressive knowledge mining, screen and verify the extracted rules through confidence evaluation, and obtain a target syndrome differentiation pattern library.
[0028] It is understandable that the execution subject of the present application can be a TCM disease diagnosis and treatment system based on machine learning, or a terminal or a server, which is not specifically limited here. The present application embodiment is described by taking the server as the execution subject as an example.
[0029] Specifically, the processing of tongue image data and complexion data is to convert visual features such as color and morphology in the image into numerical features through image semantic segmentation and feature extraction. Taking the tongue image as an example, the tongue area is segmented, the tongue color, tongue coating distribution and other features are extracted, the color information is converted from RGB space to HSV space for quantification, and the texture features of the tongue are analyzed by grayscale co-occurrence matrix to obtain tongue image feature data. The processing of sound data is to convert the audio signal into time-frequency domain through spectrogram analysis to extract characteristic parameters such as frequency and sound intensity of the sound. For the sound features in the four diagnosis of traditional Chinese medicine, the main focus is on the breath, loudness and other features of the patient's voice. The spectrogram is generated by short-time Fourier transform, and the Mel frequency cepstral coefficient features are further extracted. The processing of symptom description data uses the professional vocabulary of traditional Chinese medicine for word segmentation and semantic annotation, and converts the symptom information described by the patient into a structured feature vector. The standardization of symptom description is achieved through the mapping of synonyms and near synonyms of traditional Chinese medicine terms. Taking the headache symptom as an example, different expressions such as "headache", "headache" and "headache" are mapped to the standard symptom terms.
[0030] The pulse waveform data is subjected to time-frequency analysis through wavelet transform to extract the parameters such as the amplitude, frequency, and waveform characteristics of the pulse wave. The palpation data is numerically quantified to convert the tactile characteristics such as pressure and temperature into numerical features. These processed four diagnostic data together constitute the semantic data set of traditional Chinese medicine. The fusion feature mapping algorithm realizes the unified expression of multi-source heterogeneous data through feature space conversion and feature fusion. The features of different modes are normalized, and then the features are projected into the feature space through nonlinear mapping. Finally, the correlation relationship between the features is constructed through feature association analysis to form a traditional Chinese medicine feature association matrix.
[0031] The deep syndrome classification network extracts the key features for syndrome discrimination from the TCM feature association matrix through multi-layer feature learning and generates syndrome feature mapping vectors. The process includes three stages: feature extraction, feature selection, and feature mapping. The network structure design fully considers the hierarchy and correlation of TCM syndromes. The treatment method serialization processing is based on the syndrome feature mapping vector, and matches and combines it through the TCM treatment method knowledge base to generate a reasonable prescription compatibility sequence. The process includes steps such as treatment method rule screening and prescription combination optimization.
[0032] The synergistic correlation calculation is used to analyze the synergistic relationship between the components in the prescription compatibility sequence, and the syndrome differentiation feature weight matrix is generated through multidimensional data coupling. Hierarchical clustering analysis further reduces the dimension of the feature weight matrix and extracts the core syndrome differentiation data indicators. The progressive knowledge mining process extracts the knowledge rules of TCM syndrome differentiation and treatment from the core syndrome differentiation data indicators through association rule analysis. After confidence evaluation and screening verification, these rules constitute the target syndrome differentiation pattern library.
[0033] For example, when dealing with a case of a patient with headache, the tongue image collected from the patient showed a pale red tongue, thin white fur, pale complexion, weak voice, and symptom descriptions including "headache like wrapping", "aversion to cold", "weak pulse" and other features. Through hierarchical semantic standardization processing, the color features of the tongue image are converted into numerical parameters of the HSV color space, and the symptom description is converted into a standardized feature vector through word segmentation and semantic mapping. The fusion feature mapping algorithm maps these multi-source data features into a unified feature space to construct a feature association matrix. The deep syndrome classification network performs pattern recognition based on the feature association matrix to obtain a feature mapping vector pointing to the solar disease syndrome. The treatment method serialization processing generates a treatment method combination based on dispersing wind and cold according to the syndrome characteristics, and then generates a prescription compatibility sequence based on Gegen Decoction. The core syndrome differentiation feature indicators are extracted through collaborative correlation calculation and hierarchical clustering analysis. Through knowledge mining and rule verification, the experience of this diagnosis and treatment process is converted into reusable syndrome differentiation rules and stored in the target syndrome differentiation pattern library.
[0034] In the embodiment of the present application, clinical data conversion is performed on the patient's four diagnostic information through hierarchical semantic standardization processing, which realizes the standardized collection and preprocessing of multi-source heterogeneous data including tongue image data, complexion data, sound data, symptom description data, pulse waveform data and palpation data, laying a reliable data foundation for subsequent data analysis; a fusion feature mapping algorithm is used for multi-dimensional feature transformation, which effectively solves the problem of feature extraction and fusion of different types of data, so that features from different diagnostic dimensions can be expressed and analyzed in a unified feature space; pattern recognition is performed through a deep syndrome classification network, The accuracy and interpretability of syndrome classification are improved, making the syndrome differentiation process more objective and standardized; prescription rules are generated based on the serialization of treatment methods, realizing the automated reasoning from syndrome to treatment methods and then to prescriptions, ensuring the rationality and standardization of prescriptions; multi-dimensional data coupling and dimensionality reduction are carried out through collaborative correlation calculation and hierarchical clustering analysis, effectively extracting the core features of the syndrome differentiation process, reducing data redundancy and improving computing efficiency; finally, through progressive knowledge mining and confidence assessment, the dynamic accumulation and optimization of syndrome differentiation and treatment knowledge is realized, which promotes the inheritance and development of Chinese medicine knowledge and improves the accuracy and efficiency of Chinese medicine diagnosis.
[0035] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0036] (1) Color components are extracted from the tongue image data and complexion data in the patient's four diagnostic information through RGB color space conversion, the extracted color components are standardized through grayscale normalization, texture features are calculated based on the standardized grayscale image through Gabor filtering, and visual feature sequences are obtained through feature vector splicing;
[0037] (2) The sound data is converted into the time-frequency domain by short-time Fourier transform, the converted spectrum is enhanced by Mel frequency conversion, the features are extracted by cepstral coefficient calculation based on the enhanced spectrum, and the auscultation feature sequence is obtained by feature noise reduction processing;
[0038] (3) The symptom description data is divided into semantic units through Chinese word segmentation, and the grammatical structure is analyzed through part-of-speech tagging based on the divided semantic units. The semantic relationship of the analyzed grammatical structure is extracted through dependency syntax parsing, and the symptom feature sequence is obtained through vector space mapping;
[0039] (4) Decomposing the pulse waveform data by wavelet basis function for multi-scale analysis, characterizing the features by energy distribution calculation according to the decomposed coefficients, extracting dynamic features from the characterized features by time-frequency joint distribution, and obtaining a pulse feature sequence by feature fusion;
[0040] (5) The palpation data is converted into numerical values by quantifying the pressure values, and the features are standardized by interval mapping according to the converted values. The spatial features of the standardized features are extracted by multi-point correlation analysis, and the palpation feature sequence is obtained by feature combination;
[0041] (6) The importance of visual feature sequences, auscultatory feature sequences, symptom feature sequences, pulse feature sequences, and palpation feature sequences were evaluated by feature weight calculation. Feature fusion was performed through weighted combination based on the evaluation results. The redundancy of the fused features was eliminated through dimensionality reduction. After feature alignment, a semantic dataset of traditional Chinese medicine was obtained.
[0042] Specifically, for tongue images and facial color data, the RGB color space contains three basic components: red, green, and blue. The basic color features of the original image are obtained by separating these three channels, such as Figure 2As shown, it is the RGB color space distribution diagram of the tongue image in the embodiment of the present application. In order to make the images collected under different light conditions comparable, the extracted RGB components are gray-scale normalized, and the pixel values of each channel are mapped to the range of 0-255. The standardized grayscale image is then subjected to texture feature extraction through a Gabor filter. The Gabor filter simulates the perceptual characteristics of the human visual system and can effectively capture the directional texture information of the image. By setting a Gabor filter group of different directions and scales, the image is convolved to obtain multi-scale and multi-directional texture features. These visual features are spliced in a predefined order to form a visual feature sequence.
[0043] Short-time Fourier transform divides long-term sound signals into small time windows, performs Fourier transform in each window, and obtains the frequency characteristics of sound signals changing with time. The obtained spectrum is converted to Mel frequency, which is closer to the auditory characteristics of the human ear and can highlight the key frequency components in the sound. In the Mel frequency domain, the timbre characteristics of the sound are extracted by cepstrum analysis, and the cepstrum coefficients reflect the harmonic structure and timbre characteristics of the sound. Finally, the influence of environmental noise is eliminated by noise reduction processing to obtain a sequence of auscultation features reflecting the patient's sound characteristics. Chinese word segmentation processing divides continuous text into the smallest semantic units. Considering the particularity of TCM professional terms, a special TCM dictionary is required to support the word segmentation process. The text after word segmentation is tagged with parts of speech to identify the grammatical function of each word, such as noun, verb, adjective, etc. Dependency syntax parsing further analyzes the dependency relationship between words and constructs a syntactic tree to understand the semantic structure of the symptom description. Finally, the text features are mapped to a high-dimensional vector space through a vector space model to form a symptom feature sequence.
[0044] The processing of pulse waveform data adopts wavelet analysis method. Wavelet basis function decomposition decomposes the pulse signal at different scales to obtain wavelet coefficients reflecting different frequency components of the signal. The energy characteristics of the pulse signal are obtained by calculating the energy distribution at each scale. Time-frequency joint distribution analysis reveals the changing law of the pulse signal in the time and frequency dimensions and captures the dynamic characteristics of the pulse. These features are integrated to form a complete pulse feature sequence. The palpation data is processed starting from the quantification of the pressure value. Through professional palpation sensor equipment, the pressure during the doctor's palpation is converted into a numerical signal. These original values are standardized through predefined interval mapping to make the palpation data of different parts comparable. Multi-point correlation analysis examines the correlation between different palpation parts and extracts spatial distribution characteristics. Finally, the features of each palpation point are combined into a palpation feature sequence.
[0045] Feature fusion is a key step in integrating five different types of feature sequences into a semantic dataset of traditional Chinese medicine. The importance weight of each feature sequence is calculated, and the weight value reflects the contribution of the feature to disease diagnosis. The features are weighted and combined according to the weights to ensure that important features have a larger proportion in the feature set. Redundant and irrelevant features are removed through dimensionality reduction technology to retain the most discriminative feature combination. Finally, feature alignment is performed to ensure the consistency of features from different sources in time and space dimensions.
[0046] For example, a patient's tongue image shows a dark red tongue with a thin white coating. The color features are extracted by RGB decomposition, and the color difference of the tongue and tongue coating is highlighted after grayscale normalization. The texture features of the tongue surface are extracted by Gabor filtering. The patient's voice is low and weak. The short-time Fourier transform shows that the energy is concentrated in the low frequency band, and the Mel frequency analysis highlights the qi deficiency characteristics of the voice. The symptom description contains keywords such as "chest tightness", "shortness of breath", and "fatigue". After word segmentation and syntactic analysis, a symptom association network with "qi deficiency" as the core is constructed. The pulse is shown as a deep and thin pulse, and wavelet analysis reveals the low amplitude and slow frequency characteristics of the pulse. Palpation examination found that the abdomen was soft and weak, and the pressure value distribution showed typical characteristics of qi deficiency. The five types of feature sequences were weighted and fused to generate a feature data set reflecting qi deficiency syndrome.
[0047] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0048] (1) Decompose the TCM semantic data set into multiple dimensions, normalize and quantize the decomposed multidimensional data space, and establish a feature mapping set based on the quantized space;
[0049] (2) Divide the feature mapping set into subspaces according to feature attributes, calculate the feature correlation coefficient based on the feature distribution in the subspace, and construct an orthogonal feature vector set;
[0050] (3) Reconstruct the feature space according to the orthogonal feature vector set, determine the feature weights by calculating the correlation entropy, and perform nonlinear transformation on the weighted features to obtain a feature combination sequence;
[0051] (4) Using the feature combination sequence to construct a feature association network, the core features are selected by calculating the node importance, and the relationship strength of the core features is quantified to obtain the feature association matrix;
[0052] (5) Mapping the feature space according to the feature correlation matrix, determining the cluster center of the mapped space by density calculation, and reconstructing the features based on the cluster center to obtain a correlation feature table;
[0053] (6) The correlation feature table is converted into a high-dimensional feature space, and the feature compression is performed according to the spatial distribution. The TCM feature correlation matrix is obtained through feature fusion operation.
[0054] Specifically, a multi-dimensional decomposition of the semantic data set of traditional Chinese medicine is performed. This involves decomposing the feature space containing multi-source heterogeneous data such as tongue image, complexion, voice, symptom description, pulse and palpation into multiple subspaces. The data in each subspace is normalized and quantified, and the data of different dimensions and scales are unified into the same numerical range to form a feature mapping set. When the feature mapping set is divided into subspaces, the features are grouped according to different attribute categories of traditional Chinese medicine syndromes (such as cold and heat, deficiency and excess, qi and blood, yin and yang, etc.). In each subspace, the correlation coefficient between the features is calculated to reflect the degree of association between the features. The orthogonal feature vector set is constructed by the principal component analysis method to eliminate the linear correlation between the features and ensure the independence of the features.
[0055] The feature space is reconstructed based on the orthogonal feature vector set, and the information correlation between features is measured by correlation entropy calculation. The higher the correlation entropy, the greater the information redundancy between features. According to the calculation result of correlation entropy, a corresponding weight value is assigned to each feature. The weighted features are nonlinearly transformed to enhance the expressive power of the features and generate a feature combination sequence.
[0056] The feature combination sequence constructs a feature association network in the form of a network graph, where nodes represent features and edges represent the association between features. The formula for calculating node importance is:
[0057]
[0058] Among them, N importance represents the importance of the node, m represents the number of associated features, α k Indicates the correlation with the k-th feature, β k represents the symptom association strength with the kth feature, γ k Indicates the contribution of the feature in syndrome judgment.
[0059] The core features are selected according to the node importance, and the relationship between these core features is quantified and calculated to construct the feature association matrix. The feature association matrix reflects the strength of association between the features of different syndromes. In the feature space mapping process, the cluster center is determined by calculating the density function of the feature distribution. The areas with higher density indicate that the features are more clustered, and the center points of these areas are used as cluster centers. Based on the cluster centers, feature reconstruction is performed to map the original feature space to the new feature space to form an association feature table.
[0060] The correlation feature table is converted into a high-dimensional feature space, and the distribution characteristics of the features in the high-dimensional space are analyzed to compress the features and delete the redundant information. After the feature fusion operation, the TCM feature correlation matrix reflecting the correlation relationship of TCM syndrome features is obtained.
[0061] For example, the four diagnostic data are multidimensionally decomposed. The tongue data shows that the tongue is light red and the fur is thin and white; the complexion data shows that the complexion is pale; the voice data shows that the voice is low and weak; the symptom description includes features such as "headache like a wrap" and "aversion to cold"; the pulse shows features such as weak pulse. These features are converted into numerical form to form the initial feature mapping set. Then the subspace is divided according to the cold and heat attributes, and the correlation coefficient between the features is calculated. For example, the correlation between "aversion to cold" and "weak pulse" is high, and they are classified into the same subspace. Through feature space reconstruction, the correlation entropy of each feature is calculated. For example, the correlation entropy between "headache like a wrap" and "aversion to cold" is high, indicating that there is a strong correlation between them. In the feature association network, these symptoms form a closely connected node group, and the node importance calculation shows that "aversion to cold" and "weak pulse" are the core features. After feature space mapping and density calculation, a clustering center is formed in the solar disease syndrome feature space. Through high-dimensional feature space transformation and compression, a feature association matrix reflecting the correlation relationship between solar disease syndrome features is obtained, which describes the correlation strength and hierarchical structure between each symptom feature.
[0062] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0063] (1) The TCM feature association matrix is grouped into symptom features, and feature classification and annotation are performed through syndrome type knowledge base mapping. The syndrome type association degree of the annotated features is calculated to obtain the syndrome type feature space;
[0064] (2) Extracting symptom combination patterns from the syndrome feature space, evaluating the association strength based on the semantic similarity between symptom combinations, and calculating the syndrome probability distribution of the association features to obtain the syndrome feature hierarchy;
[0065] (3) The feature importance of the syndrome type feature hierarchy is sorted according to the primary and secondary relationship of the syndrome type, and the features are screened using the syndrome type discrimination threshold. The syndrome type discrimination sequence is constructed based on the screening results;
[0066] (4) Combining the syndrome discrimination sequence according to the TCM syndrome differentiation rules, verifying the combination pattern through syndrome association rules, and quantifying the syndrome association of the verified feature combinations to obtain the syndrome feature set;
[0067] (5) Perform multi-level decomposition on the syndrome feature set, perform feature clustering by calculating syndrome similarity, and extract syndrome features based on cluster centers to obtain a syndrome representation matrix;
[0068] (6) The syndrome representation matrix is transformed through the syndrome mapping function, and the transformed features are fused through syndrome correlation calculation to obtain the syndrome feature mapping vector.
[0069] Specifically, the TCM feature association matrix is grouped according to the type and nature of the symptoms. The symptom feature grouping includes head and face symptom group, chest and abdomen symptom group, limb symptom group, etc. The features in each symptom group are mapped and annotated through the syndrome knowledge base. The annotation process establishes a correspondence between the symptom features and the syndrome knowledge, such as annotating symptoms such as "headache", "fever", and "chills" as solar disease-related features. The annotated features are used to calculate the syndrome correlation degree and construct the syndrome feature space.
[0070] The process of extracting symptom combination patterns in the syndrome feature space is achieved by calculating the semantic similarity and association strength between symptoms. Its mathematical expression is:
[0071]
[0072] Where: P pattern represents the intensity value of the symptom combination pattern; n represents the number of symptom characteristics; m represents the number of syndrome types; λ i represents the weight coefficient of the ith symptom; ω ij represents the contribution of the i-th symptom to the j-th syndrome; ρ j represents the prior probability of the jth syndrome; δ ij Represents the semantic similarity between the i-th symptom and the j-th syndrome.
[0073] For the construction of syndrome feature hierarchy, the features are ranked according to the primary and secondary relationship of syndrome. The primary syndrome features have higher diagnostic value, and the secondary syndrome features serve as auxiliary diagnosis basis. The syndrome discrimination sequence is constructed by setting the syndrome discrimination threshold to filter features. The features in the discrimination sequence are combined according to the TCM syndrome differentiation rules to verify the degree of conformity of the feature combination with the syndrome rules. The syndrome association quantification is performed on the verified feature combination to obtain the syndrome feature set. The syndrome feature set is decomposed at multiple levels and the features are divided according to the diagnostic value at different levels. Feature clustering is performed by syndrome similarity calculation to find symptom groups with similar syndrome features. Typical syndrome features are extracted according to the cluster center to form a syndrome representation matrix.
[0074] The syndrome type representation matrix is converted to a new feature space through the syndrome type mapping function, the syndrome type correlation is calculated for the converted features, and feature fusion is performed to obtain the syndrome type feature mapping vector. The vector contains the key feature information of syndrome type diagnosis.
[0075] For example, for a group of patient data showing headache, fever, and chills, these symptom features are grouped according to the manifestation site and nature. Headache belongs to the head and face symptom group, and fever and chills belong to the systemic symptom group. Through the syndrome knowledge base mapping, these symptoms are marked as solar disease related features. When calculating the syndrome association, it is found that there is a strong correlation between the three symptoms of headache, fever, and chills, and they all point to the solar disease syndrome. In the symptom combination pattern extraction, the semantic similarity of these three symptoms is calculated, and it is found that they have a high co-occurrence probability under the solar disease syndrome. Through the syndrome probability distribution calculation, the distribution weight of this group of symptom features in the solar disease syndrome is determined. According to the primary and secondary relationship of the syndrome, fever and chills as the main symptoms of solar disease have a higher feature importance, and headache as a secondary syndrome has a relatively low weight. After feature combination verification, this group of symptoms fully meets the syndrome differentiation rules of solar disease. After multi-level decomposition, it is found that fever and chills form a highly correlated feature cluster, and headache also has a significant correlation with it. Through syndrome type mapping and feature fusion, a feature mapping vector pointing to the syndrome type of Taiyang disease is generated, which reflects the contribution of symptom characteristics to syndrome type judgment.
[0076] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0077] (1) Decomposing the syndrome feature mapping vector according to the treatment method classification, classifying the features through the treatment method knowledge association, and calculating the treatment method relevance of the classified features to obtain the treatment method feature sequence;
[0078] (2) Extract the core treatment method combination from the treatment method feature sequence, match the drug attributes according to the conversion relationship between the treatment methods, and calculate the prescription correlation degree of the matching results to obtain the prescription treatment method hierarchy;
[0079] (3) Extract prescription rules at the prescription treatment method level, verify the combination through drug compatibility, and establish a drug combination map based on the verification results to obtain a compatibility rule set;
[0080] (4) Classifying drug effects using the compatibility rule set, screening drugs based on prescription association strength, and evaluating prescription combinations based on the screening results to obtain a prescription combination sequence;
[0081] (5) Verifying the compatibility relationship of the prescription combination sequence, optimizing the combination through drug interaction calculation, and reconstructing the prescription sequence according to the optimization results to obtain a prescription sequence table;
[0082] (6) The prescription sequence table is sequenced by calculating the prescription similarity, and the integration result is sequence optimized by verifying the prescription composition rules to obtain the prescription compatibility sequence.
[0083] Specifically, the syndrome feature mapping vector is decomposed according to different treatment methods such as dispersing, clearing heat, and warming and tonifying, and the features are classified by the treatment method knowledge association. The treatment method relevance is calculated for the classified features, and the calculation formula is:
[0084]
[0085] Where: R therapy represents the relevance of treatment methods; k represents the number of syndrome characteristics; l represents the number of treatment method categories; represents the weight of the i-th syndrome feature; ψ ij represents the strength of association between the ith feature and the jth treatment method; ξ j Represents the basic weight of the j-th treatment method.
[0086] The formula for calculating the prescription correlation is:
[0087]
[0088] Among them: F correlation represents the prescription correlation; p represents the number of treatment methods; q represents the number of prescription categories; η x represents the weight of the xth treatment method; μ xy represents the matching degree between the xth treatment method and the yth prescription; σ y Represents the basic efficacy weight of the y-th prescription.
[0089] The prescription rules are extracted at the level of prescription treatment, mainly examining the compatibility relationship between drugs, including mutual dependence, mutual use, mutual fear, etc., and the drug combination map is established through combination verification, such as Figure 3 As shown, it is a schematic diagram of the drug combination map in the embodiment of the present application. A compatibility rule set is formed. The drug effects are classified according to the compatibility rule set, and the drugs are classified according to the medication principle of monarch, minister, assistant and envoy. Drug screening is performed by the prescription association strength to evaluate the rationality of the prescription combination. The compatibility relationship of the prescription combination sequence is verified, and the interaction between drugs, including synergistic, antagonistic and other relationships, is combined and optimized by calculation, and the prescription sequence is reconstructed to form a prescription sequence table. Finally, the sequence is integrated by prescription similarity calculation, and the integration result is verified and optimized by prescription rule to obtain the prescription compatibility sequence.
[0090] Taking Taiyang disease as an example, the characteristics of the treatment method of dispersing wind and cold are extracted from its syndrome feature mapping vector. Through the association of treatment knowledge, it is found that the treatment method of dispersing wind and cold is closely related to the treatment methods of dispelling exterior pathogens and dispelling cold. When calculating the correlation of treatment methods, the weight of the treatment method of dispersing wind and cold is the highest, followed by the treatment method of dispelling exterior pathogens. In the process of drug attribute matching, drugs with the effect of dispersing wind and cold, such as ephedra and cinnamon twig, are found. Through the calculation of prescription association, it is found that the association of Mahuang Decoction and Guizhi Decoction is high. When extracting prescription rules, the relationship between ephedra and almond, and the relationship between cinnamon twig and white peony root are investigated. In the classification of drug effects, ephedra and cinnamon twig are listed as monarch drugs, almond and white peony root are listed as minister drugs, and ginger and jujube are listed as adjuvant drugs. Through the verification of compatibility relationship, the rationality of these drug combinations is confirmed. For example, the compatibility of ephedra, almond and gypsum can dispel wind and cold without hurting the body. The generated prescription compatibility sequence contains complete prescription information and provides a standardized medication scheme for the treatment of Taiyang disease.
[0091] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0092] (1) The prescription compatibility sequence is feature decomposed according to the drug combination relationship, the prescription association strength is analyzed by collaborative similarity calculation, and the analysis results are weighted and quantified to obtain the initial association matrix;
[0093] (2) Perform multidimensional data mapping on the initial correlation matrix through feature cross calculation, calculate the synergy coefficient based on the mapping results, and assign weights to the calculation results to obtain a feature coupling map;
[0094] (3) The characteristic coupling map is evaluated for characteristic weight by calculating the correlation strength, the evaluation results are transformed into a matrix, and the weights are normalized according to the transformation results to obtain the dialectical characteristic weight matrix;
[0095] (4) Hierarchical grouping of the syndrome differentiation feature weight matrix, feature clustering by distance calculation, and center point extraction based on the clustering results to obtain a feature cluster sequence;
[0096] (5) Compressing the feature cluster sequence by dimensionality reduction, reconstructing the compressed results, and performing data fusion based on the reconstructed results to obtain a reduced-dimensionality feature set;
[0097] (6) Feature screening is performed on the reduced dimension feature set through correlation calculation, data integration is performed based on the screening results, and core dialectical data indicators are obtained through feature importance evaluation.
[0098] Specifically, the data processing process of the prescription compatibility sequence starts with the characteristic decomposition of the drug combination relationship, and the calculation formula for the prescription association strength analysis is:
[0099]
[0100] Where: C herb represents the strength of association of drug combination; h represents the number of drugs; c represents the number of compatibility types; υ a represents the utility weight of the a-th drug; κ ab Indicates the matching degree between the a-th drug and the b-th compatibility type; τ b Indicates the basic weight of the bth compatibility type. Each drug in the prescription compatibility sequence is feature decomposed according to the combination relationship of monarch, minister, assistant and envoy, and the similarity of the synergistic effect of each drug combination is calculated to obtain the initial association matrix.
[0101] The initial correlation matrix is mapped to multidimensional data through feature cross calculation, the synergistic effect between different drug combinations is evaluated, and the weights are assigned after calculating the synergistic coefficient to form a feature coupling map. The feature coupling map reflects the interaction strength and direction between drug combinations. The feature coupling map is used to calculate the correlation strength, evaluate the weight contribution of different features, and standardize the weights through matrix transformation to obtain the dialectical feature weight matrix. The matrix contains the relative importance of each dialectical feature.
[0102] The dialectical feature weight matrix is hierarchically grouped to cluster components with similar features. The similarity between features is evaluated by distance calculation to form feature clusters. The center points of each feature cluster are extracted according to the clustering results to generate a feature cluster sequence. The feature cluster sequence is compressed by dimensionality reduction technology to remove redundant information, reconstruct the compressed features, and fuse relevant feature information to form a reduced dimension feature set. Finally, the reduced dimension feature set is subjected to correlation calculation and feature screening. After data integration and importance evaluation, the core dialectical data indicators are obtained.
[0103] Taking the Taiyang disease syndrome as an example, the prescription compatibility sequence is analyzed. Taking Mahuang Decoction as an example, the drugs such as ephedra, cinnamon twig, and almond are feature decomposed according to the combination relationship. Through the calculation of synergistic similarity, it is found that ephedra and almond have a high compatibility association strength, and this relationship is reflected in the initial association matrix. The initial association matrix is subjected to feature cross analysis to evaluate the synergistic effect between the drugs in Mahuang Decoction. The dispersing effect of ephedra and the lung-clearing effect of almond form a synergy, and this relationship is manifested as a high coupling intensity in the feature coupling map.
[0104] The weights of each drug feature are determined by calculating the strength of association. Ephedra has the highest weight as the main drug, followed by apricot kernel as the secondary drug, and ginger and jujube as the adjuvant drugs have lower weights. These weights are normalized to form a standardized syndrome differentiation feature weight matrix. In the hierarchical grouping process, dispersing drugs (such as ephedra and cassia twig) form one feature cluster, and warming and tonic drugs (such as ginger and jujube) form another feature cluster. The central features of each cluster are determined by distance calculation to generate a feature cluster sequence. After dimensionality reduction and feature reconstruction, the drug features with similar effects are merged. Through correlation calculation and data integration, core syndrome differentiation data indicators such as dispersing wind and cold, promoting lung function and relieving asthma are extracted, which accurately reflect the treatment principles and medication characteristics of solar disease.
[0105] In a specific embodiment, the process of executing step S106 may specifically include the following steps:
[0106] (1) Analyze the syndrome type distribution characteristics of the core syndrome differentiation data indicators, calculate the characteristic values by grouping the syndrome type data, and quantify the associated data according to the distribution law of the syndrome type characteristics to obtain the initial syndrome type rule table;
[0107] (2) Calculate data association based on the initial syndrome rule table, quantitatively evaluate the interaction strength between syndrome rules, and construct a syndrome association network through hierarchical mapping to obtain a rule mapping sequence;
[0108] (3) Perform multi-level decomposition of the rule mapping sequence, progressively screen the rules by using the syndrome type association strength threshold, and generate a rule verification sequence based on the matching degree of the syndrome type rules;
[0109] (4) The rule verification sequence is reconstructed according to the syndrome feature weights, the syndrome combination analysis is performed by calculating the correlation between the rules, and the confidence of the analysis results is evaluated to obtain the core rule set;
[0110] (5) Cross-validate the core rule set, integrate the rules by analyzing the logical relationship between syndrome type rules, and optimize the rules based on the integration results to obtain a syndrome differentiation rule map;
[0111] (6) The dialectical rule map is organized through knowledge structuring, the syndrome type association of the organized rules is quantified, and the rule mapping is performed according to the quantification results to obtain the target dialectical pattern library.
[0112] Specifically, the core syndrome differentiation data indicators are analyzed for the distribution characteristics of syndrome types. The process groups the syndrome types according to their main manifestations, such as cold syndrome group, heat syndrome group, deficiency syndrome group, excess syndrome group, etc., and quantifies the characteristic values in each group. The distribution law of syndrome characteristics is determined by counting the frequency and intensity of each syndrome characteristic in each syndrome type, forming an initial syndrome rule table. The data in the initial syndrome rule table are further analyzed for correlation, and the process focuses on the interaction between syndrome rules. The interaction strength is evaluated by the co-occurrence frequency and conditional probability between syndrome characteristics, such as the co-occurrence relationship between fever and aversion to cold in solar disease, and the correlation between chest tightness and shortness of breath in qi deficiency syndrome. These associations are constructed into a multi-level syndrome association network through hierarchical mapping to form a rule mapping sequence.
[0113] During the multi-level decomposition of the rule mapping sequence, a threshold for the strength of syndrome association is set to screen out rules with a higher degree of association than the threshold. These rules are then verified by calculating the matching degree of the syndrome pattern. The matching degree calculation takes into account factors such as the primary and secondary relationship of symptoms and the order of occurrence. Through progressive screening, a rule verification sequence is generated. The rule verification sequence is reconstructed according to the weight of the syndrome characteristics, and the weight value reflects the importance of each feature in the syndrome judgment. The syndrome combination analysis is performed by calculating the correlation between rules to evaluate the effectiveness of different rule combinations. The confidence of the analysis results is evaluated, and the rule combination with high credibility is screened out to form a core rule set.
[0114] The cross-validation process of the core rule set focuses on the logical relationship between rules, including inclusion relationship, transformation relationship, etc. The rules are integrated through logical relationship analysis, similar rules are merged, and conflicting rules are eliminated. The rules are optimized according to the integration results to construct a syndrome differentiation rule map. Finally, the syndrome differentiation rule map is organized through knowledge structuring, and the rules are systematically organized according to syndrome type classification, symptom characteristics, treatment principles and other dimensions. The syndrome type association is quantified for the organized rules, the association relationship between the rules is established, and the target syndrome differentiation pattern library is formed through rule mapping.
[0115] Taking the syndrome type of Taiyang disease as an example, the features related to Taiyang disease were extracted from the core syndrome differentiation data indicators for analysis. These features were grouped according to the manifestation site and nature, such as head symptom group (headache), systemic symptom group (fever, aversion to cold), pulse group (floating pulse), etc. Through eigenvalue calculation, it was found that fever and aversion to cold had a high co-occurrence frequency in Taiyang disease, and the mutual correlation strength was high. These relationships were recorded in the initial syndrome rule table. In the process of data correlation analysis, it was found that there was a close interaction relationship between the three symptoms of fever, aversion to cold, and headache. Through hierarchical mapping, these symptoms were hierarchically organized according to the main syndrome (fever, aversion to cold) and the secondary syndrome (headache), and the syndrome association network was constructed. In the association network, the correlation strength between fever and aversion to cold was the highest, and the correlation strength with headache was second.
[0116] When the rule mapping sequence is decomposed at multiple levels, a higher threshold of association strength is set to retain the core symptom combinations such as fever and aversion to cold. In the process of rule verification, the degree of conformity of these symptom combinations with the syndrome of Taiyang disease is verified by matching calculation. The verification results show that when fever and aversion to cold occur at the same time, the credibility of judging Taiyang disease is the highest. In the process of rule reconstruction, weights are assigned according to the importance of symptoms. Fever and aversion to cold are given higher weights as the main symptoms, and headache is given a relatively lower weight as the secondary symptom. Through syndrome combination analysis, it is confirmed that the combination pattern of symptoms is of great significance for the diagnosis of Taiyang disease. In the process of rule cross-validation, the logical consistency of the rules is further verified, and the key points for distinguishing the syndrome law of Taiyang disease from other six meridian diseases are found. The verified rules are integrated into the syndrome differentiation rule map to form a complete Taiyang disease diagnosis knowledge system. Through knowledge structuring processing, these rules are organized according to the dimensions of symptom characteristics, diagnostic points, treatment principles, etc. to form a Taiyang disease syndrome differentiation model and store it in the target syndrome differentiation model library.
[0117] In a specific embodiment, the process of performing the step of analyzing the syndrome type distribution characteristics of the core syndrome differentiation data indicators may specifically include the following steps:
[0118] (1) The core syndrome differentiation data indicators are classified according to syndrome type categories, and grouped and marked by calculating the correlation between syndrome types to obtain syndrome type data groups;
[0119] (2) Perform frequency statistical analysis on the syndrome data set, extract distribution rules by calculating the frequency of occurrence of syndrome characteristics, and obtain the syndrome distribution matrix;
[0120] (3) The syndrome distribution matrix is subjected to feature correlation analysis through covariance calculation, and the feature correlation coefficient is subjected to threshold screening to obtain a feature association table;
[0121] (4) Perform data quantification conversion according to the feature association table, perform numerical processing through syndrome feature weight calculation, and obtain the syndrome feature vector;
[0122] (5) Normalizing the syndrome feature vectors, standardizing the data through distribution feature mapping, and obtaining a regular feature sequence;
[0123] (6) The rule feature sequence is transformed into a syndrome type rule to perform feature combination, and rules are extracted from the combination results to obtain an initial syndrome type rule table.
[0124] Specifically, the core syndrome differentiation data indicators are classified according to different syndrome types (such as cold syndrome, heat syndrome, deficiency syndrome, excess syndrome, etc.). The strength of the relationship between syndromes is determined by calculating the correlation between different syndromes. The correlation calculation involves multiple dimensions such as the co-occurrence frequency of symptoms, transformation relationship, and degree of mutual influence. According to the results of the correlation calculation, the data are grouped and marked to form a syndrome data group containing relevant syndrome characteristics. The characteristics in the syndrome data group are subjected to frequency statistical analysis, focusing on the frequency of each syndrome feature in different syndromes. The frequency calculation includes the frequency of occurrence of a single symptom and the co-occurrence frequency of multiple symptom combinations. Through this frequency calculation, the distribution law of syndrome characteristics is extracted and a syndrome distribution matrix is constructed. The distribution matrix records the distribution of each syndrome feature in each syndrome.
[0125] The syndrome distribution matrix was subjected to feature correlation analysis after covariance calculation. Covariance calculation reflects the degree of correlation between different features. The larger the covariance value, the stronger the correlation between the features. Feature screening was performed by setting the correlation coefficient threshold, and feature pairs with significant correlation were retained to form a feature association table, which recorded the strength of the correlation between the features. During the data quantification conversion process of the feature association table, an initial weight was assigned to each feature. The weight calculation took into account factors such as the importance, frequency of occurrence, and correlation strength of the feature in syndrome diagnosis. Through this numerical processing, the qualitative feature association relationship was converted into a quantitative syndrome feature vector.
[0126] Normalization of syndrome feature vectors unifies feature values of different dimensions into the same numerical range. Data standardization is performed through distribution feature mapping to ensure comparability between different features. The standardized data forms a regular feature sequence, and each feature in the sequence has a unified numerical representation. The regular feature sequence is combined through syndrome rule conversion. The combination process is based on the theoretical knowledge of traditional Chinese medicine, combining related features together to form syndrome rules. Rules are extracted from the combination results to obtain the initial syndrome rule table, which contains a complete set of syndrome diagnosis rules.
[0127] Taking the syndrome differentiation of Taiyang disease as an example, the relevant core syndrome differentiation data indicators are classified. Fever, aversion to cold, etc. belong to the category of superficial syndrome, headache belongs to the category of head and facial symptoms, and pulse belongs to the category of pulse diagnosis. Through correlation calculation, it is found that the symptoms in the superficial syndrome category have a strong correlation with the head and facial symptoms, and these features are grouped and marked as the Taiyang disease syndrome data group. In the frequency statistical analysis, the frequency of occurrence of symptoms such as fever and aversion to cold in the Taiyang disease syndrome is calculated. Statistics show that fever and aversion to cold have the highest co-occurrence frequency, and this distribution pattern is recorded in the syndrome distribution matrix. Through covariance calculation, it is found that there is a strong correlation between fever and aversion to cold, and this correlation is recorded in the feature association table.
[0128] In the process of data quantification conversion, higher weights are given to the main symptoms such as fever and aversion to cold, and relatively lower weights are given to secondary symptoms such as headache. These weight values constitute the syndrome feature vector. After normalization, all feature values are standardized to the range of 0-1 to form a regular feature sequence. The features in the regular feature sequence are combined according to the theory of traditional Chinese medicine. For example, the combination of fever and aversion to cold constitutes the main diagnostic basis for solar disease, and headache is used as an auxiliary diagnostic feature. These combination rules are extracted to form the initial syndrome rule table.
[0129] The above describes the TCM disease diagnosis and treatment method based on machine learning in the embodiment of the present application. The following describes the TCM disease diagnosis and treatment system based on machine learning in the embodiment of the present application. Figure 4 In the embodiment of the present application, an embodiment of a TCM disease diagnosis and treatment system based on machine learning includes:
[0130] The acquisition module 201 is used to convert clinical data through hierarchical semantic standardization processing according to the patient's four diagnostic information to obtain a TCM semantic data set, wherein the patient's four diagnostic information includes tongue image data, complexion data, voice data, symptom description data, pulse waveform data and palpation data;
[0131] The transformation module 202 is used to perform multi-dimensional feature transformation on the TCM semantic data set through a fusion feature mapping algorithm to obtain a TCM feature association matrix;
[0132] Identification module 203, used to perform pattern recognition through a deep syndrome classification network according to the TCM feature association matrix to obtain a syndrome feature mapping vector;
[0133] A generating module 204 is used to generate prescription rules based on the syndrome characteristic mapping vector by serializing the treatment method to obtain a prescription compatibility sequence;
[0134] A coupling module 205 is used to perform multi-dimensional data coupling according to the prescription compatibility sequence through collaborative correlation calculation to obtain a syndrome differentiation feature weight matrix, and perform data dimension reduction on the syndrome differentiation feature weight matrix through hierarchical clustering analysis to obtain core syndrome differentiation data indicators;
[0135] The extraction module 206 is used to extract association rules from the core syndrome differentiation data indicators through progressive knowledge mining, and to screen and verify the extracted rules through confidence evaluation to obtain a target syndrome differentiation pattern library.
[0136] Through the collaborative cooperation of the above components, the clinical data conversion of the patient's four diagnostic information is carried out through hierarchical semantic standardization processing, which realizes the standardized collection and preprocessing of multi-source heterogeneous data including tongue image data, complexion data, sound data, symptom description data, pulse waveform data and palpation data, laying a reliable data foundation for subsequent data analysis; the fusion feature mapping algorithm is used for multi-dimensional feature transformation, which effectively solves the problem of feature extraction and fusion of different types of data, so that features from different diagnostic dimensions can be expressed and analyzed in a unified feature space; the deep syndrome classification network is used for modeling The system can identify syndromes and improve the accuracy and explainability of syndrome classification, making the syndrome differentiation process more objective and standardized; generate prescription rules based on the serialization of treatment methods, realize the automated reasoning from syndromes to treatment methods and then to prescriptions, and ensure the rationality and standardization of prescriptions; couple and reduce multidimensional data through collaborative correlation calculation and hierarchical clustering analysis, effectively extract the core features of the syndrome differentiation process, reduce data redundancy, and improve computing efficiency; finally, through progressive knowledge mining and confidence assessment, realize the dynamic accumulation and optimization of syndrome differentiation and treatment knowledge, promote the inheritance and development of TCM knowledge, and improve the accuracy and efficiency of TCM diagnosis.
[0137] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the method for disease differentiation and treatment in TCM based on machine learning.
[0138] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.
[0139] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for disease diagnosis and treatment in traditional Chinese medicine based on machine learning, characterized in that: The TCM disease diagnosis and treatment method based on machine learning includes: According to the patient's four diagnostic information, clinical data conversion is performed through hierarchical semantic standardization processing to obtain a TCM semantic data set, wherein the patient's four diagnostic information includes tongue image data, complexion data, voice data, symptom description data, pulse waveform data and palpation data; Performing multi-dimensional feature transformation on the TCM semantic data set by a fusion feature mapping algorithm to obtain a TCM feature association matrix; Performing pattern recognition through a deep syndrome classification network according to the TCM feature association matrix to obtain a syndrome feature mapping vector; Based on the syndrome characteristic mapping vector, formula composition rules are generated by serializing the treatment method to obtain a formula compatibility sequence; According to the prescription compatibility sequence, multi-dimensional data coupling is performed through collaborative correlation calculation to obtain a syndrome differentiation feature weight matrix, and the syndrome differentiation feature weight matrix is subjected to data dimension reduction through hierarchical clustering analysis to obtain core syndrome differentiation data indicators; The core syndrome differentiation data indicators are used to extract association rules through progressive knowledge mining, and the extracted rules are screened and verified through confidence evaluation to obtain a target syndrome differentiation pattern library.
2. The method for disease differentiation and treatment based on machine learning in traditional Chinese medicine according to claim 1, characterized in that: The clinical data conversion is performed through hierarchical semantic standardization processing according to the patient's four diagnostic information to obtain a TCM semantic data set, wherein the patient's four diagnostic information includes tongue image data, complexion data, voice data, symptom description data, pulse waveform data and palpation data, including: According to the tongue image data and complexion data in the patient's four diagnostic information, color components are extracted through RGB color space conversion, the extracted color components are standardized through grayscale normalization, texture features are calculated through Gabor filtering based on the standardized grayscale image, and a visual feature sequence is obtained through feature vector splicing; The sound data is converted into time-frequency domain by short-time Fourier transform, the converted spectrum is feature enhanced by Mel frequency conversion, features are extracted by cepstral coefficient calculation based on the enhanced spectrum, and auscultation feature sequence is obtained by feature noise reduction processing; The symptom description data is divided into semantic units through Chinese word segmentation, grammatical structure analysis is performed through part-of-speech tagging according to the divided semantic units, semantic relationship extraction is performed on the analyzed grammatical structure through dependency syntax parsing, and symptom feature sequence is obtained through vector space mapping; Decomposing the pulse waveform data by wavelet basis function for multi-scale analysis, characterizing the features by energy distribution calculation according to the decomposed coefficients, extracting dynamic features from the characterized features by time-frequency joint distribution, and obtaining a pulse feature sequence by feature fusion; The palpation data is converted into numerical values by quantifying the pressure value, and features are standardized by interval mapping according to the converted values, and spatial features are extracted from the standardized features by multi-point correlation analysis, and a palpation feature sequence is obtained by feature combination; The importance of the visual feature sequence, auscultation feature sequence, symptom feature sequence, pulse feature sequence and palpation feature sequence is evaluated respectively by feature weight calculation, and feature fusion is performed through weighted combination according to the evaluation results. The redundancy of the fused features is eliminated through dimensionality reduction, and the TCM semantic data set is obtained after feature alignment.
3. The method for disease differentiation and treatment based on machine learning in traditional Chinese medicine according to claim 1, characterized in that: The multi-dimensional feature transformation of the TCM semantic data set is performed by the fusion feature mapping algorithm to obtain a TCM feature association matrix, including: Decomposing the TCM semantic data set into multiple dimensions, normalizing and quantizing the decomposed multidimensional data space, and establishing a feature mapping set according to the quantized space; Dividing the feature mapping set into subspaces according to feature attributes, calculating feature correlation coefficients according to feature distribution in the subspaces, and constructing an orthogonal feature vector set; Reconstructing the feature space according to the orthogonal feature vector set, determining feature weights by correlation entropy calculation, and performing nonlinear transformation on the weighted features to obtain a feature combination sequence; Using the feature combination sequence to construct a feature association network, screening core features through node importance calculation, and quantifying the relationship strength of the core features to obtain a feature association matrix; Performing feature space mapping according to the feature association matrix, determining cluster centers of the mapped space by density calculation, and reconstructing features based on the cluster centers to obtain an associated feature table; The associated feature table is converted into a high-dimensional feature space, feature compression is performed according to the spatial distribution, and the TCM feature association matrix is obtained through feature fusion operation.
4. The method for disease differentiation and treatment based on machine learning in traditional Chinese medicine according to claim 1, characterized in that: The method of performing pattern recognition through a deep syndrome classification network according to the TCM feature association matrix to obtain a syndrome feature mapping vector includes: The TCM feature association matrix is grouped into symptom features, feature classification and annotation are performed through syndrome type knowledge base mapping, and syndrome type association degree is calculated for the annotated features to obtain syndrome type feature space; Extracting symptom combination patterns from the syndrome feature space, evaluating association strength according to semantic similarity between symptom combinations, and calculating syndrome probability distribution of association features to obtain syndrome feature hierarchy; The syndrome feature hierarchy is sorted by feature importance according to the primary and secondary relationship of the syndrome, features are screened by syndrome discrimination threshold, and a syndrome discrimination sequence is constructed according to the screening results; Combining the syndrome discrimination sequence according to the syndrome differentiation rules of traditional Chinese medicine, verifying the combination mode through syndrome association rules, and quantifying the syndrome association of the verified feature combinations to obtain a syndrome feature set; Decomposing the syndrome feature set at multiple levels, clustering the features by calculating the syndrome similarity, extracting syndrome features according to cluster centers to obtain a syndrome representation matrix; The syndrome type representation matrix is subjected to feature conversion through a syndrome type mapping function, and the converted features are subjected to feature fusion through syndrome type association calculation to obtain the syndrome type feature mapping vector.
5. The method for disease differentiation and treatment based on machine learning in traditional Chinese medicine according to claim 1, characterized in that: The method of generating prescription rules based on the syndrome characteristic mapping vector by serializing the treatment method to obtain a prescription compatibility sequence includes: Decomposing the syndrome feature mapping vector according to the treatment method classification, classifying the features through the treatment method knowledge association, and calculating the treatment method relevance of the classified features to obtain the treatment method feature sequence; Extracting a core treatment method combination from the treatment method feature sequence, matching drug attributes according to the conversion relationship between treatment methods, and calculating the prescription correlation degree of the matching results to obtain the prescription treatment method hierarchy; Extracting prescription rules from the prescription treatment method level, performing combination verification through drug compatibility relationship, and establishing a drug combination map according to the verification result to obtain a compatibility rule set; Classify the drug effects of the compatibility rule set, screen the drugs according to the prescription association strength, and evaluate the prescription combination according to the screening results to obtain a prescription combination sequence; Verifying the compatibility relationship of the prescription combination sequence, optimizing the combination through drug interaction calculation, and reconstructing the prescription sequence according to the optimization result to obtain a prescription sequence table; The prescription sequence table is sequenced and integrated by prescription similarity calculation, and the integration result is sequence optimized by formula composition rule verification to obtain the prescription compatibility sequence.
6. The method for disease differentiation and treatment based on machine learning in traditional Chinese medicine according to claim 1, characterized in that: The multi-dimensional data coupling is performed according to the prescription compatibility sequence through collaborative correlation calculation to obtain a syndrome differentiation feature weight matrix, and the syndrome differentiation feature weight matrix is subjected to data dimension reduction through hierarchical clustering analysis to obtain core syndrome differentiation data indicators, including: Decomposing the prescription compatibility sequence according to the drug combination relationship, analyzing the prescription association strength by collaborative similarity calculation, and weighting and quantifying the analysis results to obtain an initial association matrix; Performing multi-dimensional data mapping on the initial correlation matrix through feature cross calculation, calculating the synergy coefficient according to the mapping result, and performing weight distribution on the calculation result to obtain a feature coupling map; The characteristic coupling map is evaluated for characteristic weight by calculating the correlation strength, the evaluation result is transformed into a matrix, and the weight is normalized according to the transformation result to obtain the dialectical characteristic weight matrix; The syndrome differentiation feature weight matrix is hierarchically grouped, feature clustering is performed by distance calculation, and center points are extracted according to the clustering results to obtain a feature cluster sequence; The feature cluster sequence is compressed by dimensionality reduction, the compressed result is reconstructed, and data fusion is performed according to the reconstructed result to obtain a reduced-dimensionality feature set; The reduced dimension feature set is subjected to feature screening through correlation calculation, data integration is performed according to the screening results, and the core dialectical data index is obtained through feature importance evaluation.
7. The method for disease differentiation and treatment based on machine learning in traditional Chinese medicine according to claim 1, characterized in that: The core syndrome differentiation data indicators are extracted through progressive knowledge mining for association rules, and the extracted rules are screened and verified through confidence evaluation to obtain a target syndrome differentiation pattern library, including: The core syndrome differentiation data indicators are analyzed for syndrome type distribution characteristics, characteristic values are calculated by syndrome type data grouping, and the associated data are quantified according to the distribution law of syndrome type characteristics to obtain an initial syndrome type rule table; Calculating data association according to the initial syndrome type rule table, quantitatively evaluating the interaction strength between syndrome type rules, and constructing a syndrome type association network through hierarchical mapping to obtain a rule mapping sequence; Decomposing the rule mapping sequence at multiple levels, progressively screening the rules by using the syndrome type association strength threshold, and generating a rule verification sequence by combining the matching degree calculation of the syndrome type rule; The rule verification sequence is reconstructed according to the syndrome characteristic weights, syndrome combination analysis is performed by calculating the correlation between rules, and the confidence of the analysis results is evaluated to obtain a core rule set; Cross-validation is performed on the core rule set, rule integration is performed through logical relationship analysis between syndrome type rules, and rule optimization is performed based on the integration results to obtain a syndrome differentiation rule map; The syndrome differentiation rule map is organized by knowledge structuring, the syndrome type association quantification is performed on the organized rules, and rule mapping is performed according to the quantification results to obtain the target syndrome differentiation pattern library.
8. The method for disease differentiation and treatment based on machine learning in traditional Chinese medicine according to claim 1, characterized in that: The core syndrome differentiation data indicators are subjected to syndrome type distribution feature analysis, characteristic value calculation is performed by syndrome type data grouping, and the associated data is quantified according to the distribution law of syndrome type characteristics to obtain an initial syndrome type rule table, including: Classify the core syndrome differentiation data indicators according to syndrome type categories, and group and mark them by calculating the correlation between syndrome types to obtain syndrome type data groups; Performing frequency statistical analysis on the syndrome data group, extracting distribution rules by calculating the frequency of occurrence of syndrome characteristics, and obtaining a syndrome distribution matrix; Perform feature correlation analysis on the syndrome type distribution matrix through covariance calculation, perform threshold screening on the feature correlation coefficient, and obtain a feature association table; Perform data quantization conversion according to the feature association table, perform numerical processing through syndrome type feature weight calculation, and obtain a syndrome type feature vector; Normalizing the syndrome characteristic vector, performing data standardization through distribution feature mapping, and obtaining a regular feature sequence; The rule feature sequence is transformed into a syndrome type rule to perform feature combination, and rule extraction is performed on the combination result to obtain the initial syndrome type rule table.
9. A TCM disease differentiation and treatment system based on machine learning, used to implement the TCM disease differentiation and treatment method based on machine learning as described in any one of claims 1 to 8, characterized in that: The TCM disease diagnosis and treatment system based on machine learning includes: The acquisition module is used to convert clinical data through hierarchical semantic standardization processing according to the patient's four diagnostic information to obtain a TCM semantic data set, wherein the patient's four diagnostic information includes tongue image data, complexion data, voice data, symptom description data, pulse waveform data and palpation data; A transformation module, used for performing multi-dimensional feature transformation on the TCM semantic data set through a fusion feature mapping algorithm to obtain a TCM feature association matrix; A recognition module, used for performing pattern recognition through a deep syndrome classification network according to the TCM feature association matrix to obtain a syndrome feature mapping vector; A generation module, used for generating prescription rules based on the syndrome characteristic mapping vector by serializing the treatment method to obtain a prescription compatibility sequence; A coupling module is used to perform multi-dimensional data coupling according to the prescription compatibility sequence through collaborative correlation calculation to obtain a syndrome differentiation feature weight matrix, and perform data dimension reduction on the syndrome differentiation feature weight matrix through hierarchical clustering analysis to obtain core syndrome differentiation data indicators; The extraction module is used to extract association rules from the core dialectical data indicators through progressive knowledge mining, screen and verify the extracted rules through confidence evaluation, and obtain a target dialectical pattern library.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, a method for disease differentiation and treatment in traditional Chinese medicine based on machine learning as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Traditional Chinese medicine prescription efficacy quantification method and system
CN120853725A
Traditional Chinese medicine tumor collaborative treatment clinical data analysis method
CN120853773A
Traditional Chinese medicine health assessment method and system based on intelligent analysis
CN121034642A