A Traditional Chinese Medicine Syndrome Differentiation Analysis Method and Device Incorporating Prior Knowledge of the Integration of Four Diagnostic Information

By integrating the four diagnosis data sets and dialectical data sets in traditional Chinese medicine dialectical analysis and using multi-core learning model to process, the problem of ignoring data differences in the existing technology is solved, and the performance and learning ability of dialectical algorithms are improved.

CN117334324BActive Publication Date: 2025-07-01GUANGDONG PHARMA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311136909.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-04
Publication Date
2025-07-01
Estimated Expiration
2043-09-04

AI Technical Summary

Technical Problem

The prior art ignores the differences between different types of data in traditional Chinese medicine dialectical analysis, affecting the performance of dialectical algorithms.

Method used

By obtaining the four-diagnostic data set and the dialectical data set, the fusion process is performed to obtain the global affinity matrix, and the multi-core learning model is used to perform multi-core learning processing on the four-diagnostic data set and the global affinity matrix, the sample coefficient matrix and the kernel weight vector are obtained, and finally map it to the output space and cluster it to obtain the results of the Chinese medicine dialectical analysis.

Benefits of technology

The performance of dialectical algorithms has been improved, and by integrating prior knowledge and multi-core learning models, the learning ability of traditional Chinese medicine dialectical decision-making and reasoning strategies has been improved, avoiding learning only the knowledge and experience of a very small number of experts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117334324B_ABST
    Figure CN117334324B_ABST
Patent Text Reader

Abstract

The present invention discloses a traditional Chinese medicine syndrome differentiation analysis method and device integrating prior knowledge of four diagnostic information. The method includes obtaining a four-diagnosis data set and a syndrome differentiation data set; performing fusion processing on the four-diagnosis data set and the syndrome differentiation data set to obtain a global affinity matrix; performing multi-kernel learning processing on the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector; mapping the sample coefficient matrix and the kernel weight vector to an output space according to the sample coefficient matrix and the kernel weight vector to obtain mapped data; and performing clustering processing on the mapped data to obtain a traditional Chinese medicine syndrome differentiation analysis result. By integrating prior knowledge into the four-diagnosis data set and performing syndrome differentiation analysis through a multi-kernel learning model, the embodiment of the present invention can achieve more accurate syndrome differentiation results and can be widely applied to the field of artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for traditional Chinese medicine syndrome differentiation and analysis by fusing four diagnostic information with prior knowledge. Background Art

[0002] Syndrome differentiation and treatment is the basic principle of TCM in understanding and treating diseases, which includes two processes: syndrome differentiation and treatment. Syndrome differentiation is to identify the nature of the disease based on the information collected by the four examinations and combined with TCM theory for comprehensive analysis. With the development of artificial intelligence technology, various machine learning methods have been applied to the syndrome differentiation of TCM to mine information from syndrome differentiation to assist clinical analysis of TCM. However, in the related technology, data items collected by different diagnostic methods are directly combined for data analysis, ignoring the differences between different types of data, which affects the performance of the syndrome differentiation algorithm. In summary, the technical problems existing in the related technology need to be solved urgently. Summary of the invention

[0003] In view of this, an embodiment of the present invention provides a method and device for TCM syndrome differentiation analysis that integrates four diagnostic information with prior knowledge, so as to improve the performance of the syndrome differentiation algorithm.

[0004] On the one hand, the present invention provides a method for TCM syndrome differentiation and analysis by integrating four diagnostic information with prior knowledge, comprising:

[0005] Obtain the four diagnostic data sets and syndrome differentiation data sets;

[0006] The four-diagnosis data set and the syndrome differentiation data set are fused to obtain a global affinity matrix;

[0007] Performing multi-core learning processing on the four-diagnosis data set and the global affinity matrix through a multi-core learning model to obtain a sample coefficient matrix and a kernel weight vector after multi-core learning;

[0008] Mapping the sample coefficient matrix and the kernel weight vector to an output space according to the sample coefficient matrix and the kernel weight vector to obtain mapping data;

[0009] The mapping data is clustered to obtain TCM syndrome differentiation analysis results.

[0010] Optionally, the fusion processing of the four-diagnosis data set and the syndrome differentiation data set to obtain a global affinity matrix includes:

[0011] Performing affinity calculation on the four diagnostic data sets according to the Gaussian kernel function to obtain an initial affinity matrix;

[0012] The initial affinity matrix is ​​fused using the syndrome differentiation data set as prior data to obtain a global affinity matrix.

[0013] Optionally, the affinity calculation process for the four diagnostic data sets according to the Gaussian kernel function to obtain an initial affinity matrix includes:

[0014] Perform normalization processing on the four diagnostic data sets to obtain a normalized data set;

[0015] Perform similarity calculation processing on the normalized data set according to the Gaussian kernel function to obtain a similarity matrix;

[0016] Perform perspective balance processing on the neighborhood information of the similarity matrix to obtain a balance matrix;

[0017] Perform multi-perspective similarity calculation processing on the balance matrix to obtain an initial affinity matrix.

[0018] Optionally, the process of fusing the initial affinity matrix with the syndrome differentiation data set as prior data to obtain a global affinity matrix includes:

[0019] Perform weighted summation processing on the syndrome differentiation data set to obtain prior data;

[0020] Fuse the prior data with the initial affinity matrix to obtain a global affinity matrix.

[0021] Optionally, the process of performing multi-kernel learning on the four diagnostic data sets and the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector after multi-kernel learning includes:

[0022] Adopt a graph embedding method to integrate the four diagnostic data sets and the global affinity matrix to form a minimum optimization problem of multi-kernel learning;

[0023] For the minimum optimization problem of multi-kernel learning, use a two-stage method for iterative optimization to calculate and obtain a sample coefficient matrix and a kernel weight vector.

[0024] Optionally, the process of using a two-stage method for iterative optimization of the minimum optimization problem of multi-kernel learning to calculate and obtain a sample coefficient matrix and a kernel weight vector includes:

[0025] Perform calculation processing on the minimum optimization problem of multi-kernel learning to obtain a first calculation expression;

[0026] Extend the first calculation expression to multi-dimensional data processing to obtain a second calculation expression;

[0027] Perform iterative update processing on the sample coefficient matrix and the kernel weight vector of the second calculation expression to obtain the sample coefficient matrix and the kernel weight vector after the iteration ends.

[0028] Optionally, performing iterative update processing on the sample coefficient matrix and the kernel weight vector of the second calculation expression to obtain the sample coefficient matrix and the kernel weight vector after the iteration ends, including:

[0029] Fix the kernel weight vector, and optimize and update the sample coefficient matrix according to the trace ratio algorithm;

[0030] Fix the sample coefficient matrix, and optimize and update the kernel weight vector according to the semi-definite programming algorithm;

[0031] Return to the step of fixing the kernel weight vector and optimizing and updating the sample coefficient matrix according to the trace ratio algorithm until the iteration end condition is satisfied, and obtain the optimized sample coefficient matrix and kernel weight vector.

[0032] On the other hand, an embodiment of the present invention further provides a traditional Chinese medicine syndrome differentiation analysis device based on fused prior knowledge, including:

[0033] A first module, configured to obtain a four-diagnosis data set and a syndrome differentiation data set;

[0034] A second module, configured to perform fusion processing on the four-diagnosis data set and the syndrome differentiation data set to obtain a global affinity matrix;

[0035] A third module, configured to perform multi-kernel learning processing on the four-diagnosis data set and the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector after multi-kernel learning;

[0036] A fourth module, configured to map the sample coefficient matrix and the kernel weight vector to an output space according to the sample coefficient matrix and the kernel weight vector to obtain mapped data;

[0037] A fifth module, configured to perform clustering processing on the mapped data to obtain a traditional Chinese medicine syndrome differentiation analysis result.

[0038] On the other hand, an embodiment of the present invention further discloses an electronic device, including a processor and a memory;

[0039] The memory is used to store a program;

[0040] The processor executes the program to implement the method as described above.

[0041] On the other hand, an embodiment of the present invention further discloses a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the method as described above.

[0042] On the other hand, an embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the foregoing method.

[0043] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects: In the embodiment of the present invention, a four-diagnosis data set and a syndrome differentiation data set are obtained; the four-diagnosis data set and the syndrome differentiation data set are fused to obtain a global affinity matrix; the syndrome differentiation data can be combined as prior data to improve the learning ability of the multi-kernel learning model and the performance of syndrome differentiation. In the embodiment of the present invention, the multi-kernel learning model is also used to perform multi-kernel learning on the four-diagnosis data set and the global affinity matrix to obtain an optimized sample coefficient matrix and a kernel weight vector; the four-diagnosis data set can be used as different input perspective information and different kernel functions, and the model is learned in a multi-kernel learning manner, so that the multi-kernel learning model learns the decision-making and reasoning strategies of traditional Chinese medicine syndrome differentiation and improves the performance of the syndrome differentiation algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0045] Figure 1 It is a flowchart of a traditional Chinese medicine syndrome differentiation analysis method integrating prior knowledge of four-diagnosis information provided by an embodiment of the present application;

[0046] Figure 2 It is a schematic structural diagram of a traditional Chinese medicine syndrome differentiation analysis device based on integrated prior knowledge provided by an embodiment of the present application;

[0047] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further describes the present application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0049] In current traditional Chinese medicine (TCM) syndrome differentiation, TCM practitioners rely on TCM theoretical knowledge and personal experience to conduct syndrome differentiation through four diagnostic methods: inspection, auscultation and olfaction, interrogation, and palpation. These four diagnostic methods of TCM can be regarded as the TCM practitioners obtaining information on the physical state of the target object from four different perspectives. Due to the complexity of TCM theoretical knowledge and the differences in personal experience, the syndrome differentiation results are somewhat subjective, so different TCM practitioners may have different syndrome differentiation results. In related technologies, various machine learning methods are applied to the syndrome differentiation typing of TCM to mine the information for syndrome differentiation to assist TCM clinical analysis. However, the related methods take the information collected by the four diagnostic methods as undifferentiated inputs, without considering the differences in data distribution and characteristics in the four different information acquisition methods of inspection, auscultation and olfaction, interrogation, and palpation, which affects the effect of syndrome differentiation.

[0050] In view of this, in the embodiments of the present application, a TCM syndrome differentiation analysis method based on integrating prior knowledge is provided. The analysis method in the embodiments of the present application can be applied to a terminal, or to a server, or can also be software running on a terminal or a server, etc. The terminal can be a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0051] Referring to Figure 1 , an embodiment of the present invention provides a TCM syndrome differentiation analysis method that integrates prior knowledge of the four diagnostic information, including:

[0052] S101. Obtain a four-diagnosis data set and a syndrome differentiation data set;

[0053] S102. Perform fusion processing on the four-diagnosis data set and the syndrome differentiation data set to obtain a global affinity matrix;

[0054] S103. Perform multi-kernel learning processing on the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector;

[0055] S104. Map the four-diagnosis data set to an output space according to the sample coefficient matrix and the kernel weight vector to obtain mapped data;

[0056] S105. Perform clustering processing on the mapped data to obtain the TCM syndrome differentiation analysis result.

[0057] In an embodiment of the present invention, a four-diagnosis data set and a syndrome differentiation data set are obtained. The four-diagnosis data set is information obtained by analyzing a target object from four different perspectives of inspection, auscultation and olfaction, inquiry, and palpation; the syndrome differentiation data set is data obtained by syndrome differentiation of the above-mentioned target object of the input data by multiple well-known veteran Chinese medicine practitioners.

[0058] The four-diagnosis data set includes sample data from multiple different perspectives. Each sample data includes M perspective type data, and its expression is as follows:

[0059]

[0060] In the formula, Ω represents the four-diagnosis data set, i represents the sample variable, and m represents the perspective variable. In this embodiment, since there is information from four perspectives of inspection, auscultation and olfaction, inquiry, and palpation in the four-diagnosis data set, the value of the number M of perspective types is 4, and there can be several features in each perspective information.

[0061] The syndrome differentiation data set is obtained by syndrome differentiation of the above-mentioned target object of the input data by multiple well-known veteran Chinese medicine practitioners, and then N t syndrome differentiation result matrices U t (t = 1,..., N t ) are obtained. N t represents the number of well-known veteran Chinese medicine practitioners. In the syndrome differentiation result of the t-th Chinese medicine practitioner, if the target object i and the target object j are differentiated into the same syndrome type, the value of U t (i, j) is 1, and if they are not the same syndrome type, it is 0, as shown in the following formula:

[0062]

[0063] Then, in the embodiment of the present invention, the global affinity matrix that fuses the prior knowledge of well-known veteran Chinese medicine practitioners and the four-diagnosis information is used as the input of the multi-kernel learning model. By integrating the prior knowledge into the multi-kernel learning model, the multi-kernel learning model performs multi-kernel learning processing on the global affinity matrix of the four-diagnosis data set to obtain a sample coefficient matrix and a kernel weight vector. Then, the four-diagnosis data is mapped to the learning representation space by using the sample coefficient matrix, and the mapped data is processed by using a clustering algorithm to obtain the result of traditional Chinese medicine syndrome differentiation analysis, providing a basis for further studying the characteristics of each syndrome type and clinical analysis. It should be noted that in this embodiment, the k-means clustering algorithm is used to cluster the mapped data, and other clustering algorithms such as hierarchical clustering and spectral clustering algorithms can also be used, which all belong to the protection scope of this solution.

[0064] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the identity or characteristics of the target object, such as the information of the target object, the behavioral data of the target object, the historical data of the target object, and the location information of the target object, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive information of the target object, the separate permission or separate consent of the target object will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiments of the present application will be obtained.

[0065] Further as an optional implementation manner, in the above step S102, the fusion processing of the four diagnostic data sets and the syndrome differentiation data sets to obtain a global affinity matrix includes:

[0066] S201. Perform affinity calculation processing on the four diagnostic data sets according to the Gaussian kernel function to obtain an initial affinity matrix;

[0067] S202. Use the syndrome differentiation data sets as prior data to perform fusion processing on the initial affinity matrix to obtain a global affinity matrix.

[0068] In the embodiments of the present invention, in order to make the information collected from each perspective of inspection, auscultation and olfaction, inquiry, and palpation in the four diagnostic data sets not lose generality, the Gaussian kernel function can be used to calculate the affinity between sample data in each perspective to obtain an initial affinity matrix. Then, the syndrome differentiation data sets are used as prior data to perform fusion processing on the initial affinity matrix to obtain a global affinity matrix. By integrating prior knowledge such as the syndrome differentiation results of Chinese medicine practitioners in the four diagnostic data sets, the embodiments of the present invention can enable the subsequent multi-kernel learning model to learn its syndrome differentiation thinking method, synthesize the wisdom of Chinese medicine experts, and improve the learning ability and syndrome differentiation performance of the model.

[0069] Further as an optional implementation manner, in the above step S201, the performing affinity calculation processing on the four diagnostic data sets according to the Gaussian kernel function to obtain an initial affinity matrix includes:

[0070] Perform normalization processing on the four diagnostic data sets to obtain a normalized data set;

[0071] Perform similarity calculation processing on the normalized data set according to the Gaussian kernel function to obtain a similarity matrix;

[0072] Perform perspective balance processing on the neighborhood information of the similarity matrix to obtain a balance matrix;

[0073] Perform multi - perspective similarity calculation on the balance matrix to obtain an initial affinity matrix.

[0074] In the embodiment of the present invention, the input four - diagnosis data set is normalized to obtain a normalized data set, that is, ‖X‖1 = 1, where X represents the four - diagnosis data set. According to the Gaussian kernel function, the similarity of the sample data in the m - th perspective of the normalized data set is calculated to obtain a similarity matrix K m . Let K m (x i ,x j ) represent the similarity between the sample pair <i, j>, and the calculation expression is as follows:

[0075]

[0076] In the formula, θ m is the kernel bandwidth parameter, which is used to control the local action range of the Gaussian kernel function.

[0077] Then, perform perspective - balance processing on the neighborhood information of the similarity matrix to obtain a balance matrix. Specifically, use α m to calculate the balance matrix to balance the contribution of different - perspective information to the neighborhood information encoded by the matrix, where α m is obtained by dividing the variance of the similarity matrix K m by the variance of the kernel with the smallest variance among all M kernels. M represents the number of Gaussian kernel functions, which corresponds to the four perspectives in the four - diagnosis data set in the embodiment of the present invention, and M = 4. Use α m to calculate the balance matrix. The calculation expression is as follows:

[0078]

[0079] (α m ≥1, m = 1, 2, …, M)

[0080] Finally, perform multi - perspective similarity calculation on the balance matrix to obtain an initial affinity matrix W’, so as to quantify the similarity between samples. The calculation expression of W’ is as follows:

[0081]

[0082] Furthermore, as an optional implementation manner, in the above step S202, the process of using the syndrome - differentiation data set as prior data to fuse the initial affinity matrix to obtain a global affinity matrix includes:

[0083] Perform weighted summation on the syndrome - differentiation data set to obtain prior data;

[0084] Fuse the prior data with the initial affinity matrix to obtain a global affinity matrix.

[0085] In the embodiment of the present invention, the prior data of the multi-kernel learning model is obtained by weighted summation of the syndrome differentiation data set, that is, the syndrome differentiation result matrix U t (t = 1, …, N t ), and the calculation expression is as follows:

[0086]

[0087] In the formula, λ t is the weight value of the syndrome differentiation result of the t-th famous traditional Chinese medicine doctor.

[0088] In the embodiment of the present invention, different weight values can be assigned according to aspects such as the professional title and years of practice of the famous traditional Chinese medicine doctor, and the sum of the weight values of the syndrome differentiation results of all famous traditional Chinese medicine doctors is 1. Therefore, the maximum value of the items in the matrix U t is 1. Without loss of generality, in the embodiment of the present invention, it can also be considered that all famous traditional Chinese medicine doctors participating in the syndrome differentiation have the same weight value. Then, λ t = 1 / N t , and corresponding settings can be made according to the actual situation. Then, take the syndrome differentiation results of the famous traditional Chinese medicine doctors as prior knowledge and integrate them into the affinity matrix constructed based on the four diagnostic data sets to obtain a global affinity matrix that fuses the prior data and the four diagnostic data sets The calculation expression is as follows:

[0089]

[0090] In the formula, δ is the prior knowledge weight coefficient of the syndrome differentiation of the famous traditional Chinese medicine doctor, which is used to adjust its role in the global affinity matrix, and its range value is [0, 1].

[0091] Further as an optional implementation manner, in the above step S103, the multi-kernel learning process of the global affinity matrix by the multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector includes:

[0092] S301. Adopt a graph embedding method to integrate the four diagnostic data sets and the global affinity matrix to form a minimum value optimization problem of multi-kernel learning;

[0093] S302. For the minimum value optimization problem of multi-kernel learning, use a two-stage method for iterative optimization to calculate and obtain a sample coefficient matrix and a kernel weight vector.

[0094] In the embodiment of the present invention, the four diagnostic data sets are subjected to feature mapping processing through an integrated kernel function to obtain mapped data samples. Among them, the integrated kernel function has m basic kernel functions The integrated kernel k is obtained in a linear combination manner, and the calculation expression for the sample pair <i,j> is as follows:

[0095]

[0096] s.t.β m ≥0

[0097] In the formula, β m represents the weight coefficient of the m-th kernel function, and M is the number of kernel functions.

[0098] In the embodiment of the present invention, the four diagnostic dataset includes four diagnostic information such as inspection, auscultation and olfaction, inquiry, and palpation, which are the physical state information of the target object obtained from different perspectives. For the four diagnostic datasets obtained from four different perspectives, different kernel functions are used for learning, so the number of kernel functions is 4. Then, the mapping data samples are subjected to graph embedding processing to obtain an embedding vector set, where graph embedding is a process of mapping graph data into low-dimensional dense vectors, providing a unified framework for dimensionality reduction algorithms, and dimensionality reduction of the global affinity data is performed to obtain the embedding vector set. Let X = [x1, x2,..., x N represent a certain sample dataset in the four diagnostic dataset, and v represent the projection vector. It can be seen that v T X = [v T x1, v T x2,..., v T x N represents the data after projection. is the global affinity matrix for fusing prior data and the four diagnostic dataset. For a certain sample dataset in the four diagnostic dataset, its optimal linear embedding vector v * The calculation expression is as follows:

[0099]

[0100]

[0101] In the formula is a diagonal matrix. The global affinity matrix and the diagonal matrix D determine the dimensionality reduction mode of graph embedding.

[0102] Let the function φ map the data from the input space to a new feature space using the integrated kernel function K. In the new feature space, the mapped sample data φ(x i ), (i = 1,..., N) are subjected to graph embedding processing, and the sample data projected into the Euclidean space is represented by v T φ(x i ), (i = 1,..., N), and its calculation expression is as follows:

[0103]

[0104] where α = [α1, …, α N is the sample coefficient vector, and β = [β1, …, β M is the vector composed of each perspective weight value. K i , (i = 1, …, N) is calculated as follows:

[0105]

[0106] where K m (i,j), (m = 1, …, M) represents the affinity between sample pairs <i,j> in the m-th perspective.

[0107] Using the above multi-kernel learning calculation expression and graph embedding technology, a one-dimensional multi-kernel learning optimization problem based on dimensionality reduction can be obtained.

[0108] Further as an optional implementation manner, in the above step S302, for the minimum value optimization problem of the multi-kernel learning, an iterative optimization is performed in a two-stage manner to calculate the sample coefficient matrix and the kernel weight vector, including:

[0109] S3301. Perform calculation processing on the minimum value optimization problem of the multi-kernel learning to obtain a first calculation expression;

[0110] S3302. Extend the first calculation expression to multi-dimensional data processing to obtain a second calculation expression;

[0111] S3303. Perform iterative update processing on the sample coefficient matrix and the kernel weight vector of the second calculation expression to obtain the sample coefficient matrix and the kernel weight vector after the iteration ends.

[0112] In the embodiment of the present invention, calculation processing is performed on the one-dimensional multi-kernel learning optimization problem based on dimensionality reduction to obtain a first calculation expression, and the first calculation expression is as follows:

[0113]

[0114]

[0115] β m ≥0, (m = 1, 2, …, M)

[0116] Then the first calculation expression is subjected to multi-dimensional data expansion processing to obtain a second calculation expression, as follows:

[0117]

[0118]

[0119] β m ≥0, (m = 1, 2, …, M)

[0120] In the formula, A represents a set of N sample coefficient vectors A = [α1, …, α N , β = [β1, …, β M is a vector composed of m perspective kernel weight values. Using the sample coefficient matrix A and the kernel weight vector β, each projection vector v i is determined by the sample coefficient vector α i and the kernel weight vector β. Since it is difficult to directly optimize the second calculation expression, a two-stage method is adopted for iterative optimization to calculate A and β. By performing iterative update processing on the sample coefficient matrix and the kernel weight vector of the second calculation expression, the sample coefficient matrix and the kernel weight vector are obtained.

[0121] Further, as an optional implementation manner, in the above step S3303, the iterative update processing of the sample coefficient matrix and the kernel weight vector of the second calculation expression to obtain the sample coefficient matrix and the kernel weight vector includes:

[0122] Fix the kernel weight vector and optimize and update the sample coefficient matrix according to the trace ratio algorithm;

[0123] Fix the sample coefficient matrix and optimize and update the kernel weight vector according to the semi-definite programming algorithm;

[0124] Return to the step of fixing the kernel weight vector and optimizing and updating the sample coefficient matrix according to the trace ratio algorithm until the iteration end condition is satisfied, and obtain the sample coefficient matrix and the kernel weight vector.

[0125] In the embodiment of the present invention, a two-stage method is adopted for iterative optimization. First, fix the kernel weight vector and optimize and update the sample coefficient matrix according to the trace ratio algorithm. Among them, the optimization problem of the second calculation expression can be expressed by the following expression:

[0126]

[0127]

[0128] Among them, Trace(*) represents the trace of the matrix. and are calculated by the following formula:

[0129]

[0130]

[0131] Among them, the optimization problem of the above expression is the trace ratio problem (Trace Ratio), which can be solved As an approximate solution of the above expression, the ITR (Iterative Trace Ratio) algorithm is used to solve it, and the updated sample coefficient matrix A is obtained.

[0132] Then, the fixed processing is performed on the updated sample coefficient matrix, and the kernel weight vector is optimized and updated according to the semi-definite programming algorithm; specifically, the problem of optimizing the second expression is changed into the following expression:

[0133]

[0134]

[0135] Among them, and are calculated by the following formula:

[0136]

[0137]

[0138] Among them, due to the constraint condition that the kernel weight vector β≥0 exists, the above expression becomes a non-convex quadratic constraint quadratic programming problem, which is an NP-hard problem that requires super-polynomial time to solve. Therefore, a variable B of size M×M is introduced, and the semi-definite programming relaxation is performed on the above expression, and the optimization problem of the above expression is transformed into a semi-definite programming problem. The calculation expression is as follows:

[0139]

[0140]

[0141]

[0142] In the formula, is a column vector, except that the m-th item value is 1 and the rest of the values are 0. B = ββ T The constraint is non-convex and is relaxed to The Lagrange algorithm (Lagrange method) can be used to solve the updated kernel weight vector β.

[0143] Finally, by repeatedly executing the above steps in a loop, A and β are continuously iteratively optimized. Specifically, by returning to the step of fixing the kernel weight vector and optimizing and updating the sample coefficient matrix according to the trace ratio algorithm until the iteration end condition is met, a sample coefficient matrix and a kernel weight vector are obtained. The iteration end condition can be set to when the algorithm converges or after reaching the maximum number of iterations. In the embodiment of the present invention, the convergence condition of the algorithm is that the absolute value of the difference between the target values of two consecutive iterations of the second calculation expression is less than 10 -5 , and the maximum number of iterations is set to 50. Using the updated and optimized sample coefficient matrix and kernel weight vector, the four diagnostic sample data is mapped to the output space, and the calculation expression is as follows:

[0144]

[0145] The output space data is obtained, and finally, the k-means method is used to cluster the data in the output space to obtain a syndrome differentiation data set based on the four diagnostic data set and integrating the syndrome differentiation experience of veteran Chinese medicine doctors, realizing a traditional Chinese medicine syndrome differentiation analysis method that integrates multi-source information of the four diagnostics and the prior knowledge of veteran Chinese medicine doctors.

[0146] In the embodiment of the present invention, by iteratively updating the sample coefficient matrix and the four diagnostic weight vector through multi-kernel learning, and then using the sample coefficient matrix and the kernel weight vector to map the four diagnostic data to the representation space, the four diagnostic data set can be modeled and learned in the way of multi-kernel learning, and the prior data can be integrated, enabling the multi-kernel learning model to directly learn the decision-making and reasoning strategies of traditional Chinese medicine syndrome differentiation from the four diagnostic data set and the syndrome differentiation results of Chinese medicine doctors, avoiding learning only the knowledge and experience of a very small number of experts, and having stronger learning ability.

[0147] Referring to Figure 2 , the embodiment of the present invention also provides a traditional Chinese medicine syndrome differentiation analysis device based on integrating prior knowledge, including:

[0148] The first module 201 is used to obtain the four diagnostic data set and the syndrome differentiation data set;

[0149] The second module 202 is used to perform fusion processing on the four diagnostic data set and the syndrome differentiation data set to obtain a global affinity matrix;

[0150] The third module 203 is used to perform multi-kernel learning processing on the four diagnostic data set and the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector after multi-kernel learning;

[0151] The fourth module 204 is used to map the sample coefficient matrix and the kernel weight vector to the output space according to the sample coefficient matrix and the kernel weight vector to obtain mapped data;

[0152] The fifth module 205 is used to perform clustering processing on the mapped data to obtain a traditional Chinese medicine syndrome differentiation analysis result.

[0153] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0154] Referring to Figure 3 , an embodiment of the present invention further provides an electronic device, including a processor 301 and a memory 302; the memory is used to store a program; the processor executes the program to implement the method as described above.

[0155] Corresponding to Figure 1 the method of, an embodiment of the present invention further provides a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the method as described above.

[0156] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute Figure 1 the method shown.

[0157] In summary, the embodiments of the present invention have the following advantages:

[0158] 1. In the embodiments of the present invention, the information collected by the four different diagnostic methods of inspection, auscultation and olfaction, inquiry, and palpation is used as information from different input perspectives and different kernel functions are adopted, and model learning is carried out in the way of multi-kernel learning, so as to improve the performance of the syndrome differentiation algorithm through the multi-kernel learning model;

[0159] 2. In the embodiments of the present invention, it is not necessary to pre-determine the syndrome differentiation of the target object. The model directly learns the decision-making and reasoning strategies of traditional Chinese medicine syndrome differentiation from the four-diagnosis data set of the target object and the syndrome differentiation results of traditional Chinese medicine doctors, avoiding learning only the knowledge and experience of a very small number of experts, and having stronger learning ability;

[0160] 3. The multi-kernel learning model in the embodiments of the present invention incorporates prior knowledge such as the syndrome differentiation results of traditional Chinese medicine doctors, and different weights can also be assigned in combination with the specialties and levels of traditional Chinese medicine doctors to learn their thinking methods of syndrome differentiation, integrate the wisdom of traditional Chinese medicine experts, and improve the learning ability of the model and the performance of syndrome differentiation.

[0161] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks may sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations described as part of a larger operation are performed independently.

[0162] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, the scope of which is determined by the full scope of the appended claims and their equivalents.

[0163] If the described functions are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence or in part that contributes to the prior art, may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0164] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.

[0165] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0166] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0167] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0168] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0169] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A traditional Chinese medicine syndrome differentiation analysis method based on prior knowledge of the integration of four diagnostic information, characterized in that, The method includes: Obtaining a four-diagnosis data set and a syndrome differentiation data set; Performing fusion processing on the four-diagnosis data set and the syndrome differentiation data set to obtain a global affinity matrix; Performing multi-kernel learning processing on the four-diagnosis data set and the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector after multi-kernel learning; Mapping the sample coefficient matrix and the kernel weight vector to an output space according to the sample coefficient matrix and the kernel weight vector to obtain mapped data; Performing clustering processing on the mapped data to obtain a traditional Chinese medicine syndrome differentiation analysis result; The performing fusion processing on the four-diagnosis data set and the syndrome differentiation data set to obtain a global affinity matrix includes: Performing affinity calculation processing on the four-diagnosis data set according to a Gaussian kernel function to obtain an initial affinity matrix; Using the syndrome differentiation data set as prior data to perform fusion processing on the initial affinity matrix to obtain a global affinity matrix; The performing multi-kernel learning processing on the four-diagnosis data set and the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector after multi-kernel learning includes: Adopting a graph embedding method to integrate the four-diagnosis data set and the global affinity matrix to form a minimum value optimization problem of multi-kernel learning; Performing iterative optimization on the minimum value optimization problem of multi-kernel learning in a two-stage manner to calculate and obtain a sample coefficient matrix and a kernel weight vector.

2. The method according to claim 1, characterized in that, The performing affinity calculation processing on the four-diagnosis data set according to a Gaussian kernel function to obtain an initial affinity matrix includes: Performing normalization processing on the four-diagnosis data set to obtain a normalized data set; Performing similarity calculation processing on the normalized data set according to a Gaussian kernel function to obtain a similarity matrix; Performing perspective balance processing on the neighborhood information of the similarity matrix to obtain a balance matrix; Performing multi-perspective similarity calculation processing on the balance matrix to obtain an initial affinity matrix.

3. The method according to claim 1, characterized in that, The using the syndrome differentiation data set as prior data to perform fusion processing on the initial affinity matrix to obtain a global affinity matrix includes: Performing weighted summation processing on the syndrome differentiation data set to obtain prior data; Performing fusion processing on the prior data and the initial affinity matrix to obtain a global affinity matrix.

4. The method according to claim 1, wherein The performing iterative optimization on the minimum value optimization problem of multi-kernel learning in a two-stage manner to calculate and obtain a sample coefficient matrix and a kernel weight vector includes: Performing calculation processing on the minimum value optimization problem of multi-kernel learning to obtain a first calculation expression; Extending the first calculation expression to multi-dimensional data processing to obtain a second calculation expression; Performing iterative update processing on the sample coefficient matrix and the kernel weight vector of the second calculation expression to obtain a sample coefficient matrix and a kernel weight vector after iteration ends.

5. The method according to claim 4, characterized in that The performing iterative update processing on the sample coefficient matrix and the kernel weight vector of the second calculation expression to obtain a sample coefficient matrix and a kernel weight vector after iteration ends includes: Performing fixed processing on the kernel weight vector and optimizing and updating the sample coefficient matrix according to a trace ratio algorithm; Perform a fixed processing on the sample coefficient matrix, and optimize and update the kernel weight vector according to the semi-definite programming algorithm; Return to the step of performing a fixed processing on the kernel weight vector and optimizing and updating the sample coefficient matrix according to the trace ratio algorithm until the iteration end condition is satisfied, and obtain the optimized sample coefficient matrix and kernel weight vector.

6. A traditional Chinese medicine syndrome differentiation analysis device based on the fusion of prior knowledge, characterized in that The device includes: A first module, configured to obtain a four-diagnosis data set and a syndrome differentiation data set; A second module, configured to perform a fusion processing on the four-diagnosis data set and the syndrome differentiation data set to obtain a global affinity matrix; A third module, configured to perform a multi-kernel learning processing on the four-diagnosis data set and the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector after multi-kernel learning; A fourth module, configured to map the sample coefficient matrix and the kernel weight vector to an output space according to the sample coefficient matrix and the kernel weight vector to obtain mapped data; A fifth module, configured to perform a clustering processing on the mapped data to obtain a traditional Chinese medicine syndrome differentiation analysis result; The second module, configured to perform a fusion processing on the four-diagnosis data set and the syndrome differentiation data set to obtain a global affinity matrix, includes: Perform an affinity calculation processing on the four-diagnosis data set according to a Gaussian kernel function to obtain an initial affinity matrix; Use the syndrome differentiation data set as prior data to perform a fusion processing on the initial affinity matrix to obtain a global affinity matrix; The third module, configured to perform a multi-kernel learning processing on the four-diagnosis data set and the global affinity matrix through a multi-kernel learning model to obtain a sample coefficient matrix and a kernel weight vector after multi-kernel learning, includes: Adopt a graph embedding method to integrate the four-diagnosis data set and the global affinity matrix to form a minimum value optimization problem of multi-kernel learning; For the minimum value optimization problem of the multi-kernel learning, perform iterative optimization in a two-stage manner to calculate and obtain a sample coefficient matrix and a kernel weight vector.

7. An electronic device, characterized in that, The electronic device includes a memory and a processor; The memory is used to store a program; The processor executes the program to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by the processor, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Self-adaptive parameter multiple kernel learning classification method based on large-scale data

    CN103678681A

  • Hyperspectral image classification method based on affinity propagation clustering and sparse multiple kernel learning

    CN105760900A