Syndrome matching method, device and equipment and storage medium
By integrating a syndrome element knowledge base and multiple machine learning models, a list of syndrome elements is generated and similarity matching is performed, which solves the problems of accuracy and interpretability in TCM syndrome matching and achieves high-precision syndrome differentiation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2022-05-20
- Publication Date
- 2026-05-08
AI Technical Summary
In existing technologies, TCM syndrome matching methods are not very accurate and lack interpretability, text similarity matching methods cannot perform semantic recognition, and embedding vector methods lack interpretable reasoning basis.
By acquiring the symptom sequences to be matched, the symptom sequences are mapped into a list of symptom elements using a symptom element knowledge base and a pre-trained symptom element mapping model. Multiple machine learning models, such as LDA topic model and variational autoencoder, are integrated to generate the final list of symptom elements. A weighted calculation is then performed using an attention model, and finally, similarity matching is performed with the known list of symptom elements.
This improved the accuracy of syndrome matching and, combined with the TCM syndrome element differentiation theory, ensured the interpretability of the matching results, thereby enhancing the accuracy and reliability of TCM syndrome differentiation.
Smart Images

Figure CN115188462B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a type matching method, apparatus, device and storage medium. Background Technology
[0002] In Traditional Chinese Medicine (TCM), a "syndrome pattern" is a characteristic (or symptom) response and summary of a certain stage in the cycle or development of a disease. In the TCM process of "disease differentiation," "syndrome differentiation" is a crucial step. Syndrome differentiation and treatment are fundamental principles of TCM. The thought process of TCM syndrome differentiation involves "identifying the location, nature, and pathogenic factors based on clinical symptoms, and then combining these factors to form a syndrome pattern." Traditional syndrome differentiation primarily relies on the physician's personal knowledge and experience.
[0003] With the development of artificial intelligence technology, the process of TCM syndrome differentiation can now be achieved through methods such as text similarity matching and deep networks. However, text similarity matching methods such as edit distance and cosine distance can only identify the distance between texts, but cannot perform semantic recognition of the text. Its ability to match TCM terminology is limited, resulting in low accuracy. In contrast, the method of similarity matching using text embedding vectors and deep networks first maps the two syndromes to be matched to embedding vectors, then inputs them into the deep network model, and finally outputs the syndrome similarity. Syndromes with higher similarity are retrieved. This method utilizes multi-dimensional text vector information, resulting in better text matching performance. However, this method of using embedding vectors lacks interpretable reasoning for applications in the TCM industry. Summary of the Invention
[0004] This application provides a syndrome matching method, apparatus, device, and storage medium to address the problems of low accuracy and lack of interpretability in existing syndrome matching methods in the field of traditional Chinese medicine.
[0005] To address the aforementioned technical problems, this application provides a method for syndrome matching, comprising: acquiring a symptom sequence to be matched; mapping the symptom sequence to be matched using a syndrome element knowledge base to obtain a first list of syndrome elements to be matched, wherein the syndrome element knowledge base is constructed based on pre-obtained empirical knowledge of syndrome elements; inputting the symptom sequence to be matched into a pre-trained syndrome element mapping model to obtain a second list of syndrome elements to be matched, wherein the syndrome element mapping model is trained based on pre-obtained sample syndromes, the corresponding list of syndrome elements, and the syndrome element knowledge base; integrating the first and second lists of syndrome elements to be matched to obtain a final list of syndrome elements to be matched; and performing similarity matching between the final list of syndrome elements to be matched and the list of known syndrome elements corresponding to each known syndrome to obtain and output the matched target known syndrome.
[0006] As a further improvement to this application, the first list of testimonial elements to be matched and the second list of testimonial elements to be matched are integrated to obtain a final list of testimonial elements to be matched, including: merging the first list of testimonial elements to be matched and the second list of testimonial elements to be matched to obtain a final list of testimonial elements to be matched.
[0007] As a further improvement of this application, integrating the first list of testimonial elements to be matched and the second list of testimonial elements to be matched to obtain the final list of testimonial elements to be matched includes: converting the first list of testimonial elements to be matched into a first vectorized representation and converting the second list of testimonial elements to be matched into a second vectorized representation; inputting the first vectorized representation and the second vectorized representation into a pre-trained attention model for weighted calculation to obtain the final vectorized representation, wherein the attention model is trained based on the sample testimonial type and the target testimonial element list of the corresponding target testimonial type; and converting the final vectorized representation into the final list of testimonial elements to be matched.
[0008] As a further improvement of this application, the attention model training method includes: obtaining sample certificate types and corresponding target certificate types, as well as a list of target certificate elements for the target certificate types; mapping the sample certificate types using a certificate element knowledge base to obtain a first list of sample certificate elements, and analyzing the sample certificate types using a certificate element mapping model to obtain a second list of sample certificate elements; converting the first and second list of sample certificate elements into vectorized representations, and constructing a query vector using the vectorized representations of the first and second list of sample certificate elements as elements; inputting the query vector into the attention model to be trained, outputting the final sample vectorized representation, and converting the final sample vectorized representation into the final list of sample certificate elements; and updating the weight parameters of the attention model in reverse based on the final list of sample certificate elements, the target certificate element list, and a pre-constructed loss function.
[0009] As a further improvement of this application, the syndrome element mapping model is implemented based on the LDA topic model; inputting the symptom sequence to be matched into the pre-trained syndrome element mapping model includes: inputting the symptom sequence to be matched into the pre-trained LDA topic model, using the LDA topic model to extract vector representations of multiple topics from the symptom sequence to be matched, the LDA topic model being trained based on pre-acquired sample syndrome types and corresponding syndrome element lists and syndrome element knowledge base; and converting the vector representations of multiple topics into a second list of syndrome elements to be matched.
[0010] As a further improvement of this application, the syndrome element mapping model is implemented based on a variational autoencoder; the input of the symptom sequence to be matched into the pre-trained syndrome element mapping model includes: inputting the symptom sequence to be matched into the pre-trained variational autoencoder to obtain the intermediate layer representation of the variational autoencoder, the variational autoencoder being trained based on pre-prepared sample syndrome types, during the training of the variational autoencoder, the input data of the input layer is the sample syndrome type, the output data of the output layer is the autoencoder output syndrome type, and the output of the intermediate layer is the syndrome element representation of the sample syndrome type; the parameters of the intermediate layer are trained by fitting the sample syndrome type and the autoencoder output syndrome type; and the intermediate layer representation is used as the second list of syndrome elements to be matched.
[0011] As a further improvement to this application, the final list of evidence elements to be matched is compared with the list of known evidence elements corresponding to each known evidence type to obtain the matched target known evidence type and output it. This includes: calculating the intersection and union of the list of evidence elements to be matched and the list of known evidence elements, where both the final list of evidence elements to be matched and the list of known evidence elements contain multiple evidence elements; calculating the ratio of the intersection to the union to obtain the similarity value between the list of evidence elements to be matched and the list of known evidence elements; and selecting the target known evidence element corresponding to the list of known evidence elements with the highest similarity value for output.
[0012] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a syndrome matching device, comprising: a first mapping module, used to acquire a symptom sequence to be matched, and to map the symptom sequence to be matched using a syndrome element knowledge base to obtain a corresponding first list of syndrome elements to be matched, wherein the syndrome element knowledge base is constructed based on pre-obtained syndrome element experience knowledge; a second mapping module, used to input the symptom sequence to be matched into a pre-trained syndrome element mapping model to obtain a second list of syndrome elements to be matched, wherein the syndrome element mapping model is trained based on pre-obtained sample syndromes, the corresponding list of syndrome elements, and the syndrome element knowledge base; an integration module, used to integrate the first list of syndrome elements to be matched and the second list of syndrome elements to be matched to obtain a final list of syndrome elements to be matched; and a matching module, used to perform similarity matching between the final list of syndrome elements to be matched and the list of known syndrome elements corresponding to each known syndrome to obtain the matched target known syndrome and output it.
[0013] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a computer device, the computer device including a processor and a memory coupled to the processor, the memory storing program instructions, and when the program instructions are executed by the processor, causing the processor to perform the steps of the type matching method described in any of the above claims.
[0014] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a storage medium storing program instructions capable of implementing the above-mentioned type matching method.
[0015] The beneficial effects of this application are as follows: The syndrome matching method of this application maps the symptom sequence to be matched to a pre-constructed syndrome element knowledge base as the first syndrome element to be matched, and then uses the syndrome element mapping model to map to obtain the second syndrome element to be matched. After integrating the first and second syndrome elements to be matched, it performs similarity matching with the known syndrome element list of known syndrome elements, thereby selecting the target known syndrome with the highest matching degree. On the one hand, it combines the syndrome element differentiation theory of traditional Chinese medicine and uses the summarized experience knowledge of traditional Chinese medicine for mapping, thereby ensuring the interpretability of the syndrome matching results in the field of traditional Chinese medicine. On the other hand, it combines the high-precision matching effect of machine learning models to improve the accuracy of syndrome matching results. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the type matching method according to an embodiment of the present invention;
[0017] Figure 2 yes Figure 1 A detailed flowchart of step S102;
[0018] Figure 3 yes Figure 1 Another detailed flowchart of step S102;
[0019] Figure 4 yes Figure 1 A detailed flowchart of step S103;
[0020] Figure 5 This is a flowchart illustrating the training process of the attention model according to an embodiment of the present invention;
[0021] Figure 6 yes Figure 1 A detailed flowchart of step S104;
[0022] Figure 7 This is a schematic diagram of the functional modules of the certificate type matching device according to an embodiment of the present invention;
[0023] Figure 8 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0024] Figure 9 This is a schematic diagram of the structure of the storage medium according to an embodiment of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0026] The terms "first," "second," and "third" in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationships and movements between components in a specific orientation (as shown in the figures). If the specific orientation changes, the directional indications also change accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0028] Figure 1 This is a flowchart illustrating the type matching method of the first embodiment of the present invention. It should be noted that if substantially the same result is obtained, the method of the present invention does not necessarily replace the original method. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, the method includes the following steps:
[0029] Step S101: Obtain the symptom sequence to be matched, and use the syndrome element knowledge base to map the symptom sequence to be matched to obtain the corresponding first list of syndrome elements to be matched. The syndrome element knowledge base is constructed based on the pre-obtained syndrome element experience knowledge.
[0030] It should be noted that syndrome type is a unique term in Traditional Chinese Medicine (TCM). It refers to a standardized expression of "syndrome elements" that reflect the spatiotemporal changes of a disease and contain pathological content such as the location of the disease, its cause, its nature, and its progression. For example, taking the common cold as an example, there are three syndrome types: wind-cold, wind-heat, and summer-dampness. For the wind-cold syndrome type, the corresponding syndrome elements include severe aversion to cold and mild fever.
[0031] Specifically, in this embodiment, the symptom sequence to be matched refers to the symptoms of the patient's current condition as determined by the doctor based on the patient's description and through diagnostic methods such as observation, auscultation, inquiry, and palpation. These symptoms include, for example, "chills," "cough," and "slight fever." The syndrome element knowledge base includes various syndrome elements derived from syndrome element differentiation and doctor's case summaries, as well as the mapping relationships between syndrome elements and various symptoms. For example, syndrome elements include "qi deficiency," "yin deficiency," and "yang deficiency." After obtaining the symptom sequence to be matched, mapping is performed based on the syndrome element knowledge base to obtain a first list of syndrome elements to be matched corresponding to the symptom sequence. This first list of syndrome elements to be matched includes multiple syndrome elements.
[0032] Step S102: Input the symptom sequence to be matched into the pre-trained syndrome element mapping model to obtain the second syndrome element list to be matched. The syndrome element mapping model is trained based on the pre-acquired sample syndrome type and the corresponding syndrome element list.
[0033] Specifically, the syndrome element mapping model is a pre-built and trained machine learning model, which is trained using sample syndrome types and corresponding syndrome element lists. The sample syndrome types include standard symptom sequence descriptions of the sample syndrome types. By inputting the state sequence to be matched into the syndrome element mapping model, the model generates a second list of syndrome elements to be matched for the symptom sequence.
[0034] Furthermore, in some embodiments, this evidence mapping model is implemented based on the LDA topic model; therefore, please refer to [link / reference needed]. Figure 2 Step S102 specifically includes:
[0035] Step S201: Input the symptom sequence to be matched into the pre-trained LDA topic model. Use the LDA topic model to extract vector representations of multiple topics from the symptom sequence to be matched. The LDA topic model is trained based on the pre-acquired sample syndrome type and corresponding syndrome element list and syndrome element knowledge base.
[0036] Step S202: Convert the vector representation of multiple topics into a second list of elements to be matched.
[0037] It's important to note that LDA (Latent Dirichlet Allocation) is a document generation model. It assumes that a text has multiple topics, each corresponding to different words. The process of constructing a text involves first selecting a topic with a certain probability, and then selecting a word within that topic with a certain probability. This generates the first word of the text. This process is repeated to generate the entire text, assuming that the words are not in any particular order. The LDA topic model works by reversing this process. It identifies the topics of a text and the corresponding words for those topics.
[0038] In this embodiment, each syndrome element in the syndrome element knowledge base is used as a topic. The LDA topic model is trained by combining pre-prepared sample syndrome types and corresponding syndrome element lists. This enables the LDA topic model to extract multiple topics from the patient's symptom sequence to be matched, and the list of multiple topics is used as the second list of syndrome elements to be matched for the symptom sequence.
[0039] Furthermore, in some embodiments, this evidence mapping model is implemented based on a variational autoencoder; therefore, please refer to [link to relevant documentation]. Figure 3 Step S102 specifically includes:
[0040] Step S301: Input the symptom sequence to be matched into the pre-trained variational autoencoder to obtain the intermediate layer representation of the variational autoencoder; the variational autoencoder is trained based on the pre-prepared sample syndrome. During the training of the variational autoencoder, the input data of the input layer is the sample syndrome, the output data of the output layer is the autoencoder output syndrome, and the output of the intermediate layer is the syndrome element representation of the sample syndrome. The parameters of the intermediate layer are trained by fitting the sample syndrome and the autoencoder output syndrome.
[0041] Step S302: Represent the intermediate layer as the second list of elements to be matched.
[0042] It's important to note that a variational autoencoder (VAE) consists of an encoder and a decoder. The encoder compresses data into a latent space to obtain latent vectors, and the decoder uses these latent vectors to reconstruct the data. Both the encoder and decoder are neural networks. During training, noisy sample data is input into the encoder for encoding, resulting in latent vectors. These latent vectors are the outputs of the intermediate layers of the VAE. The decoder then decodes and reconstructs the output data based on these latent vectors. By comparing the differences between the output data and the sample data, a loss function is used to train the entire VAE, gradually reducing the difference between the sample data and the output data until a preset requirement is met.
[0043] In this embodiment, when constructing the variational autoencoder, the output of the intermediate layer of the variational autoencoder is used as the vector representation of the symptom list. By inputting the sample symptom pattern with added noise into the variational autoencoder, the autoencoder output symptom pattern is output. The sample symptom pattern and the autoencoder output symptom pattern are then fitted to train the variational autoencoder. After the variational autoencoder is trained, the symptom sequence to be matched is input into the variational autoencoder, and the intermediate layer representation of the variational autoencoder is output, which is used as the second symptom list to be matched.
[0044] Step S103: Integrate the first list of elements to be matched and the second list of elements to be matched to obtain the final list of elements to be matched.
[0045] In step S103, after obtaining the first list of syndrome elements to be matched and the second list of syndrome elements to be matched, the first list of syndrome elements to be matched and the second list of syndrome elements to be matched are integrated to obtain the final list of syndrome elements to be matched. This final list of syndrome elements to be matched includes both syndrome element knowledge summarized from traditional Chinese medicine experience and syndrome element knowledge obtained based on machine learning methods. Specifically, the methods for integrating the first list of syndrome elements to be matched and the second list of syndrome elements to be matched include, but are not limited to, splicing and fusing the first list of syndrome elements to be matched and the second list of syndrome elements to be matched, or performing a weighted fusion of the first list of syndrome elements to be matched and the second list of syndrome elements to be matched.
[0046] Furthermore, in some embodiments, step S103 specifically includes: merging the first list of elements to be matched and the second list of elements to be matched to obtain the final list of elements to be matched.
[0047] For example, if the first list of elements to be matched includes [A, B, C, D], and the second list of elements to be matched includes [E, F, G, H], merging the first and second lists of elements to be matched yields [A, B, C, D, E, F, G, H]. It should be noted that when the first and second lists of elements to be matched contain the same element, only one element needs to be retained.
[0048] Furthermore, in other embodiments, such as Figure 4 As shown, step S103 specifically includes:
[0049] Step S401: Convert the first list of evidence elements to be matched into a first vectorized representation, and convert the second list of evidence elements to be matched into a second vectorized representation.
[0050] Specifically, both the first and second lists of evidence elements to be matched are non-numerical representations, while the subsequent evidence element integration process requires numerical data. Therefore, both the first and second lists of evidence elements to be matched are converted into vector representations, thereby obtaining the first vectorized representation of the first list of evidence elements to be matched and the second vectorized representation of the second list of evidence elements to be matched.
[0051] The process of converting the list of evidence elements to be matched into a vector representation may include: performing feature encoding on each evidence element in the list of evidence elements to be matched, specifically, representing the evidence element list with a vector of fixed length, the fixed length being the number of all evidence elements in the evidence element knowledge base, and corresponding each evidence element with each component of the vector one-to-one according to the arrangement order of evidence elements in the evidence element knowledge base, and then using one-hot encoding, marking the component corresponding to the evidence element appearing in the list of evidence elements to be matched with 1, and marking the component corresponding to the other evidence element that does not appear with 0, thereby obtaining the vectorized representation of the list of evidence elements to be matched. For example, assuming the fixed length of the vector is 10, and the evidence elements in the evidence element knowledge base are [A, B, C, D, E, F, G, H, I, J], if the first evidence element list to be matched is [A, B, C, D], then its vectorized representation is [1,1,1,1,0,0,0,0,0,0]. If the second evidence element list to be matched is [A, B, C, I, J], then its vectorized representation is [1,1,1,0,0,0,0,0,1,1].
[0052] Step S402: Input the first vectorized representation and the second vectorized representation into the pre-trained attention model for weighted calculation to obtain the final vectorized representation. The attention model is trained based on the sample certificate type and the target certificate element list of the corresponding target certificate type.
[0053] Specifically, after obtaining the first and second vectorized representations, a query vector is constructed using these representations. This query vector is then input into a pre-trained attention model to perform a weighted calculation on the first and second vectorized representations, resulting in the final vectorized representation. This attention model is trained based on a list of target elements for the sample certificate type and the corresponding target certificate type.
[0054] For further details, please refer to Figure 5 The training method for attention models includes the following steps:
[0055] Step S501: Obtain the sample certificate type and the corresponding target certificate type, as well as the target certificate element list of the target certificate type.
[0056] Step S502: Use the evidence element knowledge base to map the sample evidence type to obtain the first sample evidence element list, and use the evidence element mapping model to analyze the sample evidence type to obtain the second sample evidence element list.
[0057] Step S503: Convert the first sample evidence list and the second sample evidence list into vectorized representations, and construct a query vector using the vectorized representations of the first sample evidence list and the second sample evidence list as elements.
[0058] Step S504: Input the query vector into the attention model to be trained, output the final sample vectorized representation, and convert the final sample vectorized representation into the final sample element list.
[0059] Step S505: Update the weight parameters of the attention model in reverse based on the final sample evidence list, the target evidence list, and the pre-built scoring function.
[0060] It should be noted that the attention model is a neural network model built on the attention mechanism. Its calculation process can be divided into two parts: 1. Calculate the attention distribution using the query vector on all input information; 2. Calculate the weighted average of each input information based on the attention distribution.
[0061] Step S203: Convert the final vectorized representation into the final list of elements to be matched.
[0062] Specifically, after obtaining the final vectorized representation, the final vectorized representation is converted into the final list of elements to be matched by combining the evidence element knowledge base. Taking the above example as an example, if the calculated final vectorized representation is [1,1,1,0,0,0,0,0,1,0], then it is converted into the final list of elements to be matched as [A, B, C, I].
[0063] In this embodiment, when integrating the first list of testimonial elements to be matched and the second list of testimonial elements to be matched, an attention mechanism is introduced to perform weighted calculations on the first list of testimonial elements to be matched and the second list of testimonial elements to be matched. This makes the integration process of the first list of testimonial elements to be matched and the second list of testimonial elements to be matched more scientific and accurate. It takes into account the possibility that the testimonial knowledge base mapping and machine learning model prediction have different degrees of impact on the accuracy of the final result in practical applications. Thus, the weighted average of the first list of testimonial elements to be matched and the second list of testimonial elements to be matched is achieved based on the attention mechanism.
[0064] Step S104: Perform similarity matching between the final list of evidence elements to be matched and the list of known evidence elements corresponding to each known evidence type to obtain the matched target known evidence type and output it.
[0065] It should be noted that the known syndrome types refer to all known syndrome types summarized based on existing TCM experience, and each known syndrome type corresponds to a list of known syndrome elements.
[0066] Specifically, after obtaining the final list of syndrome elements to be matched for the symptom sequence to be matched, the final list of syndrome elements to be matched is used to perform similarity matching with the list of known syndrome elements for each known syndrome type, thereby identifying the list of target known syndrome elements with the highest similarity to the symptom sequence to be matched, and then identifying the target known syndrome type corresponding to the list of target known syndrome elements. The target known syndrome type is then output as the syndrome type corresponding to the symptom sequence to be matched, thereby assisting TCM doctors in quickly identifying the patient's disease condition and making corresponding diagnosis and treatment strategies.
[0067] It should be noted that in this embodiment, the list of known elements corresponding to each known element type is obtained in advance, and the process of obtaining it is the same as the process of obtaining the final list of elements to be matched. That is, it also needs to be generated using the element mapping knowledge base and the element mapping model. The process is the same as the process of generating the final list of elements to be matched, and will not be described again here.
[0068] Furthermore, in some embodiments, such as Figure 6 As shown, step S104 specifically includes:
[0069] Step S601: Calculate the intersection and union of the list of evidence elements to be matched and the list of known evidence elements. Finally, both the list of evidence elements to be matched and the list of known evidence elements include multiple evidence elements.
[0070] Step S602: Calculate the ratio of intersection to union to obtain the similarity value between the list of evidence elements to be matched and the list of known evidence elements.
[0071] Step S603: Select the target known evidence element corresponding to the list of known evidence elements with the highest similarity value and output it.
[0072] Specifically, the formula for calculating the similarity value is as follows:
[0073]
[0074] Among them, similarity n The similarity value is for syndrome. a For the final list of elements to be matched, syndrome b Given a list of known elements of a known syndrome type, max((len(syndrome)) a ),len(syndrome b )) represents the largest set of evidence elements between the final list of evidence elements to be matched and the list of known evidence elements, which is the union of the final list of evidence elements to be matched and the list of known evidence elements.
[0075] The syndrome matching method of this invention maps the symptom sequence to be matched to a pre-constructed syndrome element knowledge base as a first syndrome element to be matched, and then uses a syndrome element mapping model to map to obtain a second syndrome element to be matched. After integrating the first and second syndrome elements to be matched, it performs similarity matching with a list of known syndrome elements of known syndrome elements, thereby selecting the target known syndrome with the highest degree of matching. On the one hand, it combines the syndrome element differentiation theory of traditional Chinese medicine and uses the summarized experience knowledge of traditional Chinese medicine for mapping, thereby ensuring the interpretability of the syndrome matching results in the field of traditional Chinese medicine. On the other hand, it combines the high-precision matching effect of machine learning models to improve the accuracy of syndrome matching results.
[0076] Figure 7 This is a schematic diagram of the functional modules of the type matching device according to an embodiment of the present invention. Figure 7 As shown, the device 70 includes a first mapping module 71, a second mapping module 72, an integration module 73, and a matching module 74.
[0077] The first mapping module 71 is used to obtain the symptom sequence to be matched, and to map the symptom sequence to be matched using the syndrome element knowledge base to obtain the corresponding first list of syndrome elements to be matched. The syndrome element knowledge base is constructed based on the pre-obtained syndrome element experience knowledge.
[0078] The second mapping module 72 is used to input the symptom sequence to be matched into the pre-trained syndrome element mapping model to obtain the second syndrome element list to be matched. The syndrome element mapping model is trained based on the pre-acquired sample syndrome type and the corresponding syndrome element list and syndrome element knowledge base.
[0079] Integration module 73 is used to integrate the first list of elements to be matched and the second list of elements to be matched to obtain the final list of elements to be matched.
[0080] The matching module 74 is used to perform similarity matching between the final list of evidence elements to be matched and the list of known evidence elements corresponding to each known evidence type, so as to obtain the matched target known evidence type and output it.
[0081] Optionally, the matching module 74 performs the operation of integrating the first list of testimonies to be matched and the second list of testimonies to be matched to obtain the final list of testimonies to be matched, including: merging the first list of testimonies to be matched and the second list of testimonies to be matched to obtain the final list of testimonies to be matched.
[0082] Optionally, the matching module 74 performs the operation of integrating the first list of testimonial elements to be matched and the second list of testimonial elements to be matched to obtain the final list of testimonial elements to be matched, including: converting the first list of testimonial elements to be matched into a first vectorized representation and converting the second list of testimonial elements to be matched into a second vectorized representation; inputting the first vectorized representation and the second vectorized representation into a pre-trained attention model for weighted calculation to obtain the final vectorized representation, wherein the attention model is trained based on the sample testimonial type and the target testimonial element list of the corresponding target testimonial type; and converting the final vectorized representation into the final list of testimonial elements to be matched.
[0083] Optionally, it also includes a training model for training the attention model, specifically including: obtaining sample certificate types and corresponding target certificate types, as well as a list of target certificate elements for the target certificate types; mapping the sample certificate types using a certificate element knowledge base to obtain a first list of sample certificate elements, and analyzing the sample certificate types using a certificate element mapping model to obtain a second list of sample certificate elements; converting the first and second list of sample certificate elements into vectorized representations, and constructing a query vector using the vectorized representations of the first and second list of sample certificate elements as elements; inputting the query vector into the attention model to be trained, outputting the final sample vectorized representation, and converting the final sample vectorized representation into the final list of sample certificate elements; and updating the weight parameters of the attention model in reverse based on the final list of sample certificate elements, the target list of certificate elements, and a pre-constructed loss function.
[0084] Optionally, the syndrome element mapping model is implemented based on the LDA topic model; the second mapping module 72 performs the operation of inputting the symptom sequence to be matched into the pre-trained syndrome element mapping model, including: inputting the symptom sequence to be matched into the pre-trained LDA topic model, using the LDA topic model to extract vector representations of multiple topics from the symptom sequence to be matched, the LDA topic model being trained based on pre-acquired sample syndrome types and corresponding syndrome element lists and syndrome element knowledge bases; and converting the vector representations of multiple topics into a second list of syndrome elements to be matched.
[0085] Optionally, the syndrome element mapping model is implemented based on a variational autoencoder; the second mapping module 72 performs the operation of inputting the symptom sequence to be matched into the pre-trained syndrome element mapping model, specifically including: inputting the symptom sequence to be matched into the pre-trained variational autoencoder to obtain the intermediate layer representation of the variational autoencoder, the variational autoencoder is trained based on pre-prepared sample syndrome types, during the training of the variational autoencoder, the input data of the input layer is the sample syndrome type, the output data of the output layer is the autoencoder output syndrome type, and the output of the intermediate layer is the syndrome element representation of the sample syndrome type; the parameters of the intermediate layer are trained by fitting the sample syndrome type and the autoencoder output syndrome type; and the intermediate layer representation is used as the second list of syndrome elements to be matched.
[0086] Optionally, the matching module 74 performs a similarity matching operation between the final list of evidence elements to be matched and the list of known evidence elements corresponding to each known evidence type, to obtain the matched target known evidence type and output it. Specifically, this includes: calculating the intersection and union of the list of evidence elements to be matched and the list of known evidence elements, where both the final list of evidence elements to be matched and the list of known evidence elements include multiple evidence elements; calculating the ratio of the intersection to the union to obtain the similarity value between the list of evidence elements to be matched and the list of known evidence elements; and selecting the target known evidence element corresponding to the list of known evidence elements with the highest similarity value for output.
[0087] For other details regarding the implementation techniques of each module in the above-described type matching device, please refer to the description of the type matching method in the above-described embodiments, which will not be repeated here.
[0088] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0089] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 8 As shown, the computer device 60 includes a processor 61 and a memory 62 coupled to the processor 61. The memory 62 stores program instructions. When the program instructions are executed by the processor 61, the processor 61 performs the steps of the type matching method described in any of the above embodiments.
[0090] The processor 61 can also be referred to as a CPU (Central Processing Unit). The processor 61 may be an integrated circuit chip with signal processing capabilities. The processor 61 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor.
[0091] See Figure 9 , Figure 9This is a schematic diagram of the structure of a storage medium according to an embodiment of the present invention. The storage medium of this embodiment stores program instructions 71 capable of implementing all the above methods. These program instructions 71 can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or computer devices such as computers, servers, mobile phones, and tablets.
[0092] In the several embodiments provided in this application, it should be understood that the disclosed computer devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0093] Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A type matching method, characterized in that, include: Obtain the symptom sequence to be matched, and map the symptom sequence to be matched using the syndrome element knowledge base to obtain the corresponding first list of syndrome elements to be matched. The syndrome element knowledge base is constructed based on the pre-obtained syndrome element experience knowledge. The symptom sequence to be matched is input into a pre-trained syndrome element mapping model to obtain a second list of syndrome elements to be matched. The syndrome element mapping model is trained based on the pre-acquired sample syndrome type, the corresponding syndrome element list, and the syndrome element knowledge base. By integrating the first list of evidence elements to be matched and the second list of evidence elements to be matched, a final list of evidence elements to be matched is obtained. The final list of evidence elements to be matched is matched with the list of known evidence elements corresponding to each known evidence type to obtain the target known evidence type and output it. The integration of the first list of evidence elements to be matched and the second list of evidence elements to be matched yields the final list of evidence elements to be matched, including: The first list of elements to be matched is converted into a first vectorized representation, and the second list of elements to be matched is converted into a second vectorized representation; The first vectorized representation and the second vectorized representation are input into a pre-trained attention model for weighted calculation to obtain the final vectorized representation. The attention model is trained based on the sample certificate type and the target certificate element list of the corresponding target certificate type. The final vectorized representation is converted into the final list of elements to be matched; Attention model training methods include: Obtain the sample certificate type and the corresponding target certificate type, as well as the target certificate element list of the target certificate type; The sample evidence type is mapped using the evidence element knowledge base to obtain a first sample evidence element list, and the sample evidence type is analyzed using the evidence element mapping model to obtain a second sample evidence element list. The first sample evidence list and the second sample evidence list are converted into vectorized representations, and the vectorized representations of the first sample evidence list and the second sample evidence list are used as elements to construct a query vector. The query vector is input into the attention model to be trained, the final sample vectorized representation is output, and the final sample vectorized representation is converted into a final sample element list; The weight parameters of the attention model are updated in reverse based on the final sample evidence list, the target evidence list, and the pre-built scoring function.
2. The certificate type matching method according to claim 1, characterized in that, The integration of the first list of evidence elements to be matched and the second list of evidence elements to be matched yields the final list of evidence elements to be matched, including: The first list of evidence elements to be matched and the second list of evidence elements to be matched are merged to obtain the final list of evidence elements to be matched.
3. The certificate type matching method according to claim 1, characterized in that, The evidence element mapping model is implemented based on the LDA topic model; The step of inputting the symptom sequence to be matched into a pre-trained symptom mapping model includes: The symptom sequence to be matched is input into a pre-trained LDA topic model. The LDA topic model is used to extract vector representations of multiple topics from the symptom sequence to be matched. The LDA topic model is trained based on pre-acquired sample syndrome types and corresponding syndrome element lists and syndrome element knowledge base. The vector representations of the multiple themes are converted into the second list of elements to be matched.
4. The certificate type matching method according to claim 1, characterized in that, The evidence element mapping model is implemented based on a variational autoencoder. The step of inputting the symptom sequence to be matched into a pre-trained symptom mapping model includes: The symptom sequence to be matched is input into a pre-trained variational autoencoder to obtain the intermediate layer representation of the variational autoencoder. The variational autoencoder is trained based on a pre-prepared sample syndrome. During the training of the variational autoencoder, the input data of the input layer is the sample syndrome, the output data of the output layer is the autoencoded output syndrome, and the output of the intermediate layer is the syndrome element representation of the sample syndrome. The parameters of the intermediate layer are trained by fitting the sample syndrome and the autoencoded output syndrome. The intermediate layer is represented as the second list of elements to be matched.
5. The certificate type matching method according to claim 1, characterized in that, The step of performing similarity matching between the final list of evidence elements to be matched and the list of known evidence elements corresponding to each known evidence type to obtain and output the matched target known evidence type includes: Calculate the intersection and union of the list of evidence elements to be matched and the list of known evidence elements, wherein both the final list of evidence elements to be matched and the list of known evidence elements include multiple evidence elements; Calculate the ratio of the intersection to the union to obtain the similarity value between the list of evidence elements to be matched and the list of known evidence elements; Output the target known evidence element corresponding to the list of known evidence elements with the highest similarity value.
6. A certificate type matching device, used to implement the certificate type matching method as described in any one of claims 1-5, characterized in that, include: The first mapping module is used to obtain the symptom sequence to be matched, and to map the symptom sequence to be matched using the syndrome element knowledge base to obtain the corresponding first list of syndrome elements to be matched. The syndrome element knowledge base is constructed based on the pre-obtained syndrome element experience knowledge. The second mapping module is used to input the symptom sequence to be matched into a pre-trained syndrome element mapping model to obtain a second list of syndrome elements to be matched. The syndrome element mapping model is trained based on the pre-acquired sample syndrome type, the corresponding syndrome element list, and the syndrome element knowledge base. An integration module is used to integrate the first list of elements to be matched and the second list of elements to be matched to obtain a final list of elements to be matched. The matching module is used to perform similarity matching between the final list of evidence elements to be matched and the list of known evidence elements corresponding to each known evidence type, so as to obtain the matched target known evidence type and output it.
7. A computer device, characterized in that, The computer device includes a processor and a memory coupled to the processor, the memory storing program instructions that, when executed by the processor, cause the processor to perform the steps of the type matching method as described in any one of claims 1-5.
8. A storage medium, characterized in that, The system stores program instructions capable of implementing the type matching method as described in any one of claims 1-5.
Citation Information
Patent Citations
Traditional Chinese medicine assistant diagnosis system based on syndrome element
CN110459321A
Typhoid fever syndrome differentiation reasoning system based on knowledge graph
CN110827990A