Method for identifying roles in acupoint combination based on named entities

Through the identification method based on named entities and the BiLSTM+CRF model, the acupuncture point compatibility relationship in medical cases is analyzed, and the problem that the existing technology cannot effectively identify acupuncture point compatibility is solved, and the intelligent and precise development of acupuncture treatment is achieved.

CN119783676BActive Publication Date: 2025-05-13CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510279466.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-13
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing technology cannot effectively use deep learning algorithms to identify acupuncture points and their compatibility relationships in medical cases, which limits the intelligent and precise development of acupuncture treatment.

Method used

The identification method based on named entities is adopted, and the literature database for acupuncture treatment is built by collecting ancient medical cases and clinical research literature, and the acupuncture treatment is identified using the BiLSTM+CRF model to analyze the dependence of the main and auxiliary roles in the acupuncture compatibility.

Benefits of technology

The role recognition of acupuncture points in the clinical effect of acupuncture has been achieved, the acupuncture points are optimized, the efficacy and reliability of acupuncture have been improved, and the intelligent and precise development of acupuncture treatment has been promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783676B_ABST
    Figure CN119783676B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying various roles in acupoint compatibility based on named entities, and relates to a method for identifying acupoint compatibility in medical records, including: S1, building a literature database D for acupuncture treatment; S2, analyzing and processing the literature database D, and characterizing each acupoint name through a self-built acupuncture acupoint compatibility data set; S3, performing hyperparameter optimization on a subset of the acupuncture acupoint compatibility dictionary based on PSO to obtain the hyperparameter configuration in the BiLSTM+CRF model, training and testing the entire literature database D, and obtaining an acupoint compatibility recognition model; S4, for any acupoint sequence and the corresponding label sequence, identifying the role of each acupoint in treating the corresponding disease through the acupoint compatibility recognition model. The present invention provides a method for identifying various roles in acupoint compatibility based on named entities, which has high accuracy, strong generalization ability, high intelligence level, and the effect of promoting the modernization of traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for identifying acupoint combinations in medical records, and more specifically, to a method for identifying roles in acupoint combinations based on named entities. Background Art

[0002] Named Entity Recognition (NER) is a task that aims to determine the boundaries of entities in text and accurately classify them. The named entity recognition task is the basis of many natural language processing (NLP) tasks, such as information extraction, question answering, information retrieval, knowledge graphs, etc., and has attracted much attention from researchers.

[0003] With the integration of medicine and information technology, the importance of named entity recognition in the field of medical research has become increasingly prominent. Its main research methods include rule-based and dictionary-based methods, statistical machine learning methods, and deep learning-based methods. The earliest named entity recognition was mainly based on rules and dictionaries. This method relies on pre-built rules and dictionaries, cannot effectively recognize named entities outside the dictionary and rules, and is difficult to adapt to different fields and languages. To solve the above problems, machine learning models have gradually replaced rule-based and dictionary-based methods. Statistical machine learning methods mainly use Markov models and conditional random field (CRF) models. Statistical machine learning methods can recognize named entities outside dictionaries and rules, but this method relies on a lot of feature engineering and requires professional knowledge.

[0004] Acupoint compatibility refers to combining different acupoints with similar effects to exert the synergistic effect between acupoints to achieve improved efficacy, which is the basis of acupuncture prescriptions. Appropriate acupoint compatibility is the key to achieving and improving the efficacy of acupuncture. Named entity recognition of acupoints is a key area of ​​information extraction in biochemical research. NER provides support for text mining of drugs, including entity relationship extraction, attribute extraction, etc. However, the existence of complex naming features in the biomedical field, such as polysemy and special characters, makes the NER task extremely challenging. In the clinical field, it is necessary to take the acupoint compatibility relationship (the acupoint compatibility relationship refers to the main and auxiliary relationship of each acupoint in a medical case) as the starting point to parse the main and auxiliary role dependency relationship in acupoint compatibility. The existing technology mainly recognizes the role status of acupoints from a large amount of clinical experience, and does not involve the use of deep learning algorithms to identify the relationship between acupoints and compatibility in medical cases, which cannot be applied to the intelligent and precise development of acupuncture treatment. Summary of the invention

[0005] An object of the present invention is to solve at least the above problems and / or disadvantages and to provide at least the advantages which will be described hereinafter.

[0006] In order to achieve these purposes and other advantages of the present invention, a method for identifying various roles in acupoint combination based on named entities is provided, comprising:

[0007] S1. Collect ancient medical records of related diseases and clinical research literature on acupuncture treatment, and build a literature database D on acupuncture treatment;

[0008] S2, based on the analysis and processing of the literature database D, each acupoint name is characterized by a self-built acupuncture point compatibility data set;

[0009] S3, based on the segmented PSO model structure optimization, a subset of the acupuncture point compatibility dictionary is optimized for hyperparameters to obtain the hyperparameter configuration in the BiLSTM+CRF model for named entity recognition, and the entire literature database D is trained and tested to obtain the acupuncture point compatibility recognition model;

[0010] S4. For any acupoint sequence And the corresponding label sequence The acupoint compatibility recognition model identifies the role of each acupoint in treating the corresponding disease through the following formula:

[0011]

[0012] In the above formula, P is the calculation score matrix of the two-layer LSTM neural network, Indicates that the i-th acupoint in the acupuncture prescription is named y i The score of the role label, σ represents the sigmoid function in logistic regression, A represents the transition probability matrix between labels, The first label of the sequence is y i , Ɛ1 and Ɛ2 represent the weights of the label probability matrix and label score respectively, and Ɛ1+Ɛ2=1.

[0013] Preferably, in S2, the analysis process includes:

[0014] S20, preprocessing the text in the literature database D;

[0015] S21, performing dependency grammar analysis on the preprocessed text through the open source toolkit LTP to extract the dependency relationship between the main and auxiliary roles of each acupoint in the acupoint combination;

[0016] S22. Analyze the preprocessed text using the open source syntactic analyzer BerkeleyParser to identify the components and phrase structures in the sentence, and segment the symptom sentence according to the analysis results to complete the phrase structure syntactic analysis;

[0017] S23. Based on the dependency relationship between the main and auxiliary roles and the results of the syntactic analysis of the phrase structure, the effect characteristics of each acupoint in the treatment of the corresponding disease are obtained to establish an acupuncture point compatibility dataset.

[0018] Preferably, in S3, the process of obtaining the hyperparameter configuration includes:

[0019] S30, in the initialization of PSO, the dimension of the particle is used to represent the parameters of each acupoint role labeling;

[0020] S31. In the first stage, the particle parameters were substituted into the BiLSTM+CRF model of acupoint role labeling for double cross validation to evaluate the fitness value of the function, update the speed and position of each particle, and complete 50 iterations;

[0021] S32. Based on the convergence analysis of each dimension of particles in the first stage, the initial range of each dimension of PSO particles in the second stage is determined, and the BiLSTM+CRF model is cross-validated 10 times and iterated ten times to select the 6 best particles and obtain the corresponding hyperparameter configuration.

[0022] Preferably, the BiLSTM+CRF model includes an LSTM unit and a Bi-LSTM unit;

[0023] The LSTM unit takes the acupuncture point compatibility dataset z as input, and constructs an enhanced dynamic neural memory unit by introducing a dynamic gating mechanism, a multi-head memory coupling mechanism, a hierarchical feature fusion strategy and acupoint functional characteristics. The update and output of the LSTM unit are characterized by the following formula:

[0024]

[0025]

[0026] In the above formula, h t is the output of the LSTM unit at time t, represents the encoded acupoint function features dynamically obtained from the self-built dictionary, C is the value of the LSTM memory unit, i t , f t , O t , , C t They represent the input gate, forget gate, output gate, candidate value of the memory unit state at the current moment, and state value respectively. W c , W i , W f , W o Respectively represent the weight matrices input to the memory unit state candidate value, input gate, forget gate, and output gate, represents K independent memory head parameter groups, represents M independent input gate parameter groups, , , , represents the learnable gating parameter vector, Z t express t Input the acupoint characteristics of the acupoint dictionary at any time. h t-1 Represents the output of the LSTM unit at the previous moment, b c , b i , b f , b o They represent the bias items input to the candidate value of the memory unit state, the input gate, the forget gate, and the output gate, respectively. U c , U i , U f , U o They represent the weight matrices from the previous hidden state to each gate, ATT(·) represents the adaptive attention mechanism based on the self-built acupoint role matrix, represents the pth memory transformation layer with differentiable wavelet transform kernel function, MLP(·) represents the multi-layer perceptron based on the self-built acupoint role matrix, Feature extraction network representing the functional characteristics of acupoints, represents the weight distribution coefficient of the kth memory head when the memory unit state is updated, represents the weight distribution coefficient of the mth input gate parameter group when the memory unit state is updated, represents the learnable parameters, represents element-wise multiplication, ⊕ represents tensor concatenation and projection operations, ⊗ represents Kronecker product expansion operations, C t-1represents the state value of the memory unit at the previous moment, σ represents the sigmoid function in logistic regression, K represents the number of independent memory heads, M represents the number of independent input gates, and P represents the number of memory transformation layers.

[0027] Preferably, the output of the Bi-LSTM unit P Obtained by the following formula:

[0028] In the above formula, , They are based on the output of LSTM units respectively h t The calculated forward hidden layer sequence and reverse layer sequence, g is the adaptive gating weight, and .

[0029] Preferably, the BiLSTM+CRF model is strengthened by transferring constraints through a CRF layer;

[0030] Among them, in the label transfer probability matrix of the CRF layer, a dynamic penalty term is set for illegal transfers by the following formula: :

[0031] .

[0032] The present invention includes at least the following beneficial effects: the present invention takes the acupoint effect as the starting point, determines the role of each acupoint in the clinical effect of acupuncture through named entity recognition, and finally analyzes the main and auxiliary role dependencies in the acupoint combination by designing an acupoint combination recognition model, so as to optimize the acupoint combination, serve the clinic, promote the development of more precise and intelligent acupuncture treatment, and improve the efficacy and reliability of acupuncture.

[0033] Other advantages, objectives and features of the present invention will be embodied in part through the following description, and in part will be understood by those skilled in the art through study and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 The organizational structure diagram of the acupoint main and auxiliary role entity naming and recognition method of the present invention;

[0035] Figure 2 It is a schematic diagram of the process of optimizing the PSO hyperparameters of the present invention. DETAILED DESCRIPTION

[0036] The present invention is further described in detail below in conjunction with the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.

[0037] 1. Literature collection and database construction

[0038] 1. Collection of ancient medical records: Through libraries, ancient book databases and other channels, collect relevant ancient medical records on the treatment of diseases such as epigastric pain, loss of appetite, fullness, vomiting, abdominal distension, abdominal pain, etc.

[0039] 2. Collection of clinical research literature: Using databases such as PubMed, CNKI, and Wanfang, we searched with keywords such as “acupuncture,” “painful diseases,” and “functional dyspepsia” to collect clinical research literature on acupuncture treatment of painful diseases, functional dyspepsia, and other diseases in the past thirty years.

[0040] 3. Database construction: Organize the collected ancient medical records and clinical research documents to establish a literature database D for acupuncture treatment. The database should include basic information of the literature (such as author, year, literature type, etc.), symptoms, etiology and pathogenesis, treatment principles, acupoints, treatment effects, etc.

[0041] 2. Acupoint Effect Feature Extraction

[0042] The feature extraction of acupoint effects starts from the dependency relationship and phrase structure, and then a self-built dictionary of acupoint compatibility is obtained, including:

[0043] 1. Dependency feature extraction: preprocess the text in the literature database D, including removing stop words, word segmentation, etc. Use the open source toolkit LTP (Language Technology Platform) to perform dependency grammar analysis on the preprocessed text to identify the main and auxiliary role dependencies in acupoint compatibility. Extract and analyze these dependencies to understand the role and relationship of different acupoints in the compatibility.

[0044] 2. Phrase structure syntactic analysis: preprocess the text in the literature database D, including removing stop words, word segmentation, etc. Use Berkeley Parser to perform phrase structure syntactic analysis on the preprocessed text to identify the components and phrase structures in the sentence. According to the results of the phrase structure syntactic analysis, segment the symptom sentence for subsequent analysis and processing.

[0045] The acupuncture point compatibility dictionary refers to extracting all acupuncture point names and compatible acupuncture points from the preprocessed text in the literature database D through text cleaning, word segmentation, and stop word removal to form an initial vocabulary, assigning a unique index to each acupuncture point name, and constructing a dictionary of compatibility relationships between acupuncture points based on the compatibility information in the text. The acupuncture point compatibility dictionary is used to identify symptoms in text sentences to improve the accuracy of symptom recognition.

[0046] 3. Analysis after feature extraction: ancient medical records and clinical research literature are integrated into a unified literature database D for comprehensive analysis and comparative research. That is, using the results of dependency and phrase structure syntactic analysis, we can deeply explore the effect characteristics of acupoints in treating specific diseases. Combining clinical practice and literature data, we can further verify and optimize the acupoint combination scheme.

[0047] 3. Acupoint role labeling

[0048] In actual applications, the acupoint role labeling process is as follows Figure 1 As shown, it mainly includes:

[0049] 1. Select the corpus for training word vectors

[0050] In the literature database D, first, the corpus for training word vectors is selected, and the training corpus comes from the literature database D;

[0051] Secondly, according to the characteristics of clinical data corpus (small amount of data), we chose to use a self-built dictionary of acupuncture point combinations to generate word vectors with a 5-dimensional dimension.

[0052] Again, the words in the sentence are concatenated with word vectors converted from additional linguistic features and sent to the next layer.

[0053] In this solution, the acupoint roles in the acupoint text are identified and segmented, and then concatenated with the word vectors converted by word embedding (a technology that converts words into low-dimensional vector representation). Then, a self-built acupoint dictionary is used to make each acupoint name have its own unique expression, which is used as the input of LSTM.

[0054] 2. Particle Swarm Optimization (PSO) based on segmentation model structure optimization: In the process of particle swarm optimization (PSO), PSO is used to optimize the hyperparameters of the BiLSTM model used for acupoint role labeling. The hyperparameter configuration obtained after selecting the best particles through PSO is then used to train and test the entire dataset using the BiLSTM architecture. In practical applications, such as Figure 2 As shown in Figure 2, the segmentation-based PSO model structure optimization mainly includes the following two stages:

[0055] Phase 1:

[0056] The first step is to initialize PSO, where each dimension of the particle represents the parameter of each acupoint role labeling that needs to be optimized.

[0057] The second step is to calculate the fitness function value of each particle. The calculation method is to bring the particle parameters into the model construction process of acupoint role annotation. The data set uses 50% of the data and adopts 2-fold cross validation to obtain the mean of the accuracy results. Sort the fitness values ​​of all particles and save the position of the best particle as the current global optimal position. When running to this step again, compare the value of the best particle in the current cycle with the historical global optimal. If the current one is better, replace it.

[0058] In the third step, the vector of the particle is calculated based on the motion vector of the particle itself, the optimal position of all particles obtained in the second step, and the optimal position in the history of the current particle.

[0059] The fourth step is to update the particle position information.

[0060] The fifth step is to determine whether the iteration stop requirement has been met. When the iteration is set to 50 times, the process stops if it is met, and goes to the second step of the first stage if it is not met.

[0061] The sixth step requires collecting and analyzing the convergence status of each dimension of the PSO particles from the previous stage to determine the initial range of each dimension of the PSO particles in the next stage.

[0062] Phase 2:

[0063] Step 7: Initialize the second-stage PSO particle swarm according to the range determined in step 6.

[0064] In the eighth step, the average accuracy calculated by 10-fold cross validation is used as the fitness value of function. Since the acupuncture data set obtained is a small sample scenario, there is too little training data and the fitness value deviation may be large. In order to quickly clarify the value range of the optimal result. The overall calculation time is controlled by adjusting the PSO particle swarm population size and the number of iterations. The 10-fold method is used to calculate the fitness value, sort the calculated fitness values, and update the global optimal position and the historical optimal position of each particle.

[0065] Step 9~Step 10: Update the particle vector and position as in steps 3 and 4.

[0066] The eleventh step is to determine whether the iteration stop condition has been reached (the stop condition is: the second stage stops after 10 iterations).

[0067] In this scheme, hyperparameter optimization is performed by using a subset of the self-built acupoint dictionary in the particle swarm optimization (PSO) process. The hyperparameter configuration obtained after selecting the 6 best particles through PSO is then used to train and test the entire self-built acupoint dictionary using the BiLSTM (Bidirectional Long Short-Term Memory, BiLSTM) architecture.

[0068] 3. Use the PSO optimized hyperparameter configuration in the Bilstm-CRF (Conditional Random Field, CRF) model for named entity recognition (NER)

[0069] In actual applications, the word at the beginning of a sentence cannot obtain information about the subsequent word data, which may cause the previous word of the text to lack necessary information when predicting labels. Bi-LSTM (BidirectionalLSTM) further combines historical information with future information, merging features from the left and right directions to improve the ability to capture the global semantics of the text, so that the results of the current word after LSTM prediction are passed forward and backward to the adjacent words respectively, and the prediction result of the current word contains the processed information of the adjacent words in the front and back directions.

[0070] The LSTM unit takes the acupuncture point compatibility dataset z as input, and constructs an enhanced dynamic neural memory unit by introducing a dynamic gating mechanism, a multi-head memory coupling mechanism, a hierarchical feature fusion strategy and acupoint functional characteristics. h is the output of the LSTM unit, C is the value of the LSTM memory unit, and z is the input data. The update and output of the relevant state in the LSTM unit are as follows:

[0071]

[0072] In the above formula, h t is the output of the LSTM unit at time t, represents the encoded acupoint function features dynamically obtained from the self-built dictionary, C is the value of the LSTM memory unit, i t , f t , O t , , C tThey represent the input gate, forget gate, output gate, candidate value of the memory unit state at the current moment, and state value respectively. W c , W i , W f , W o Respectively represent the weight matrices input to the memory unit state candidate value, input gate, forget gate, and output gate, represents K independent memory head parameter groups, represents M independent input gate parameter groups, , , , represents the learnable gating parameter vector, Z t express t Input the acupoint characteristics of the acupoint dictionary at any time. h t-1 Represents the output of the LSTM unit at the previous moment, b c , b i , b f , b o They represent the bias items input to the candidate value of the memory unit state, the input gate, the forget gate, and the output gate, respectively. U c , U i , U f , U o They represent the weight matrices from the previous hidden state to each gate, ATT(·) represents the adaptive attention mechanism based on the self-built acupoint role matrix, represents the pth memory transformation layer with differentiable wavelet transform kernel function, MLP(·) represents the multi-layer perceptron based on the self-built acupoint role matrix, Feature extraction network representing the functional characteristics of acupoints, represents the weight distribution coefficient of the kth memory head when the memory unit state is updated, represents the weight distribution coefficient of the mth input gate parameter group when the memory unit state is updated, Represents a learnable parameter, which is used to input the gate of historical information and is part of the dynamic gating mechanism. represents element-wise multiplication, ⊕ represents tensor concatenation and projection operations, ⊗ represents Kronecker product expansion operations, C t-1represents the state value of the memory unit at the previous moment, σ represents the sigmoid function in logistic regression, K represents the number of independent memory heads, M represents the number of independent input gates, and P represents the number of memory transformation layers.

[0073] In the problem of sequence labeling, future information and historical information are equally helpful for sequence prediction. Therefore, Bi-LSTM is introduced. Its basic idea is to capture the role labeling information of acupoint roles from the positive and negative directions by constructing two hidden layers. That is, Bi-LSTM first calculates the forward hidden layer sequence , then calculate the reverse layer sequence , and finally the forward sequence With reverse sequence Splice to get the calculated score matrix output P :

[0074]

[0075] In the above formula, Θ is the element-wise multiplication, which is used to dynamically adjust the forward and backward information contribution to improve the ability to capture complex compatibility relationships, σ represents the sigmoid function in logistic regression, g is the adaptive gating weight, and .

[0076] In order to obtain deeper semantic information and consider the balance between training time and annotation effect, a two-layer LSTM framework and a one-layer Bi-LSTM framework are used, and the output of the second LSTM layer is used as the input of the corresponding node of the third Bi-LSTM neural network layer.

[0077] 4. For an input acupoint role sequence ,in, Indicates the corresponding acupuncture point name, and the corresponding label sequence ,in, They respectively represent the main (main acupoint) or auxiliary (accessory acupoint) role of the acupoint in treating the corresponding disease.

[0078] P is the calculation score matrix of the two-layer LSTM neural network, P ij represents the score of the jth role of the ith acupoint in the acupuncture prescription, and σ represents the sigmoid function in logistic regression. Let A represent the transition probability matrix between labels, then the element A mm It represents the probability that label m will be transferred to label n at the next moment. The elements that are unlikely to be transferred are assigned a value of -3000000, and the remaining transfers are obtained during model training.

[0079] In addition, The first label of the sequence is y iThe probability of , Ɛ1, Ɛ2 represent the weights of the label probability matrix and the label score respectively, let Ɛ1=Ɛ2=0.5, so the score of the label sequence is defined as:

[0080]

[0081] In the above formula, P is the calculation score matrix of the two-layer LSTM neural network, Indicates that the i-th acupoint in the acupuncture prescription is named the y-th acupoint i The score of the role label, A represents the transition probability matrix between labels, Indicates that the first label of the sequence is y i , and Ɛ1+Ɛ2=1.

[0082] According to the probability results corresponding to the label probability transfer matrix, the role of the acupoint in treating a certain disease can be identified.

[0083] Example: Identification of the role of acupoints in the treatment of epigastric pain in ancient medical records

[0084] The original text of the medical case is: "The patient suffered from recurrent epigastric pain, accompanied by fullness and vomiting. Zusanli, Zhongwan, Neiguan and Gongsun were used as the main points, supplemented by Liangmen and Taichong."

[0085] 1. Literature collection and database construction

[0086] 1. Medical record entry The above medical records are entered into the literature database D and stored in a structured manner as follows:

[0087] Symptoms: epigastric pain, fullness, vomiting

[0088] Acupoints: Zusanli, Zhongwan, Neiguan, Gongsun, Liangmen, Taichong

[0089] Tags: Main acupoints (Zusanli, Zhongwan, Neiguan, Gongsun), auxiliary acupoints (Liangmen, Taichong)

[0090] Literature information: Source: Acupuncture and Moxibustion Jia Yi Jing, author Huangfu Mi, dated around 256 AD.

[0091] 2. Data Integration

[0092] Combined with other literature (such as the modern clinical research "Analysis of acupoint combinations for acupuncture treatment of functional dyspepsia"), acupoint combinations with similar symptoms (such as the high-frequency combination of "Zhongwan + Zusanli") are associated with database D.

[0093] 2. Acupoint Effect Feature Extraction

[0094] 1Dependency Analysis (LTP Tool)

[0095] Analyzing the sentence structure, we find that "take... as the main" indicates that the main acupoint is the direct object of the action, and "assisted by" suggests that the auxiliary acupoint is a supplementary component.

[0096] Extract dependency relationship: Subject-predicate relationship: Take → Zusanli, Zhongwan, Neiguan, Gongsun; Modification relationship: Assisted by → Liangmen, Taichong

[0097] Conclusion: The main acupoints are directly related through the core verb "take", and the auxiliary acupoints are modified by "assist with".

[0098] 2. Phrase structure analysis (Berkeley Parser)

[0099] The punctuation is: "Take [Zusanli, Zhongwan, Neiguan, Gongsun] as the main points" and "Assist with [Liangmen, Taichong]".

[0100] Identify noun phrases (NP): main acupoint NP (Zusanli, etc.), auxiliary acupoint NP (Liangmen, etc.).

[0101] 3. Matching dictionary

[0102] In the self-built dictionary, "Zusanli-Zhongwan" is marked as a high-frequency combination for "harmony of spleen and stomach", and "Neiguan-Gongsun" is a classic pairing for "expanding chest and regulating qi".

[0103] Matching results: The combination of main acupoints conforms to the known compatibility pattern for treating epigastric pain, and the acupoint combination "Liangmen-Taichong" is an auxiliary solution for soothing the liver and regulating qi.

[0104] 3. Acupoint role labeling (BiLSTM-CRF model based on PSO optimization)

[0105] 1. Data preprocessing and word vector generation

[0106] Word segmentation results: [Zusanli, Zhongwan, Neiguan, Gongsun, Liangmen, Taichong]

[0107] Generate 5-dimensional word vector (example):

[0108] Zusanli: [0.8, 0.2, 0.5, 0.1, 0.3] (encoding the “strengthening the spleen and stomach” feature)

[0109] Beam gate: [0.3, 0.6, 0.1, 0.4, 0.2] (encoding the “local analgesia” feature)

[0110] Concatenate linguistic features (such as "primary" and "secondary" labels) and input them into BiLSTM.

[0111] 2. PSO Optimizes Hyperparameters

[0112] Phase 1: Optimize the parameter range in 50 iterations to determine the learning rate (a=0.01), hidden layer dimension (b=128), and Dropout rate (c=0.1).

[0113] Phase 2: After 10 iterations, the optimal parameters were locked, and the accuracy of 10-fold cross-validation reached 92%.

[0114] 3. BiLSTM-CRF model prediction

[0115] Input sequence: [Zusanli, Zhongwan, Neiguan, Gongsun, Liangmen, Taichong], load ∆bacup from self-built dictionary, and enhance gated calculation.

[0116] BiLSTM output score matrix:

[0117] Main acupoints with high scores: Zusanli (0.9), Zhongwan (0.85), Neiguan (0.88), Gongsun (0.82)

[0118] Low score for acupoint pairing: Liangmen (0.45), Taichong (0.4)

[0119] CRF layer constraints: set illegal transfer penalties (such as A=-∞ for auxiliary acupoint→main acupoint), calculate legal path scores, and force the “main acupoint” label to be continuous to avoid jump labeling.

[0120] 4. Role recognition results

[0121] Main acupoints: Zusanli, Zhongwan, Neiguan, Gongsun

[0122] Acupoints: Liangmen, Taichong

[0123] Basis: "Take... as the main point" in the dependency relationship directly refers to the main acupoints; the combination of main acupoints in the compatibility dictionary is the core solution for treating epigastric pain;

[0124] The model output score is consistent with the CRF transition probability.

[0125] 4. Result Verification and Optimization

[0126] Clinical verification: Compared with modern research (CNKI literature), the main acupoint combination "Zusanli + Zhongwan" was used in more than 80% of epigastric pain cases, which is consistent with the model results.

[0127] Optimization direction: If "Taichong" is mistakenly marked as a supporting acupoint (it is actually the main acupoint in some cases of liver depression-type epigastric pain), it is necessary to add the pathological characteristics of "liver depression" to the dictionary and retrain the model.

[0128] The effects of this embodiment are:

[0129] 1. Domain adaptation: Injecting TCM prior knowledge through the ∆bacup mechanism improves the accuracy of acupoint recognition;

[0130] 2. Compatibility logic guarantee: CRF constraints reduce the error rate of illegal label transfer (such as primary acupoint → auxiliary acupoint);

[0131] 3. Efficient optimization: Segmented PSO shortens the hyperparameter search time while maintaining parameter quality.

[0132] The above solution is only an illustration of a preferred embodiment, but is not limited thereto. When implementing the present invention, appropriate replacement and / or modification can be performed according to user needs.

[0133] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily realized. Therefore, without departing from the general concept defined by the claims and equivalent scope, the present invention is not limited to the specific details and the illustrations shown and described here.

Claims

1. A method for identifying various roles in acupoint combination based on named entities, characterized in that: include: S1. Collect ancient medical records of related diseases and clinical research literature on acupuncture treatment, and build a literature database D on acupuncture treatment; S2, based on the analysis and processing of the literature database D, each acupoint name is characterized by a self-built acupuncture point compatibility data set; S3, based on the segmented PSO model structure optimization, a subset of the acupuncture point compatibility dictionary is optimized for hyperparameters to obtain the hyperparameter configuration in the BiLSTM+CRF model for named entity recognition, and the entire literature database D is trained and tested to obtain the acupuncture point compatibility recognition model; S4. For any acupoint sequence And the corresponding label sequence The acupoint compatibility recognition model identifies the role of each acupoint in treating the corresponding disease through the following formula: In the above formula, P is the calculation score matrix of the two-layer LSTM neural network, Indicates that the i-th acupoint in the acupuncture prescription is named y i The score of the role label, σ represents the sigmoid function in logistic regression, A represents the transition probability matrix between labels, The first label of the sequence is y i , Ɛ1 and Ɛ2 represent the weights of the label probability matrix and label score, and Ɛ1+Ɛ2=1; The BiLSTM+CRF model includes an LSTM unit and a Bi-LSTM unit; The LSTM unit takes the acupuncture point compatibility dataset z as input, and constructs an enhanced dynamic neural memory unit by introducing a dynamic gating mechanism, a multi-head memory coupling mechanism, a hierarchical feature fusion strategy and acupoint functional characteristics. The update and output of the LSTM unit are characterized by the following formula: In the above formula, h t is the output of the LSTM unit at time t, represents the encoded acupoint function features dynamically obtained from the self-built dictionary, C is the value of the LSTM memory unit, i t , f t , O t , , C t They represent the input gate, forget gate, output gate, candidate value of the memory unit state at the current moment, and state value respectively. W c , W i , W f , W o Respectively represent the weight matrices input to the memory unit state candidate value, input gate, forget gate, and output gate, represents K independent memory head parameter groups, represents M independent input gate parameter groups, , , , represents the learnable gating parameter vector, Z t express t Input the acupoint characteristics of the acupoint dictionary at any time. h t-1 Represents the output of the LSTM unit at the previous moment, b c , b i , b f , b o They represent the bias items input to the candidate value of the memory unit state, the input gate, the forget gate, and the output gate, respectively. U c , U i , U f , U o They represent the weight matrices from the previous hidden state to each gate, ATT(·) represents the adaptive attention mechanism based on the self-built acupoint role matrix, represents the pth memory transformation layer with differentiable wavelet transform kernel function, MLP(·) represents the multi-layer perceptron based on the self-built acupoint role matrix, Feature extraction network representing the functional characteristics of acupoints, represents the weight distribution coefficient of the kth memory head when the memory unit state is updated, represents the weight distribution coefficient of the mth input gate parameter group when the memory unit state is updated, represents the learnable parameters, represents element-wise multiplication, ⊕ represents tensor concatenation and projection operations, ⊗ represents Kronecker product expansion operations, C t-1 represents the state value of the memory unit at the previous moment, σ represents the sigmoid function in logistic regression, K represents the number of independent memory heads, M represents the number of independent input gates, and P represents the number of memory transformation layers.

2. The method for identifying roles in acupoint combination based on named entities according to claim 1, characterized in that: In S2, the analysis process includes: S20, preprocessing the text in the literature database D; S21, performing dependency grammar analysis on the preprocessed text through the open source toolkit LTP to extract the dependency relationship between the main and auxiliary roles of each acupoint in the acupoint combination; S22. Analyze the preprocessed text using the open source syntactic analyzer Berkeley Parser to identify the components and phrase structures in the sentence, and segment the symptom sentences based on the analysis results to complete the phrase structure syntactic analysis; S23. Based on the dependency relationship between the main and auxiliary roles and the results of the syntactic analysis of the phrase structure, the effect characteristics of each acupoint in the treatment of the corresponding disease are obtained to establish an acupuncture point compatibility dataset.

3. The method for identifying roles in acupoint combination based on named entities as claimed in claim 1, characterized in that: In S3, the process of obtaining hyperparameter configuration includes: S30, in the initialization of PSO, the dimension of the particle is used to represent the parameters of each acupoint role labeling; S31. In the first stage, the particle parameters were substituted into the BiLSTM+CRF model of acupoint role labeling for double cross validation to evaluate the fitness value of the function, update the speed and position of each particle, and complete 50 iterations; S32. Based on the convergence analysis of each dimension of particles in the first stage, the initial range of each dimension of PSO particles in the second stage is determined, and the BiLSTM+CRF model is cross-validated 10 times and iterated ten times to select the 6 best particles and obtain the corresponding hyperparameter configuration.

4. The method for identifying roles in acupoint combination based on named entities according to claim 1, characterized in that: The output of the Bi-LSTM unit P Obtained by the following formula: In the above formula, , They are based on the output of LSTM units respectively h t The calculated forward hidden layer sequence and reverse layer sequence, g is the adaptive gating weight, and .

5. The method for identifying roles in acupoint combination based on named entities according to claim 1, characterized in that: The BiLSTM+CRF model is strengthened by CRF layer transfer constraints; Among them, in the label transfer probability matrix of the CRF layer, a dynamic penalty term is set for illegal transfers by the following formula: : 。

Citation Information

Patent Citations

  • A software multi-error positioning method based on particle swarm optimization and a processing device

    CN109885471A

  • Driver road rage emotion detection method based on machine learning

    CN116189267A