A moxibustion decision-making method and system based on Bert pre-training and deep reinforcement learning

By using the BERT pre-trained model and the deep reinforcement learning algorithm DQN, the problem of moxibustion robots being unable to provide personalized treatment was solved, realizing personalized and intelligent moxibustion treatment plans and improving treatment effectiveness and efficiency.

CN119694493BActive Publication Date: 2026-02-06HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411970946.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2026-02-06
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing moxibustion robots are unable to provide personalized treatment plans based on the individual patient's case, resulting in poor treatment outcomes.

Method used

A BERT pre-trained model is used for intent recognition and entity recognition of moxibustion-related data. The deep reinforcement learning algorithm DQN is combined for agent decision training to build personalized moxibustion treatment plans.

Benefits of technology

The moxibustion robot can provide personalized and intelligent treatment plans based on the patient's actual medical information, thereby improving treatment effectiveness and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119694493B_ABST
    Figure CN119694493B_ABST
Patent Text Reader

Abstract

The application discloses a moxibustion decision-making method and system based on Bert pre-training and deep reinforcement learning, belongs to the field of artificial intelligence technology and the medical industry, and comprises the following steps: step 1, data collection, collecting data related to moxibustion therapy, including case data, medical literature and moxibustion operation records; the collected data is subjected to data cleaning and preprocessing, noise is removed, missing values are processed, and data is standardized; step 2, according to user data information and a pre-trained Bert network model, state representation of moxibustion diagnosis and treatment related information; step 3, based on the obtained state representation, input into a deep reinforcement learning network model, and decision learning of an intelligent agent is carried out. The application serves patients themselves, realizes the autonomy, intelligence and individualization of moxibustion diagnosis and treatment, and better promotes the development of moxibustion diagnosis and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of artificial intelligence technology and medical industry, and particularly relates to a moxibustion decision-making method and system based on Bert pre-training and deep reinforcement learning. BACKGROUND

[0002] With the deep cultivation of artificial intelligence technology in the medical field and the invention of mechanical arms, sensing devices and other equipment, moxibustion diagnosis and treatment technology has been well developed, and moxibustion robots have gradually entered thousands of households. The tasks that originally needed to be completed by moxibustion experts can gradually be completed by robots, reducing the difficulty of diagnosis and treatment operations.

[0003] Among them, the application number CN202010302192.7 proposes a traditional Chinese medicine moxibustion system and method, which simulates the moxibustion operation process through sensors and mechanical arms, and constructs a Markov model for moxibustion decision-making according to the temperature data of the skin acupoints of the moxibustion within a certain time.

[0004] Among them, the application number CN202210356305.0 provides a traditional Chinese medicine dynamic diagnosis and treatment scheme optimization method and system based on deep reinforcement learning, which combines a reinforcement learning model to train a diagnosis and treatment optimization model and obtain the best patient diagnosis and treatment scheme.

[0005] Traditional moxibustion methods require experts to hold moxa sticks and repeatedly adjust the distance between the moxa sticks and the acupoint skin, which is very inefficient. With the development of technology, new moxibustion devices and moxibustion robots have appeared, which are simple to operate and have high safety, but cannot be targeted to acupoints "prescription according to symptoms", so the treatment effect is not good. Therefore, according to the individual case of the patient, a personalized diagnosis and treatment scheme is recommended, and reinforcement learning is combined to let the moxibustion robot learn the best strategy for moxibustion, which can well improve the treatment effect. SUMMARY

[0006] The application uses a Bert model to perform intent recognition and entity recognition on user input moxibustion-related questions, constructs a state representation of moxibustion diagnosis and treatment, and uses a deep reinforcement learning algorithm DQN (Deep Q Network) for agent decision-making training, so that the moxibustion robot can learn the best strategy for moxibustion and provide personalized moxibustion diagnosis and treatment schemes. The above application is different in design idea and functional module.

[0007] Specifically, the application provides the following technical solutions:

[0008] A moxibustion decision-making method based on Bert pre-training and deep reinforcement learning, comprising the following steps:

[0009] Step 1: Data collection. Collect data related to moxibustion therapy, including case data, medical literature, and moxibustion operation records; perform data cleaning and preprocessing on the collected data, including noise removal, handling missing values, and data standardization.

[0010] Step 2: Based on user data and a pre-trained BERT network model, obtain the state representation of moxibustion diagnosis and treatment related information;

[0011] Step 3: Based on the obtained state representation, input it into the deep reinforcement learning network model to perform decision learning for the agent.

[0012] This invention also proposes an moxibustion decision-making system based on BERT pre-training and deep reinforcement learning, comprising:

[0013] Moxibustion data acquisition module: used to acquire data from moxibustion diagnosis and treatment;

[0014] BERT model training module: used to extract effective features for subsequent model training and decision-making;

[0015] Moxibustion treatment decision-making module: used to continuously interact with the environment and adjust parameters, enabling the agent to learn the optimal moxibustion treatment decision.

[0016] The present invention has the following beneficial effects:

[0017] The moxibustion method and system based on the BERT model and deep reinforcement learning described above can calculate the treatment plan required by the user based on the patient's actual case information. The treatment plan is then fed into a DQN for learning feedback, allowing the agent to obtain the corresponding treatment plan. Finally, the treatment plan is evaluated and fed back, ultimately serving the patient and realizing the autonomy, intelligence, and personalization of moxibustion treatment, thus better promoting the development of moxibustion treatment. Attached Figure Description

[0018] Figure 1 This is a flowchart of the method in an embodiment of the present invention;

[0019] Figure 2 This is a schematic diagram of the structure of the moxibustion decision system based on BERT pre-training and deep reinforcement learning in an embodiment of the present invention.

[0020] The attached diagram is labeled as follows: Moxibustion data acquisition module 100, BERT model training module 200, and Moxibustion diagnosis and treatment decision module 300. Detailed Implementation

[0021] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other. In order to achieve the above purpose, the technical scheme of the present application is as follows.

[0022] Figure 1 A moxibustion decision-making flowchart based on Bert pre-training and deep reinforcement learning. The method comprises:

[0023] Step 1: Data collection, collect data related to moxibustion therapy, and convert the collected data into a form acceptable to the Bert model; specifically including:

[0024] S1: Obtain the clinical moxibustion diagnosis and treatment data of the user;

[0025] The collected data includes open source data sets, doctor-patient diagnosis and treatment conversation records, and moxibustion-related clinical data: including medical records, symptom descriptions, moxibustion plans, etc.

[0026] S11: According to the information obtained in S1, convert the text data into a form acceptable to the Bert model. Preprocess the text data such as user requirements, symptom descriptions and moxibustion plans, use the word segmentation tool jieba library for word segmentation, then mark and encode the word segmentation results, and convert them into the input format of the BERT model, that is, convert each sentence into a vector sequence, and add a special identifier "[CLS]" before each text.

[0027] Step 2: Obtain the state representation of moxibustion diagnosis and treatment related information according to the user data information and the pre-trained Bert network model; specifically including:

[0028] S2: Use the pre-trained Bert model to extract features, and input the preprocessed text data into the Bert model to obtain the output of the BERT model, including the word vector representation of each word and the sentence-level representation. Then use a small amount of labeled data for fine-tuning to adapt to the intent recognition and entity recognition tasks of moxibustion diagnosis and treatment.

[0029] S21: After obtaining the representation vector of the text, use the Softmax classifier to calculate the probability distribution of the text belonging to each category, and calculate the classification loss (Loss) according to the probability distribution, and finally update the parameters of the entire model according to the loss. Use the cross-entropy loss function to calculate the classification loss, and the specific formula is as follows:

[0030] ,

[0031] In the formula, M represents all categories of text classification, represents the probability that the classifier predicts the text o as the c category. When the c category is the true category of the text o, the value is 1, and otherwise the value is 0.

[0032] S22: Fuse the extracted vector and the output of the Bert model to construct a complete state representation. Use splicing, weighted averaging, etc. to fuse different features to form a comprehensive state representation. The state representation of moxibustion diagnosis and treatment related information includes user's demand, disease description, etc. and the current moxibustion scheme and treatment progress. The specific introduction is as follows:

[0033] User demand: including user's moxibustion purpose, symptom description, pain degree, etc. These information is through the Bert model to carry on the intent recognition and entity recognition, will user's natural language input into the machine can understand the representation.

[0034] Disease description: including patient's medical history, past treatment, disease development, etc. These information is encoded through the Bert model, which converts text data into vector representation.

[0035] Current moxibustion scheme: including the current selected moxibustion acupoint, moxibustion time, temperature, etc. These information is represented by discrete One-hot encoding.

[0036] Treatment progress: including the improvement degree of disease, the feedback of patients, etc. These information is evaluated by setting three indexes, which are (1) disease improvement degree index: by comparing the symptom improvement of patients before and after treatment to evaluate the treatment progress, a symptom improvement degree scoring system is set, using 0-10 score, 0 represents no improvement, 10 represents complete improvement. According to the change of patient's symptoms, give the corresponding score, then calculate the average score to evaluate the treatment progress. (2) Pain degree evaluation: for the improvement of pain symptoms, use the pain visual analog scale (Visual Analog Scale, VAS) tool, let the patient give a 0-10 score according to his own feeling, 0 represents no pain, 10 represents the most severe pain. By comparing the pain scores before and after treatment, the treatment progress is evaluated. (3) Patient satisfaction evaluation: by designing a satisfaction questionnaire, including the evaluation of patients on treatment effect, pain relief degree, life quality improvement, etc. The score of patient satisfaction is calculated to evaluate the treatment progress.

[0037] Step 3: Based on the obtained state representation, input into the deep reinforcement learning network model for intelligent agent decision learning. Specifically, S3: Establish a moxibustion diagnosis and treatment decision system based on deep reinforcement learning:

[0038] In the moxibustion decision-making process, the state representation output by the Bert model is used as the input state for the agent (moxibustion robot) decision-making. According to the application scenario and requirements of the agent, the action space of the agent is defined. The action space includes selecting different moxibustion acupoints, adjusting moxibustion time and temperature, etc. Each action is encoded as a discrete action label. Taking the action space definition of a heat-sensitive moxa as an example, the corresponding label data is designed: actions = {

[0039] "moxa acupoint selection": ["baihui acupoint", "fengchi acupoint", "dazhui acupoint", "yongquan acupoint", "taichong acupoint"],

[0040] "moxa time adjustment": ["5 minutes", "10 minutes", "15 minutes", "20 minutes", "25 minutes"],

[0041] "moxa temperature adjustment": ["low temperature", "medium-low temperature", "medium temperature", "medium-high temperature", "high temperature"]

[0042] }。

[0043] Through interaction with the environment, the robot selects the best action according to the current state to optimize the moxibustion strategy. According to the current state and the selected action, the state transition is simulated, and the reward function is defined. The reward function is based on the improvement of the disease, patient feedback and doctor evaluation to comprehensively evaluate the three indicators to guide the reinforcement learning model to learn good decision-making strategies. The disease improvement index is valued between 0 and 1, where 0 means no improvement in the disease, and 1 means complete improvement of the disease, with a weight of 0.4. The patient feedback index is valued between 0 and 1, where 0 means the patient is not satisfied with the moxibustion effect, and 1 means the patient is very satisfied with the moxibustion effect, with a weight of 0.3; The doctor evaluation index is valued between 0 and 1, where 0 means the doctor thinks the moxibustion effect is very poor, and 1 means the doctor thinks the moxibustion effect is very good, with a weight of 0.3.

[0044] A deep reinforcement learning algorithm, DQN, is used to establish a neural network model based on moxibustion decision-making. The input of the model is the state representation, and the output is the corresponding action of the agent. A convolutional neural network (CNN) is used as the network structure of DQN. Using existing state transition data and reward functions, the model is trained through reinforcement learning algorithm. Experience Replay is used to improve the training effect and stability. During the training process, the model will try different actions and feedback according to the reward function, gradually optimizing the decision-making strategy. In practical application, according to the current state input into the trained reinforcement learning model, the model will output an action, that is, the decision that the moxibustion diagnosis and treatment system should take. According to the action output by the model, the system can make the agent perform the corresponding moxibustion operation or adjustment.

[0045] In addition, a moxibustion diagnosis and treatment system based on Bert model and deep reinforcement learning is also provided.

[0046] Figure 2 is an embodiment of a moxibustion decision-making system structure based on Bert pre-training and deep reinforcement learning. The system includes a moxibustion data acquisition module 100, a Bert model training module 200, and a moxibustion diagnosis and treatment decision-making module 300.

[0047] The moxibustion data acquisition module 100 is used to acquire data for moxibustion diagnosis and treatment.

[0048] The user logs in to the system, fills out the user's personal CRF table, and collects the patient's relevant data, including the patient's basic information (name, age, gender, medical history, etc.); symptom description (main symptoms, pain description, other symptoms, etc.), and stores the above information in text form and uploads it to the cloud platform. The collected data set is preprocessed, such as cleaning, denoising and normalization, to ensure the accuracy and consistency of the data.

[0049] The Bert model training module 200 is used to extract effective features for subsequent model training and decision-making.

[0050] The pre-trained Bert model encodes and processes the collected data to extract main features, predicts the moxibustion effect according to the patient's characteristics, and determines the specific operation scheme of moxibustion, such as selecting acupoints, moxibustion methods, and heat levels, etc. It generates as output data of the Bert model, which is used as the initial state representation for agent decision-making input to the next module.

[0051] The moxibustion diagnosis and treatment decision-making module 300 is used to train the agent (moxibustion robot) to learn relevant decision-making operations.

[0052] By accepting the predicted state in the Bert model training module 200 as the initial state in the agent decision-making process, combining the pre-set action space and reward function, inputting the deep reinforcement learning-based network model training agent, the agent will output an action, i.e. the decision that the moxibustion robot should take, and through continuous interaction with the environment and updating of model parameters, the moxibustion robot can learn the best strategy for moxibustion and improve the personalized moxibustion diagnosis and treatment scheme.

Claims

1. A moxibustion decision system based on Bert pre-training and deep reinforcement learning, characterized in that, The method comprises the following steps: The moxibustion data collection module is used for data collection, collecting data related to moxibustion therapy, and converting the collected data into an input form acceptable to the Bert model; The Bert model training module obtains the state representation of the moxibustion diagnosis and treatment related information according to the user data information and the pre-trained Bert network model; the pre-trained Bert model is used for feature extraction, and the preprocessed text data is input into the Bert model to obtain the output of the Bert model, including the word vector representation of each word and the sentence level representation, and then the labeled data is used for fine tuning to adapt to the intent recognition and entity recognition tasks of moxibustion diagnosis and treatment; the extracted vector and classification features and the output of the Bert model are fused to construct a complete state representation, different features are fused by using splicing and weighted average to form a comprehensive state representation, and the state representation of the moxibustion diagnosis and treatment related information includes user demand, disease description information, and the current moxibustion scheme and treatment progress; the treatment progress includes: disease improvement degree index, pain degree evaluation, and patient satisfaction evaluation; The moxibustion diagnosis and treatment decision module inputs the obtained state representation into a deep reinforcement learning network model for agent decision learning; the action space of the agent is defined according to the application scene and requirements of the agent, and the action space includes selecting different moxibustion acupoints, adjusting moxibustion time and temperature, and encoding each action into a discrete action label; Through interaction with the environment and parameter adjustment, the agent selects the best action according to the current state to optimize the moxibustion strategy; the state transition is simulated according to the current state and the selected action, and a reward function is defined, which comprehensively evaluates the disease improvement degree, patient feedback and doctor evaluation to guide the reinforcement learning model to learn a good decision strategy.

2. The moxibustion decision system based on Bert pre-training and deep reinforcement learning according to claim 1, characterized in that, The collected data includes open source data sets, specifically including: doctor-patient diagnosis and treatment dialogue records, and moxibustion related clinical data: including user demand, medical history, symptom description, and moxibustion scheme.

3. The moxibustion decision system based on Bert pre-training and deep reinforcement learning according to claim 2, characterized in that, The preprocessing specifically includes: using the jieba library for word segmentation, then marking and encoding the word segmentation results, and converting them into the input format of the Bert model, that is, converting each sentence into a vector sequence, and adding an identifier "[CLS]" before each text.

4. The moxibustion decision system based on Bert pre-training and deep reinforcement learning according to claim 3, characterized in that, The Bert model training module specifically includes: extracting data features using a pre-trained Bert model to obtain text representation vectors; classifying the text representation vectors through Softmax to extract corresponding classification features; and fusing the extracted vectors and classification features with the output of the Bert model to construct a complete state representation.

5. The moxibustion decision system based on Bert pre-training and deep reinforcement learning according to claim 4, characterized in that, The Bert model training module further includes: After obtaining the text representation vectors, a Softmax classifier is used to calculate the probability distribution of the text belonging to each class, and the classification loss Loss is calculated according to the probability distribution, and finally the loss is back propagated to update the parameters of the entire model, and the cross-entropy loss function is used to calculate the classification loss, and the specific formula is as follows: , In the formula, M represents all categories of text classification, represents the probability that the classifier predicts the text o to be of class c, when c is the true class of the text o, has a value of 1, and otherwise has a value of 0.

6. The moxibustion decision system based on Bert pre-training and deep reinforcement learning according to claim 5, characterized in that, User needs, disease description, and current moxibustion scheme and treatment progress also include: User needs include user's moxibustion purpose, symptom description, and pain degree; these information is through Bert model for intent recognition and entity recognition, the natural language input of user is converted into machine understandable representation; Disease description includes patient's medical history, past treatment, and disease development; these information is encoded by Bert model, and text data is converted into vector representation; Current moxibustion scheme includes currently selected moxibustion acupoint, moxibustion time, and temperature, which are represented by discrete One-hot encoding; Disease improvement degree index is: by comparing the symptom improvement of patients before and after treatment, a symptom improvement degree scoring system is set, using 0-10 score, 0 means no improvement, 10 means complete improvement; according to the change of patient's symptoms, give the corresponding score, then calculate the average score to evaluate the treatment progress; pain degree evaluation: for the improvement of pain symptoms, use the pain visual analog scale VAS tool, let the patient give a 0-10 score according to his own feeling, 0 means no pain, 10 means the most severe pain, by comparing the pain score before and after treatment, evaluate the treatment progress; patient satisfaction evaluation: by designing a satisfaction questionnaire, including patient's evaluation of treatment effect, pain relief degree, and life quality improvement, statistical patient satisfaction score to evaluate the treatment progress.

7. The moxibustion decision system based on Bert pre-training and deep reinforcement learning according to claim 6, characterized in that, In the moxibustion diagnosis and treatment decision module, in the moxibustion decision process, the state representation output by the Bert model is used as the input state of the agent decision, the value range of the disease improvement degree index is 0 to 1, where 0 means no improvement of the disease, and 1 means complete improvement of the disease, the weight is set to 0.4, the value range of the patient feedback index is 0 to 1, where 0 means the patient is not satisfied with the moxibustion effect, and 1 means the patient is very satisfied with the moxibustion effect, the weight is set to 0.3; the value range of the doctor evaluation index is 0 to 1, where 0 means the doctor thinks the moxibustion effect is very poor, and 1 means the doctor thinks the moxibustion effect is very good, the weight index is set to 0.

3.

8. The moxibustion decision system based on Bert pre-training and deep reinforcement learning according to claim 7, characterized in that, The agent uses deep reinforcement learning algorithm DQN to establish a neural network model based on moxibustion decision, the input of the model is state representation, and the output is the action corresponding to the agent, convolutional neural network is used as the network structure of DQN, existing state transition data and reward function are used, reinforcement learning algorithm is used to train the model, experience replay is used to improve the training effect and stability, during the training process, the model will try different actions, and feedback according to the reward function, gradually optimize the decision strategy, in actual application, according to the current state input into the trained reinforcement learning model, the model outputs an action, that is, the decision that the moxibustion diagnosis and treatment system should take, according to the action output by the model, the agent carries out corresponding moxibustion operation or adjustment.

Citation Information

Patent Citations

  • Traditional Chinese medicine moxibustion system and method thereof

    CN111449942A

  • Traditional Chinese medicine dynamic diagnosis and treatment scheme optimization method and system based on deep reinforcement learning

    CN114783571A

  • Control method and device of intelligent moxibustion robot and storage medium

    CN116549284A

  • Massage robot control system optimization method and system based on AI driving

    CN118544372A