A method and device for a traditional Chinese medicine chat robot based on a pre-trained language model

Through the traditional Chinese medicine chat robot based on the pre-trained language model, people's lack of knowledge about traditional Chinese medicine has been solved, high-quality responses to traditional Chinese medicine health and life problems and convenient traditional Chinese medicine consulting services have been achieved, and popularization of traditional Chinese medicine knowledge and scientific research development have been promoted.

CN118916458BActive Publication Date: 2025-05-13YANSHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410951645.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-05-13
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

People lack knowledge of traditional Chinese medicine, and existing chat robot technology cannot meet users' needs for personalized and intelligent dialogue.

Method used

A Chinese medicine chat robot is constructed based on a pre-trained language model, and by constructing a Chinese medicine data set, training the pre-trained model, processing user input information, and generating Chinese medicine knowledge answers or suggestions.

Benefits of technology

It has achieved high-quality response to users' health and life problems in traditional Chinese medicine, popularized traditional Chinese medicine knowledge, provided convenient and efficient traditional Chinese medicine consulting services, optimized model training and efficiency, and promoted the development of traditional Chinese medicine scientific research and product.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118916458B_ABST
    Figure CN118916458B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for a traditional Chinese medicine chat robot based on a pre-trained language model, which belongs to the field of natural language processing and artificial intelligence technology, and includes four modules: data collection and pre-processing, model training, model reasoning, and interaction. The present invention first collects original corpus data of sensitive words and traditional Chinese medicine consultation dialogues, then pre-processes the data to construct a relevant traditional Chinese medicine data set, then transfers the sorted data set to the pre-trained model for model training and saves it, and deploys the trained model to an actual system for reasoning and answering traditional Chinese medicine related questions raised by users. Finally, the interaction module designs an interface that can perform multiple rounds of interaction with users, and displays corresponding traditional Chinese medicine knowledge and suggestions generated by the robot according to the user's questions. The present invention improves the user's satisfaction and convenience in traditional Chinese medicine inquiries, and brings a brand-new solution to the field of traditional Chinese medicine health consultation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing and artificial intelligence technology, and specifically relates to a method and device for a traditional Chinese medicine chat robot based on a pre-trained language model. Background Art

[0002] Traditional Chinese medicine has a long history and rich cultural connotations, and is an indispensable cultural treasure of our country. The ancient Chinese people's cognition and treatment of diseases have experienced a long period of historical accumulation and development, forming a unique system of traditional Chinese medicine. Chinese medicine has always served people's healthy lives and made indelible contributions to ancient and contemporary disease prevention and control.

[0003] Nowadays, the Internet, information communication and artificial intelligence technologies are developing rapidly. As a new way to communicate with computing devices, human-computer dialogue systems have natural and convenient characteristics and are considered to be a new generation of future interaction paradigms. More and more users are beginning to use intelligent chatbots to interact with computers, which will further develop the technology of intelligent chatbots. Therefore, under the current technological background, we can seize this opportunity to invent a Chinese medicine chatbot to fill the gap in people's access to Chinese medicine knowledge. People can use mobile phones or computers to ask questions or symptoms of discomfort, and the robot can give corresponding answers or suggestions on Chinese medicine knowledge, which not only enriches people's knowledge of Chinese medicine but also serves their healthy life.

[0004] From the perspective of application technology, intelligent chatbots can be divided into goal-driven and non-goal-driven, retrieval-based and generative. However, traditional retrieval-based chatbots can no longer meet users' needs for personalized and intelligent conversations. Therefore, intelligent chatbots based on deep learning and pre-trained language models have emerged. With the Transformer model proposed by Google and the GPT series of models launched by the OpenAI team, pre-trained language models have become a research hotspot in the field of natural language processing. By training on large-scale unlabeled corpora, pre-trained language models can learn the statistical characteristics and semantic information of language and show excellent performance in various natural language processing tasks. Inspired by this technology, researchers began to apply pre-trained language models to intelligent chatbots. This method builds an open-domain, automatically generated response chatbot, which enables the robot to understand users' questions more accurately and generate fluent and natural answers. Summary of the invention

[0005] In order to solve the problem of people's lack of knowledge of traditional Chinese medicine, facilitate the service of people's healthy life, and overcome the shortcomings of existing chat robot technology, the present invention provides a method and device of a traditional Chinese medicine chat robot based on a pre-trained language model. The present invention provides a method and device of a traditional Chinese medicine chat robot based on a pre-trained language model, which is used to respond to users' health life problems with high quality based on relevant knowledge of traditional Chinese medicine.

[0006] The technical solution adopted by the method and device of a traditional Chinese medicine chat robot based on a pre-trained language model of the present invention is:

[0007] A method for a traditional Chinese medicine chat robot based on a pre-trained language model, comprising the following steps:

[0008] S1. Build relevant TCM datasets;

[0009] S2, training the pre-trained model based on the traditional Chinese medicine dataset;

[0010] S3: Process the current user's input information based on the trained model, obtain the corresponding output vector, splice and generate content, and realize model reasoning;

[0011] S4. Display the inference results in the form of a web page, interact with users, and give Chinese medicine advice.

[0012] A further improvement of the solution of the present invention is that the specific steps of step S1 are as follows:

[0013] S1.1, traverse the collected raw data files and make judgments, dividing the data into two categories: questions and answers;

[0014] S1.2. Remove incomplete questions and simple answers, use existing sensitive word identification tools to detect data, and discard any sensitive words;

[0015] S1.3, remove HTML tags, repeated punctuation and tag symbols in questions and answers; if the length of the cleaned data is less than a preset threshold, it is discarded; where HTML stands for Hypertext Markup Language;

[0016] S1.4. Save the cleaned data as a TCM dataset in the form of question-answer pairs in the format required for model training.

[0017] A further improvement of the solution of the present invention is that the specific steps of step S2 are as follows:

[0018] S2.1. Use the data loading tool to pass the TCM dataset organized into question-answer pairs into the pre-training model;

[0019] S2.2. Set the parameters required for model training;

[0020] S2.3, optimize model parameters through loss function;

[0021] S2.4. Set the random seed to ensure that the model effect can be reproduced under the same device and the repeatability of the experimental results;

[0022] S2.5, load the model parameters that have been successfully pre-trained and adjust them according to the new data set to meet the needs of specific tasks;

[0023] S2.6. Load the data required for training and perform mask operations on the target sentences in the original text dataset; record and convert the relevant information into a tensor form that can be used by the model;

[0024] S2.7. Set the parameters of the Adamw optimizer and start model training. Divide the original text dataset into 32 batches, traverse each batch of data, and calculate its loss value.

[0025] S2.8, the loss value is transmitted back, and the model parameters are updated through the optimizer, the gradient is calculated through the back propagation algorithm, and the model parameters are updated through the optimizer;

[0026] S2.9. After each batch of data training is completed, save the model.

[0027] A further improvement of the solution of the present invention is that the step S2.3 of optimizing the model parameters by using the loss function comprises the following steps:

[0028] S2.3.1. Set the model initialization function and define the components required for model training;

[0029] S2.3.2. Load the weights and configure the forward propagation function of the model, which includes passing the input data through the encoder part and obtaining the sequence output, whose dimension is [batch_size, sequence_length, hidden_size];

[0030] Among them, batch_size is the batch size, sequence_length is the sequence length, and hidden_size is the number of hidden memory units;

[0031] S2.3.3. Obtain the vector of the masked position, calculate the loss value of the word or words replaced with the special token, and use the calculated loss value in the optimization process of the model parameters.

[0032] A further improvement of the solution of the present invention is that the overall architecture of the model trained in step S2 can be expressed as the following formula:

[0033]

[0034] The Encoder i is the i-th TransformerEncoder layer, and They represent the output tensors of the i-th and i-1-th encoding layers respectively;

[0035] Each encoder layer contains two sub-layers: Multi-head Self-Attention (MSA) and Feedforward Neural Network (FFN). The outputs of these two sub-layers are connected through residuals and input into a layer normalization (LayerNorm) layer, which is specifically expressed as follows:

[0036] y = text{LayerNorm}(x + text{SubLayer}(x)) (2)

[0037] Where y represents the output tensor, x represents the input tensor, text{LayerNorm} represents the layer normalization layer, and text{SubLayer} represents the sublayer (i.e., MSA or FFN) layer;

[0038] The FFN (feedforward neural network) layer is expressed as follows:

[0039] FFN(x)=max(0,xW1+b1)W2+b2 (3)

[0040] Among them, the input tensor of the FFN layer is x, W1 and W2 in the FFN layer are the weight parameters of the first and second hidden linear layers respectively, and b1 and b2 represent the bias parameters of the first and second hidden linear layers respectively.

[0041] A further improvement of the solution of the present invention is that the specific steps of step S3 are as follows:

[0042] S3.1, set the parameters required for model inference;

[0043] S3.2. Load the trained model and place it on the corresponding device;

[0044] S3.3, load the sensitive word filter to ensure that the output does not contain any sensitive information;

[0045] S3.4. Detect the content input by the user. If it contains sensitive words, directly reply with a specific sentence, otherwise proceed to the next step;

[0046] S3.5, vectorize the sentence input by the user and obtain the corresponding sequence;

[0047] S3.6, traverse according to the maximum length limit of the output sequence of the model, obtain the output vector according to the current input, and select the words with high and low probability in the vocabulary;

[0048] S3.7. When the "[SEP]" sign appears in the generated content, it indicates that the generation is completed. All generated content is spliced ​​and integrated, and it is checked whether sensitive words appear. If there are sensitive words, a specific sentence is replied; otherwise, the content generated by the model is replied.

[0049] A device including a method for a traditional Chinese medicine chat robot based on a pre-trained language model, comprising:

[0050] Data collection and preprocessing module: collect original corpus data of TCM consultation dialogues and TCM knowledge by reading ancient and modern documents; preprocess the data, clean and organize the original data, traverse the content to divide it into questions and answers, organize it into the required data format, and build relevant TCM data sets;

[0051] Model training module: pass the data set organized into question-answer pairs into the pre-trained model, set appropriate parameters, optimize the model parameters through the loss function, load the required data, start the training task, and calculate the loss value in real time to update the model parameters according to the optimization algorithm;

[0052] Model inference module: loads the trained model, configures sensitive word filters to process user input content, obtains the corresponding output vector by processing the current user input, splices and generates content, and implements model inference, that is, generates the model's response or reply to the user input;

[0053] Interaction module: The inference results are displayed in the form of a web page. Multiple rounds of interaction can be carried out with users to accurately answer user questions and provide TCM advice.

[0054] Among them, the output end of the data collection and preprocessing module is connected to the input end of the model training module, the output end of the model training module is connected to the input end of the model reasoning module, and the output end of the model reasoning module is connected to the input end of the interaction module.

[0055] The beneficial effects of the present invention are:

[0056] 1. Popularize TCM knowledge: By building a special TCM data set and training based on a pre-trained model, the TCM question-and-answer robot can provide users with knowledge about TCM theory, diagnostic methods, treatment techniques, etc., helping to fill the problem of modern people's lack of knowledge about TCM.

[0057] 2. Optimized model training and efficiency, and provided convenient and efficient TCM consultation services: The loss function was used to optimize the model parameters, and the text was processed and generated through a multi-layer Transformer encoder to ensure the quality and fluency of the generated content. At the same time, by setting random seeds and using the Adamw optimizer, the stability and efficiency of the model were improved. The TCM Q&A robot can provide TCM knowledge and consulting services anytime and anywhere, provide auxiliary suggestions and guidance, help users understand their condition, solve TCM health problems, and save users the time cost of seeking medical advice.

[0058] 3. Provide a friendly user interaction interface: This invention uses large-scale unlabeled text data for training, learns the statistical characteristics and semantic information of the language, accurately understands the user's intention, and displays the model reasoning results through an intuitive and friendly web interface to interact with the user. This design enables users to easily obtain Chinese medicine advice, and also promotes the convenience and operability of users' long-term health management.

[0059] 4. Promote TCM research and product development: By analyzing the conversations between users and chatbots, it can provide valuable information and feedback for TCM research, contribute to the development and innovation of the TCM field, and tap into users’ potential needs, providing useful reference for product design and market research. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a flow chart of a conversation method of a device of a traditional Chinese medicine chat robot based on a pre-trained language model provided by the present invention;

[0061] Figure 2 It is a model training framework diagram of a method for a traditional Chinese medicine chat robot based on a pre-trained language model provided by the present invention;

[0062] Figure 3 and Figure 4 It is a display diagram of the interactive interface of the present invention. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in combination with specific implementation methods and with reference to the accompanying drawings. In the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.

[0064] The present invention specifically comprises the following steps:

[0065] S1. Construct relevant TCM datasets.

[0066] Specifically, step S1 includes the following steps: S1.1, traverse the collected original data files and make judgments to divide the data into two categories: questions and answers.

[0067] S1.2. Remove incomplete questions and simple answers, use existing sensitive word identification tools to detect data, and discard any sensitive words.

[0068] S1.3. Remove HTML tags, repeated punctuation and tag symbols from questions and answers; if the length of the cleaned data is less than a preset threshold, it will be discarded; where HTML stands for Hypertext Markup Language.

[0069] S1.4. Save the cleaned data as a TCM dataset in the form of question-answer pairs in the format required for model training.

[0070] S2. Train the pre-trained model based on the traditional Chinese medicine dataset.

[0071] Specifically, step S2 includes the following steps: S2.1. Use a data loading tool to pass the TCM dataset organized into question-answer pairs into the pre-training model.

[0072] S2.2. Set the parameters required for model training.

[0073] S2.3. Optimize model parameters through loss function.

[0074] The above step S2.3 specifically includes the following steps: S2.3.1, setting the initialization function of the model and defining the various components required for model training.

[0075] S2.3.2. Load the weights and configure the forward propagation function of the model, which includes passing the input data through the encoder part and obtaining the sequence output, whose dimension is [batch_size, sequence_length, hidden_size].

[0076] Among them, batch_size is the batch size, sequence_length is the sequence length, and hidden_size is the number of hidden memory units.

[0077] S2.3.3. Obtain the vector of the masked position, calculate the loss value of replacing certain words or phrases in the input text with special tokens, and use the calculated loss value in the optimization process of the model parameters.

[0078] S2.4. Set the random seed to ensure that the model effect can be reproduced under the same device and to ensure the repeatability of the experimental results.

[0079] S2.5. Load the pre-trained model parameters and adjust them according to the new data set to meet the needs of specific tasks.

[0080] S2.6. Load the data required for training and perform mask operations on the target sentences in the original text dataset; record and convert the relevant information into a tensor form that can be used by the model.

[0081] S2.7. Set the parameters of the Adamw optimizer and start model training. Divide the original text dataset into 32 batches, traverse each batch of data, and calculate its loss value.

[0082] S2.8. The loss value is fed back and the model parameters are updated through the optimizer. The gradient is calculated through the back-propagation algorithm and the model parameters are updated through the optimizer.

[0083] S2.9. After each batch of data training is completed, save the model.

[0084] The overall architecture of the model obtained in S2.9 can be expressed as follows:

[0085]

[0086] The Encoder i is the i-th TransformerEncoder layer, and They represent the output tensors of the i-th and i-1-th encoding layers respectively;

[0087] Each encoder layer contains two sub-layers: Multi-head Self-Attention (MSA) and Feedforward Neural Network (FFN). The outputs of these two sub-layers are connected through residuals and input into a layer normalization (LayerNorm) layer, which is specifically expressed as follows:

[0088] y = text{LayerNorm}(x + text{SubLayer}(x)) (2)

[0089] Where y represents the output tensor, x represents the input tensor, text{LayerNorm} represents the layer normalization layer, and text{SubLayer} represents the sublayer (i.e., MSA or FFN) layer;

[0090] The FFN (feedforward neural network) layer is expressed as follows:

[0091] FFN(x)=max(0,xW1+b1)W2+b2 (3)

[0092] Among them, the input tensor of the FFN layer is x, W1 and W2 in the FFN layer are the weight parameters of the first and second hidden linear layers respectively, and b1 and b2 represent the bias parameters of the first and second hidden linear layers respectively.

[0093] S3. Process the current user’s input information based on the trained model, obtain the corresponding output vector, splice and generate content, and realize model reasoning.

[0094] Specifically, step S3 includes the following steps: S3.1, setting parameters required for model reasoning.

[0095] S3.2. Load the trained model and place it on the corresponding device.

[0096] S3.3. Load the sensitive word filter to ensure that the output does not contain any sensitive information.

[0097] S3.4. Check the content input by the user. If it contains sensitive words, directly reply with a specific sentence, otherwise proceed to the next step.

[0098] S3.5. Vectorize the statements input by the user and obtain the corresponding sequence.

[0099] S3.6. Traverse according to the maximum length limit of the model's output sequence, obtain the output vector based on the current input, and select subwords with higher and lower probabilities in the vocabulary.

[0100] S3.7. When the "[SEP]" sign appears in the generated content, it indicates that the generation is completed. All generated content is spliced ​​and integrated, and it is checked whether sensitive words appear. If there are sensitive words, a specific sentence is replied; otherwise, the content generated by the model is replied.

[0101] S4. Display the inference results in the form of a web page, interact with users, and give Chinese medicine advice.

[0102] according to Figure 1 It can be seen that the present invention provides a device of a traditional Chinese medicine chat robot based on a pre-trained language model, which includes a data collection module, a data preprocessing module, a model training module, a model reasoning module and an interaction module connected in sequence.

[0103] The data collection and preprocessing module in the above device is used to collect original corpus data of TCM consultation dialogues and TCM knowledge by reading ancient and modern documents; preprocess the data, clean and organize the original data, traverse the content to divide it into questions and answers, organize it into the required data format, and construct a relevant TCM data set.

[0104] The model training module in the above-mentioned device is used to pass the data set organized into the form of question and answer pairs into the pre-trained model, set appropriate parameters, and optimize and adjust the model parameters through the loss function; by training on a large-scale text corpus, a general representation is obtained, and then the model parameters of the pre-trained model are adjusted on the traditional Chinese medicine question and answer data set, the required data is loaded, the training task is started, and the loss value is calculated in real time so as to update the model parameters according to the optimization algorithm.

[0105] The above adjustment operations include using Transformer to train on a large-scale text corpus so that it can learn the general representation of language, and applying this pre-trained model to the TCM question-answering task. By adjusting the model parameters on the TCM question-answering dataset, the performance and generalization ability of the model on the task can be improved, which helps to speed up the convergence of the model on the task and improve its performance. Among them, Transformer represents the pre-trained model.

[0106] The model inference module in the above device is used to load the trained model, configure the sensitive word filter to process the user input content, obtain the corresponding output vector by processing the current user input, splice and generate content, and realize model inference, that is, generate the model's response or reply to the user input.

[0107] The interactive module in the above device is used to display the reasoning results in the form of a web page, and can conduct multiple rounds of interaction with the user, accurately answer user questions, and give traditional Chinese medicine suggestions.

[0108] refer to Figure 3 and Figure 4 The present invention can interact with the user in multiple rounds and respond with appropriate answers based on the user's sentences. For example:

[0109] [User]: What is the reason for frequent cold hands and feet and insomnia and nightmares?

[0110] [Chatbot]: According to traditional Chinese medicine, cold hands and feet are related to insufficient Yang energy.

[0111] [User]: What Chinese medicines can I take for conditioning?

[0112] [Chatbot]: You can take some warming and tonic Chinese medicines, such as cinnamon, aconite, evodia rutaecarpa, cloves, galangal, etc.

[0113] [User]: Any dietary recommendations?

[0114] [Chatbot]: Warm foods such as lamb, beef, chicken, as well as chestnuts, wolfberries, yam, and dates can improve the symptoms of cold hands and feet.

[0115] [User]: What should we pay attention to in our daily life?

[0116] [Chatbot]: Warm foods such as lamb, beef, chicken, as well as chestnuts, wolfberries, yam, and dates can improve the symptoms of cold hands and feet.

[0117] like Figure 2 The model training framework diagram shown in the figure has the following steps:

[0118] Step 1: Extract the medical knowledge of traditional Chinese medicine from ancient and modern Chinese medicine literature such as "Compendium of Materia Medica" and "Treatise on Febrile and Miscellaneous Diseases" to ensure the accuracy and standardization of the corpus package.

[0119] Step 2: Use the already trained large language model Tromsformer, and adjust it to let the large language model learn professional knowledge of traditional Chinese medicine, so as to realize question and answer of traditional Chinese medicine knowledge.

[0120] Step 3: Through simple tests, input questions and have the robot answer them to verify whether the TCM robot has learned TCM knowledge and whether the answers are accurate.

[0121] The following is a comparison of the answers given by the present invention and other large language models such as Wenxinyiyan and Chatgpt:

[0122] Table 1 Comparison of answers to the question “What is the cause of angular cheilitis?”

[0123]

[0124] Table 2 Comparison of answers to the question “What are the benefits of purslane?”

[0125]

[0126]

[0127] Table 3 Comparison of answers to the question “How to treat insomnia and dreaminess?”

[0128]

[0129] According to the comparison, it can be seen that the answer of the present invention is more concise and accurate, without too much redundancy, and is more inclined to the knowledge in the field of traditional Chinese medicine.

[0130] In the above embodiments, the present invention provides a method and device for a traditional Chinese medicine chat robot based on a pre-trained language model. Through the traditional Chinese medicine question-and-answer robot, the present invention can enable users to easily acquire knowledge about traditional Chinese medicine theory, diagnostic methods, treatment techniques, etc., thereby promoting the dissemination and popularization of traditional Chinese medicine knowledge.

[0131] The above-described embodiments are merely descriptions of preferred implementations of the present invention, and are not intended to limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various modifications and improvements made by ordinary persons in the art to the technical solution of the present invention should fall within the protection scope of the present invention, and the technical contents for which protection is sought in the present invention have been fully recorded in the claims.

Claims

1. A method for a traditional Chinese medicine chat robot based on a pre-trained language model, characterized in that: The following steps are involved: S1. Build relevant TCM datasets; S2, training the pre-trained model based on the traditional Chinese medicine dataset; The specific steps of step S2 are as follows: S2.

1. Use the data loading tool to pass the TCM dataset organized into question-answer pairs into the pre-training model; S2.

2. Set the parameters required for model training; S2.3, optimize model parameters through loss function; S2.

4. Set the random seed to ensure that the model effect can be reproduced under the same device and the repeatability of the experimental results; S2.5, load the model parameters that have been successfully pre-trained and adjust them according to the new data set to meet the needs of specific tasks; S2.

6. Load the data required for training and perform mask operations on the target sentences in the original text dataset; record and convert the relevant information into a tensor form that can be used by the model; S2.

7. Set the parameters of the Adamw optimizer and start model training. Divide the original text dataset into 32 batches, traverse each batch of data, and calculate its loss value. S2.8, the loss value is transmitted back, and the model parameters are updated through the optimizer, the gradient is calculated through the back propagation algorithm, and the model parameters are updated through the optimizer; S2.

9. After each batch of data is trained, save the model; The overall architecture of the model trained in step S2 can be expressed as the following formula: The Encoder i is the i-th Transformer Encoder layer, and They represent the output tensors of the i-th and i-1-th encoding layers respectively; Each encoder layer contains two sub-layers: multi-head self-attention and feed-forward neural network (FFN). The outputs of these two sub-layers are connected through residuals and input into a layer normalization layer, which is specifically expressed as follows: y = text{LayerNorm}(x + text{SubLayer}(x)) (2) Among them, y represents the output tensor, x represents the input tensor, text{LayerNorm} represents the layer normalization layer, and text{SubLayer} represents the sublayer; The FFN layer is expressed as follows: FFN(x)=max(0,xW1+b1)W2+b2 (3) Among them, the input tensor of the FFN layer is x, W1 and W2 in the FFN layer are the weight parameters of the first and second hidden linear layers respectively, and b1 and b2 represent the bias parameters of the first and second hidden linear layers respectively; S3: Process the current user's input information based on the trained model, obtain the corresponding output vector, splice and generate content, and realize model reasoning; S4. Display the inference results in the form of a web page, interact with users, and give Chinese medicine advice.

2. The method of a Chinese medicine chat robot based on a pre-trained language model according to claim 1, characterized in that: The specific steps of step S1 are as follows: S1.1, traverse the collected raw data files and make judgments, dividing the data into two categories: questions and answers; S1.

2. Remove incomplete questions and simple answers, use existing sensitive word identification tools to detect data, and discard any sensitive words; S1.3, remove HTML tags, repeated punctuation and tag symbols in questions and answers; if the length of the cleaned data is less than a preset threshold, it is discarded; where HTML stands for Hypertext Markup Language; S1.

4. Save the cleaned data as a TCM dataset in the form of question-answer pairs in the format required for model training.

3. The method of a Chinese medicine chat robot based on a pre-trained language model according to claim 1, characterized in that: The step S2.3 performs model parameter optimization through the loss function and comprises the following steps: S2.3.

1. Set the model initialization function and define the components required for model training; S2.3.

2. Load the weights and configure the forward propagation function of the model, which includes passing the input data through the encoder part and obtaining the sequence output, whose dimension is [batch_size, sequence_length, hidden_size]; Among them, batch_size is the batch size, sequence_length is the sequence length, and hidden_size is the number of hidden memory units; S2.3.

3. Obtain the vector of the masked position, calculate the loss value of the word or words replaced with the special token, and use the calculated loss value in the optimization process of the model parameters.

4. The method of a Chinese medicine chat robot based on a pre-trained language model according to claim 1, characterized in that: The specific steps of step S3 are as follows: S3.1, set the parameters required for model inference; S3.

2. Load the trained model and place it on the corresponding device; S3.3, load the sensitive word filter to ensure that the output does not contain any sensitive information; S3.

4. Detect the content input by the user. If it contains sensitive words, directly reply with a specific sentence, otherwise proceed to the next step; S3.5, vectorize the sentence input by the user and obtain the corresponding sequence; S3.6, traverse according to the maximum length limit of the output sequence of the model, obtain the output vector according to the current input, and select the words with high and low probability in the vocabulary; S3.

7. When the "[SEP]" mark appears in the generated content, it indicates that the generation is completed. All the generated content is spliced ​​and integrated, and it is checked whether sensitive words appear. If sensitive words are found, a specific sentence is replied; Otherwise reply with the content generated by the model.

5. A device comprising the method of a Chinese medicine chat robot based on a pre-trained language model as described in any one of claims 1 to 4, characterized in that: include: Data collection and preprocessing module: collect original corpus data of TCM consultation dialogues and TCM knowledge by reading ancient and modern documents; Preprocess the data, clean and organize the raw data, traverse the content to divide it into questions and answers, organize it into the required data format, and build relevant TCM data sets; Model training module: pass the data set organized into question-answer pairs into the pre-trained model, set appropriate parameters, optimize the model parameters through the loss function, load the required data, start the training task, and calculate the loss value in real time to update the model parameters according to the optimization algorithm; Model inference module: loads the trained model, configures sensitive word filters to process user input, obtains the corresponding output vector by processing the current user input, splices and generates content, and implements model inference, that is, generates the model's response or reply to the user input; Interaction module: The inference results are displayed in the form of a web page. Multiple rounds of interaction can be carried out with users to accurately answer user questions and provide TCM advice. Among them, the output end of the data collection and preprocessing module is connected to the input end of the model training module, the output end of the model training module is connected to the input end of the model reasoning module, and the output end of the model reasoning module is connected to the input end of the interaction module.

Citation Information

Patent Citations

  • Medical question and answer reply method and system based on doctor feedback and reinforcement learning

    CN116383364A

  • Medical question-answering system based on improved named entity recognition and construction method thereof

    CN116719913A