Elderly health management instruction data set construction method based on large language model
By combining the information extraction ability and voting method of the large language model, the problem of lack of field specialization and data quality control of the existing technology middle-aged and elderly health management instruction data sets is solved, and a high-quality elderly health management instruction data set is built to support the accurate data provision of intelligent question-and-answer questions and answers for elderly health management.
Patent Information
- Application Number
- CN202510067758.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
When constructing the data set of elderly health management instructions, the existing large language model lacks domain specialization and data quality control, resulting in the generated data not meeting the needs of the domain and insufficient accuracy and relevance.
An information extraction method based on a large language model is adopted, combined with semantic similarity judgment and voting method, an elderly health management instruction data set is constructed. Specific steps include data crawling, cleaning, information extraction, semantic similarity judgment, manual evaluation and multiple indicator voting screening based on prior knowledge.
A high-quality, field-specific data set of elderly health management instructions has been built, which improves the accuracy and relevance of the data and supports the accurate data provision of intelligent questions and answers for elderly health management.
Smart Images

Figure CN120067420A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to technologies such as large language models, information extraction, machine learning, and ensemble learning. Specifically, it relates to a method for constructing an elderly health management instruction dataset based on large language model information extraction and voting method. Background Art
[0002] With the intensification of population aging, the health management needs of the elderly population have become one of the core issues in the field of public health. The health challenges faced by the elderly population include the long-term management of multiple chronic diseases, complication prevention, and mental health maintenance, which not only affect the quality of personal life but also have a profound impact on families and society. The elderly's demand for obtaining health information is increasing day by day. However, due to the imbalance between the supply and demand of traditional medical resources and the limited channels for the elderly to obtain information, at the same time, false health information is more likely to mislead the elderly population, resulting in adverse consequences. Therefore, elderly health management faces many challenges. The rapid development of large models provides a new opportunity for the elderly health management Q&A system. By providing personalized and convenient consultation and recommendation services, it can help the elderly and their caregivers obtain relevant health knowledge in a timely manner and optimize the health management process. However, due to the lack of a dedicated instruction dataset for elderly health management, it is difficult for large models to capture effective keywords during training and inference, resulting in the output information not conforming to the actual situation and even generating hallucinated information, affecting the accuracy and interpretability of online consultation answers. Therefore, how to design a high-quality instruction dataset focusing on the elderly population and efficiently train the model to provide scientific guidance for elderly health management is a current necessity in the aging society.
[0003] Large language models such as GPT-4 have attracted wide attention due to their excellent conversation and generation capabilities. They can not only understand complex language structures and grasp subtle meanings, interact with users naturally and smoothly, but also demonstrate powerful information extraction capabilities in natural language processing tasks, and can generate coherent and highly creative texts. They are widely used in fields such as text understanding, knowledge mining, and question answering generation, providing important support for the construction of instruction data. However, existing research still faces some challenges when applying large language models to construct instruction data. First, the domain specialization degree of instruction data generation is insufficient. Most existing models are trained based on general domain data and are difficult to meet the precise needs of specific domains (such as elderly health management). Second, the data quality is uneven, and there is still room for improvement in the logical consistency and semantic accuracy of automatically generated data pairs. Therefore, how to combine domain knowledge with the extraction capabilities of large language models to construct high-quality, domain-specific instruction datasets has become a hot topic and a difficult point in current research.
[0004] In view of the above problems, this patent proposes a method for constructing a dataset of elderly health management instructions based on information extraction from large language models and the voting method. This method uses the information extraction method based on large language models to generate data that meets the domain requirements by leveraging its efficient understanding and parsing capabilities; uses semantic similarity judgment to filter out data with large semantic differences, poor accuracy, and low domain relevance; after manually evaluating 10,000 pieces of data, trains a classification model based on the voting method and prior knowledge, and conducts multi-index voting evaluation on the remaining data to further screen and form a final high-quality dataset of elderly health management instructions, providing accurate data support for optimizing intelligent question answering for elderly health management. Summary of the Invention
[0005] The technical problem solved by the present invention is: using the powerful information extraction ability of large language models to initially generate data pairs that conform to the domain characteristics, and using semantic similarity, manual evaluation, and the voting method based on prior knowledge to further screen and filter the data to construct a dataset of elderly health management instructions.
[0006] The technical solution of the present invention is: The present invention proposes a method for constructing a dataset of elderly health management instructions based on information extraction from large language models and the voting method. This method first obtains high-quality, elderly health management Q&A data and unsupervised data (i.e., unstructured text data) through data cleaning and filtering based on multiple data acquisition methods such as web crawlers; in the information extraction stage of large language models, design prompt words to knowledge-guide instruction data related to background knowledge; prompt the large language model to judge the semantic similarity between the Q&A data formed by information extraction and the answers generated by different base models, and delete content with large semantic differences, poor accuracy, and low domain relevance; after manually evaluating 10,000 pieces of data, train a classification model based on the voting method supported by the evaluation data and prior knowledge, conduct multi-index voting on the remaining instruction data, and screen out data that meets the high-quality standards to construct a dataset of elderly health management instructions with high accuracy and domain relevance. The specific steps are as follows:
[0007] (1) Crawl authoritative websites to obtain elderly health management Q&A data and unsupervised data (i.e., unstructured text data), and store the data. The steps include:
[0008] a. Data crawling. Based on the web crawler strategy, browse and crawl websites such as DXY (www.dxy.com), XYWY (www.xywy.com), and the Chinese government website (www.gov.cn) to obtain elderly health management Q&A data and unsupervised data (i.e., unstructured text data), including: authoritative data such as professional knowledge of elderly diseases, common sense questions, health care, prevention and nursing, and medical insurance service policies;
[0009] b. Data storage. The Q&A data pairs are directly stored in JSON format; the obtained unsupervised data is
[0010] logically divided into each piece of data with "@@@@@" as the delimiter and stored as a txt file.
[0011] (2) Utilize the powerful information extraction ability of the large language model to guide the generation of Q&A data that meets the domain requirements. The specific steps include:
[0012] a. Select the base model ChatGLM, input the obtained unsupervised data as background knowledge, and use prompt engineering to design prompts to prompt the large language model to generate instruction questions with strong domain characteristics and close relationship with the background knowledge. The prompt design is as follows:
[0013] Table 1 Design of prompts for generating instruction questions
[0014]
[0015] b. Design answer generation prompts to prompt the large language model to generate corresponding answers based on the background knowledge and the generated instruction questions. The prompt design is as follows:
[0016] Table 2 Design of prompts for generating answers
[0017]
[0018]
[0019]
[0020] K represents unsupervised knowledge, Q represents the generated questions, and A represents the answers generated by the model's reading comprehension.
[0021] (3) Perform data filtering and cleaning based on regular expressions. The specific steps include:
[0022] a. Filter and remove duplicate data;
[0023] b. Based on regular expressions, eliminate data that violates the instruction constraints, has incomplete or empty answers, or has incorrect formats.
[0024] (4) Screen instruction data based on semantic similarity evaluation. The specific steps include:
[0025] a. Call other base models and prompt the models to generate different answers based on the background knowledge and the instruction questions;
[0027] b. Design prompt words for semantic similarity judgment to prompt the large language model to perform semantic similarity judgment on different answers, and delete instructions with large semantic differences, poor accuracy, and low domain relevance.
[0028] Data. The design of the prompt words for semantic similarity judgment is as follows:
[0029] Table 3 Design of Prompt Words for Semantic Similarity Judgment
[0030]
[0031]
[0032] (5) Manually evaluate pairs of instruction data. The specific steps include:
[0033] a. Based on the filtered dataset above, randomly select 10,000 pieces of data for manual annotation. Among them, the label value of the instruction data closely related to elderly health management is set to 0, and the label value of the data with low or no relevance is set to 1.
[0034] (6) Based on the data manually evaluated, train a classification model to screen and evaluate the remaining instruction data to form the final elderly health management instruction dataset. The specific steps include:
[0035] a. Convert the data into the DataFrame (DF) format
[0036] b. Use TF-IDF for feature extraction. Among them, the term frequency (TF) represents the probability that a keyword appears in the text, and the inverse document frequency (IDF) is obtained by dividing the total number of documents by the number of documents containing the word, and then taking the logarithm of the resulting quotient. The specific formula is as follows:
[0037] TF-IDF = TF * IDF
[0038] Among them:
[0039] That is:
[0040] n i,j is the number of times the word appears in the document, and the denominator is the total number of times all words appear in the document.
[0041] That is:
[0042] |D| represents the total number of documents in the corpus. |{j:t i ∈d j}| represents the number of documents containing the word t i (which is n i,j(number of non - zero files). If the term is not in the corpus, it causes the denominator to be 0. Therefore, generally, 1 + |{j:t i ∈d j}| is used.
[0043] c. According to the manual evaluation criteria, define an expert rule model and implement the predict_proba method required by the SoftVotingClassifier in the model. Soft voting takes the average of the probabilities that all model prediction samples belong to a certain category as the criterion, and the corresponding category with the highest probability is the final prediction result. The predict_proba method is used to calculate the probability that each sample belongs to each category and returns a two - dimensional array, where each row represents a sample and each column represents a category, and the value of the array represents the probability that the sample belongs to the corresponding category;
[0044] d. Define the parameters of four models: random forest, SVC, decision tree, and XGBoost. Use the grid optimization algorithm to search for the optimal parameters during the model training process, and use the optimal parameters to replace the original model after K - fold cross - validation (3 - fold cross - validation is used in the present invention). The K - fold cross - validation formula is as follows:
[0045]
[0046] where M is the model, D train,i is the training set for the i - th iteration, D test,i is the test set for the i - th iteration, and Acc is the accuracy performance metric.
[0047] e. Add the expert rule model and train the soft voting classifier;
[0048] f. Iteratively train and optimize the model, and encapsulate and save the trained optimal model;
[0049] g. Apply the encapsulated model to classify unevaluated data, filter and remove the instruction data pairs with label value 1, and form the final elderly health management instruction dataset.
[0050] The advantages of the present invention compared with the prior art are as follows:
[0051] 1. It provides a method for constructing an elderly health management instruction dataset based on large - language model information extraction and voting method. Compared with the existing technologies, the instruction data constructed by this method is targeted and specialized, and better meets the needs of specific fields.
[0052] 2. Innovatively combines information extraction with large - language models. Based on the information extraction method of large - language models, design prompt words to guide knowledge for unstructured elderly health management texts, effectively improving the quality and domain relevance of instruction data generation.
[0053] 3. Use regular expressions and semantic similarity judgment mechanisms to eliminate Q&A content with large semantic differences, poor accuracy, and low domain relevance, and optimize data quality in multiple dimensions.
[0054] 4. Based on manual evaluation and prior knowledge, combined with the voting method, perform multiple-index screening on the data. Through the training and evaluation of the classification model, ensure that the data meets the standards of high accuracy and domain relevance, overcoming the limitations of traditional single evaluation methods.
[0055] 5. Compared with traditional methods, the present invention combines automated information extraction and the voting method, greatly reducing the workload of manual intervention, and at the same time significantly improving the efficiency and quality of dataset construction.
[0056] The present invention has carried out technological innovations in various aspects such as data acquisition, information extraction, and data screening, significantly improving the construction efficiency and quality of the elderly health management instruction dataset, and providing strong support for promoting the intelligent development of the elderly health management field. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In conjunction with the accompanying drawings, the present invention will be better understood from the following detailed description of the embodiments of the present invention. Figure 1 It is the overall step flow chart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0058] The present invention includes the following steps:
[0059] (1) Crawl authoritative websites to obtain elderly health management Q&A data and unsupervised data (i.e., unstructured text data), and store the data. The steps include:
[0060] a. Data crawling. Based on the web crawler strategy, browse and crawl websites such as Dingxiang Doctor Network (www.dxy.com), XYwy Network (www.xywy.com), and the Chinese government website (www.gov.cn) to obtain elderly health management Q&A data and unsupervised data (i.e., unstructured text data), including: authoritative data such as professional knowledge of elderly diseases, common sense questions, health care, prevention and nursing, and medical insurance service policies;
[0061] b. Data storage. The Q&A data pairs are directly stored in json format; the obtained unsupervised data is
[0062] logically divided into each piece of data with "@@@@@" as the delimiter and stored as a txt file.
[0063] (2) Utilize the powerful information extraction ability of the large language model to generate Q&A data that meets the domain requirements under the guidance of knowledge. The specific steps include:
[0064] a. Select the base model ChatGLM, input the obtained unsupervised data as background knowledge, and use prompt engineering to design prompts to prompt the large language model to generate instruction questions with strong domain characteristics and close relevance to the background knowledge. The prompt design is as follows:
[0065] Table 1 Design of Prompts for Generating Instruction Questions
[0066]
[0067]
[0068] b. Design prompts for answer generation, prompting the large language model to generate corresponding answers based on the background knowledge and the generated instructions
[0069] The prompt design is as follows:
[0070] Table 2 Design of Prompts for Generating Answers
[0071]
[0072]
[0073] K represents unsupervised knowledge, Q represents the questions generated by the instructions, and A represents the answers generated by the model's reading comprehension.
[0074] (3) Perform data filtering and cleaning based on regular expressions. The specific steps include:
[0075] a. Filter and remove duplicate data;
[0076] b. Based on regular expressions, eliminate data that violates the instruction constraints, has incomplete or empty answers, or incorrect formats.
[0077] (4) Screen instruction data based on semantic similarity evaluation. The specific steps include:
[0078] a. Call other base models, prompting the model to generate different answers based on the background knowledge and the instruction questions;
[0080] b. Design prompts for semantic similarity judgment, prompting the large language model to judge the semantic similarity of different answers, and delete instruction data with large semantic differences, poor accuracy, and low domain relevance. The prompt design for semantic similarity judgment is as follows:
[0081] Table 3 Design of Prompts for Semantic Similarity Judgment
[0082]
[0083] (5) Manually evaluate pairs of instruction data. The specific steps include:
[0084] a. Based on the above filtered dataset, randomly select 10,000 pieces of data and manually annotate them. Set the label value of the instruction data closely related to elderly health management to 0, and set the label value of the data with low or no relevance to 1.
[0085] (6) Based on the data evaluated manually, train a classification model to screen and evaluate the remaining instruction data, and form the final elderly health management instruction dataset. The specific steps include:
[0086] a. Convert the data into the DataFrame (DF) format
[0087] b. Use TF-IDF for feature extraction. Among them, the term frequency (TF) represents the probability that a keyword appears in the text, and the inverse document frequency (IDF) is obtained by dividing the total number of documents by the number of documents containing the word, and then taking the logarithm of the obtained quotient. The specific formula is as follows:
[0088] TF-IDF = TF * IDF
[0089] Among them:
[0090] That is:
[0091] n i,j is the number of times the word appears in the document, and the denominator is the total number of times all words appear in the document.
[0092] That is:
[0093] |D| represents the total number of documents in the corpus. |{hj:t i ∈d j}| represents the number of documents containing the word t i (that is, the number of documents where n i,j ≠0). If the denominator is 0 because the word is not in the corpus, so generally 1 + |{j:t i ∈d j}| is used.
[0094] c. According to the manual evaluation criteria, define an expert rule model and implement the predict_proba method required by the SoftVotingClassifier in the model. Soft voting takes the average of the probabilities of all model predictions for a certain category of samples as the standard, and the category corresponding to the highest probability is the final prediction result. The predict_proba method is used to calculate the probability of each sample belonging to each category and returns a two-dimensional array, where each row represents a sample and each column represents a category, and the value of the array represents the probability of the sample belonging to the corresponding category;
[0095] d. Define the parameters of four models: random forest, SVC, decision tree, and XGBoost. Use the grid optimization algorithm to search for the optimal parameters during the model training process, and replace the original model with the optimal parameters after K-fold cross-validation (in this invention, 3-fold cross-validation is adopted). The formula for K-fold cross-validation is as follows:
[0096]
[0097] where M is the model, D train,i is the training set for the i-th iteration, and D test,i is the test set for the i-th iteration, and Acc is the accuracy performance metric.
[0098] e. Add the expert rule model and train the soft voting classifier;
[0099] f. Iteratively train and optimize the model, and encapsulate and save the trained optimal model;
[0100] g. Apply the encapsulated model to classify the unevaluated data, filter and clear the instruction data pairs with the label value of 1, and form the final elderly health management instruction dataset.
[0101] The above has introduced the embodiments of the present invention in detail. In this article, specific implementation manners are used to elaborate on the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for constructing an elderly health management instruction dataset based on a large language model, characterized in that Here are the steps: Step (1), crawl authoritative websites, obtain elderly health management question and answer data and unsupervised data, and store the data; Step (2), using the powerful information extraction capability of the large language model, knowledge guides the generation of question-answering data that meets domain requirements; Step (3), filtering and cleaning data based on regular expressions; Step (4), filtering instruction data based on semantic similarity evaluation; Step (5), manually evaluating the instruction data pair; Step (6), based on the manually evaluated data, the classification model is trained to screen and evaluate the remaining instruction data to form the final elderly health management instruction data set.
2. The method for constructing an elderly health management instruction dataset based on a large language model according to claim 1, characterized in that: In the step (1), crawling the web and performing data processing includes: (1) Crawl websites to obtain elderly health management question-and-answer data and unsupervised data, including: professional knowledge of elderly diseases, common sense questions, health care, preventive care, and medical insurance service policies; (2) The crawled question-answer data pairs are stored in json format, and the unsupervised data are divided into each data with the "@@@@@" delimiter and stored in txt files.
3. The method for constructing an elderly health management instruction dataset based on a large language model according to claim 1, characterized in that: In step (2), the large language model information extraction and knowledge-guided data generation include: (1) Select a base model, use the acquired unsupervised data as background knowledge input, and use prompt engineering to design prompt words to prompt the large language model to generate instruction problems that have strong domain characteristics and are closely related to background knowledge; (2) Design answer generation prompts to prompt the large language model to generate corresponding answers based on background knowledge and the generated instruction questions.
4. The method for constructing an elderly health management instruction dataset based on a large language model according to claim 1, characterized in that: In the step (3), Data filtering and cleaning based on regular expressions include: (1) Filter and remove duplicate data; (2) Based on regular expressions, eliminate data that violates indicated constraints, has incomplete or empty answers, or has incorrect formats.
5. The method for constructing an elderly health management instruction dataset based on a large language model according to claim 1, characterized in that: In the step (4), screening instruction data based on semantic similarity evaluation includes: (1) Call other base models to prompt them to generate different answers based on background knowledge and instruction questions; (2) Design semantic similarity judgment prompts to prompt the large language model to make semantic similarity judgments on different answers and delete instruction data with large semantic differences, poor accuracy, and low domain relevance.
6. The method for constructing an elderly health management instruction dataset based on a large language model according to claim 1, characterized in that: In the step (5), the manual evaluation instruction data pair includes: (1) Based on the above filtered data set, 10,000 data were randomly selected and manually labeled, where the label value of instruction data closely related to elderly health management was set to 0, and the label value of data with low correlation or irrelevant was set to 1.
7. The method for constructing an elderly health management instruction dataset based on a large language model according to claim 1, characterized in that: In step (6), the training data classification and screening model includes: (1) Data is converted into DataFrame format. DataFrame format is a data structure used to store and manipulate tabular data. It is organized in rows and columns, making it easy to perform various data analysis and processing tasks. (2) Perform feature extraction; (3) Based on the manual evaluation criteria, define the expert rule model and implement the predict_proba method required for soft voting in the model; the predict_proba method is used to calculate the probability that each sample belongs to each category and returns a two-dimensional array, where each row represents a sample and each column represents a category. The value of the array represents the probability that the sample belongs to the corresponding category; (4) Define the parameters of four models: random forest, SVC, decision tree, and XGBoost. Use the grid optimization algorithm to search for the optimal parameters during model training and use the optimal parameters to replace the original model. (5) Add the expert rule model and train the soft voting classifier; (6) Iterate and optimize the model, and encapsulate and save the trained optimal model; (7) Apply the encapsulated model to classify the un-evaluated data, filter out the instruction data pairs with label value 1, and form the final elderly health management instruction data set.
Citation Information
Cited By
Data enhancement method, device and system based on combination of three-dimensional scene and language data
CN120671074A
Data augmentation method, device and system based on combination of three-dimensional scene and language data
CN120671074B
Digital entertainment interaction method and device based on intelligent agent and storage medium
CN120932640A
Adaptive instruction induction method, system and equipment based on large language model and storage medium
CN121072788A