Intelligent Question-Answering Processing Method and Device for the Elderly, Storage Medium, Computer Equipment

An intelligent question answering system integrates and transforms diverse health knowledge sources into a multi-source database, addressing resource limitations and terminology barriers to enhance elderly users' health consultation experience.

CN119739836BActive Publication Date: 2025-07-15HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510216473.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-15
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

Traditional health consultation methods have limited medical resources for the elderly, poor consultation channels, and it is difficult to meet the health knowledge needs of the elderly. In addition, the elderly’s questions are colloquial and there are obstacles to understanding professional terms.

Method used

Integrate health knowledge of the elderly scattered in different sources and forms, build a multi-source database, conduct intelligent Q&A through multi-source databases, and generate colloquial answers.

Benefits of technology

It improves the user experience of the elderly, reduces the barriers to professional terminology, and provides more accurate and convenient health knowledge Q&A services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739836B_ABST
    Figure CN119739836B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology, and discloses a method and device for processing intelligent questions and answers for the elderly, a storage medium, and a computer device. The method includes: converting the original medical field documents into various forms of text to construct a multi-source database, where the forms of text include term and explanation form text, question and answer pair form text, and hierarchical summary form text, and the text content in each form of text is colloquial language; receiving a medical question raised by an elderly user, and obtaining a question sentence vector based on the medical question; generating an answer statement corresponding to the medical question based on the multi-source database and the question sentence vector, and returning the rewritten answer statement after performing colloquial rewriting on the answer statement. By effectively integrating the health knowledge of the elderly scattered in different sources and forms and making colloquial improvements, a multi-source database is formed. Through the multi-source database for intelligent question and answer, it can effectively handle colloquial questions and the obstacle of understanding professional terms, and improve the user experience of the elderly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and device for intelligent question-answering processing for the elderly, a storage medium, and a computer device. Background Art

[0002] As the global population continues to age, the health of the elderly has become a focus of widespread concern in all sectors of society. The health needs of the elderly are growing, and the demand for health consultation and knowledge answers is becoming more urgent. However, traditional health consultation methods often have many inconveniences, such as limited medical resources and poor consultation channels, which make it difficult to meet the actual needs of the elderly. Summary of the invention

[0003] In view of this, the present application provides a method and device for intelligent question and answer processing for the elderly, a storage medium, and a computer device. By effectively integrating the health knowledge of the elderly scattered in different sources and forms and improving it in colloquial terms, a multi-source database is formed. Intelligent question and answering through the multi-source database can effectively deal with oral questions and obstacles to understanding professional terms, thereby improving the user experience of the elderly.

[0004] According to one aspect of the present application, a method for processing intelligent questions and answers for the elderly is provided, the method comprising:

[0005] Collect multiple original medical documents containing health knowledge about the elderly;

[0006] For any original medical field document, convert the original medical field document into multiple forms of text, wherein the form text includes terminology and explanation form text, question-answer form text and hierarchical summary form text, and the text content in each form text is colloquial language;

[0007] Construct a multi-source database based on multiple forms of text corresponding to multiple original medical documents;

[0008] receiving a medical question raised by an elderly user, and obtaining a question sentence vector based on the medical question;

[0009] Based on the multi-source database and the question sentence vector, an answer sentence corresponding to the medical question is generated, and the answer sentence is rewritten in colloquial language and the rewritten answer sentence is returned.

[0010] According to another aspect of the present application, a device for intelligent question-answering processing for the elderly is provided, the device comprising:

[0011] The medical document collection module is used to collect multiple original medical documents containing health knowledge for the elderly;

[0012] A formal text conversion module is used to convert any original medical field document into multiple forms of text. Among them, the forms of text include term and explanation form text, Q&A pair form text, and hierarchical summary form text, and the text content in each form of text is colloquial language;

[0013] A multi-source database construction module is used to construct a multi-source database based on multiple forms of text corresponding to multiple original medical field documents;

[0014] A medical problem processing module is used to receive a medical problem proposed by an elderly user and obtain a problem sentence vector based on the medical problem;

[0015] An answer statement generation and return module is used to generate an answer statement corresponding to the medical problem based on the multi-source database and the problem sentence vector, and return the rewritten answer statement after performing colloquial rewriting on the answer statement.

[0016] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the above-mentioned intelligent Q&A processing method for the elderly is implemented.

[0017] According to still another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the above-mentioned intelligent Q&A processing method for the elderly is implemented.

[0018] By means of the above technical solutions, an intelligent Q&A processing method and device, storage medium, and computer device for the elderly provided by the present application can effectively integrate the health knowledge of the elderly scattered in different sources and forms and perform colloquial improvement to form a multi-source database. Through the multi-source database for intelligent Q&A, it can effectively handle colloquial questions and the obstacle of understanding professional terms, and improve the user experience of the elderly.

[0019] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. Brief Description of the Drawings

[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0021] Figure 1 A flowchart showing an intelligent Q&A processing method for the elderly provided by an embodiment of the present application is shown;

[0022] Figure 2 shows a schematic flowchart of a medical field document collection method provided by an embodiment of the present application;

[0023] Figure 3 shows a schematic flowchart of a multi-source database construction method provided by an embodiment of the present application;

[0024] Figure 4 shows a schematic flowchart of a form text conversion method provided by an embodiment of the present application;

[0025] Figure 5 shows a schematic flowchart of a medical problem processing method provided by an embodiment of the present application;

[0026] Figure 6 shows a schematic flowchart of an answer statement generation method provided by an embodiment of the present application;

[0027] Figure 7 shows a schematic flowchart of another answer statement generation method provided by an embodiment of the present application;

[0028] Figure 8 shows a schematic structural diagram of an intelligent question and answer processing device for the elderly provided by an embodiment of the present application. Detailed implementation manners

[0029] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0030] In this embodiment, an intelligent question and answer processing method for the elderly is provided. As Figure 1 shown, the method includes:

[0031] Step 101, collecting a plurality of original medical field documents containing health knowledge of the elderly.

[0032] Currently, when the elderly conduct medical question and answer through the elderly health knowledge Q&A robot, they face a series of challenges. (The elderly health knowledge Q&A robot is a new type of intelligent auxiliary tool that can provide health consultation and knowledge answering services for the elderly), mainly including:

[0033] Knowledge fragmentation: Knowledge related to the health of the elderly is scattered in various sources such as e-books, official account articles, blogs, etc. These data have diverse forms and inconsistent structures, resulting in difficulties in knowledge integration.

[0034] User questions are colloquial: As the main user group, the questions asked by the elderly tend to be colloquial and life-oriented, showing significant differences in expression styles compared to written professional literature.

[0035] Barriers of professional terms: The elderly generally lack knowledge of professional medical terms, and a large number of professional nouns are filled in the existing health knowledge materials, which increases the difficulty for the elderly to understand and use these resources.

[0036] Limitations of the retrieval system: The answers that exactly match the questions may not necessarily exist in the databases of traditional retrieval systems, and the answers are not natural. The answers of large language models are natural, but they are prone to incorrect answers. For example, for the RAG+ (Retrieval-Augmented Generation) large language model, the recall rate and accuracy rate are relatively low.

[0037] In the above embodiments of the present application, it can be applied to the health knowledge Q&A robot for the elderly. Through technical means such as offline corpus collation by the large model, structured storage, multi-source matching, and large model re-ranking, the data utilization rate can be improved, the answer coverage and accuracy can be enhanced, so as to provide more accurate and convenient health knowledge Q&A services for the elderly. Specifically, based on different sources (e-books, official accounts, blogs, etc.), multiple original medical field documents containing the health knowledge of the elderly are collected to prepare for the establishment of a multi-source database in the future.

[0038] Optionally, in step 101 of collecting multiple original medical field documents containing the health knowledge of the elderly, with reference to Figure 2 as shown, it specifically includes:

[0039] Step 1011, collecting multiple original candidate medical field documents containing the health knowledge of the elderly.

[0040] Step 1012, for any original candidate medical field document, according to the preset evaluation index and the index weight corresponding to the preset evaluation index, obtain the evaluation value of the original candidate medical field document, and obtain the original medical field document corresponding to the evaluation value greater than the preset evaluation threshold, where the preset evaluation index includes the document collection source, the document release time, and the authority of the document source.

[0041] In the above embodiments of the present application, original candidate medical field documents containing health knowledge of the elderly are widely collected through various channels. These channels may include: academic databases such as PubMed and CNKI, which contain a large number of medical papers and research reports; government and medical institution websites such as the World Health Organization and the official websites of major hospitals, which publish authoritative health guidelines and policy documents; professional medical websites and forums such as Medscape and Dingxiangyuan, which gather exchanges and discussions among many medical professionals and patients. It can also be collected from channels such as e-books, official accounts, and blogs to expand the collection scope and enrich the original data.

[0042] Next, evaluation indicators are set. The preset evaluation indicators are used to evaluate the quality and reliability of the original candidate medical field documents. According to actual needs, the following evaluation indicators and indicator weights can be set: document collection source (30%), document release time (20%). Considering the timeliness of the document, newer documents often contain more cutting-edge and accurate information. Document source authority (50%), evaluating the authority and credibility of the document release institution. Documents published by government agencies, well-known medical institutions, etc. are usually more authoritative.

[0043] For each original candidate medical field document, a value is obtained by calculating according to the preset evaluation indicators and the corresponding indicator weights. The specific steps are as follows:

[0044] Score the document collection source: For example, papers in academic databases get 30 points (full score), documents published on government and medical institution websites get 25 points, and documents on professional medical websites and forums get 20 points.

[0045] Score the document release time: Score according to the gap between the document release time and the current time. For example, those released within the past year get 20 points, those released within one to three years get 15 points, and those released more than three years ago get 10 points.

[0046] Score the document source authority: Documents published by government agencies get 50 points (full score), those published by well-known medical institutions get 45 points, authoritative papers in academic databases get 40 points, and other reliable sources get 35 points.

[0047] For example, a research paper on the management of diabetes in the elderly is collected. Assuming that the document collection source is an academic database, the publication time is within the past year, and the authority of the document source is an authoritative paper in the academic database, the evaluation value is calculated as follows: Evaluation value = 30 (document collection source) × 30% + 20 (publication time) × 20% + 40 (authority of document source) × 50% = 27 + 4 + 20 = 51 (points). A preset evaluation threshold is set, such as 50 points, and then the original medical field documents with evaluation values greater than the preset evaluation threshold are screened out. These documents are high-quality and reliable original medical field documents containing health knowledge of the elderly.

[0048] In the above example, the evaluation value of the research paper on the management of diabetes in the elderly is 51 points, which is higher than the preset evaluation threshold of 50 points. Therefore, the research paper on the management of diabetes in the elderly is selected as the original medical field document. Through this process, multiple high-quality original medical field documents containing health knowledge of the elderly can be effectively collected and screened, improving the subsequent answer matching accuracy.

[0049] Step 102, for any original medical field document, convert the original medical field document into multiple forms of text, where the form of text includes a form of text with terms and explanations, a form of text with question-and-answer pairs, and a form of text with hierarchical summaries, and the text content in each form of text is colloquial language.

[0050] Step 103, based on the multiple forms of text corresponding to each of the multiple original medical field documents, construct a multi-source database.

[0051] Next, as Figure 3 shown, for any original medical field document, preprocess the original medical field document, including extracting medical term entity words and their explanatory notes, questions (questions that can be asked) and their answers (answer sentences), and fragment content and summaries (summarized content) of the document, so as to convert the original medical field document into multiple forms of text, the form of text includes a form of text with terms and explanations, a form of text with question-and-answer pairs, and a form of text with hierarchical summaries, and construct a multi-source database based on the multiple forms of text corresponding to each of the multiple original medical field documents, in preparation for the subsequent generation of answer sentences.

[0052] Specifically, the text content in each formal text is colloquial language so that the elderly health knowledge Q&A robot can better "understand" the colloquial questions of the elderly and generate colloquial answers that the elderly can understand. Specifically, after obtaining the formal text, natural language processing techniques can be used to analyze and process the formal text. For example, tools such as word segmentation and part-of-speech tagging can be used to preprocess the text, and then the content in the formal text can be improved in terms of colloquialization by combining the rules and strategies of colloquial conversion. After the improvement is completed, manual review and correction can also be carried out to ensure that the improved text meets the requirements of colloquialization. It is also possible to learn how to better improve by searching for examples of colloquial conversion. These examples can come from resources such as daily conversations, colloquial textbooks, or colloquial writing guides. It should be noted that when performing colloquial conversion, appropriate adjustments need to be made according to the target audience and context.

[0053] Optionally, in step 102, for any original medical field document, when converting the original medical field document into multiple formal texts, with reference to Figure 4 as shown, specifically including:

[0054] Step 1021, for any original medical field document, perform word segmentation on the original medical field document, determine multiple medical term entity words among the segmented words, obtain the respective explanatory notes corresponding to each medical term entity word, construct term-explanation pairs based on the medical term entity words and the explanatory notes corresponding to the medical term entity words, after performing colloquial rewriting on the term-explanation pairs, based on the rewritten term-explanation pairs, obtain the term-and-explanation formal text corresponding to the original medical field document.

[0055] Step 1022, for any original medical field document, generate answer statements that can be used to answer in the original medical field document and the question statements that can be answered by the answer statements based on a preset medical field Q&A model, obtain question-and-answer pairs based on the question statements and the answer statements corresponding to the question statements, after performing colloquial rewriting on the question-and-answer pairs, based on the rewritten question-and-answer pairs, obtain the question-and-answer pair formal text corresponding to the original medical field document, where the preset medical field Q&A model is established based on LLaMA-13B.

[0056] Step 1023, for any original medical field document, segment the original medical field document into multiple fragment contents based on a text segmentation algorithm, respectively summarize each fragment content based on a text summarization algorithm to obtain the respective summary content corresponding to each fragment content, obtain fragment-summary content pairs based on the fragment content and the summary content corresponding to the fragment content, after performing colloquial rewriting on the fragment-summary content pairs, based on the rewritten fragment-summary content pairs, obtain the hierarchical summary formal text corresponding to the original medical field document.

[0057] In the above embodiments of the present application, specifically, the formal text includes term and explanation formal text, Q&A pair formal text, and hierarchical summary formal text.

[0058] For the term and explanation formal text, the text content included in the original medical field document is, for example: "Heart disease is a common circulatory system disease, mainly manifested as symptoms such as palpitations and chest pain." The original medical field document is segmented into words: "Heart disease, is, a, common, of, circulatory, system, disease, mainly, manifested, as, palpitations, chest pain, etc., symptoms." Among the segmented words, multiple medical term entity words are determined, such as: "Heart disease", "Circulatory system", "Palpitations", "Chest pain". The explanatory notes for each medical term entity word are obtained: "Heart disease" refers to a disease caused by abnormal heart structure or function; "Circulatory system" refers to the system composed of the cardiovascular system and the lymphatic system; "Palpitations" refers to the feeling of a rapid, strong or irregular heartbeat; "Chest pain" refers to the symptom of chest pain. Then, term-explanation pairs are constructed: ["Heart disease - A disease caused by abnormal heart structure or function", "Circulatory system - The system composed of the cardiovascular system and the lymphatic system", "Palpitations - The feeling of a rapid, strong or irregular heartbeat", "Chest pain - The symptom of chest pain"]. After the term-explanation pairs are rewritten in a colloquial way, it becomes: "Heart disease, well, it means there's something wrong with the heart, either the structure is incorrect or the function is not good. The circulatory system, that is the system in our body composed of the cardiovascular and lymphatic systems. Palpitations, that is when you feel your heart beating very fast, very strongly, or irregularly. Chest pain, that is pain in the chest area." The finally generated term and explanation formal text can be: "Heart disease is a common disease, which means there's something wrong with the heart, maybe the structure is incorrect or the function is not good. It is mainly manifested as palpitations, that is when you feel your heart beating very fast, very strongly, or irregularly, and chest pain, that is pain in the chest area."

[0059] For Q&A pair-formatted texts, for original medical field documents such as: "Diabetes is a chronic disease, mainly manifested as high blood sugar. It is usually caused by insufficient insulin secretion or reduced responsiveness of cells to insulin. Prolonged high blood sugar may lead to damage to tissues such as blood vessels and nerves." Using a pre-set medical field Q&A model, the content of the original document can be analyzed, and multiple possible answer statements can be generated, as well as the question that these answer statements can answer. For example: Answer 1: Chronic disease. Answer 2: High blood sugar. Answer 3: Usually caused by insufficient insulin secretion or reduced responsiveness of cells to insulin. Answer 4: May lead to damage to tissues such as blood vessels and nerves. These answers can be used to answer: Question 1, What type of disease is diabetes? Question 2, What are the main manifestations of diabetes? Question 3, What causes diabetes? Question 4, What consequences will prolonged high blood sugar cause? For each question, combine the question and the answer to form a Q&A pair. For example: Q&A pair 1: Question, What type of disease is diabetes? Answer, Chronic disease. To make the Q&A pairs easier to understand and communicate, the Q&A pairs can be rewritten in a colloquial way. For example, the rewritten Q&A pair 1: Question, What kind of disease does diabetes belong to? Answer, It is a chronic disease. Finally, combine the rewritten Q&A pairs in a certain format (such as a list, a paragraph, etc.) to form a Q&A pair-formatted text.

[0060] Specifically, the medical field Q&A model can be built based on LLaMA-13B. LLaMA-13B is a large natural language processing model with powerful language understanding and generation capabilities. It is one of the Llama series of models released by the Meta AI team and contains 13 billion parameters. LLaMA-13B adopts the Transformer Decoder architecture and has been improved in many aspects to optimize performance. It can perform incremental pre-training on large-scale corpora, thus having a wide range of knowledge backgrounds. LLaMA-13B has high efficiency and scalability and can be applied to various specific fields. Through customized training for domain knowledge, it can improve the understanding and Q&A capabilities within the domain. For example, in the medical field, the MedicalGPT (Medical Generative Pre-trained Transformer, Chinese full name is medical generative pre-trained transducer) model built based on LLaMA-13B can accurately understand and answer medical-related questions through multiple stages of training such as secondary pre-training, supervised fine-tuning, reward modeling, and reinforcement learning training, providing a convenient communication bridge for doctors and patients.

[0061] For hierarchical summary-form texts, original medical field documents such as: "Diabetes is a chronic metabolic disease that mainly affects blood sugar levels in the body. It is divided into two types: type 1 diabetes and type 2 diabetes. Type 1 diabetes usually occurs in children and adolescents and is caused by absolute deficiency of insulin secretion. Type 2 diabetes is more common and occurs mostly in adults, caused by relative deficiency of insulin secretion or weakened body response to insulin (insulin resistance). Symptoms of diabetes include polyuria, thirst, weight loss, etc. If not treated promptly, diabetes can cause various complications such as cardiovascular diseases, retinopathy, neuropathy, and kidney diseases."

[0062] First, use a text segmentation algorithm to split the original document into multiple fragment contents. These fragments can be split based on sentences, paragraphs, or topics. In this example, for instance, splitting by paragraphs, a total of 6 fragments can be obtained, and two of these fragments, for example:

[0063] Fragment 1: Diabetes is a chronic metabolic disease that mainly affects blood sugar levels in the body.

[0064] Fragment 2: It is divided into two types: type 1 diabetes and type 2 diabetes.

[0065] Next, use a text summarization algorithm to summarize each fragment content. These summaries should concisely convey the main information of the fragments. For example: Summary of Fragment 1 (generalized content): Diabetes affects blood sugar levels. Summary of Fragment 2 (generalized content): Diabetes is divided into type 1 and type 2.

[0066] Combine each fragment content with its corresponding summary (generalized content) to form fragment-summary pairs. For example: Fragment-summary pair 1: Fragment content: Diabetes is a chronic metabolic disease that mainly affects blood sugar levels in the body. Generalized content: Diabetes affects blood sugar levels.

[0067] To make the fragment-summary pairs easier to understand and communicate, they can be rewritten in a colloquial way. For example: Rewritten fragment-summary pair 1: Fragment content (colloquial): Diabetes, well, it's a chronic disease that affects blood sugar levels. Generalized content (colloquial): That is to say, diabetes is related to blood sugar levels.

[0068] Finally, combine the rewritten fragment-summary pairs in a logical order to form a hierarchical summary-form text, for example, it can be organized according to the structure or theme of the original document.

[0069] For this reason, since the health knowledge of the elderly is scattered in various documents with different formats and structures, preprocessing and structured storage can unify the data format, facilitate subsequent retrieval and matching, and improve the availability of the data.

[0070] Step 104: Receive the medical questions raised by elderly users, and based on the medical questions, obtain question sentence vectors.

[0071] Step 105: Based on the multi-source database and the question sentence vectors, generate answer statements corresponding to the medical questions, and return the rewritten answer statements after performing colloquial rewriting on the answer statements.

[0072] Next, receive the medical questions raised by elderly users, obtain question sentence vectors, generate answer statements corresponding to the medical questions based on the multi-source database and the question sentence vectors, and return the rewritten answer statements after performing colloquial rewriting on the answer statements.

[0073] Optionally, in Step 104 where the medical questions raised by elderly users are received and question sentence vectors are obtained based on the medical questions, with reference to Figure 5 , specifically including:

[0074] Step 1041: Receive the medical questions raised by elderly users, perform word segmentation on the medical questions to obtain multiple question words, perform vectorization operations on each question word respectively to obtain question word vectors, and perform combined encoding on the question word vectors of each question word to obtain question sentence vectors.

[0075] In the above embodiments of the present application, when receiving the medical questions raised by elderly users, segment the medical questions to obtain multiple question words. After obtaining the question word vectors corresponding to each question word, use the Transformer decoder architecture to perform encoding and combination on the question word vectors to obtain question sentence vectors. Specifically, for example, if the medical question raised by the elderly user received is "I have always been dizzy recently. Is my hypertension acting up again? What should I do?", analyze the received medical question and extract the key information therein, such as performing preprocessing steps such as word segmentation and removing stop words on the question. For the above example question, the following key information (i.e., question words) can be analyzed: dizziness, hypertension, what to do.

[0076] After obtaining the question words, a preset word vector model can be used to match the question word vectors corresponding to each question word. Specifically, a suitable word vector model can be selected, such as Word2Vec, GloVe, etc. These models have been trained on large-scale corpora and include word vectors of multiple words, and can accurately represent the semantic information of words.

[0077] That is, after splitting the sentence into several words A, B, C, vectors a, b, c are obtained respectively, and encoding is performed by combining vectors a, b, c (for example, using the Transformer decoder for combined encoding) to obtain a vector x, which is used to represent this question sentence.

[0078] Optionally, in step 105, based on the multi-source database and the problem sentence vector, an answer statement corresponding to the medical problem is generated, and after the answer statement is rewritten into a colloquial form, the rewritten answer statement is returned, with reference to Figure 6 As shown, it specifically includes

[0079] Step 1051: In the multi-source database, obtain the medical term entity word vectors corresponding to the medical term entity words in the term and explanation form text, the question word vectors corresponding to each question word segmented from the askable questions in the question and answer pair form text, and the answer statement word vectors corresponding to each answer statement word segmented from the answer statements, and the segment content word vectors corresponding to each segment content word segmented from the segment content in the hierarchical summary form text.

[0080] In the above embodiment of the present application, in the multi-source database, obtain the medical term entity word vectors corresponding to the medical term entity words in the term and explanation form text, the question word vectors corresponding to each question word segmented from the askable questions in the question and answer pair form text, and the answer statement word vectors corresponding to each answer statement word segmented from the answer statements, and the segment content word vectors corresponding to each segment content word segmented from the segment content in the hierarchical summary form text. Specifically, the word vectors can be obtained through vectorization operations.

[0081] Step 1052: After adding a term explanation label to the medical term entity word vector, adding a Q&A label to the question word vector and the answer statement word vector, and adding a segment summary content label to the segment content word vector, construct an answer word vector index library based on the word vectors after adding labels.

[0082] Next, based on the medical term entity word vector, the question word vector, the answer statement word vector, and the segment content word vector, construct an answer word vector index library to prepare for subsequent answer matching. In particular, the answer word vector index library is divided into three types of labels, namely, the term explanation label, the Q&A label, and the segment summary content label, and different types of word vectors are added with different labels.

[0083] Step 1053: For the word vectors under any one of the labels in the answer word vector index library, calculate the similarity between each word vector and the problem sentence vector respectively, and based on the calculated similarity, determine multiple initial candidate answer word vectors under the type of label.

[0084] Step 1054, when the medical term entity word vector is determined as the initial candidate answer word vector, determine the candidate answer word vector based on the term explanation pair corresponding to the medical term entity word vector; when the question word vector is determined as the initial candidate answer word vector, determine the candidate answer word vector based on the Q&A pair corresponding to the question word vector; when the answer sentence word vector is determined as the initial candidate answer word vector, determine the candidate answer word vector based on the Q&A pair corresponding to the answer sentence word vector; when the fragment content word vector is determined as the initial candidate answer word vector, determine the candidate answer word vector based on the fragment summary content pair corresponding to the fragment content word vector.

[0085] Suppose the question raised by the elderly user is: "What is diabetes?" First, perform word segmentation on this question to obtain a word sequence such as [what / is / diabetes], and then convert these words into corresponding word vectors. A pre-trained language model (such as Word2Vec, BERT, etc.) can be used to obtain the word embedding representation (word vector), and then these word vectors are encoded and combined to obtain the question sentence vector.

[0086] For each word vector under any label in the answer word vector index library, for example, for various word vectors labeled with term explanations, calculate the similarity between each word vector and the question sentence vector respectively. Similarity calculation can be performed by calculating methods such as cosine similarity and Euclidean distance. Cosine similarity is a commonly used method, and its value range is [-1, 1]. The larger the value, the more similar the two vectors are. The similarity calculation formula is, for example:

[0087]

[0088] where v1 and v2 represent the word vector and the question sentence vector respectively, v1·v2 represents the inner product (dot product) of vector v1 and vector v2, and ||v1|| and ||v2|| represent the norms (Euclidean norms) of vector v1 and vector v2 respectively.

[0089] Next, select several (which can be set to other thresholds such as 3, 5, etc.) word vectors with the highest similarity as the candidate answer word vectors. Specifically:

[0090] For the medical term entity word vector: If the medical term entity word vector is determined as the initial candidate answer word vector (such as the word vector of "insulin"), then find the explanation of this term (insulin) through the term explanation pair where insulin is located, and use the word vectors corresponding to the medical term entity word and the explanation respectively as the candidate answer word vectors for subsequent answer generation.

[0091] For the word vectors of questions that can be asked: If the initial candidate answer word vector corresponds to a question that can be asked (such as "What are the symptoms of diabetes?"), then answers can be retrieved from the question-and-answer pairs in the multi-source database based on this question, and the word vectors corresponding to the question and the answer respectively are jointly used as the candidate answer word vectors for generating answers subsequently.

[0092] For the word vectors of answer sentences: If the initial candidate answer word vector directly corresponds to the answer to a question (such as "Diabetes is a chronic disease"), then questions can be retrieved from the question-and-answer pairs in the multi-source database based on this answer, and the word vectors corresponding to the answer and the question respectively are jointly used as the candidate answer word vectors for generating answers subsequently.

[0093] For the word vectors of fragment content: If the initial candidate answer word vector is a word vector of fragment content, then the word vectors corresponding to the fragment summary content pair where the fragment content corresponding to the word vector of fragment content is located can be used to determine the candidate answer word vectors for generating answers subsequently.

[0094] For example, find several question-and-answer pairs (question-and-answer pairs) highly relevant to "diabetes", such as "What are the symptoms of diabetes? (Answer: Polyuria, thirst, weight loss, etc.)" and "What are the treatment methods of diabetes? (Answer: Medication, diet control, exercise, etc.)". Then, according to the specific context and intention of the medical questions of the elderly users, the most appropriate answer can be selected from these candidate answers and presented to the elderly users. This process combines multiple technologies such as information retrieval, word vector matching, and semantic understanding to achieve accurate and efficient medical Q&A services.

[0095] Step 1055: Obtain the similarity scoring scores between the candidate answer word vectors corresponding to various tags and the question sentence vector respectively through a preset similarity scoring model, sort the candidate answer word vectors according to the obtained similarity scoring scores, and determine the target answer word vector based on the sorting result.

[0096] Step 1056: Fill the target answer word corresponding to the target answer word vector into a preset answer generation template to obtain the answer sentence corresponding to the medical question, and return the rewritten answer sentence after performing colloquial rewriting on the answer sentence, where the preset answer generation template includes multiple colloquially rewritten connecting words for connecting the answer words.

[0097] Suppose an elderly user asks a medical question: "How to prevent diabetes?", and the candidate answer word vectors are, for example, ["Control diet", "Increase exercise", "Regularly check blood sugar", "Avoid high-sugar foods"], etc.

[0098] To determine which candidate answers are more relevant and should be presented to elderly users first, the similarity scoring scores between the candidate answer word vectors corresponding to various tags and the question sentence vector can be obtained again through a preset similarity scoring model. To this end, by re-scoring and sorting the three candidate answer word vectors, the accuracy of the generated answer statements can be improved. After being trained, the preset similarity scoring model can output the similarity scoring scores of each candidate answer word vector by inputting the candidate answer word vector to be scored and the question sentence vector.

[0099] Next, sort the candidate answer word vectors according to the obtained similarity scoring scores, and determine the target answer word vectors corresponding to the scores ranked in the top preset number (such as 3 or other thresholds). Specifically, for example, take the top three with the highest scores. Multiple candidate answer word vectors may have the same score, so all candidate answer vectors with the highest top three scores are used as target answer word vectors. To convert the target answer word vectors into answer statements that are easy for elderly users to understand, a preset answer generation template can be used to generate answer statements. This preset answer generation template contains multiple conjunctions connecting the answer words, and these conjunctions use colloquial language such as "you need to...", "you also need to...", "and", etc., to make the answer more natural and fluent.

[0100] For example, assuming the determined target answer word vectors are ["control diet", "increase exercise", "regularly check blood sugar"], the generated answer statement may be: "To prevent diabetes, you need to control your diet, increase exercise, and regularly check your blood sugar."

[0101] Another example, the preset answer generation template is for example: "The [symptom] you mentioned may be due to [reason 1], [reason 2] or [reason 3]. It is recommended that you [suggested action]."

[0102] Among them, [symptom], [reason 1], [reason 2], [reason 3] and [suggested action] are slots to be filled.

[0103] Fill the terms or phrases (target answers) corresponding to the target answer word vectors into the corresponding slots in the template. For example, if the target answer word vectors are v{arrhythmia}, v{anemia} and v{anxiety}, the corresponding terms "arrhythmia", "anemia" and "anxiety" can be filled into the template.

[0104] After filling the slots, to ensure that the answer statement is colloquial, it can be achieved by selecting colloquial conjunctions (such as "because", "or", "it is recommended that you") and appropriate sentence structures. For example, the generated answer statement may be:

[0105] "The palpitation you mentioned may be caused by arrhythmia, anemia, or anxiety. It is recommended that you seek medical attention in a timely manner to determine the specific cause."

[0106] Next, rewrite the answer statement in a more colloquial way. For example, you can use tools such as word segmentation and part-of-speech tagging to preprocess the text, and then combine the rules and strategies of colloquial conversion to improve the content in the formal text. Specifically, you can use a more natural, friendly, and daily-conversation-suitable language to replace written words with colloquial ones. For example, "may be caused by" can be changed to "may be", and "it is recommended to seek medical attention in a timely manner" can be changed to "it's better to go to the hospital for a check". This makes the sentence more fluent and easier to understand. Also, for example, you can split a sentence with a causal relationship into two independent sentences to make the expression clearer. At the same time, add modal particles such as "ah", "ba", etc. to make the conversation more natural. Finally, generate a colloquial answer statement to return to the elderly user. The generated colloquial answer statement is, for example:

[0107] Your palpitation may be due to arrhythmia, anemia, or anxiety. It's better to go to the hospital for a check to find out the specific cause.

[0108] Also, for example, you can set up a prompt template, fill in the question into the prompt template, and input this prompt into a colloquial rewriting large model for rewriting. Among them, the prompt template is, for example:

[0109]

Task

[0110]

Example

[0111]

Answer to be rewritten

[0112]

Rewritten answer

[0113] Through this process, rewrite the original answer statement into a more colloquial, natural, and friendly expression, which is more suitable for use in daily conversations. In this way, the elderly user can understand their health status in a more understandable way and take corresponding actions. In this way, it is possible to screen out the most relevant and useful information from the candidate answers in combination with the medical problems raised by the elderly users and present it to the elderly users in a natural and fluent manner.

[0114] By applying the technical solution of this embodiment, the original medical field documents of the health knowledge of the elderly, which are scattered in different sources and forms, are converted into various forms of text and effectively integrated, and then transformed into structured data to construct a multi-source database, which can improve the retrieval efficiency and understanding ability. In addition, when faced with the questions of the elderly, it can effectively match the written and terminological content in professional literature. This reduces the barrier of professional terms encountered by the elderly when using the question-answering robot, and improves the understanding and acceptance of health knowledge. It also improves the retrieval accuracy of the question-answering robot for the questions of the elderly and generates high-quality and highly similar answers. By making full use of the health data of the elderly from different channels, including medical term entity words, questions, answers, document fragments and abstracts, etc., the coverage and accuracy of the question-answering system of the question-answering robot are improved. Finally, a health knowledge question-answering system that can better serve the elderly population is provided, improving the effectiveness in practical applications and the satisfaction of elderly users.

[0115] In a specific embodiment, for example Figure 7 As shown, according to the medical questions raised by the elderly users, the question sentence vector q is obtained, and the similarity is calculated separately with the word vectors of each label in the answer word vector index library. The topk candidate answer word vectors that are most matched for the question word vector q under each label are selected, and then the topk candidate answer word vectors corresponding to each label are scored for similarity, and re-sorted according to the scores, and the top3 target answer word vectors that are most relevant are selected. Specifically, a prompt word (connector) can also be designed to combine the questions of the elderly users with the top3 candidate answer word vectors to generate an answer statement, and the answer statement is returned after being rewritten into a colloquial form. Through the structured storage, multi-source matching and re-sorting mechanisms, the data utilization rate can be improved, matching clues can be found from multiple perspectives, and the most accurate information can be screened out through the large model to generate accurate answers.

[0116] Specifically, the medical questions raised by the elderly users are processed by the semantic model to obtain the question sentence vector q, and the [medical term entity word + explanatory note, term explanation pair] is processed by the semantic model to obtain the medical term word vector E. The [questionable question + answer statement, question-answer pair] is processed by the semantic model to obtain the questionable question word vector Q. The answer statement in the [questionable question + answer statement, question-answer pair] is processed by the semantic model to obtain the answer statement word vector A. The fragment content in the [fragment content + summary content, hierarchical summary] is processed by the semantic model to obtain the fragment content word vector D.

[0117] Build an answer word vector index library for E, Q, A, and D. The answer word vector index library can quickly convert the natural language questions of elderly users into vector form, match them with the word vectors in the multi-source database, and improve the efficiency and accuracy of retrieval. Calculate the similarity between the question sentence vector and the word vectors corresponding to various tags in the answer word vector index library, and select the top k candidate answer word vectors that are closest to q under each tag. If E enters the candidate, place the corresponding entity and explanatory description into the candidate. If Q enters the candidate, add the corresponding A to the candidate. If A enters the candidate, add the corresponding Q to the candidate. If D enters the candidate, add the corresponding abstract (summarized content) to the candidate. Through multi-source matching, different sources of data can be fully utilized to improve the coverage and diversity of answers.

[0118] Perform similarity scoring on the top k candidates under each tag respectively, and reorder them according to the scores to select the most relevant top 3 candidates, which are the goals. Reordering can further screen out the candidates most relevant to the questions of elderly users, that is, the goals, and improve the accuracy of answers.

[0119] Design prompt words, add q and the top 3 goals (E / Q / A / D) to the prompt words to obtain the answer statement. The steps of generating the answer statement ensure the coherence and accuracy of the answer, and can provide personalized answers according to the context. For this reason, in the face of problems such as scattered knowledge, colloquial questions from elderly users, and obstacles of professional terms. Through these technical solutions, the integration degree and utilization rate of data can be improved, the accuracy of retrieval and matching can be enhanced, and finally, health knowledge answers more in line with the needs of the elderly can be generated, thereby improving the experience and service quality of elderly users.

[0120] By applying the technical solutions of this embodiment, through structured storage and multi-source matching technologies, the efficiency and accuracy of the elderly health knowledge Q&A system are improved. Scattered data is effectively integrated, and the data utilization rate is increased. At the same time, by understanding colloquial questions and reducing obstacles of professional terms, the experience of elderly users is enhanced. In addition, the application of large models optimizes information screening and improves the accuracy of answers. It not only saves resources but also promotes the update and iteration of knowledge, improves the intelligent level in the Q&A process, and provides more convenient and accurate health consultation services for the elderly.

[0121] Further, as Figure 1 a specific implementation of the method, an elderly intelligent Q&A processing device is provided in an embodiment of the present application, as Figure 8 shown. The device includes:

[0122] A medical field document collection module 201 for collecting multiple original medical field documents containing elderly health knowledge;

[0123] A formal text conversion module 202, which is used to convert any original medical field document into multiple formal texts. Among them, the formal texts include term and explanation formal texts, question-and-answer pair formal texts, and hierarchical summary formal texts, and the text content in each formal text is colloquial language;

[0124] A multi-source database construction module 203, which is used to construct a multi-source database based on multiple formal texts corresponding to multiple original medical field documents;

[0125] A medical problem processing module 204, which is used to receive a medical problem proposed by an elderly user and obtain a problem sentence vector based on the medical problem;

[0126] An answer statement generation and return module 205, which is used to generate an answer statement corresponding to the medical problem based on the multi-source database and the problem sentence vector, and return the rewritten answer statement after performing colloquial rewriting on the answer statement.

[0127] Optionally, the formal text conversion module 202 is further used for:

[0128] Perform word segmentation on the original medical field document, determine multiple medical term entity words among the segmented words, obtain the corresponding explanatory notes for each medical term entity word, construct a term-explanation pair based on the medical term entity word and the explanatory note corresponding to the medical term entity word, and after performing colloquial rewriting on the term-explanation pair, obtain the term and explanation formal text corresponding to the original medical field document based on the rewritten term-explanation pair;

[0129] Generate answer statements that can be used to answer in the original medical field document and the questions that can be asked that the answer statements can answer based on a preset medical field question-answering model, obtain question-and-answer pairs based on the questions that can be asked and the answer statements corresponding to the questions that can be asked, and after performing colloquial rewriting on the question-and-answer pairs, obtain the question-and-answer pair formal text corresponding to the original medical field document, where the preset medical field question-answering model is established based on LLaMA-13B;

[0130] Segment the original medical field document into multiple fragment contents based on a text segmentation algorithm, respectively summarize each fragment content based on a text summarization algorithm to obtain the summary content corresponding to each fragment content, obtain a fragment-summary content pair based on the fragment content and the summary content corresponding to the fragment content, and after performing colloquial rewriting on the fragment-summary content pair, obtain the hierarchical summary formal text corresponding to the original medical field document based on the rewritten fragment-summary content pair.

[0131] Optionally, the medical problem processing module 204 is further used for:

[0132] Segment the medical problem to obtain multiple problem words. Perform vectorization operations on each problem word respectively to obtain problem word vectors. Combine and encode the problem word vectors of each problem word to obtain a problem sentence vector.

[0133] Optionally, the answer statement generation and return module 205 is further configured to:

[0134] In a multi-source database, obtain the medical term entity word vectors corresponding to the medical term entities in the term and explanation form text, the question word vectors corresponding to each question word segmented from the question that can be asked in the question-and-answer pair form text, and the answer statement word vectors corresponding to each answer statement word segmented from the answer statement, and the segment content word vectors corresponding to each segment content word segmented from the segment content in the hierarchical summary form text;

[0135] After adding a term explanation label to the medical term entity word vector, adding a question-and-answer label to the question word vector and the answer statement word vector, and adding a segment summary content label to the segment content word vector, construct an answer word vector index library based on each word vector after adding the label;

[0136] Generate an answer statement corresponding to the medical problem based on the answer word vector index library and the question sentence vector.

[0137] Optionally, the answer statement generation and return module 205 is further configured to:

[0138] For the word vectors under any one label in the answer word vector index library, calculate the similarity between each word vector and the question sentence vector respectively. Based on the calculated similarity, determine multiple initial candidate answer word vectors under the label;

[0139] When the medical term entity word vector is determined as an initial candidate answer word vector, determine the candidate answer word vector based on the term explanation corresponding to the medical term entity word vector. When the question word vector that can be asked is determined as an initial candidate answer word vector, determine the candidate answer word vector based on the question-and-answer pair corresponding to the question word vector that can be asked. When the answer statement word vector is determined as an initial candidate answer word vector, determine the candidate answer word vector based on the question-and-answer pair corresponding to the answer statement word vector. When the segment content word vector is determined as an initial candidate answer word vector, determine the candidate answer word vector based on the segment summary content pair corresponding to the segment content word vector;

[0140] Determine the target answer word vector based on the candidate answer word vectors of each label, and obtain an answer statement corresponding to the medical problem based on the target answer word vector.

[0141] Optionally, the response statement generation and return module 205 is further configured to:

[0142] Obtain the similarity scoring scores between the candidate answer word vectors corresponding to various tags and the question sentence vector respectively through a preset similarity scoring model, sort the candidate answer word vectors according to the obtained similarity scoring scores, and determine the target answer word vector based on the sorting result;

[0143] Fill the target answer word corresponding to the target answer word vector into a preset answer generation template to obtain a response statement corresponding to the medical question, where the preset answer generation template includes a plurality of connectives rewritten in a colloquial manner, and the connectives are used to connect the answer words.

[0144] Optionally, the medical field document collection module 201 is further configured to:

[0145] Collect a plurality of original candidate medical field documents containing health knowledge for the elderly;

[0146] For any original candidate medical field document, obtain the evaluation value of the original candidate medical field document according to a preset evaluation index and the index weight corresponding to the preset evaluation index, and obtain the original medical field document corresponding to the evaluation value greater than the preset evaluation threshold, where the preset evaluation index includes the document collection source, the document release time, and the authority of the document source.

[0147] It should be noted that for other corresponding descriptions of each functional unit involved in the elderly intelligent question-answering processing device provided in the embodiments of the present application, reference can be made to Figure 1 、 Figure 2 、 Figures 4 to 6 the corresponding descriptions in the method, which will not be elaborated here.

[0148] Based on the above as Figure 1 、 Figure 2 、 Figures 4 to 6 shown in the method, correspondingly, the embodiments of the present application further provide a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the elderly intelligent question-answering processing method as shown in the above as Figure 1 、 Figure 2 、 Figures 4 to 6 shown.

[0149] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in various implementation scenarios of the present application.

[0150] Based on the methods described above such as Figure 1 , Figure 2 , Figures 4 to 6 shown, and the virtual device embodiments shown in Figure 8 in order to achieve the above object, an embodiment of the present application further provides a computer device, specifically, it may be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the methods for intelligent question answering processing of the elderly as described above such as Figure 1 , Figure 2 , Figures 4 to 6 shown.

[0151] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.

[0152] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not limit the computer device, and it may include more or fewer components, or combine certain components, or have different component arrangements.

[0153] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing and storing the hardware and software resources of the computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, and communication between other hardware and software in the entity device.

[0154] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can be implemented through hardware. By effectively integrating the elderly health knowledge scattered in different sources and forms and making oral improvement, a multi-source database is formed. Through the multi-source database for intelligent question answering, it can effectively handle oral questions and the understanding obstacles of professional terms, and improve the user experience of the elderly.

[0155] Those skilled in the art can understand that the attached drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the attached drawings are not necessarily essential for implementing this application. Those skilled in the art can understand that the modules in the devices in the implementation scenario can be distributed in the devices in the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more devices different from this implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0156] The above serial numbers of this application are only for description and do not represent the advantages or disadvantages of the implementation scenario. The above disclosure is only several specific implementation scenarios of this application. However, this application is not limited thereto, and any changes that can be made by those skilled in the art should fall within the protection scope of this application.

Claims

1. An intelligent question-and-answer processing method for the elderly, characterized in that, The method includes: Collecting multiple original medical field documents containing health knowledge for the elderly; For any original medical field document, converting the original medical field document into multiple forms of text, where the form of text includes term and explanation form text, question-and-answer pair form text, and hierarchical summary form text, and the text content in each form of text is colloquial language; Based on the multiple forms of text corresponding to each of the multiple original medical field documents, constructing a multi-source database; Receiving a medical question raised by an elderly user, and obtaining a question sentence vector based on the medical question; In the multi-source database, obtaining the medical term entity word vectors corresponding to the medical term entities in the term and explanation form text, the questionable question word vectors corresponding to the respective questionable question words segmented from the questionable questions in the question-and-answer pair form text, the answer sentence word vectors corresponding to the respective answer sentence words segmented from the answer sentences, and the fragment content word vectors corresponding to the respective fragment content words segmented from the fragment content in the hierarchical summary form text; After adding a term explanation label to the medical term entity word vector, adding a question-and-answer label to the questionable question word vector and the answer sentence word vector, and adding a fragment summary content label to the fragment content word vector, constructing an answer word vector index library based on the word vectors after adding the labels; For the word vectors under any one of the labels in the answer word vector index library, calculating the similarity between each word vector and the question sentence vector respectively, and determining multiple initial candidate answer word vectors under the label based on the calculated similarity; When the medical term entity word vector is determined as the initial candidate answer word vector, determining the candidate answer word vector based on the term explanation pair corresponding to the medical term entity word vector; when the questionable question word vector is determined as the initial candidate answer word vector, determining the candidate answer word vector based on the question-and-answer pair corresponding to the questionable question vector; when the answer sentence word vector is determined as the initial candidate answer word vector, determining the candidate answer word vector based on the question-and-answer pair corresponding to the answer sentence word vector; when the fragment content word vector is determined as the initial candidate answer word vector, determining the candidate answer word vector based on the fragment summary content pair corresponding to the fragment content word vector; Obtaining the similarity scoring scores between the candidate answer word vectors corresponding to each of the labels and the question sentence vector through a preset similarity scoring model, sorting each candidate answer word vector according to the obtained similarity scoring scores, and determining the target answer word vector based on the sorting result; Filling the target answer word corresponding to the target answer word vector into a preset answer generation template to obtain an answer statement corresponding to the medical question, where the preset answer generation template includes multiple colloquially rewritten conjunctions for connecting answer words; Returning the rewritten answer statement after colloquially rewriting the answer statement.

2. The method according to claim 1, characterized in that, Converting the original medical field document into term and explanation form text includes: Perform word segmentation on the original medical field document, determine multiple medical term entity words among the segmented words, obtain the corresponding explanatory descriptions for each medical term entity word, construct term-explanation pairs based on the medical term entity words and their corresponding explanatory descriptions, and after performing colloquial rewriting on the term-explanation pairs, obtain the term-and-explanation form text corresponding to the original medical field document based on the rewritten term-explanation pairs; Convert the original medical field document into a question-and-answer pair form text, including: Generate answer statements that can be used to answer in the original medical field document and the question statements that can be answered by the answer statements based on a preset medical field question-and-answer model, obtain question-and-answer pairs based on the question statements and the answer statements corresponding to the question statements, and after performing colloquial rewriting on the question-and-answer pairs, obtain the question-and-answer pair form text corresponding to the original medical field document, where the preset medical field question-and-answer model is established based on LLaMA-13B; Convert the original medical field document into a hierarchical summary form text, including: Segment the original medical field document into multiple fragment contents based on a text segmentation algorithm, respectively summarize each fragment content based on a text summarization algorithm to obtain the summary content corresponding to each fragment content, obtain fragment-summary content pairs based on the fragment content and the summary content corresponding to the fragment content, and after performing colloquial rewriting on the fragment-summary content pairs, obtain the hierarchical summary form text corresponding to the original medical field document based on the rewritten fragment-summary content pairs.

3. The method according to claim 1, characterized in that, The obtaining the question sentence vector based on the medical question includes: Perform word segmentation on the medical question to obtain multiple question words, perform vectorization operations on each question word respectively to obtain question word vectors, and perform combined encoding on the question word vectors of each question word to obtain a question sentence vector.

4. The method according to claim 1, characterized in that, The collecting multiple original medical field documents containing elderly health knowledge includes: Collect multiple original candidate medical field documents containing elderly health knowledge; For any original candidate medical field document, obtain the evaluation value of the original candidate medical field document according to a preset evaluation index and the index weight corresponding to the preset evaluation index, and obtain the original medical field document corresponding to the evaluation value greater than the preset evaluation threshold, where the preset evaluation index includes the document collection source, the document release time, and the authority of the document source.

5. An intelligent question-answering processing device for the elderly, characterized in that, The device includes: A medical field document collection module for collecting multiple original medical field documents containing elderly health knowledge; A form text conversion module for converting the original medical field document into multiple form texts for any original medical field document, where the form texts include term-and-explanation form text, question-and-answer pair form text, and hierarchical summary form text, and the text content in each form text is colloquial language; A multi-source database construction module for constructing a multi-source database based on the multiple form texts corresponding to the multiple original medical field documents; A medical problem processing module, configured to receive medical problems raised by elderly users, and obtain problem sentence vectors based on the medical problems; An answer statement generation and return module, configured to obtain medical term entity word vectors corresponding to medical term entities in the term and explanation form text, question word vectors corresponding to each question word segmented from the question in the question-and-answer pair form text, and answer statement word vectors corresponding to each answer statement word segmented from the answer statement, and segment content word vectors corresponding to each segment content word segmented from the segment content in the hierarchical summary form text in a multi-source database; adding a term explanation label to the medical term entity word vectors, adding a question-and-answer label to the question word vectors and the answer statement word vectors, and adding a segment summary content label to the segment content word vectors, and then constructing an answer word vector index library based on the word vectors with added labels; for the word vectors under any one of the labels in the answer word vector index library, calculate the similarity between each word vector and the question sentence vector respectively, and determine multiple initial candidate answer word vectors under the label based on the calculated similarity; when the medical term entity word vector is determined as an initial candidate answer word vector, determine the candidate answer word vector based on the term explanation corresponding to the medical term entity word vector, when the question word vector is determined as an initial candidate answer word vector, determine the candidate answer word vector based on the question-and-answer pair corresponding to the question word vector, when the answer statement word vector is determined as an initial candidate answer word vector, determine the candidate answer word vector based on the question-and-answer pair corresponding to the answer statement word vector, and when the segment content word vector is determined as an initial candidate answer word vector, determine the candidate answer word vector based on the segment summary content corresponding to the segment content word vector; obtain the similarity scoring scores between the candidate answer word vectors corresponding to each label and the question sentence vector respectively through a preset similarity scoring model, sort the candidate answer word vectors according to the obtained similarity scoring scores, and determine the target answer word vector based on the sorting result; fill the target answer word corresponding to the target answer word vector into a preset answer generation template to obtain an answer statement corresponding to the medical problem, wherein the preset answer generation template includes a plurality of connectives rewritten in a colloquial form, and the connectives are used to connect answer words, and return the rewritten answer statement after colloquializing the answer statement.

6. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for processing intelligent question and answer for the elderly according to any one of claims 1 to 4.

7. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for processing intelligent question and answer for the elderly according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • A spoken medical question and answer method and system

    CN109684445A

  • Medical auxiliary question and answer method and system based on knowledge calibration and retrieval enhancement

    CN117573843A