Multi-dimensional disease diet recommendation method and computer program product
Patent Information
- Application Number
- CN202510035177.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The disease dietary recommendation model in the prior art has low generalization ability and is difficult to cope with complex medical and nutritional information.
Using a multi-dimensional disease diet recommendation method, a combination of a disease diet knowledge base with a hierarchical storage structure and a large language model is obtained, user query instructions are obtained and topic prompts are generated, and layered search and response generation are performed.
It improves the generalization ability in different situations, enhances the search accuracy of the disease dietary knowledge base and the understanding of disease and semantic associations by large language models, thereby improving the accuracy of recommended dietary results.
Smart Images

Figure CN119943284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of nutritional health treatment principles and disease dietary health technology, and in particular to a multi-dimensional disease dietary recommendation method and a computer program product. Background Art
[0002] For many chronic diseases (such as diabetes, hypertension, heart disease, etc.), dietary control is a key part of disease management. Traditional disease dietary recommendations mainly rely on rules and statistical models, lacking certain personalization and precision. With the rapid development of artificial intelligence (AI) and the health field, AI can develop personalized meal plans based on personal dietary preferences, restrictions and health goals. Compared with traditional disease dietary recommendations, it has improved personalization and precision, and can implement precise medical strategies based on individual differences of patients.
[0003] For example, the Chinese invention patent document (CN117725186B) discloses a method and device for generating dialogue samples for a specific disease diet plan, which combines a professional knowledge base with multiple large models to generate dialogue samples, quickly generate a large number of professional field dialogue samples, and screen effective dialogue samples through similarity judgment for fine-tuning training of the large model. It includes: inputting the first prompt word into the first type of large model to generate sample questions related to the specific disease diet plan; inputting the sample question as the second prompt word into multiple second type large models, and using each second type large model to generate a sample answer corresponding to the sample question; performing text similarity recognition between the sample answers generated by the multiple second type large models, if the text similarity between the two reaches or exceeds the preset threshold, the sample answer is judged to be preliminarily valid; the sample answers generated by the multiple second type large models are input as the third prompt word into the third type large model, and the third type large model is used to judge whether the meaning of the sample answers is the same according to the third prompt word; if the sample answer is preliminarily valid, and the third type large model determines that the sample answers have the same meaning, then the sample question and the sample answer are saved as dialogue samples.
[0004] However, this method is designed for specific diseases, the generalization ability of the model is low, and its applicability in different diseases and different patient groups is also low. In addition, the dual judgment of text similarity and semantic similarity to screen effective samples still has certain limitations when processing complex medical and nutritional information. Summary of the invention
[0005] The present invention provides a multi-dimensional disease diet recommendation method and a computer program product to solve the problems in the prior art that the disease diet recommendation model has low generalization ability and is difficult to cope with complex medical and nutritional information.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is to provide a multi-dimensional disease diet recommendation method, the multi-dimensional disease diet recommendation method comprising: pre-building a disease diet knowledge base, wherein the disease diet knowledge base is configured as a hierarchical storage structure, each level storing different data categories; Obtain a user query instruction, and extract the subject information of the user query instruction to generate a subject prompter; perform a hierarchical search in the disease diet knowledge base according to the semantic clues in the subject prompter, until finding nodes and / or documents related to the user query instruction.
[0007] The topic prompter is optimized according to the topic level loss of hierarchical retrieval of the disease diet knowledge base; the user query instruction and the node and / or the document content are integrated and input into a large language model, and a response is generated.
[0008] The technical solution provided by the present invention has the following beneficial effects compared with the prior art: The pre-built multi-layer storage structure of the knowledge base can better understand and process the complex information related to various diseases and diets, thereby improving the generalization ability in different situations. Among them, the extraction of the subject information of the user query instruction to generate a subject prompter can improve the subsequent retrieval accuracy in the disease diet knowledge base and the understanding of the disease and semantic association of the large language model, thereby improving the accuracy of the response (i.e., the recommended diet result).
[0009] Among them, by efficiently clarifying, deduplicating and optimizing data from multiple sources, a pre-built disease diet knowledge base can solve the current problem of insufficient knowledge in the field of disease diet. Compared with the current deep learning methods that rely too much on structured databases and fixed rules, the above pre-built disease diet knowledge base has significantly improved knowledge update efficiency and retrieval speed.
[0010] In addition, by enhancing the multi-level demand parsing of large language models, we can better understand the questions input by users, make the content generated by subsequent responses more complete and accurate, and reduce dependence on expert knowledge.
[0011] In some embodiments, obtaining a user query instruction and extracting subject information of the user query instruction to generate a subject prompter includes: The user query instruction is converted into a tokenized input embedding; topic information is extracted from the tokenized input embedding and a topic tag embedding is generated; the contextual fine semantics of the topic tag embedding is obtained; the contextual fine semantics is dynamically modeled through a state space model to generate a topic state embedding with a time series; the topic state embedding is integrated with the contextual fine semantics and a topic prompter is generated.
[0012] In some embodiments, the pre-built disease diet knowledge base includes: Acquire clinical reference data, the clinical reference data including expert doctor advice data and clinical reference book text data; delete the directory, chapter names and subheadings in the clinical reference book text data, and integrate the expert doctor advice data, which is recorded as the first data set.
[0013] Obtain chronic disease recommended dietary data text, and organize them into chronic disease recommended dietary text documents one by one; integrate the chronic disease recommended dietary text documents, and delete the chapter names and subheadings of the chronic disease recommended dietary text documents, and record them as the second data set.
[0014] The Prompt command is used to automatically clean and screen the first data set and the second data set through the large language model, and then the data sets are integrated into a third data set.
[0015] In some embodiments, the pre-built disease diet knowledge base further comprises: The third data set is semantically vectorized, and the semantic similarity between the text vectors of the third data set is calculated; the third data set is deduplicated, and the third data set is optimized using a hierarchical clustering algorithm. The third data set is vectorized, and vectorized and stored using a built-in function to construct a disease diet knowledge base.
[0016] In some embodiments, the integrating the user query instruction and the node and / or the document content, inputting into a large language model, and generating a response includes: The large language model is pre-connected to an external Internet knowledge base; context information is integrated according to the node and / or the document content, and the context information is encoded as an embedding; an output sequence embedding is generated according to the tokenized input embedding of the user query instruction and the subject tag embedding; the large language model generates a response according to the embedded context information, the output sequence embedding and the external Internet knowledge base.
[0017] In some embodiments, the integrating the user query instruction and the node and / or the document content, inputting into a large language model, and generating a response includes: When the user query instruction is a voice sample, the user input instruction is converted into text data in real time by an external Whisper module; the text answer generated by the large language model according to the user input instruction is converted into a voice response through the FastSpeech2 module; when the user query instruction is a text sample, the large language model generates a text response according to the user input instruction.
[0018] In some implementations, the large language model is deployed online through the VLLM so that users can access it through a Web page.
[0019] In some embodiments, the large language model is CLLaMA2-13B, wherein the large language model CLLaMA2-13B is deployed with a topic selection state space model, and the topic selection state space model is used to perform: Performing hierarchical retrieval in the disease diet knowledge base according to the semantic clues in the subject prompter until finding nodes and / or documents related to the user query instruction; Performing hierarchical retrieval in the disease diet knowledge base according to the semantic clues in the subject prompter until finding nodes and / or documents related to the user query instruction; The topic prompter is optimized according to the topic level loss of hierarchical retrieval of the disease diet knowledge base.
[0020] In some embodiments, the optimizing the topic prompter according to the topic hierarchy loss of hierarchical search of the disease diet knowledge base comprises: A pre-built hierarchical index includes a first-level index structure and a second-level index structure, wherein the first-level index structure is used to store field classifications, and the second-level index structure is used to store specific entries; the topic selection state space model optimizes the topic indicator according to the hierarchical loss generated by the hierarchical index in the retrieval enhancement generation framework.
[0021] In some embodiments, the present application also provides a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the above-mentioned multi-dimensional disease diet recommendation method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work, among which: Figure 1 This is a schematic diagram of a process of an embodiment of a multi-dimensional disease diet recommendation method provided by the present invention. Figure 1 ; Figure 2 This is a schematic diagram of a process of an embodiment of a multi-dimensional disease diet recommendation method provided by the present invention. Figure 2 ; Figure 3It is a structural block diagram of an embodiment of a multi-dimensional disease diet recommendation method provided by the present invention; Figure 4 This is a schematic diagram of a process of an embodiment of a multi-dimensional disease diet recommendation method provided by the present invention. Figure 3 ; Figure 5 This is a prompt instruction schematic diagram of an embodiment of a multi-dimensional disease diet recommendation method provided by the present invention; Figure 6 This is a schematic diagram of a process of an embodiment of a multi-dimensional disease diet recommendation method provided by the present invention. Figure 4 ; Figure 7 It is a multimodal interactive logic diagram of an embodiment of a multidimensional disease diet recommendation method provided by the present invention. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0024] See also Figure 1 As shown, Figure 1 The following is a flow chart showing an embodiment of a multi-dimensional disease dietary recommendation method provided by the present application. Figure 1 .
[0025] In some embodiments, the multidimensional disease dietary recommendation method comprises: Step S100, pre-constructing a disease diet knowledge base.
[0026] Among them, the disease diet knowledge base is configured as a hierarchical storage structure, and each level stores different data categories. By efficiently clarifying, deduplicating and optimizing data from multiple sources, the speed and accuracy of knowledge retrieval can be greatly improved. Exemplarily, the above-mentioned disease diet knowledge base is a vector knowledge base, which can provide personalized disease diet recommendations by collecting data related to disease diet (such as academic literature, clinical data from tertiary hospitals, documents recommended by the National Health Commission, etc.) in combination with the user's personal information, physical condition and medical history development, and taste information.
[0027] Compared with the current deep learning methods that rely too much on structured databases and fixed rules, the above-mentioned pre-built disease diet knowledge base has significantly improved the knowledge updating efficiency and retrieval speed.
[0028] See also Figures 2 to 3 As shown, Figure 2The following is a flow chart showing an embodiment of a multi-dimensional disease dietary recommendation method provided by the present application. Figure 2 ; Figure 3 A structural block diagram of an embodiment of a multi-dimensional disease diet recommendation method provided by the present application is shown.
[0029] Step S200: Obtain a user query instruction and extract the subject information of the user query instruction to generate a subject prompter. Figure 2 As shown, step S200 includes: Step S210, converting the user query instruction into a tokenized input embedding.
[0030] After receiving a user query command, the embedding layer of natural language processing can be used to convert the text of the user query command into a numerical form that the model can understand, such as tokenization, part-of-speech tagging, and named entity recognition.
[0031] Step S220, extracting topic information from the tokenized input embedding and generating a topic tag embedding; and step S230, obtaining contextual fine semantics of the topic tag embedding.
[0032] The topic tag embedding is obtained through the hierarchical topic retrieval enhancement generation module. Specifically, the tokenized input embedding is linearly transformed to generate the corresponding embedding vector. Then, a layer of convolutional network is used to obtain detailed context information, so as to obtain the local relationship between words and phrases in the text, thereby improving the understanding of user query instructions.
[0033] Step S240, dynamically modeling the contextual fine semantics through a state space model to generate a topic state embedding with a time series.
[0034] Using the topic selection state space model to analyze the changes of topics over time can capture and simulate the temporal dynamic changes or complex dependencies in the text, and further enhance the ability to understand the user query instruction text.
[0035] Step S250, integrating topic state embedding with contextual fine semantics and generating a topic prompter.
[0036] The topic prompter integrates contextual information and topic information to guide subsequent hierarchical retrieval and response generation, so as to generate more accurate disease dietary recommendation answers.
[0037] Step S300, performing hierarchical search in the disease diet knowledge base according to the semantic clues in the subject prompter, until finding nodes and / or documents related to the user's query instruction.
[0038] The topic selection state space model performs layer-by-layer screening based on the semantic clues in the above-mentioned topic prompter. Exemplarily, the disease diet knowledge base is divided into multiple layers by the topic selection state space model. Exemplarily, the first layer is the domain classification (such as disease, ingredients, recipes, etc.), and the second layer is the specific entries (such as disease = hypertension; recipe = low-sodium breakfast), that is, the recipes corresponding to specific diseases are recommended (such as low-sodium breakfast recipes corresponding to hypertension).
[0039] Step S400, optimizing the topic prompter according to the topic level loss of hierarchical retrieval of the disease diet knowledge base.
[0040] Combination Figure 3 As shown in the figure, by pre-building a multi-layer knowledge base, the HierarchicalTopic Loss represents the loss function used to optimize the model parameters in the disease diet knowledge base stored in a hierarchical structure to better capture the multi-level topics implicit in the document. By continuously adjusting the relationship between documents and topics in the disease diet knowledge base, the retrieval query is optimized.
[0041] That is, through multi-level demand analysis, we can better understand the questions input by users and compare them with disease types to further understand user questions. Then, through the layered topic retrieval enhancement generation module, we can supplement the knowledge and semantic level capabilities, so that the content generated by subsequent responses is more complete and accurate.
[0042] Step S500: Integrate user query instructions and node and / or document content, input into a large language model, and generate a response.
[0043] Exemplarily, the large language model used in this application is the CLLaMA2-13B model, which is a large language model based on the Transformer architecture and has 13 billion parameters.
[0044] Among them, the present application deploys a topic selection state space model for the large language model CLLaMA2-13B, and the topic selection state space model is used to execute: executable step S200, obtaining user query instructions, and extracting the topic information of the user query instructions to generate a topic prompter; step S300, performing a hierarchical search in the disease diet knowledge base according to the semantic clues in the topic prompter, until the nodes and / or documents related to the user query instructions are searched; step S400, performing a hierarchical search in the disease diet knowledge base according to the semantic clues in the topic prompter, until the nodes and / or documents related to the user query instructions are searched; and step S500, optimizing the topic prompter according to the topic hierarchy loss of the hierarchical search of the disease diet knowledge base.
[0045] The above steps can greatly improve the ability to understand disease diet terms and semantic associations, thereby improving the accuracy and personalization of recommendation results, while reducing dependence on expert knowledge. In addition, the above disease diet knowledge base is hierarchically retrieved, which can reduce the cost of structured storage data and speed up the update speed.
[0046] In addition, the above steps can perform multi-task processing, such as being able to handle multiple tasks such as disease question answering, dietary recommendations, etc. For example, the large language model is pre-connected to an external Internet knowledge base, making it scalable and generalizable, able to adapt to the needs of different tasks, and obtain the latest data for response. Figure 3 As shown, corresponding dietary recommendations can be generated based on dietary recommendations and disease nutritional treatment principles. See also Figures 4 to 5 As shown, Figure 4 The following is a flow chart showing an embodiment of a multi-dimensional disease dietary recommendation method provided by the present application. Figure 3 ; Figure 5 A prompt instruction schematic diagram of an embodiment of a multi-dimensional disease diet recommendation method provided by the present application is shown.
[0047] In some embodiments, in combination Figure 4 As shown, step S100, pre-building a disease diet knowledge base, includes: Step S110, obtaining clinical reference data, which includes expert doctor advice data and clinical reference book text data; Step S120, deleting the directory, chapter names and subheadings in the reference book text data, and integrating the expert doctor advice data, which is recorded as the first data set.
[0048] Exemplarily, the accuracy of the first data set is determined by communicating with expert doctors. Specifically, the reference data text data is passed through Python code, the table of contents, chapter names and subheadings in the reference book text data are deleted to be consistent with subsequent data, and they are sorted out one by one to ensure accuracy.
[0049] Step S130, obtaining the chronic disease recommended dietary data text, and organizing them one by one into a chronic disease recommended dietary text document; Step S140, integrating the chronic disease recommended dietary text document, and deleting the chapter names and subheadings of the chronic disease recommended dietary text document, and recording them as the second data set.
[0050] Exemplarily, the recommended diet for chronic diseases PDF is downloaded from the official website of the National Health Commission, and it is sorted item by item according to the title and subsection to ensure accuracy. After sorting, the chapter names and subheadings are deleted to keep the first data set consistent.
[0051] Step S150: Use the Prompt command to automatically clean and filter the first data set and the second data set through the large language model, and integrate them into a third data set.
[0052] For example, in combination Figure 5 As shown, the ChatGPT model is used to guide the cleaning of the first and second data sets through the Prompt command. For example, a role of a disease diet expert and a clinical doctor is set, and the ChatGPT model is required to determine whether the given context contains content related to disease diet or clinical nutrition. If it does, the answer is based on the context content. In some application scenarios, the third database also needs to be manually reviewed to ensure the accuracy of the data.
[0053] Combination Figure 6 As shown, Figure 6 The following is a flow chart showing an embodiment of a multi-dimensional disease dietary recommendation method provided by the present application. Figure 4 .
[0054] In some embodiments, step S100, pre-building a disease diet knowledge base, further includes: Step S160, semantic vectorization is performed on the third data set, and semantic similarity between text vectors of the third data set is calculated. Exemplarily, an instruction-optimized embedding model (such as a gte-Qwen1.5-7B-instruct model) can be used to perform semantic similarity on the third data set (such as calculating cosine similarity between semantic vectors).
[0055] Step S170, de-duplication of the third data set is performed, and the third data set is optimized using a hierarchical clustering algorithm. Exemplarily, the third data set is de-duplicated by using a cosine distance to remove similar or repeated text data to reduce redundant information of the third data set. Exemplarily, the above operations such as deleting titles can ensure the standardization of the third data set, thereby ensuring that the scales of different features do not affect the clustering grouping results of the hierarchical clustering algorithm.
[0056] Step S180, vectorizing the third data set and storing the vectors through built-in functions to construct a disease diet knowledge base.
[0057] Exemplarily, the text is converted into a semantic vector (for example, using a BGE model or other model that embeds text into words), and a first-layer domain vector database is constructed through FAISS.
[0058] In an embodiment of the present application, a public instruction dataset can be used to fine-tune the general instructions of the large language model to give the large language model (CLLaMA2-13B model) the ability to understand and execute instructions, so as to improve the application flexibility of the large language model in disease diet recommendations.
[0059] In some embodiments, step S300, integrating the user query instruction and the node and / or document content, inputting into the large language model, and generating a response, includes: Integrate the node and / or document content into context information and encode the context information into embedding; generate output sequence embedding based on the tokenized input embedding of the user query instruction and the topic tag embedding; the large language model generates a response based on the embedded context information, the output sequence embedding and the external Internet knowledge base.
[0060] In an embodiment of the present application, the CLLaMA2-13B model obtains the tokenized input embedding and the topic tag embedding and then generates the output sequence embedding, wherein the output sequence embedding contains the semantic information generated by the large language model, which enables the large language model to have a multi-level demand parsing capability, thereby facilitating the large language model to better understand the user input questions and generate more complete and accurate responses. Exemplarily, the integrated context information is converted into a numerical vector through embedding technology (such as Word2Vec, BERT or GPT), and then the information from different sources (such as a disease diet recommendation knowledge base and an external Internet knowledge base) is integrated through an attention mechanism or weighted average.
[0061] See also Figure 7 As shown, Figure 7 A multimodal interaction logic diagram of an embodiment of a multidimensional disease diet recommendation method provided by the present application is shown.
[0062] In some embodiments, step S500, integrating the user query instruction and the node and / or document content, inputting into the large language model, and generating a response, includes: When the user query command is a voice sample, the user input command is converted into text data in real time by the external Whisper module; the text answer generated by the large language model based on the user input command is converted into a voice response through the FastSpeech2 module; when the user query command is a text sample, the large language model generates a text response based on the user input command.
[0063] In the embodiment of the present application, two interaction modes are set to realize text dialogue and voice dialogue. For example, in the voice dialogue mode, the user input command is converted into text data in real time by the external Whisper module, and the text data is passed to the optimized large language model (CLLaMA2-13B model), and the corresponding information entry is queried from the disease diet database, and a text answer is generated. The FastSpeech2 module converts the text answer into a voice response and outputs it to the user in voice, realizing the function of a voice assistant.
[0064] In the text dialogue mode, users can directly input user query commands into the optimized large language model (CLLaMA2-13B model) for communication, thus ensuring text interaction capabilities.
[0065] In some application scenarios, the above-mentioned voice mode and text mode can be converted according to user selection and user input in the same conversation. The above-mentioned multimodal interaction can process and integrate multiple modal information and perform multimodal tasks, such as text interaction mode and voice interaction mode, so as to provide users with a more diverse and flexible intelligent experience. The above-mentioned multi-dimensional disease diet recommendation method considers information from multiple dimensions to make it more in line with the needs of the elderly group for intelligent services.
[0066] For example, since chronic diseases account for a considerable proportion of the elderly population, and the prevalence and comorbidity of chronic diseases increase with age, the voice mode can also be combined with aging-friendly language input and output modules, such as dialect and accent recognition, to improve the accuracy of voice input.
[0067] In some implementations, the large language model is deployed online through the VLLM so that users can access it through a Web page.
[0068] In the embodiment of the present application, by deploying the adjusted large language model online, a more convenient access method can be provided to users, that is, users can interact with it without downloading software. In addition, online deployment enables the service to cover a wider user group, support large-scale users to use it at the same time, and improve user experience.
[0069] In some embodiments, step S200, optimizing the topic prompter according to the topic level loss of hierarchical retrieval of the disease diet knowledge base, comprises: A pre-built hierarchical index includes a first-level index structure and a second-level index structure. The first-level index structure is used to store field classifications, and the second-level index structure is used to store specific entries. The topic selection state space model optimizes the topic indicator according to the hierarchical loss generated by the hierarchical index in the retrieval enhancement generation framework.
[0070] In some embodiments, the present application also provides a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the above-mentioned multi-dimensional disease diet recommendation method is implemented.
[0071] It should be understood by those skilled in the art that the embodiments of the present invention can be provided as methods, systems or computer program products. Those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.
[0072] The above description is only an implementation mode of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly used in other related technical fields, should be included in the protection scope of the present invention.
Claims
1. A multi-dimensional disease diet recommendation method, characterized in that: include: Pre-constructing a disease diet knowledge base, wherein the disease diet knowledge base is configured as a hierarchical storage structure, each level storing different data categories; Obtaining a user query instruction, and extracting subject information of the user query instruction to generate a subject prompter; Performing hierarchical retrieval in the disease diet knowledge base according to the semantic clues in the subject prompter until finding nodes and / or documents related to the user query instruction; Optimizing the topic prompter according to the topic level loss of hierarchical retrieval of the disease diet knowledge base; The user query instruction and the node and / or the document content are integrated, input into a large language model, and a response is generated.
2. The multidimensional disease dietary recommendation method according to claim 1, characterized in that: The obtaining of the user query instruction and extracting the subject information of the user query instruction to generate a subject prompter includes: Converting the user query into a tokenized input embedding; extracting topic information from the tokenized input embedding and generating a topic token embedding; Obtaining contextual fine semantics of the topic tag embedding; Dynamically modeling the contextual fine semantics through a state space model to generate a topic state embedding with a time series; The topic state embedding is fused with the contextual fine semantics and a topic prompter is generated.
3. The multi-dimensional disease dietary recommendation method according to claim 1, characterized in that: The pre-built disease diet knowledge base includes: Acquiring clinical reference data, wherein the clinical reference data includes expert doctor advice data and clinical reference book text data; Deleting the table of contents, chapter names and subheadings in the clinical reference book text data, and integrating the expert doctor advice data, which is recorded as the first data set; Obtain the recommended dietary data text for chronic diseases, and organize them into a recommended dietary text document for chronic diseases; Integrate the chronic disease recommended diet text document, and delete the chapter names and subheadings of the chronic disease recommended diet text document, and record it as a second data set; The Prompt command is used to automatically clean and screen the first data set and the second data set through the large language model, and then the data sets are integrated into a third data set.
4. The multi-dimensional disease dietary recommendation method according to claim 3, characterized in that: The pre-built disease diet knowledge base also includes: Performing semantic vectorization on the third data set, and calculating semantic similarity between text vectors of the third data set; De-duplication of the third data set, and optimization of the third data set using a hierarchical clustering algorithm; The third data set is vectorized and vectorized through built-in functions to construct a disease diet knowledge base.
5. The multi-dimensional disease dietary recommendation method according to claim 1, characterized in that: The integrating the user query instruction and the node and / or the document content, inputting into the large language model, and generating a response includes: The large language model is pre-connected to an external Internet knowledge base; Integrate into context information according to the node and / or the document content, and encode the context information into an embedding; Generating an output sequence embedding based on the tokenized input embedding of the user query instruction and the topic tag embedding; The large language model generates a response based on the embedded context information, the output sequence embedding and the external Internet knowledge base.
6. The multi-dimensional disease dietary recommendation method according to claim 1, characterized in that: The integrating the user query instruction and the node and / or the document content, inputting into the large language model, and generating a response includes: When the user query instruction is a voice sample, the user input instruction is converted into text data in real time by an external Whisper module; the text answer generated by the large language model according to the user input instruction is converted into a voice response through the FastSpeech2 module; When the user query instruction is a text sample, the large language model generates a text response according to the user input instruction.
7. The multidimensional disease dietary recommendation method according to any one of claims 1 to 6, characterized in that: The large language model is deployed online through VLLM so that users can access it through a Web page.
8. The multidimensional disease dietary recommendation method according to any one of claims 1 to 6, characterized in that: The large language model is CLLaMA2-13B, wherein the large language model CLLaMA2-13B is deployed with a topic selection state space model, and the topic selection state space model is used to perform: Performing hierarchical retrieval in the disease diet knowledge base according to the semantic clues in the subject prompter until finding nodes and / or documents related to the user query instruction; Performing hierarchical retrieval in the disease diet knowledge base according to the semantic clues in the subject prompter until finding nodes and / or documents related to the user query instruction; The topic prompter is optimized according to the topic level loss of hierarchical retrieval of the disease diet knowledge base.
9. The multi-dimensional disease dietary recommendation method according to claim 8, characterized in that: The step of optimizing the topic prompter according to the topic level loss of hierarchical retrieval of the disease diet knowledge base comprises: Pre-constructing a hierarchical index, including a primary index structure and a secondary index structure, wherein the primary index structure is used to store domain classifications, and the secondary index structure is used to store specific entries; The topic selection state space model optimizes the topic indicator in the retrieval enhancement generation framework according to the hierarchical loss generated by the hierarchical index.
10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the multidimensional disease dietary recommendation method according to any one of claims 1 to 9.
Citation Information
Patent Citations
A method and device for generating a dialogue sample of a diet plan for a specific disease
CN117725186B
Cited By
Healthy diet knowledge accurate retrieval engine based on large language model
CN120632073A