Knowledge-injection-based data processing methods, terminal devices, and readable storage media
By introducing a dynamic knowledge injection method into the TCM big language model and utilizing multiple retrieval methods to obtain real-time TCM research results, the problem of knowledge lag in the TCM big language model was solved, and the accuracy and reliability of TCM consultation were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINESE MEDICINE GUANGDONG LABORATORY
- Filing Date
- 2025-09-29
- Publication Date
- 2026-06-02
AI Technical Summary
Because the Large Language Model for Traditional Chinese Medicine (LLM) relies on static data for pre-training, it is difficult to obtain the latest research results and policy changes in traditional Chinese medicine in a timely manner, resulting in knowledge lag and low accuracy of output results.
By introducing a dynamic knowledge injection method into the TCM big data model, we can obtain real-time TCM research results related to patient consultations from different channels using various retrieval methods, including online retrieval, database retrieval, and vector retrieval. Combined with historical interaction information, we can generate accurate TCM diagnosis and treatment suggestions.
It improves the accuracy and reliability of TCM consultations, ensuring that the output results are based on real-time and authoritative TCM research findings, and significantly enhances applicability and professionalism.
Smart Images

Figure CN122135922A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence and medical health technology, and in particular relates to a data processing method, terminal device and readable storage medium based on knowledge injection. Background Technology
[0002] Traditional Chinese medicine, a medical treasure passed down for thousands of years, plays a significant role in disease prevention, diagnosis and treatment, and health maintenance. With the development of artificial intelligence, large language models (LLM) based on deep learning have brought possibilities for the intelligent application of traditional Chinese medicine.
[0003] Traditional Chinese medicine (TCM) knowledge is highly specialized and complex. LLM training mainly relies on pre-training with existing large-scale static data, making it difficult to obtain the latest TCM research results, clinical guidelines, or policy changes in a timely manner. This results in a knowledge lag problem, making it difficult for the LLM knowledge system to fit the unique theoretical framework of TCM. This can easily lead to "illusion" output, meaning that the accuracy of the output results is low. Summary of the Invention
[0004] This application provides a data processing method, terminal device, and readable storage medium based on knowledge injection, which can significantly improve the accuracy and reliability of the output results of TCM auxiliary diagnosis and treatment models, as well as their applicability in TCM application scenarios, based on real-time and authoritative TCM research results.
[0005] In a first aspect, embodiments of this application provide a data processing method based on knowledge injection, including: Obtain the first interactive information input by the patient; wherein, the first interactive information is the consultation information of the patient corresponding to the disease; First knowledge information is obtained based on the first interactive information; wherein, the first knowledge information includes real-time research results of traditional Chinese medicine related to the first interactive information; Output the output result corresponding to the first interactive information based on the first knowledge information.
[0006] In this embodiment, patient consultation information regarding their symptoms (first interaction information) is obtained; subsequently, based on this consultation information, the latest relevant TCM research findings (first knowledge information) are obtained; finally, combining this knowledge information, the corresponding consultation result is output. By using this latest and relevant professional knowledge (first knowledge information), the output result is strictly based on real-time and authoritative TCM research findings, ensuring that the output result is supported by clear professional knowledge, significantly improving the accuracy, reliability, and applicability of the answer in TCM application scenarios.
[0007] In one possible implementation of the first aspect, obtaining the first knowledge information based on the first interaction information includes: Obtain the patient's historical interaction information; wherein, the historical interaction information is the consultation information prior to the first interaction. The first interaction information is supplemented based on historical interaction information to obtain the supplemented second interaction information; First knowledge information is obtained based on the second interaction information.
[0008] In this embodiment of the application, supplementing the current consultation content with historical consultation information can make the acquired knowledge information more in line with the patient's complete needs, thereby improving the accuracy and pertinence of TCM consultation responses.
[0009] In one possible implementation of the first aspect, obtaining the first knowledge information based on the second interaction information includes: Perform security checks on the second interactive information; After passing the security check, the information requirement type corresponding to the second interactive information is identified to obtain the first requirement information; If the first requirement information meets the preset conditions, the first knowledge information is obtained based on the second interaction information.
[0010] In this embodiment of the application, the patient consultation is first tested for safety, and then the type of demand is confirmed to be compliant. This ensures that the TCM knowledge information obtained subsequently is safe and compliant with regulations, so that the final response is both safe and in line with the legitimate demand.
[0011] In one possible implementation of the first aspect, obtaining the first knowledge information based on the second interaction information includes: The first sub-information is obtained by performing an online retrieval of the second interactive information; The second interactive information is retrieved from the database to obtain the second sub-information; The third sub-information is obtained by performing vector retrieval on the second interactive information; The first knowledge information is obtained based on the first sub-information, the second sub-information, and the third sub-information.
[0012] In this embodiment, TCM knowledge is obtained through three methods: online retrieval, database retrieval, and vector retrieval. The results are then integrated to form the first knowledge information. This allows the obtained knowledge to cover diverse information (online), authoritative theories (database), and precise needs (vector), ultimately improving the comprehensiveness and accuracy of TCM knowledge information.
[0013] In one possible implementation of the first aspect, the second interactive information is retrieved online to obtain the first sub-information, including: Extract at least one first keyword corresponding to the second interaction information; Based on the first keyword, an online search is performed to obtain a first preset number of web page information; Extract the first text fragment that matches the second interactive information from each webpage; The first sub-information is obtained from multiple first text fragments.
[0014] In this embodiment of the application, by extracting the first keyword of the second interactive information and searching online, the first text fragment matching the preset number of web page information is selected, and then the first sub-information is obtained accordingly. This method can quickly locate the content closely related to the user's inquiry from a massive amount of network resources, making the obtained information diverse and in line with the needs, and greatly improving the breadth and accuracy of the information.
[0015] In one possible implementation of the first aspect, the second interactive information is retrieved from a database to obtain the second sub-information, including: Obtain a first preset database; wherein the first preset database contains multiple second text fragments and at least one second keyword corresponding to each second text fragment; wherein the second text fragments are text fragment information obtained by text segmentation of TCM professional books; Extract at least one third keyword corresponding to the second interaction information; Each third keyword is matched against multiple second keywords in a pre-defined database to obtain the second keywords that match each third keyword. The second text fragments corresponding to each of the multiple second keywords are deduplicated to obtain the second sub-information.
[0016] In this embodiment, a pre-defined database with keywords is constructed based on excerpts from TCM professional books. Then, by extracting the third keyword from the user's consultation, matching the database keywords, and deduplicating and integrating relevant excerpts, authoritative and non-redundant TCM professional knowledge can be efficiently obtained, ensuring the professionalism and accuracy of the second sub-information.
[0017] In one possible implementation of the first aspect, vector retrieval is performed on the second interactive information to obtain the third sub-information, including: Obtain a second preset database; wherein the second preset database contains multiple first vector data; wherein the first vector data is vector data obtained after vector transformation of the third text fragment; wherein the third text fragment is text fragment information obtained by text segmentation of TCM professional books; Convert the second interactive information into second vector data; Calculate the similarity between the second vector data and each first vector data in the second preset database; The third text fragment corresponding to the first vector data with a similarity higher than a preset threshold is identified as the third sub-information.
[0018] In this embodiment, by converting fragments from TCM professional books into vectors to construct a second preset database, and then converting user inquiries into vectors and matching them with knowledge fragments corresponding to highly similar vectors in the database, the limitations of keywords can be overcome, and semantically relevant TCM professional knowledge can be accurately located, ensuring the accuracy and relevance of the third sub-information.
[0019] In one possible implementation of the first aspect, the first knowledge information is obtained based on the first sub-information, the second sub-information, and the third sub-information, including: Arrange the first sub-information, the second sub-information, and the third sub-information in descending order of their relevance to the second interactive information to obtain the first knowledge information after sequential arrangement. Obtain a second preset number of second knowledge information from the first knowledge information in the order of arrangement; The output result corresponding to the second interactive information is obtained based on the second knowledge information.
[0020] In this embodiment of the application, by sorting the three types of sub-information according to their relevance to the user's consultation, extracting the core knowledge and organizing it into the output results, the final TCM consultation response can highlight the key points and be logically clear, ensuring the core of the information while improving the user's reading and comprehension efficiency.
[0021] In a second aspect, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a knowledge injection-based data processing method as described in any of the first aspects above.
[0022] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a knowledge injection-based data processing method as described in any of the first aspects above.
[0023] Fourthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the knowledge injection-based data processing method described in any of the first aspects above.
[0024] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart illustrating the data processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the process for obtaining injection information provided in the embodiments of this application. Figure 1 ; Figure 3 This is a schematic diagram of the process for obtaining injection information provided in the embodiments of this application. Figure 2 ; Figure 4 This is a schematic diagram of the process for obtaining injection information provided in the embodiments of this application. Figure 3 ; Figure 5 This is a schematic diagram of the process for obtaining the first sub-information provided in an embodiment of this application; Figure 6 This is a schematic diagram of the process for obtaining the second sub-information provided in an embodiment of this application; Figure 7 This is a schematic diagram of the process for obtaining third sub-information provided in an embodiment of this application; Figure 8 This is a schematic diagram of the overall structure of the data processing method provided in the embodiments of this application; Figure 9 This is a schematic diagram of the query optimization module structure provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of the knowledge retrieval module provided in an embodiment of this application; Figure 11 This is a schematic diagram of the structure of the knowledge fusion module provided in an embodiment of this application; Figure 12 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation
[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0028] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0029] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0030] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0031] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0033] Traditional Chinese medicine, a medical treasure passed down for thousands of years, plays a significant role in disease prevention, diagnosis and treatment, and health maintenance. With the development of artificial intelligence, large language models (LLM) based on deep learning have brought possibilities for the intelligent application of traditional Chinese medicine.
[0034] Traditional Chinese medicine (TCM) knowledge is highly specialized and complex. LLM training mainly relies on pre-training with existing large-scale static data, making it difficult to obtain the latest TCM research results, clinical guidelines, or policy changes in a timely manner. This results in a knowledge lag problem, making it difficult for the LLM knowledge system to fit the unique theoretical framework of TCM. This can easily lead to "illusion" output, meaning that the accuracy of the output results is low.
[0035] To address the aforementioned technical issues, this application incorporates a dynamic knowledge injection method into the original large-scale model in the field of traditional Chinese medicine. Before the model generates an answer, knowledge related to the question is dynamically acquired from different channels through various retrieval methods. This knowledge is then input into the model along with the original question, thereby improving the accuracy and reliability of the answer.
[0036] See Figure 1 This is a flowchart illustrating a data processing method based on knowledge injection provided in an embodiment of this application. It is intended as an example and not a limitation. The method may include the following steps: S101, Obtain the first interactive information input by the patient; wherein, the first interactive information is the consultation information of the patient corresponding to the disease.
[0037] In this embodiment of the application, the system first collects the information input by the patient (this is the "first interaction information"), and this information must be limited to the scope of "the patient's TCM-related consultation information for their own or related objects' symptoms, such as consultation on TCM diagnosis, treatment methods, prescription recommendations, etc. for a certain symptom, and does not include content unrelated to the symptom.
[0038] S102, Obtain first knowledge information based on the first interactive information; wherein, the first knowledge information includes real-time research results of traditional Chinese medicine related to the first interactive information.
[0039] In this embodiment of the application, "first interactive information" refers to the patient's consultation information (such as "asthma treatment"). "First knowledge information" refers to the set of relevant knowledge that the system selects and extracts from knowledge sources in the field of traditional Chinese medicine around the consultation, and must include "real-time research results of traditional Chinese medicine directly related to the disease" (such as the latest clinical research conclusions on asthma published in recent years).
[0040] This type of real-time research information not only ensures that the knowledge is relevant to the patients' consultation needs, but also solves the problem of "knowledge lag" in traditional large models, providing timely and professional support for generating accurate answers.
[0041] S103, output the output result corresponding to the first interaction information based on the first knowledge information.
[0042] In this embodiment, "outputting the output result corresponding to the first interactive information based on the first knowledge information" is essentially the process by which the TCM big data model combines "professional knowledge" with "patient needs" to generate the final answer. This process does not directly generate an answer based on the patient's original consultation, but rather uses "first knowledge information" as the sole core basis to ensure that the answer strictly relies on real-time, authoritative knowledge in the field of TCM, avoiding the problems of "illusory output" or "knowledge lag" in traditional big data models.
[0043] The generated "output results" must precisely correspond to the "first interactive information," meaning the answer must be entirely focused on the patient's consultation regarding their condition (e.g., if the patient inquires about "asthma management," the output results should focus on TCM management methods for asthma, solutions supported by the latest research, etc.), without deviating from the patient's core needs. Furthermore, the output results must conform to TCM professional standards and be timely and practical (e.g., "For elderly patients with long-term asthma, TCM often uses the method of warming the lungs and resolving phlegm. According to the latest research in the 2023 'Guidelines for the Diagnosis and Treatment of Respiratory Diseases in Traditional Chinese Medicine,' modifications can be made to the classic formula Xiao Qinglong Tang..."). This directly addresses the patient's consultation needs regarding their condition.
[0044] In the above method, patient consultation information regarding their symptoms is obtained (first interaction information); subsequently, based on this consultation information, the latest relevant TCM research findings are obtained (first knowledge information); finally, these knowledge information are combined to output the corresponding consultation results. By using this latest and relevant professional knowledge (first knowledge information), the output results are strictly based on real-time and authoritative TCM research findings, ensuring that the output results are supported by clear professional knowledge, significantly improving the accuracy, reliability, and applicability of the answers in TCM application scenarios.
[0045] In one embodiment, see Figure 2 This is a schematic diagram of the process for obtaining injection information provided in the embodiments of this application. Figure 1 ,like Figure 2 As shown, step S102 includes: S201, Obtain the patient's historical interaction information; wherein, the historical interaction information is the consultation information prior to the first interaction information.
[0046] In this embodiment, after obtaining the consultation information (first interaction information) input by the patient, it is necessary to collect the consultation content that the patient had already entered when interacting with the system before this input of the "first interaction information" (i.e., the current symptom consultation information). This historical interaction information must be related to the symptom consultation and be earlier than the first interaction information in time. Its purpose is to provide data support for the subsequent "using historical sessions to rewrite the current problem and make up for information gaps", ensuring that the first interaction information is more complete and clear.
[0047] S202, supplement the first interaction information based on the historical interaction information to obtain the supplemented second interaction information.
[0048] In this embodiment of the application, "historical interaction information" refers to the disease-related consultation content provided by the patient when interacting with the system before inputting "first interaction information" (i.e., current symptom consultation information); "first interaction information" refers to the original symptom consultation currently input by the patient (which may contain missing information, such as only saying "the symptoms are more severe" without specifying the specific meaning of "symptoms").
[0049] The system retrieves key information from historical interactions (such as the patient's previous mention of "dry mouth and sore throat") and adds it to the first interaction, filling the information gaps in the original input and ultimately forming the "second interaction." For example, "These symptoms have worsened" is supplemented to "The dry mouth and sore throat symptoms you mentioned before are now worse," ensuring that the second interaction can still be clearly understood even outside the original dialogue context. This lays the foundation for the subsequent knowledge retrieval module to accurately acquire TCM-related knowledge.
[0050] S203, Obtain the first knowledge information based on the second interactive information.
[0051] In the embodiments of this application, the "second interactive information" is the patient's symptom consultation information supplemented by historical interactive information (such as supplementing the original vague "the symptoms have become more severe" to "the dry mouth and sore throat symptoms you mentioned before have become more severe"). It has completeness and clarity, and can be understood independently of the dialogue context. It is a precise "demand-oriented" knowledge acquisition.
[0052] After obtaining the second interactive information, the system will use it as the core basis to filter and extract relevant knowledge in the field of traditional Chinese medicine through a multi-channel retrieval and integration mechanism, and finally form the "first knowledge information".
[0053] In the above methods, supplementing the current consultation content with historical consultation information can make the acquired knowledge and information more aligned with the patient's complete needs, thereby improving the accuracy and relevance of TCM consultation responses.
[0054] In one embodiment, see Figure 3 This is a schematic diagram of the process for obtaining injection information provided in the embodiments of this application. Figure 2 ,like Figure 3 As shown, step S203 includes: S301, perform security checks on the second interactive information.
[0055] In this embodiment, the second interactive information is patient consultation information supplemented by historical interactive information (which already possesses completeness and independent comprehensibility). When the system performs security checks on this information, it primarily achieves this through two dimensions: "sensitive word filtering" and "intent recognition compliance assessment," including: Sensitive word filtering detection: The system calls upon the collected "sensitive word database" to comprehensively scan the text content of the second interaction information to check for sensitive words, prohibited words, or interfering information unrelated to the disease consultation. If any illegal content is detected (such as involving prohibited Chinese medicine names, malicious treatment requests, or non-medical related illegal expressions), it is judged as "not passing" and the system directly "refuses to answer," preventing the second interaction information from entering the subsequent knowledge retrieval stage; if no illegal content is detected and the focus is only on the disease consultation (such as "You mentioned dry mouth and sore throat symptoms before, which are now more severe, and you want to know how to treat them with traditional Chinese medicine"), then this round of detection passes.
[0056] Intent recognition compliance assessment: After sensitive word filtering passes, the system further detects the core intent of the second interaction information. On the one hand, it determines whether it falls within the reasonable service scope of the TCM model (i.e., whether it is a TCM-related consultation intent such as disease diagnosis, prescription recommendation, or health guidance); on the other hand, it assesses whether the intent complies with TCM diagnosis and treatment safety standards (e.g., no excessive medical consultation, no unreasonable demands beyond the scope of TCM theory). If the intent does not comply with the standards (e.g., "wanting to quickly cure hypertension through TCM folk remedies, regardless of risks"), it is judged as "not passing," triggering a refusal to answer; only when the intent focuses on legitimate and compliant TCM disease consultation is the second interaction information security test confirmed to have passed, allowing entry into the subsequent "obtaining first knowledge information based on the second interaction information" stage.
[0057] S302, after the security check is passed, the information requirement type corresponding to the second interactive information is identified to obtain the first requirement information.
[0058] In this embodiment, the second interactive information is complete disease consultation information supplemented by historical interactive information and passing security checks. It already possesses "compliance + completeness," but the patient's core demand type, i.e., "information demand type," needs further clarification. The system, combined with service scenarios in the field of Traditional Chinese Medicine (TCM), deconstructs and analyzes the textual semantics and core demands of the second interactive information to determine its demand category. These categories all revolve around TCM disease-related services, and typical demand types include, but are not limited to: Disease diagnosis type of requirement: For example, if the second interactive information is "recurrent diarrhea symptoms after supplementation, what disease does traditional Chinese medicine judge it to be", the requirement type is "disease diagnosis".
[0059] Requests for prescription recommendations: such as "My dry mouth and sore throat are getting worse, and I want to know how to treat it with traditional Chinese medicine". If the core request is for a specific prescription, the request type is "prescription recommendation".
[0060] Health and wellness guidance needs: such as "How can elderly people with hypertension prevent recurrence through diet after treatment with traditional Chinese medicine?" The need type is "health and wellness guidance".
[0061] Through the above identification, the system clarifies the core demand type corresponding to the second interactive information as "first demand information" (such as "recommendation of traditional Chinese medicine treatment for worsening dry mouth and sore throat"). The core function of the first demand information is to provide precise guidance for the subsequent "obtaining first knowledge information based on the second interactive information," ensuring that the subsequently generated first knowledge information can accurately support the patient's needs.
[0062] S303, if the first requirement information meets the preset conditions, the first knowledge information is obtained based on the second interaction information.
[0063] In this embodiment, the first requirement information must fall within the preset service scope of the TCM big data model, that is, it must revolve around TCM-related needs such as "disease diagnosis, prescription recommendation, and health guidance," excluding non-TCM treatment needs (such as "market price inquiry for Chinese medicinal materials"). It must meet the safety standards of TCM treatment and must not involve excessive medical treatment or illegal requests (such as "seeking risk-free and quick cures for chronic diseases," which are beyond the scope of reasonable treatment and are judged as not meeting the preset conditions).
[0064] Only if both of the above conditions are met simultaneously will the first requirement information be confirmed as meeting the preset conditions, allowing entry into the subsequent knowledge acquisition stage; if it does not meet the conditions, the system will trigger a refusal to answer and will not conduct knowledge retrieval.
[0065] The above method involves first conducting a safety test on the patient's consultation and then confirming whether the type of request is compliant. This ensures that the TCM knowledge information obtained subsequently is safe and compliant with regulations, thus making the final response both safe and in line with legitimate needs.
[0066] In one embodiment, see Figure 4 This is a schematic diagram of the process for obtaining injection information provided in the embodiments of this application. Figure 3 ,like Figure 4 As shown, step S303 includes: S401, perform online retrieval of the second interactive information to obtain the first sub-information.
[0067] In this embodiment, online retrieval refers to web retrieval. The system inputs the question (first interactive information) into the Baidu Search API for online retrieval. This is because online information is abundant and updated promptly, allowing access to the latest TCM research findings, clinical experience sharing, and other materials, which helps provide users with cutting-edge knowledge. See steps S501-S504 for details.
[0068] In one embodiment, see Figure 5 This is a schematic diagram of the process for obtaining the first sub-information provided in an embodiment of this application, such as... Figure 5 As shown, step S401 includes: S501, extract at least one first keyword corresponding to the second interaction information.
[0069] In this application embodiment, the "second interactive information" is complete disease consultation information supplemented by historical interactive information and security checks (such as "the symptoms of dry mouth and sore throat you mentioned before are now more severe, and you want to know about TCM treatment methods"), which has completeness and clarity; the "first keyword" is a word extracted from this type of information that can accurately represent the core needs, and it needs to revolve around the two core elements of "disease" and "consultation intent" to ensure that it can support subsequent knowledge retrieval.
[0070] For example, for "consultation on TCM treatment for worsening dry mouth and sore throat", "dry mouth and sore throat" and "TCM treatment" can be extracted as the first keywords (at least one, usually 3-5 core keywords). This avoids the problems of low efficiency and results deviating from the core caused by using complete information retrieval, and ensures that TCM knowledge highly related to the second interactive information can be quickly obtained, which is in line with the technical goal of "improving the accuracy and efficiency of knowledge retrieval" in the document.
[0071] S502, perform an online search based on the first keyword to obtain a first preset number of web page information.
[0072] In this embodiment, the system selects Baidu Search as the online search tool and performs a search by inputting the first keyword through the Baidu Search API. This is because Baidu possesses a vast amount of web page data, enabling the acquisition of rich information resources. The system typically retrieves the first 100 search results (a first preset number) as the web page information obtained in this search. This setting ensures both the richness of the information obtained and, to a certain extent, controls the amount of data, avoiding excessive invalid information from interfering with subsequent processing.
[0073] S503, extract the first text fragment that matches the second interactive information from each webpage information.
[0074] In this embodiment, after the system obtains the webpage, it parses the webpage content and extracts text fragments highly related to the first keyword. During this process, information unrelated to the keyword, such as advertisements and irrelevant topic content, is filtered out, retaining only information closely related to "TCM treatment for insomnia," such as TCM's dialectical analysis of insomnia and specific treatment methods.
[0075] S504, obtain the first sub-information based on multiple first text fragments.
[0076] In this embodiment of the application, obtaining the first sub-information from multiple first text fragments typically involves steps such as text integration, deduplication, semantic understanding, and feature extraction. The specific details are as follows: Text integration: This involves merging multiple primary text fragments to form a relatively complete text. The fragments can be directly connected sequentially. For example, if there are three primary text fragments: "Traditional Chinese medicine believes that insomnia may be caused by deficiency of both the heart and spleen," "Symptoms of deficiency of both the heart and spleen also include loss of appetite," and "For insomnia caused by deficiency of both the heart and spleen, Gui Pi Tang can be used for treatment," then the integrated text would be: "Traditional Chinese medicine believes that insomnia may be caused by deficiency of both the heart and spleen, and symptoms of deficiency of both the heart and spleen also include loss of appetite. For insomnia caused by deficiency of both the heart and spleen, Gui Pi Tang can be used for treatment."
[0077] Deduplication: Remove duplicate content from the integrated text to avoid information redundancy. For example, if two first text segments both mention "a common symptom of insomnia is difficulty falling asleep," only one instance will be retained.
[0078] Semantic understanding and analysis: Natural language processing techniques are used to perform semantic analysis on the integrated text. Lexical analysis, syntactic analysis, and semantic role labeling can be employed to understand the meaning of each word and sentence in the text, as well as the relationships between them. For example, analysis reveals that in the phrase "Gui Pi Tang for treating insomnia," "Gui Pi Tang" is a specific formula for treating "insomnia," demonstrating a relationship between the two: a treatment method and a symptom.
[0079] Feature Extraction and Summarization: Based on the results of semantic analysis, key information and features are extracted to form the first sub-information. Core information such as symptoms, causes, and treatments can be extracted. For example, in the above example, "insomnia," "deficiency of both heart and spleen," and "Gui Pi Tang" can be extracted as key content of the first sub-information. Alternatively, the text can be summarized and generalized according to specific needs to obtain a more concise first sub-information, such as "Traditional Chinese medicine believes that deficiency of both heart and spleen can lead to insomnia, which can be treated with Gui Pi Tang."
[0080] In addition, text classification techniques can be used to categorize multiple first text fragments according to different themes or categories, and then key information can be extracted for each category to form the first sub-information. Alternatively, machine learning models can be used to train and predict multiple first text fragments to obtain key information and summaries about these text fragments, which can then be used as the first sub-information.
[0081] In the above method, by extracting the first keyword of the second interactive information and searching online, the first text fragment matching the preset number of web page information is selected, and then the first sub-information is obtained. This method can quickly locate content closely related to the user's inquiry from a massive amount of online resources, making the obtained information diverse and in line with the needs, and greatly improving the breadth and accuracy of the information.
[0082] S402, perform a database retrieval on the second interactive information to obtain the second sub-information.
[0083] In this embodiment, based on the complete symptom consultation (second interactive information, such as "TCM treatment consultation for worsening dry mouth and sore throat") after supplementation and security testing, matching information is retrieved from the pre-built TCM professional database of the system. The final TCM basic professional knowledge directly related to the consultation need is the second sub-information. See steps S601-S604 for details.
[0084] In one embodiment, see Figure 6 This is a schematic diagram of the process for obtaining the second sub-information provided in an embodiment of this application, such as... Figure 6 As shown, step S402 includes: S601, Obtain a first preset database; wherein the first preset database contains multiple second text fragments and at least one second keyword corresponding to each second text fragment; wherein the second text fragments are text fragment information obtained by text segmentation of TCM professional books.
[0085] In this embodiment, the system first obtains a pre-built "first preset database". The core content of this database consists of two parts: first, "multiple second text fragments". These fragments are not random text, but fragments containing TCM professional knowledge obtained by text segmentation of TCM professional books (such as classic texts, treatment guidelines, etc.) (e.g., excerpts from the "Differentiation and Treatment of Dry Mouth and Sore Throat" chapter in a TCM book); second, "at least one second keyword" matched for each second text fragment (e.g., matching the fragment "Differentiation and Treatment of Dry Mouth and Sore Throat" with keywords such as "dry mouth and sore throat" and "differentiation of syndromes"). This facilitates the subsequent quick retrieval of corresponding TCM professional knowledge fragments through keywords, providing an accurate and authoritative knowledge source for the database retrieval of second interactive information (obtaining second sub-information).
[0086] Specifically, sqlite3 is used to build a text retrieval database, and the segmented text (second text fragment) and its corresponding keywords (second keywords) are stored in the database (first preset database) to construct a structured TCM knowledge base for subsequent retrieval.
[0087] S602, extract at least one third keyword corresponding to the second interaction information.
[0088] In this embodiment, the system extracts at least one "third keyword" from the second interactive information. These keywords include not only entity words directly appearing in the information (such as "elderly," "cough," and "TCM cough remedy"), but also synonyms and near-synonyms of these entity words (such as "cough asthma" and "coughing disease" for "cough," and "TCM cough remedy" for "TCM cough remedy"). The aim is to expand the coverage of subsequent database searches, ensuring accurate matching of relevant TCM professional knowledge fragments in the first preset database, and providing a more comprehensive search basis for generating the second sub-information.
[0089] S603, match each third keyword with multiple second keywords in the preset database; obtain the second keyword that matches each third keyword.
[0090] In this embodiment, the system compares each third keyword extracted from the second interactive information (e.g., "cough," "asthma," "TCM cough relief") with all the second keywords (i.e., tags corresponding to fragments from TCM professional books, such as "cough," "coughing disease," "TCM cough relief methods") in the first preset database. Through semantic similarity or word matching rules, the system finds the second keywords that correspond to each third keyword—for example, "cough" matches "cough" and "coughing disease" in the database, and "TCM cough relief" matches "TCM cough relief methods." Finally, the system obtains the second keywords in the database corresponding to each third keyword, laying the foundation for subsequent location of relevant TCM knowledge fragments (second text fragments).
[0091] S604, deduplicate the second text fragments corresponding to the multiple second keywords to obtain the second sub-information.
[0092] In this embodiment, after the keyword matching described above, each successfully matched second keyword (from the first preset database) corresponds to a second text fragment carrying TCM professional knowledge (such as "cough diagnosis and treatment with Xing Su San" and "cough treatment requires lung-clearing"). However, these fragments may contain duplicate content (for example, different second keywords correspond to fragments with similar expressions such as "cough diagnosis"). At this time, the system will perform deduplication on these scattered second text fragments, removing duplicate or highly similar content. The final integrated set of TCM professional knowledge without redundancy is the second sub-information, which provides authoritative and concise basic theoretical support for subsequent fusion with the first sub-information.
[0093] The above method constructs a pre-defined database with keywords based on excerpts from TCM professional books, and then extracts the third keyword from the user's consultation to match the database keywords, deduplicates and integrates relevant excerpts. This can efficiently obtain authoritative and non-redundant TCM professional knowledge, ensuring the professionalism and accuracy of the second sub-information.
[0094] S403, perform vector retrieval on the second interactive information to obtain the third sub-information.
[0095] In this embodiment of the application, vector retrieval first converts the patient's clear symptom consultation (second interactive information) into "data code" (vector) that the computer can recognize, and then uses this "code" to find the professional knowledge in the TCM knowledge base that is closest in meaning to it. The knowledge found is the third sub-information, specifically referring to steps S701-S704.
[0096] In one embodiment, see Figure 7 This is a schematic diagram of the process for obtaining third sub-information provided in an embodiment of this application, such as... Figure 7 As shown, step S403 includes: S701, Obtain the second preset database; wherein the second preset database contains multiple first vector data; wherein the first vector data is vector data obtained after vector transformation of the third text fragment; wherein the third text fragment is text fragment information obtained by text segmentation of TCM professional books.
[0097] In this embodiment, the second preset database refers to a pre-built database specifically for vector retrieval, which is called by the system. It serves as the core knowledge base for semantic matching in the TCM big data model, complementing the "first preset database" used for keyword retrieval mentioned earlier. The core content stored in this database is not text, but "first vector data," which is the transformation of TCM knowledge in textual form into numerical sequences (vectors) that can be calculated and compared by a computer. These vectors accurately convey the "semantic meaning" of TCM knowledge, rather than simply the surface information of the text.
[0098] The first vector data is obtained by segmenting the text of 100 professional books on traditional Chinese medicine (such as classics like "Huangdi Neijing" and "Shanghan Lun" and modern TCM diagnosis and treatment guidelines). The resulting fragmented professional knowledge (such as "the key points of differentiation of wind-cold cough" or "the lung-warming effect of ginger in dietary therapy") is then transformed into a vector using the Embedding Model in BCEmbedding. The resulting numerical sequence is the "first vector data" and is stored in the second preset database.
[0099] S702, convert the second interactive information into second vector data.
[0100] In this embodiment, the "second interactive information" is a complete TCM symptom consultation that has been supplemented (e.g., supplementing symptom details by combining historical dialogues) and security checks have been performed. "Converting it into second vector data" means that the system uses BCEmbedding to transform the "semantic meaning" of this text into an ordered sequence of numbers (i.e., a vector).
[0101] Simply put, it involves translating the patient's "symptom consultation" into a "digital code" (second vector data) that the computer can "understand and compare," laying the foundation for using this vector to match similar TCM knowledge vectors in the "second preset database."
[0102] S703, calculate the similarity between the second vector data and each first vector data in the second preset database.
[0103] In this embodiment of the application, the similarity between the "second vector data" (i.e., the second interactive information, such as the numerical sequence transformed from "the elderly with a long-term cough due to wind-cold seeking physiotherapy") and all the "first vector data" (i.e. the numerical sequence transformed from fragments of traditional Chinese medicine books, such as the vector corresponding to "moxibustion can be used for cough due to wind-cold at the Feishu acupoint") in the "second preset database" is calculated.
[0104] The system uses a pre-defined algorithm (such as the implicit vector similarity calculation rules in the document) to compare the numerical features of each "second vector data" with those of each "first vector data" one by one. For example, it calculates the "distance" between two vectors (the closer the distance, the higher the similarity) to determine whether the two represent similar semantics. For instance, the vector for "treatment of chronic cough due to wind-cold in the elderly" will have a high similarity with vectors related to "moxibustion for cough due to wind-cold" and "treatment of cough due to cold in the elderly," but a low similarity with the vector for "medication for cough due to wind-heat."
[0105] S704, the third text segment corresponding to the first vector data with a similarity higher than a preset threshold is determined as the third sub-information.
[0106] In this embodiment of the application, each "first vector data" in the second preset database corresponds to an original "third text fragment" (i.e., a knowledge fragment cut from a professional TCM book, such as "moxibustion can be applied to the Feishu acupoint for cough due to wind-cold"). The similarity between the "second vector data" (the vector converted from user consultation) and all "first vector data" in the database has been calculated previously.
[0107] Next, the filtering rules are executed: the system sets a "preset threshold" (e.g., similarity ≥ 0.8, the specific value is determined by model training), and filters out all "first vector data" with similarity exceeding this threshold. Then, the "third text fragments" corresponding to these filtered vector data are found. These fragments are TCM professional knowledge that highly matches the semantics of the user's needs and are finally identified as "third sub-information".
[0108] In the above method, by converting fragments of TCM professional books into vectors to construct a second preset database, and then converting user inquiries into vectors and matching them with knowledge fragments corresponding to highly similar vectors in the database, it is possible to overcome the limitations of keywords, accurately locate semantically relevant TCM professional knowledge, and ensure the accuracy and relevance of the third sub-information.
[0109] S404, obtain the first knowledge information based on the first sub-information, the second sub-information, and the third sub-information.
[0110] In this embodiment, the first sub-information comes from online web page searches, focusing on the latest or diverse treatment suggestions; the second sub-information comes from searches of professional TCM databases, focusing on classical texts and theories; and the third sub-information comes from vector searches, focusing on semantically relevant personalized knowledge. Integrating these three sources ensures that the first piece of knowledge is both authoritative and comprehensive, while also meeting the specific needs of the user.
[0111] The above methods acquire TCM knowledge through online retrieval, database retrieval, and vector retrieval, and then integrate them to form primary knowledge information. This allows the acquired knowledge to cover diverse information (online), authoritative theories (database), and precise needs (vector), ultimately improving the comprehensiveness and accuracy of TCM knowledge information.
[0112] In one embodiment, step S404 includes: Arrange the first, second, and third sub-information in descending order of their relevance to the second interactive information to obtain the first knowledge information in the ordered order.
[0113] In this embodiment, the relevance of the first sub-information (online webpage knowledge), the second sub-information (classical database knowledge), and the third sub-information (vector matching knowledge) to the user's symptom consultation (second interactive information) is first determined. Then, the Reranker Model in BCEmbedding is used to rearrange these three types of knowledge according to "relevance from high to low". The resulting ordered set of knowledge is the first knowledge information after sequential arrangement, which can make the subsequent replies to the user more in line with the core needs, with the key content placed first.
[0114] The above method sorts the three types of sub-information according to their relevance to the user's consultation, extracts the core knowledge, and organizes it into the output results. This makes the final TCM consultation response highlight the key points and make the logic clear, ensuring the core of the information while improving the user's reading and comprehension efficiency.
[0115] In one embodiment, step S103 includes: Obtain a second preset number of second knowledge information from the first knowledge information in the order of arrangement; obtain the output result corresponding to the second interactive information based on the second knowledge information.
[0116] In this embodiment of the application, in the "first knowledge information sorted by relevance", according to the pre-set "second preset quantity" (such as taking the first 5 or the first 3, the specific quantity is determined according to the response length requirements), the core knowledge fragments obtained are the "second knowledge information" - which is equivalent to only keeping the most critical and relevant content to the user's needs and eliminating secondary information.
[0117] The second knowledge information is combined with the second interaction information through a carefully designed Prompt template, and this is passed to the large LLM model so that the model can answer the question based on the knowledge. Finally, the model's answer to the question is output and the corresponding output result is output.
[0118] See Figure 8 This is a schematic diagram of the overall structure of the knowledge injection-based data processing method provided in the embodiments of this application, as shown below. Figure 8 As shown, it includes a query optimization module, a knowledge retrieval module, and a knowledge fusion module.
[0119] After obtaining the question (the first interactive information), the system first uses the preprocessing module, namely the query optimization module, to process the question, including contextual analysis and keyword detection. Then, the processed question is given to the knowledge retrieval module for knowledge retrieval. Finally, the retrieved knowledge and the question are processed together by the post-processing module, namely the knowledge fusion module, to obtain the final answer.
[0120] See Figure 9 This is a schematic diagram of the query optimization module structure provided in an embodiment of this application, as shown below. Figure 9 As shown, the specific optimization process includes: After receiving input from the user or patient, the system first rewrites the question using historical conversations. The main purpose of rewriting is to fill in the information gaps in the user's expression, enhance the completeness and clarity of the question, and ensure that the question remains independently comprehensible even after being removed from the original dialogue context, so as to facilitate the subsequent knowledge retrieval module to obtain relevant information.
[0121] Next, the collected sensitive word database is used to filter and check the questions to ensure they do not contain sensitive words, prohibited words, or other distracting information, thus reducing potential risks. Following this, the intent recognition module determines the intent of the question. This process aims to deeply understand the user's potential intent, determine its task category (such as disease diagnosis, medication recommendation, health guidance, etc.), and simultaneously assess whether the question's intent complies with system usage guidelines and security standards. Once the question has passed all the above checks, the information is sent to the knowledge retrieval module for knowledge retrieval. Questions that fail keyword filtering and intent recognition will be rejected.
[0122] See Figure 10 This is a schematic diagram of the structure of the knowledge retrieval module provided in an embodiment of this application, as shown below. Figure 10 As shown, the specific recall process includes: After obtaining the information from the preprocessing module, the system first retrieves knowledge information highly relevant to the question from multiple data sources through network retrieval, database retrieval, and vector retrieval. After integrating and rearranging the retrieved knowledge, the most effective knowledge for answering the question is selected and fed to the model. The model then uses the retrieved relevant knowledge to answer the question, enabling it to combine real-time and authoritative knowledge in the field of traditional Chinese medicine to provide the answer.
[0123] Web search During the online search process, the system first inputs the question into the Baidu Search API for online retrieval and retrieves the top 100 search results. Then, the system parses the webpage content and uses a model to extract text fragments highly relevant to the question, serving as knowledge extracted from individual webpages. For each retrieved webpage, the system extracts relevant knowledge and uses it as external knowledge for recall. During the webpage retrieval process, the system does not use the complete question for searching; instead, it first uses the model to extract 3-5 core keywords from the question and then uses only these keywords for searching. This is to avoid the search engine assigning excessive weight to secondary information in the question, thus preventing the search results from deviating from the core of the question.
[0124] Database retrieval For database retrieval, we collected and organized over 100 professional books in the field of Traditional Chinese Medicine (TCM), converted them to text format (txt), and stored them locally. Then, we used LangChain's text segmenter to segment the book text and extracted 10 core keywords from each text segment using a large model. Next, we used sqlite3 to build a text retrieval database, storing the segmented text and its corresponding keywords to construct a structured TCM knowledge base for subsequent retrieval. After a question is input, the system uses a model to extract entity words and their synonyms and near-synonyms from the question, and uses these related words for database queries. Finally, the system integrates and deduplicates the retrieved relevant text to ensure efficient and accurate knowledge retrieval, providing a foundation for subsequent knowledge rearrangement.
[0125] 3) Vector retrieval In the vector retrieval process, we also used the collected TCM professional books and segmented the text. Then, we used the Embedding Model in BCEmbedding to convert each text segment into vectors and stored them in the vector database. After obtaining the question, the system also used BCEmbedding to vectorize the question and retrieved the several texts most similar to the question through similarity retrieval as effective knowledge sources.
[0126] Knowledge Reordering During the knowledge reordering stage, the system uses the Reranker Model from BCEmbedding to reorder the recalled relevant knowledge items, ensuring that the most relevant content is prioritized for the main model. Finally, the system selects the top-ranked knowledge items, merges them as the final extracted knowledge, and inputs it along with the question into the model, requiring it to answer strictly based on the extracted knowledge. This avoids answer bias caused by incomplete knowledge, and also provides more reference information to the model when some knowledge contains errors, improving its robustness and the reliability of the answer. See Figure 11 This is a structural diagram of the knowledge fusion module provided in an embodiment of this application, as shown below. Figure 11 As shown, it specifically includes: In the knowledge fusion module, the system combines the relevant knowledge retrieved by the knowledge retrieval module with the questions through a carefully designed Prompt template, and then passes this to the model so that the model can answer the questions based on the knowledge. Finally, the system outputs the model's answer to the questions.
[0127] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0128] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0129] Figure 12 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. For example... Figure 12 As shown, the terminal device 12 of this embodiment includes: at least one processor 120 ( Figure 12 (Only one is shown in the diagram) a processor, a memory 121, and a computer program 122 stored in the memory 121 and executable on at least one processor 120. When the processor 120 executes the computer program 122, it implements the steps in any of the above-described embodiments of the knowledge injection-based data processing methods.
[0130] The terminal device can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. This terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 12 This is merely an example of terminal device 12 and does not constitute a limitation on terminal device 12. It may include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0131] The processor 120 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0132] In some embodiments, memory 121 may be an internal storage unit of terminal device 12, such as a hard disk or memory of terminal device 12. In other embodiments, memory 121 may be an external storage device of terminal device 12, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on terminal device 12. Furthermore, memory 121 may include both internal and external storage units of terminal device 12. Memory 121 is used to store operating system, application programs, boot loader, data, and other programs, such as program code of computer programs. Memory 121 may also be used to temporarily store data that has been output or will be output.
[0133] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.
[0134] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.
[0135] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0136] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0137] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0138] In the embodiments provided in this application, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A data processing method based on knowledge injection, characterized in that, The method includes: Obtain the first interactive information input by the patient; wherein, the first interactive information is consultation information for the patient's corresponding ailment; First knowledge information is obtained based on the first interactive information; wherein, the first knowledge information includes real-time research results of traditional Chinese medicine related to the first interactive information; Output the output result corresponding to the first interaction information based on the first knowledge information.
2. The data processing method based on knowledge injection as described in claim 1, characterized in that, The step of obtaining the first knowledge information based on the first interaction information includes: Obtain the patient's historical interaction information; wherein, the historical interaction information is the consultation information prior to the first interaction information; The first interaction information is supplemented based on the historical interaction information to obtain the supplemented second interaction information; The first knowledge information is obtained based on the second interaction information.
3. The data processing method based on knowledge injection as described in claim 2, characterized in that, The step of obtaining the first knowledge information based on the second interaction information includes: Perform security checks on the second interactive information; After the security check is passed, the information request type corresponding to the second interactive information is identified to obtain the first request information; If the first requirement information meets the preset conditions, the first knowledge information is obtained based on the second interaction information.
4. The data processing method based on knowledge injection as described in claim 3, characterized in that, The step of obtaining the first knowledge information based on the second interaction information includes: The second interactive information is retrieved online to obtain the first sub-information; The second interactive information is retrieved from the database to obtain the second sub-information; Vector retrieval is performed on the second interactive information to obtain the third sub-information; The first knowledge information is obtained based on the first sub-information, the second sub-information, and the third sub-information.
5. The data processing method based on knowledge injection as described in claim 4, characterized in that, The online retrieval of the second interactive information to obtain the first sub-information includes: Extract at least one first keyword corresponding to the second interaction information; Based on the first keyword, an online search is performed to obtain a first preset number of web page information; Extract a first text fragment that matches the second interactive information from each of the web page information; The first sub-information is obtained from multiple first text fragments.
6. The data processing method based on knowledge injection as described in claim 4, characterized in that, The step of performing a database retrieval on the second interactive information to obtain the second sub-information includes: Obtain a first preset database; wherein the first preset database contains a plurality of second text fragments and at least one second keyword corresponding to each second text fragment; wherein the second text fragments are text fragment information obtained by text segmentation of TCM professional books; Extract at least one third keyword corresponding to the second interaction information; Each of the third keywords is matched with a plurality of second keywords in the preset database; to obtain the second keyword that matches each of the third keywords; The second text fragments corresponding to each of the multiple second keywords are deduplicated to obtain the second sub-information.
7. The data processing method based on knowledge injection as described in claim 6, characterized in that, The step of performing vector retrieval on the second interactive information to obtain the third sub-information includes: Obtain a second preset database; wherein the second preset database contains multiple first vector data; wherein the first vector data is vector data obtained by vector transformation of the third text fragment; wherein the third text fragment is text fragment information obtained by text segmentation of TCM professional books; The second interactive information is converted into second vector data; Calculate the similarity between the second vector data and each of the first vector data in the second preset database; The third text fragment corresponding to the first vector data whose similarity is higher than a preset threshold is determined as the third sub-information.
8. The data processing method based on knowledge injection as described in claim 7, characterized in that, The step of obtaining the first knowledge information based on the first sub-information, the second sub-information, and the third sub-information includes: The first sub-information, the second sub-information, and the third sub-information are arranged in descending order of their relevance to the second interaction information to obtain the first knowledge information in the ordered order.
9. The data processing method based on knowledge injection as described in claim 8, characterized in that, The step of outputting the output result corresponding to the first interaction information based on the first knowledge information includes: Obtain a second preset number of second knowledge information from the first knowledge information in the order of arrangement; The output result corresponding to the second interactive information is obtained based on the second knowledge information.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 9.