Data processing method and device, storage medium and electronic equipment
By receiving input text, extracting keywords and vectorizing them, and using the knowledge base to construct structured prompt words, the problem of inaccurate answers of large language models is solved, and more accurate answer results are achieved.
Patent Information
- Application Number
- CN202510718688.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-15
AI Technical Summary
When generating answers, large language models cause inaccurate or deviate from the topic because structured prompt words cannot fully cover the information or training data that does not involve specific areas.
By receiving input text, extracting keywords and vectorizing them, using the knowledge base to build structured prompt words, enriching the knowledge data set of the big model, and ensuring that they generate accurate answers.
It improves the accuracy of the answers of the big model, provides an accurate and reliable target knowledge data set, and enhances the accuracy of the answer results of the big model.
Smart Images

Figure CN120492638A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, device, storage medium and electronic device. Background Art
[0002] Large language models (LLMs) have made significant progress in natural language processing over the past few years. These models, trained using deep learning techniques, are capable of performing well on a variety of tasks, such as text generation, machine translation, and question-answering. Large language models are typically trained on large corpora, enabling them to cover a wide range of knowledge domains and possess powerful language understanding and generation capabilities.
[0003] However, the quality of the output is still affected by the training data and the structured prompts used. If the structured prompts provided by the user don't fully cover the required information, or if the dataset used to train the large model doesn't cover the user's desired domain, the large language model may generate inaccurate or off-topic responses. Summary of the Invention
[0004] The technical problem to be solved by this application is to provide a data processing method, device, storage medium and electronic device that can improve the accuracy of the answer results of large models. The specific solution is as follows:
[0005] A data processing method, comprising:
[0006] Receive input text;
[0007] Extract keywords from the input text to obtain a keyword set corresponding to the input text;
[0008] Performing vectorization processing on the keyword set and the input text to obtain a word vector for each keyword in the keyword set and a text vector for the input text;
[0009] Querying a knowledge base according to each of the keywords in the keyword set, the word vector of each of the keywords, and the text vector to obtain a target knowledge data set;
[0010] Constructing a large model of structured prompt words based on the input text and the target knowledge dataset;
[0011] The structured prompt words are input into the large model to obtain result data corresponding to the input text.
[0012] In the above method, optionally, extracting keywords from the input text to obtain a keyword set corresponding to the input text includes:
[0013] Performing keyword extraction on the input text to obtain keywords of the input text;
[0014] The keywords of the input text are expanded using a knowledge graph database to obtain a keyword set corresponding to the input text, where the keyword set includes each keyword of the input text and at least one associated keyword of the keyword.
[0015] Optionally, the method of querying a knowledge base based on each keyword in the keyword set, the word vector of each keyword, and the text vector to obtain a target knowledge dataset includes:
[0016] Querying a first knowledge index library according to the word vector of each keyword in the keyword set and the text vector to obtain a first knowledge data set corresponding to the keyword set and a second knowledge data set corresponding to the text vector;
[0017] Querying a second knowledge index library according to each keyword in the keyword set to obtain a third knowledge data set corresponding to the keyword set;
[0018] A target knowledge dataset is obtained based on the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset.
[0019] In the above method, optionally, obtaining a target knowledge dataset based on the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset includes:
[0020] In a case where there is an intersection among the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset, determining the intersection among the first knowledge dataset, the second knowledge dataset, and the third dataset as a target dataset;
[0021] When there is no intersection between the first knowledge data set, the second knowledge data set and the third knowledge data set, multiple target knowledge data are selected from each of the knowledge data according to the similarity between each knowledge data in the first knowledge data set, the second knowledge data set and the third data set and the input text; and the selected multiple target knowledge data are combined into a target data set.
[0022] Optionally, in the above method, the step of selecting a plurality of target knowledge data from each of the first knowledge data set, the second knowledge data set, and the third knowledge data set based on the similarity between the input text and each of the knowledge data sets includes:
[0023] Determining an edit distance between each knowledge data in the first knowledge data set, the second knowledge data set, and the third knowledge data set and the input text, wherein the edit distance between each knowledge data and the input text is the minimum number of edits required to convert the knowledge data into the input text;
[0024] determining the similarity between each piece of knowledge data and the input text according to the edit distance between each piece of knowledge data and the input text;
[0025] Based on the similarity between each of the knowledge data and the input text, a plurality of target knowledge data are selected from each of the knowledge data.
[0026] Optionally, in the above method, the step of constructing a large model of structured prompt words based on the input text and the target knowledge dataset includes:
[0027] Performing intent analysis on the input text to construct entities and entity relationships in the input text;
[0028] Generate a query statement based on entities and entity relationships in the input text;
[0029] Querying the medical record database according to the query statement to obtain the medical record data set corresponding to the input text;
[0030] A structured prompt word of a large model is constructed according to the input text, entities and entity relationships in the input text, the target knowledge dataset, and the medical record dataset.
[0031] Optionally, in the above method, the step of constructing a large model of structured prompt words based on the input text and the target knowledge dataset includes:
[0032] Obtaining description information and question content in the input text;
[0033] The structured prompt words of the large model are constructed according to the input text, the description information, the question content and the target knowledge data set.
[0034] A data processing device, comprising:
[0035] A receiving unit, configured to receive input text;
[0036] A keyword extraction unit, configured to extract keywords from the input text to obtain a keyword set corresponding to the input text;
[0037] A first execution unit is configured to perform vectorization processing on the keyword set and the input text to obtain a word vector for each keyword in the keyword set and a text vector for the input text;
[0038] A query unit, configured to query a knowledge base according to each keyword in the keyword set, a word vector of each keyword, and the text vector to obtain a target knowledge data set;
[0039] A construction unit, configured to construct structured prompt words of a large model according to the input text and the target knowledge dataset;
[0040] The second execution unit is used to input the structured prompt words into the large model to obtain result data corresponding to the input text.
[0041] A storage medium includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the data processing method as described above.
[0042] An electronic device includes a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to perform the above-mentioned data processing method.
[0043] The present application provides a data processing method, device, storage medium and electronic device, the method comprising: receiving an input text; performing keyword extraction on the input text to obtain a keyword set corresponding to the input text; performing vectorization processing on the keyword set and the input text to obtain a word vector of each keyword in the keyword set and a text vector of the input text; querying a knowledge base based on each keyword in the keyword set, the word vector of each keyword and the text vector to obtain a target knowledge data set; constructing a structured prompt word of a large model based on the input text and the target knowledge data set; inputting the structured prompt word into the large model to obtain result data corresponding to the input text. By applying the method provided in the embodiment of the present application, the structured prompt word of the large model can be constructed based on the input text and the target knowledge data set, thereby effectively enriching the structured prompt word of the large model, and providing an accurate and reliable target knowledge data set for the large model, which can effectively improve the accuracy of the answer result of the large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0045] Figure 1 A flow chart of a data processing method provided in this application;
[0046] Figure 2 An example diagram of a medical knowledge graph database provided for this application;
[0047] Figure 3 A flowchart of a process for obtaining a target knowledge dataset provided in this application;
[0048] Figure 4 An example diagram of a construction example of an entity provided in this application;
[0049] Figure 5 An example diagram of a knowledge model provided for this application;
[0050] Figure 6 An example diagram of another knowledge model provided in this application;
[0051] Figure 7 An example diagram of the result data provided for this application;
[0052] Figure 8 A schematic diagram of the structure of a data processing device provided in this application;
[0053] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0055] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0056] The embodiment of the present application provides a data processing method, which can be applied to electronic devices, such as smart phones, tablet devices, smart wearable devices, etc. The method flow chart is as follows: Figure 1 As shown, specifically including:
[0057] S101: receiving input text.
[0058] In this embodiment, the input text may be various types of text, for example, a medical consultation text.
[0059] Optionally, the user can input text through the text box in the front-end interface as needed, so that the electronic device receives the user's input text.
[0060] S102: Extract keywords from the input text to obtain a keyword set corresponding to the input text.
[0061] In this embodiment, the keyword set includes at least each keyword in the input text.
[0062] Optionally, keyword extraction can be performed on the input text using a deep learning-based entity recognition algorithm, which can include a bidirectional long short-term memory network (BiLSTM) and a conditional random field (CRF) method.
[0063] S103: Perform vectorization processing on the keyword set and the input text to obtain a word vector for each keyword in the keyword set and a text vector for the input text.
[0064] In this embodiment, the keyword set and the input text can be vectorized through the word embedding model, and each keyword in the keyword set and the input text can be mapped to a vector space of a preset dimension to obtain the word vector of each keyword and the text vector of the input text.
[0065] S104: Query the knowledge base according to each keyword in the keyword set, the word vector of each keyword, and the text vector to obtain a target knowledge dataset.
[0066] In this embodiment, the knowledge base may include a first knowledge index base and a second knowledge index base. The first knowledge index base may be a vector knowledge base, and the second knowledge index base may be a text knowledge base. The first knowledge index base can be queried through the word vectors and text vectors of each keyword, and the second knowledge index base can be queried through each keyword.
[0067] Optionally, the target knowledge data set includes knowledge data related to the input text.
[0068] S105: Constructing structured prompt words of a large model based on the input text and the target knowledge dataset.
[0069] In this embodiment, a structured prompt word template may be obtained, and then a large model of structured prompt words may be generated according to the input text, the target data set, and the structured prompt word template.
[0070] Optionally, structured prompt words are used to instruct the large model to generate a response corresponding to the input text based at least on the target knowledge dataset.
[0071] S106: Input the structured prompt words into the large model to obtain result data corresponding to the input text.
[0072] In this embodiment, the result data can be output as reply data to the input text.
[0073] By applying the method provided in the embodiment of the present application, it is possible to construct structured prompt words of a large model based on the input text and the target knowledge data set, thereby effectively enriching the structured prompt words of the large model, and providing the large model with an accurate and reliable target knowledge data set, which can effectively improve the accuracy of the answer results of the large model.
[0074] In an embodiment provided by the present application, based on the above solution, optionally, keyword extraction is performed on the input text to obtain a keyword set corresponding to the input text, including:
[0075] Perform keyword extraction on the input text to obtain the keywords of the input text;
[0076] The keywords of the input text are expanded using the knowledge graph database to obtain a keyword set corresponding to the input text. The keyword set includes each keyword of the input text and at least one associated keyword of the keyword.
[0077] In this embodiment, the knowledge graph database can record various entities and the relationships between entities.
[0078] In some embodiments, the knowledge graph database may be a medical knowledge graph database. The knowledge graph database may be constructed based on a graph schema. The schema may be generated from medical entities (e.g., diseases, symptoms, and medications), and edges connecting multiple medical entities, where edges represent relationships between entities (e.g., the association between diseases and symptoms, the relationship between medications and disease treatments, required examinations for a disease, required tests for a disease, etc.). An example schema is as follows:
[0079] name: string @index(term) .
[0080] symptom: uid @count .
[0081] description: string @index(term) .
[0082] medicine: uid @count .
[0083] relate_disease: uid @count .
[0084] like Figure 2 As shown in the figure, a medical knowledge graph database can be built based on the schema. The medical knowledge graph database can include medical entities such as "coronary heart disease", "dyspnea", "syndrome X", "pulmonary embolism", "blood routine test" and "syncope", and entity relationships such as "symptoms", "differential diagnosis" and "test".
[0085] In this embodiment, the graph database can be queried by keywords. If the related keywords of the keyword are found, a keyword set is formed based on each keyword and the related keywords of each keyword to achieve keyword expansion of the input text.
[0086] In one embodiment provided in the present application, based on the above solution, optionally, the process of querying the knowledge base according to each keyword in the keyword set, the word vector of each keyword, and the text vector to obtain the target knowledge data set is as follows: Figure 3 As shown, including:
[0087] S301: Query a first knowledge index library according to the word vector and text vector of each keyword in the keyword set to obtain a first knowledge data set corresponding to the keyword set and a second knowledge data set corresponding to the text vector.
[0088] In this embodiment, a knowledge model may be constructed first, and then a multidimensional knowledge base may be constructed based on the knowledge model, and then a first knowledge index base and a second knowledge index base may be constructed based on the knowledge base.
[0089] In this embodiment, the process of constructing the initial knowledge base can be to perform entity recognition on the data in the data source, then extract entity attribute information of the data in the data source, and then extract the entity relationships between the entities in the data source, thereby constructing the initial knowledge base based on the entities, entity attributes and entity relationships between the entities.
[0090] Optionally, the data source may include medical record data and literature data, etc.
[0091] In this embodiment, entity classification can be performed based on medical records, documents, etc. to obtain entity categories. Named entity recognition can then be performed on the entity categories through rule-based entity recognition and machine learning-based entity recognition. Rule-based entity recognition can be implemented using at least one of the following methods: feature dictionaries, word segmentation, part-of-speech tagging, and regular expressions. Machine learning-based entity recognition can use deep learning-based entity recognition algorithms such as bidirectional long short-term memory networks (BiLSTMs) and conditional random fields (CRFs).
[0092] In this embodiment, entity attribute information can be extracted from data by combining rules with machine learning models. For example, a rule-based model can be used to process guidance data and structured data in textbooks within the data source. For entity attribute information described in free text within the data source, a translation-based embedding model can be used to predict entity attributes.
[0093] In this embodiment, a relationship extraction method based on machine learning can be used to extract entity relationships from text data in the data source. Specifically, the deep learning model Bert+BiLSTM model can be used to model the data, and then the entity recognition and relationship extraction method can be used to extract the semantic relationship between two or more entities from the text to complete the entity relationship extraction. Figure 4 The figure shows an example diagram of the construction example of one of the entities "hypertension" in the disease ontology.
[0094] Optionally, when constructing the first knowledge index, a word embedding model can be used to map each field in the knowledge base into a vector in a high-dimensional space. For example, the description "Hypertension is a cardiovascular syndrome characterized by elevated systemic arterial pressure" is converted into a series of numeric vectors. The index is then created, with the field type set to dense_vector. Cosine similarity can be used as the vector calculation function to measure the similarity between the query vector and the document vector.
[0095] Specifically, the cosine similarity is calculated by cosineSimilarity(params.queryVector, doc['my_vector_field']) + 1.0. Adding 1.0 here makes the score non-negative.
[0096] The first knowledge index library obtained by applying the embodiment of the present application can reflect the direct relationship between data and the semantic similarity between sentences. For example, the vectors of "headache" and "headache" are very close, while the vectors of "headache" and "stomachache" are far apart, which clearly shows their similarities or differences in medical concepts.
[0097] S302: Query the second knowledge index database according to each keyword in the keyword set to obtain a third knowledge data set corresponding to the keyword set.
[0098] In this embodiment, the construction process of the second knowledge index library can be as follows: using the knowledge model field names, define the names of the fields in the index; based on the type definitions of unstructured fields (e.g., disease ontology_overview) and structured fields (e.g., disease ontology_clinical manifestations_symptoms) in the model, design the data types of the fields in the schema, as shown in Table 1:
[0099] Table 1
[0100] Serial number Type definitions for fields in the model The data type of the field in the schema 1 Num Float 2 Date date(format:yyyy-MM-dd HH:mm:ss) 3 Boolean Keyword 4 Enum Keyword 5 massenum multi-field(keyword,ik_max_word) 6 Text ik_max_word 7 object Properties 8 position Geo
[0101] In this example, if the field type is defined as object, if the object is non-repeatable, the parent-child relationship is directly set using properties. If the object is repeatable, the parent-child relationship is also established using properties, and the field type is set to nested. For text fields, a word segmenter is required to split the text. The IK Analyzer can be used to load knowledge data for more accurate word segmentation and indexing.
[0102] By applying the method provided in the embodiments of this application, knowledge base data can be indexed into a search engine according to ontology dimensions, and an efficient secondary knowledge index can be constructed using inverted indexing technology, thereby improving search accuracy and efficiency. This not only ensures the rationality and flexibility of the data structure, but also enhances the accuracy of text processing, facilitating better management and retrieval of complex medical knowledge data.
[0103] S303: Obtain a target knowledge dataset based on the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset.
[0104] In this embodiment, the first knowledge data set, the second data set and the third data set can all be used as target knowledge data sets, or target knowledge data similar to the input text can be selected from the first knowledge data set, the second knowledge data set and the third knowledge data set as the target knowledge data set, or the intersection of at least any two of the first knowledge data set, the second knowledge data set and the third knowledge data set can be used as the target knowledge data set.
[0105] In an embodiment provided in the present application, based on the above solution, optionally, obtaining a target knowledge dataset based on the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset includes:
[0106] In the case where there is an intersection among the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset, determining the intersection among the first knowledge dataset, the second knowledge dataset, and the third dataset as a target dataset;
[0107] When there is no intersection between the first knowledge data set, the second knowledge data set and the third knowledge data set, multiple target knowledge data are selected from each knowledge data according to the similarity between each knowledge data in the first knowledge data set, the second knowledge data set and the third knowledge data set and the input text; and the selected multiple target knowledge data are combined into a target data set.
[0108] In this embodiment, it is possible to detect whether there is an intersection among the first knowledge data set, the second knowledge data set, and the third knowledge data set. If there is an intersection among the three, the intersection of the three can be determined as the target data set. If there is no intersection among the three, the similarity between each knowledge data in the first knowledge data set, the second knowledge data set, and the third data set and the input text can be determined. Based on the similarity between each knowledge data and the input text, multiple target knowledge data sets are selected, and then the target data set is composed based on the target knowledge data sets.
[0109] In one embodiment provided in the present application, based on the above solution, optionally, based on the similarity between each knowledge data in the first knowledge data set, the second knowledge data set, and the third knowledge data set and the input text, multiple target knowledge data are selected from each knowledge data, including:
[0110] Determine an edit distance between each piece of knowledge data in the first knowledge data set, the second knowledge data set, and the third knowledge data set and the input text, wherein the edit distance between each piece of knowledge data and the input text is the minimum number of edits required to convert the knowledge data into the input text;
[0111] Determine the similarity between each piece of knowledge data and the input text according to the edit distance between each piece of knowledge data and the input text;
[0112] Based on the similarity between each piece of knowledge data and the input text, a plurality of target knowledge data are selected from each piece of knowledge data.
[0113] In this embodiment, if the edit distance between the knowledge data and the input text is longer, the similarity between the knowledge data and the input text is determined to be smaller; if the edit distance between the knowledge data and the input text is shorter, the similarity between the knowledge data and the input text is determined to be higher.
[0114] Optionally, a preset number of target knowledge data may be selected from each knowledge data in descending order of similarity between each knowledge data and the input text.
[0115] In one embodiment provided in the present application, based on the above solution, optionally, a structured prompt word of a large model is constructed according to the input text and the target knowledge dataset, including:
[0116] Perform intent analysis on the input text to construct entities and entity relationships in the input text;
[0117] Generate query statements based on entities and entity relationships in the input text;
[0118] Query the medical record database according to the query statement to obtain the medical record data set corresponding to the input text;
[0119] Build structured prompt words for the large model based on the input text, entities and entity relationships in the input text, target knowledge dataset, and medical record dataset.
[0120] In this embodiment, the medical record data can be a medical record database. Clinical data from various medical data sources can be obtained and integrated based on a common data model. Based on this common data model, ETL tools are used to map and associate visit times / IDs, integrating data from various systems (such as the HIS information management system, CIS clinical information system, EMR electronic medical record, LIS laboratory information, PACS imaging, and RIS radiology). A rule dictionary is used to convert field values and automatically normalize terms to achieve standardized processing of structured fields. For large amounts of unstructured medical text, supervised learning algorithms such as LSTM and TextCNN are used for text classification, BiLSTM combined with CRF algorithms for named entity recognition, and a Bert-BiLSTM-BiLSTM method are used for entity relationship extraction. By converting medical text into structured data, a medical record database is constructed. This database can capture information about the patient's complete treatment process, including the medical record homepage, admission records, laboratory test reports, examination reports, discharge records, surgical records, and daily medical records, thereby fully supporting medical data analysis and applications.
[0121] In this embodiment, intent analysis can be performed on the input text to identify and construct medical entities and their relationships. Natural language processing techniques are used to parse the user's query, identifying key medical entities (such as disease names, symptoms, and treatment plans) and inter-entity relationships (such as causal and concomitant relationships). A structured query statement is then generated based on the extracted information. This query statement searches the integrated multi-source medical record database to retrieve medical records relevant to the user's query. Finally, the query generates a result set of relevant medical records, providing the user with the medical information support they need.
[0122] In one embodiment provided in the present application, based on the above solution, optionally, a structured prompt word of a large model is constructed according to the input text and the target knowledge dataset, including:
[0123] Get the description information and question content in the input text;
[0124] Build structured prompt words for a large model based on input text, description information, question content, and target knowledge dataset.
[0125] In this embodiment, the description information may include symptom description information, and the question content may be the user's question information.
[0126] The data processing method provided in this application can be applied in different technical fields, for example, it can be applied in the medical field, and the following examples are given to illustrate in detail:
[0127] The first step is to build a multi-dimensional medical knowledge model, construct a knowledge base entity library based on the model, design a schema based on the model, and build a text index library (second knowledge index library), a vector index library (first knowledge index library), and a graph database based on the data in the knowledge base.
[0128] In this embodiment, based on the medical data storage standards and terminology standards, the data model is abstracted and a general model structure that conforms to the characteristics of the ontology is constructed. All ontology information is analyzed, sorted and integrated, and dozens of ontologies are constructed. For the above ontologies, concepts such as relationships, attributes, and relationship attributes are designed. Figure 5 and Figure 6 shown.
[0129] The second step is to build a common data model, integrate and structure the patient's multi-source heterogeneous medical record data based on the model, and build a medical record database.
[0130] The third step is to perform keyword mining and vectorization on the input free text (input text), and obtain the traceability document set (knowledge data set) by searching in the index library.
[0131] In this embodiment, information extraction is performed on the free text input for medical knowledge search, and a deep learning-based entity recognition algorithm - Bidirectional Long Short-Term Memory Network (BiLSTM) and Conditional Random Field (CRF) method is used to extract a keyword list from the text.
[0132] The free text input could be: "I often feel discomfort in my head, manifested as paroxysmal or persistent headaches and dizziness, sometimes accompanied by nausea and vomiting. Is this high blood pressure? How should I treat it?"
[0133] Optionally, the list of medical keywords extracted from the free text can be as follows:
[0134] [Headache, dizziness, nausea, vomiting, high blood pressure]
[0135] In this embodiment, after extracting the medical keyword list, the keywords are extended based on the knowledge graph database. The extended keyword list is as follows:
[0136] [Headache, dizziness, nausea, vomiting, hypertension, grade 1 hypertension, benign hypertension]
[0137] In this embodiment, a word embedding model is used to map the expanded medical keyword list and free text into vectors in a high-dimensional space.
[0138] The vector representation of the medical keyword list extracted from the information is as follows:
[0139] Head discomfort: [-0.2703271, 0.38279012, ..., -0.29274252, -0.24937081]
[0140] Headache: [-0.22879271, 0.43286988,...,-0.21742335,0.16286954]
[0141] Dizziness: [-0.24912262, 0.40769795,...,-0.26663426,0.1063835]
[0142] Nausea: [0.51697373, -0.01454506, ..., 0.1063835, -0.2986216]
[0143] Vomit: [0.16286954,-0.20245396,...,1.1556625,-0.112049]
[0144] Hypertension: [-0.29274252,-0.24937081,...,0.7212287,0.0751707]
[0145] Hypertension stage 1: [0.01726123, 0.1450473, ... , 0.16286954 , -0.20245396]
[0146] Benign hypertension: [-0.28274352,-0.24835081,...,0.7213257,0.0561707]
[0147] The vector representation of free text is as follows:
[0148] I often feel headaches, dizziness, or paroxysmal headaches, sometimes accompanied by nausea and vomiting. Is this high blood pressure? How should I treat it? :[[-0.2703271,0.38279012,-0.29274252,...,-0.24937081,0.7212287,0.0751707],[0.01726123,0.1450473,0.16286954 ,... ,-0.20245396,1.1556625,-0.112049],[0.51697373,-0.01454506,0.1063835,...,-0.2986216,0.69151103,0.13124703]]
[0149] In this embodiment, after obtaining the vectors of each keyword and the vectors of the free text, the vectors of each keyword are used to search in the knowledge vector index library to obtain the source document set 1 (the first knowledge data set). The source document set 1 may include the following contents:
[0150] “Document id: hypertension;
[0151] Document id: dizziness;
[0152] Document id: High Altitude Hypertension".
[0153] Optionally, the vector of the input free text is used to search in the knowledge vector index library to obtain a traceability document set 2, including the following content:
[0154] “Document id: hypertension;
[0155] Document id: dizziness;
[0156] Document id: benign hypertension".
[0157] Optionally, keywords in the keyword list are used to search in the text knowledge index library in the form of or. The index expression may be: [head discomfort or headache or dizziness or nausea or vomiting or hypertension]. The traceability document set 3 obtained by the search may include the following content:
[0158] “Document id: hypertension;
[0159] Document id: dizziness;
[0160] Document id: Hypertension stage 1".
[0161] In this embodiment, a candidate result set (target knowledge dataset) can be generated through traceability document set 1, traceability document set 2 and traceability document set 3, a prompt can be constructed based on the candidate result set and the free text of the question, and an answer can be given based on the constructed prompt.
[0162] Optionally, the intersection of traceability document set 1, traceability document set 2, and traceability document set 3 can be used as a candidate result set. The intersection can be expressed as follows:
[0163] “Document id: hypertension;
[0164] Document id: dizzy".
[0165] In some embodiments, if there is no intersection between the traceability document set 1, the traceability document set 2 and the traceability document set 3, the union of the traceability document set 1, the traceability document set 2 and the traceability document set 3 is calculated, and the similarity is calculated with the free text; the top N knowledge document data with higher similarity in the union are selected as the candidate result set.
[0166] Optionally, assume that the current source documents do not have any intersection, and the union is as follows:
[0167] “Document id: hypertension;
[0168] Document id: high altitude hypertension;
[0169] Document id: Hypertension stage 1;
[0170] Document id: benign hypertension;
[0171] Document id: dizzy".
[0172] In this case, we can calculate the text similarity between each knowledge document in the union and the free text. Specifically, we can calculate the edit distance between the free text "I often feel discomfort in my head, manifested as paroxysmal or persistent headaches, dizziness, sometimes accompanied by nausea and vomiting. Is it hypertension? What treatment is needed?" and each knowledge document in the union. The edit distance refers to the minimum number of edit operations required to replace one of the two texts with the other. The greater the distance, the lower the similarity. We can obtain the three texts in the union with the highest similarity to the free text, as follows:
[0173] “Document id: hypertension;
[0174] Document id: dizziness;
[0175] Document id: Hypertension stage 1".
[0176] In this embodiment, intent analysis can also be performed based on the input free text, medical entities and entity relationships in the free text can be constructed, query statements can be generated, and searches can be performed in the integrated medical record library to obtain relevant medical record result sets.
[0177] The input free text is structured according to the ontology in the medical knowledge module, and the generated candidate result set documents and medical record result set documents are input into the large model using triple quotation marks as the introduced external knowledge document set to enhance the answering ability of the large model, enabling it to generate more detailed and accurate answers.
[0178] In this embodiment, a structured prompt can be constructed. The structured prompt consists of four main parts: the original free text, the medical knowledge ontology description, the question, and the external knowledge document set. The description of the medical knowledge symptom ontology is further subdivided into symptoms of the head part and accompanying symptoms. The content of the structured prompt is as follows:
[0179] {
[0180] "Free text": "I often feel headache, which manifests as paroxysmal or persistent headache, dizziness, and sometimes nausea and vomiting. Is it hypertension? How should it be treated?"
[0181] "Symptom Description": {
[0182] "Head discomfort": {
[0183] "Symptoms": ["paroxysmal headache", "persistent headache", "dizziness"],
[0184] "Accompanying symptoms": ["nausea", "vomiting"]
[0185] }
[0186] },
[0187] "Ask": [
[0188] "Is it high blood pressure?"
[0189] Treatment recommendations
[0190] ],
[0191] "External Knowledge Document Set":{"name":"Hypertension","obj":{"people":["Male","Female"],"etiology_and_risk_factors":{"description":[{"description":"The causes of essential hypertension..."}]},"lab":["Biochemical panel","Sodium ion determination"],"exam":["Electrocardiogram","24-hour blood pressure monitoring","Echocardiography"]...}"""{"name":"Dizziness","obj":{"definition":"Dizziness is a common functional brain disorder...","symptom_part_attribution":["Head"],"people":["Male","Female","All age groups" ]}],"alias":["fainting","dizziness","dizziness","dizziness","dizziness"]}],"inducements":[{"name":["drinking","staying up late","falling","fatigue","bending head for a long time","trauma","activity","turning head and neck","posture change","tension and anxiety"]}],"properties":[{"category":"feeling","name":["bloating"]}],"aggravation_factor":[{"name":["standing or sitting for a long time","riding in a car","bending head","posture change","posture change","side-lying position"]}],"occurrence_time":[{"name":["before meals","in the morning","before bedtime"]}],"clinical_significance":" There are many types of diseases that cause dizziness..."}}"""{"Medical record homepage":"","Admission record":"Past medical history: Hypertension was discovered 9 days ago, with the highest being 150 / 90 mmHg. Currently, he is taking 20 mg of nifedipine controlled-release tablets daily for treatment..."、"Test report":""、"Inspection report":""、"Discharge record":""、"Surgery record":""、"Daily medical history record":""}"""{"Medical record homepage":"","Admission record":"Present medical history: 3 years ago, the patient experienced shortness of breath while walking or climbing stairs quickly, which was relieved after a few minutes of rest. He had no chest pain..."、"Homepage diagnosis":""、"Inspection report":""、"Daily ward round record":""、"Surgery record":""、"Daily medical history record":""}"""
[0192] In this embodiment, the constructed structured prompt is used to input into the large model to generate the final result text. Figure 7 shown.
[0193] By applying the method provided in the embodiments of this application, a structured expression of prompts is implemented using a JSON-like data structure, effectively improving the readability of prompts and making them more consistent with human expression habits and the cognitive patterns of large models, thereby reducing the difficulty of understanding for both humans and large models. Structured prompts greatly improve the efficiency of semantic cognition, making problem understanding and processing more efficient and accurate.
[0194] and Figure 1 Corresponding to the method, the embodiment of the present application further provides a data processing device for Figure 1 The specific implementation of the method is shown in the following diagram: Figure 8 As shown, including:
[0195] Receiving unit 801, for receiving input text;
[0196] The keyword extraction unit 802 is used to extract keywords from the input text to obtain a keyword set corresponding to the input text;
[0197] A first execution unit 803 is configured to perform vectorization processing on the keyword set and the input text to obtain a word vector for each keyword in the keyword set and a text vector for the input text;
[0198] A query unit 804 is configured to query the knowledge base based on each keyword in the keyword set, the word vector of each keyword, and the text vector to obtain a target knowledge dataset;
[0199] A construction unit 805 is used to construct structured prompt words of a large model based on the input text and the target knowledge dataset;
[0200] The second execution unit 806 is used to input the structured prompt words into the large model to obtain result data corresponding to the input text.
[0201] In an embodiment provided in the present application, based on the above solution, optionally, the keyword extraction unit 802 includes:
[0202] The extraction subunit is used to extract keywords from the input text to obtain keywords of the input text;
[0203] The expansion subunit is used to expand the keywords of the input text using the knowledge graph database to obtain a keyword set corresponding to the input text. The keyword set includes each keyword of the input text and associated keywords of at least one keyword.
[0204] In an embodiment provided in the present application, based on the above solution, optionally, the query unit 404 includes:
[0205] A first query subunit is configured to query a first knowledge index library based on a word vector and a text vector of each keyword in the keyword set to obtain a first knowledge data set corresponding to the keyword set and a second knowledge data set corresponding to the text vector;
[0206] A second query subunit is configured to query the second knowledge index library according to each keyword in the keyword set to obtain a third knowledge data set corresponding to the keyword set;
[0207] The execution subunit is configured to obtain a target knowledge dataset based on the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset.
[0208] In an embodiment provided in the present application, based on the above solution, optionally, the execution subunit includes:
[0209] a first execution module, configured to determine the intersection of the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset as a target dataset when there is an intersection among the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset;
[0210] The second execution module is used to select multiple target knowledge data from each knowledge data according to the similarity between each knowledge data in the first knowledge data set, the second knowledge data set and the third knowledge data set and the input text when there is no intersection between the first knowledge data set, the second knowledge data set and the third knowledge data set; and form the target data set with the selected multiple target knowledge data.
[0211] In an embodiment provided in the present application, based on the above solution, optionally, the second execution module includes:
[0212] A first determining submodule is configured to determine an edit distance between each piece of knowledge data in the first knowledge data set, the second knowledge data set, and the third knowledge data set and the input text, wherein the edit distance between each piece of knowledge data and the input text is the minimum number of edits required to convert the knowledge data into the input text;
[0213] A second determination submodule is configured to determine the similarity between each piece of knowledge data and the input text based on the edit distance between each piece of knowledge data and the input text;
[0214] The selection submodule is used to select a plurality of target knowledge data from each knowledge data based on the similarity between each knowledge data and the input text.
[0215] In an embodiment provided in the present application, based on the above solution, optionally, the construction unit 405 includes:
[0216] The analysis subunit is used to perform intent analysis on the input text to construct entities and entity relationships in the input text;
[0217] A generation subunit, used to generate query statements based on entities and entity relationships in the input text;
[0218] The query subunit is used to query the medical record database according to the query statement and obtain the medical record data set corresponding to the input text;
[0219] The first construction subunit is used to construct structured prompt words of a large model based on the input text, entities and entity relationships in the input text, the target knowledge dataset and the medical record dataset.
[0220] In an embodiment provided in the present application, based on the above solution, optionally, the construction unit 405 includes:
[0221] The acquisition subunit is used to obtain the description information and question content in the input text;
[0222] The second construction subunit is used to construct structured prompt words of a large model based on the input text, description information, question content and target knowledge data set.
[0223] The specific principles and execution processes of each unit and module in the data processing device disclosed in the above embodiment of the present application are the same as the data processing method disclosed in the above embodiment of the present application. Please refer to the corresponding parts of the data processing method provided in the above embodiment of the present application, and no further details will be given here.
[0224] An embodiment of the present application further provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the above-mentioned data processing method.
[0225] The present application also provides an electronic device, the structure of which is shown in FIG. Figure 9 As shown, it specifically includes a memory 901 and one or more instructions 902, wherein one or more instructions 902 are stored in the memory 901 and are configured to be executed by one or more processors 903 to perform the above-mentioned data processing method.
[0226] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0227] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0228] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0229] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0230] The above is a detailed introduction to a data processing method provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A data processing method, characterized in that: include: Receive input text; Extract keywords from the input text to obtain a keyword set corresponding to the input text; Performing vectorization processing on the keyword set and the input text to obtain a word vector for each keyword in the keyword set and a text vector for the input text; Querying a knowledge base according to each of the keywords in the keyword set, the word vector of each of the keywords, and the text vector to obtain a target knowledge data set; Constructing a large model of structured prompt words based on the input text and the target knowledge dataset; The structured prompt words are input into the large model to obtain result data corresponding to the input text.
2. The method according to claim 1, characterized in that The step of extracting keywords from the input text to obtain a keyword set corresponding to the input text includes: Performing keyword extraction on the input text to obtain keywords of the input text; The keywords of the input text are expanded using a knowledge graph database to obtain a keyword set corresponding to the input text, where the keyword set includes each keyword of the input text and at least one associated keyword of the keyword.
3. The method according to claim 1, characterized in that The step of querying a knowledge base according to each keyword in the keyword set, the word vector of each keyword, and the text vector to obtain a target knowledge data set includes: Querying a first knowledge index library according to the word vector of each keyword in the keyword set and the text vector to obtain a first knowledge data set corresponding to the keyword set and a second knowledge data set corresponding to the text vector; Querying a second knowledge index library according to each keyword in the keyword set to obtain a third knowledge data set corresponding to the keyword set; A target knowledge dataset is obtained based on the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset.
4. The method according to claim 3, characterized in that The obtaining of a target knowledge dataset based on the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset includes: In a case where there is an intersection among the first knowledge dataset, the second knowledge dataset, and the third knowledge dataset, determining the intersection among the first knowledge dataset, the second knowledge dataset, and the third dataset as a target dataset; When there is no intersection between the first knowledge data set, the second knowledge data set and the third knowledge data set, multiple target knowledge data are selected from each of the knowledge data according to the similarity between each knowledge data in the first knowledge data set, the second knowledge data set and the third data set and the input text; and the selected multiple target knowledge data are combined into a target data set.
5. The method according to claim 4, characterized in that The selecting a plurality of target knowledge data from each of the knowledge data in the first knowledge data set, the second knowledge data set, and the third knowledge data set according to the similarity between the input text and each of the knowledge data includes: Determining an edit distance between each knowledge data in the first knowledge data set, the second knowledge data set, and the third knowledge data set and the input text, wherein the edit distance between each knowledge data and the input text is the minimum number of edits required to convert the knowledge data into the input text; determining the similarity between each piece of knowledge data and the input text according to the edit distance between each piece of knowledge data and the input text; Based on the similarity between each of the knowledge data and the input text, a plurality of target knowledge data are selected from each of the knowledge data.
6. The method according to claim 1, characterized in that The step of constructing a large model of structured prompt words based on the input text and the target knowledge dataset includes: Performing intent analysis on the input text to construct entities and entity relationships in the input text; Generate a query statement based on entities and entity relationships in the input text; Querying the medical record database according to the query statement to obtain the medical record data set corresponding to the input text; A structured prompt word of a large model is constructed according to the input text, entities and entity relationships in the input text, the target knowledge dataset, and the medical record dataset.
7. The method according to claim 1, characterized in that The step of constructing a large model of structured prompt words based on the input text and the target knowledge dataset includes: Obtaining description information and question content in the input text; The structured prompt words of the large model are constructed according to the input text, the description information, the question content and the target knowledge data set.
8. A data processing device, characterized in that: include: A receiving unit, configured to receive input text; A keyword extraction unit, configured to extract keywords from the input text to obtain a keyword set corresponding to the input text; A first execution unit is configured to perform vectorization processing on the keyword set and the input text to obtain a word vector for each keyword in the keyword set and a text vector for the input text; A query unit, configured to query a knowledge base according to each keyword in the keyword set, a word vector of each keyword, and the text vector to obtain a target knowledge data set; A construction unit, configured to construct structured prompt words of a large model according to the input text and the target knowledge dataset; The second execution unit is used to input the structured prompt words into the large model to obtain result data corresponding to the input text.
9. A storage medium, characterized in that: The storage medium includes storage instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the data processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The device comprises a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to perform the data processing method according to any one of claims 1 to 7.