Text processing method, device, and storage medium
By semantic analysis and retrieval recall of the query text input by the large language model, prompt information is constructed and inputted to the second largest language model, the problem of low accuracy in generating reply texts is solved, and the accuracy of reply texts is improved.
Patent Information
- Application Number
- CN202510457589.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-14
AI Technical Summary
When generating reply text, the large language model has interfering factors in the input query text, which makes it impossible to accurately understand the semantics of the query text, thereby reducing the accuracy of the reply text.
Through the first major language model, the target user query text is semantic analysis, hypothetical answer text is generated, and the query text and hypothetical answer text are searched and recalled separately, the target recall text collection and prompt information are constructed, and the second major language model is input to generate the target answer text.
By refining the semantic expression of query text input to the large language model, the accuracy of the reply text generated by the large language model is improved, and the user's needs for accurate and reliable information are met.
Smart Images

Figure CN119990142B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a text processing method, device, and storage medium. Background Art
[0002] Currently, large language models can generate corresponding response texts for the input query text through their powerful reasoning capabilities. In related technologies, query texts can be generated according to actual application requirements and input into the large language model to generate the response text for the query text. However, due to the presence of many interference factors in the query text input into the large language model, the large language model cannot accurately understand the semantics of the query text, resulting in a low accuracy of the finally generated response text.
[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a text processing method, device, and storage medium, aiming to solve the technical problem of the low accuracy of the response text generated by the large language model.
[0005] To achieve the above object, this application proposes a text processing method, which includes:
[0006] Performing semantic analysis on the target user query text through a first large language model to obtain a hypothetical answer text corresponding to the target user query text;
[0007] Performing retrieval and recall on the target user query text and the hypothetical answer text respectively to obtain a first recall text set corresponding to the target user query text and a second recall text set corresponding to the hypothetical answer text;
[0008] Rearranging the passage texts in the first recall text set and the passage texts in the second recall text set together to obtain a target recall text set;
[0009] Constructing a prompt message according to the target recall text set and the target user query text;
[0010] Inputting the prompt message into a second large language model to obtain a target answer text corresponding to the target user query text.
[0011] In one embodiment, the step of performing retrieval and recall on the target user query text and the hypothetical answer text respectively to obtain a first recall text set corresponding to the target user query text and a second recall text set corresponding to the hypothetical answer text includes:
[0012] Perform vector retrieval and word retrieval on the target user's query text to obtain each first paragraph text similar to the target user's query text, and combine each first paragraph text to obtain a first recall text set; and,
[0013] Perform word retrieval on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text, and combine each second paragraph text to obtain a second recall text set.
[0014] In one embodiment, the step of performing vector retrieval and word retrieval on the target user's query text to obtain each first paragraph text similar to the target user's query text includes:
[0015] Use a preset embedding model to vectorize the target user's query text to obtain a first vector representation of the target user's query text;
[0016] Retrieve multiple preset vector representations in a preset vector database whose similarity rankings with the first vector representation are before a first preset ranking, and use the preset paragraph text associated with each preset vector representation as the first paragraph text; and,
[0017] Use a word segmentation tool to perform word segmentation on the target user's query text to obtain first entry data;
[0018] Retrieve multiple first preset entry data in a first preset index thesaurus whose similarity rankings with the first entry data are before a second preset ranking, and use the preset paragraph text associated with each first preset entry data as the first paragraph text.
[0019] In one embodiment, the text processing method further includes:
[0020] Parse the document data parsing sample in the knowledge database into a set of paragraph text data;
[0021] Perform vectorization on the set of paragraph text data to obtain a set of paragraph text data vectors, and use the set of paragraph text data vectors to construct a preset vector database; and,
[0022] Perform word segmentation on the set of paragraph text data to obtain a set of paragraph text data entries, and use the set of paragraph text data entries to construct a first preset index thesaurus.
[0023] In one embodiment, the step of performing word retrieval on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text includes:
[0024] Use a word segmentation tool to perform word segmentation on the hypothetical answer text to obtain second entry data;
[0025] Retrieve multiple second preset entry data in the second preset index word library whose similarity rankings with the second entry data are before the third preset ranking, and use the preset paragraph text associated with each second preset entry data as the second paragraph text.
[0026] In one embodiment, the step of retrieving multiple second preset entry data in the second preset index word library whose similarity rankings with the second entry data are before the third preset ranking includes:
[0027] According to the second entry data and each second preset entry data, determine the inverse document frequency calculation results corresponding to each word in the second entry data, and the word frequency calculation results corresponding to each word relative to each second preset entry data;
[0028] According to the inverse document frequency calculation results corresponding to each word, and the word frequency calculation results corresponding to each word relative to each second preset entry data, obtain the scores corresponding to each word relative to each second preset entry data;
[0029] According to the sum of the scores corresponding to each word relative to each second preset entry data, obtain the correlation scores of each second preset entry data;
[0030] Based on the correlation scores, sort all the second preset entry data in descending order, and select multiple second preset entry data whose rankings are before the third preset ranking.
[0031] In one embodiment, the step of re-ranking the paragraph texts in the first recall text set and the paragraph texts in the second recall text set together to obtain the target recall text set includes:
[0032] Extract features from each paragraph text in the first recall text set and each paragraph text in the second recall text set respectively to obtain the lexical features, syntactic features, and semantic features of each paragraph text;
[0033] According to the lexical features, syntactic features, and semantic features of each paragraph text, and the weight coefficients corresponding to each feature, perform weighted calculation to obtain the feature scores corresponding to each paragraph text;
[0034] Select multiple paragraph texts whose feature scores are greater than the preset feature score from all the paragraph texts;
[0035] Combine the selected multiple paragraph texts to obtain the target recall text set.
[0036] In one embodiment, before the step of performing semantic analysis on the target user query text through the first large language model to obtain the hypothetical answer text corresponding to the target user query text, it further includes:
[0037] Obtain the current user's query text;
[0038] Determine whether there are historical Q&A pairs related to the current user's query text in the historical records;
[0039] If there are historical Q&A pairs related to the current user's query text in the historical records, generate a target user query text based on the third large language model, the historical Q&A pairs, and the current user's query text;
[0040] If there are no historical Q&A pairs related to the current user's query text in the historical records, use the current user's query text as the target user query text.
[0041] In addition, to achieve the above object, the present application also proposes a text processing device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the text processing method as described above.
[0042] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, it implements the steps of the text processing method as described above.
[0043] The present application performs semantic analysis on the target user query text through the first large language model to obtain a hypothetical answer text corresponding to the target user query text; performs retrieval and recall on the target user query text and the hypothetical answer text respectively to obtain a first recall text set corresponding to the target user query text and a second recall text set corresponding to the hypothetical answer text; rearranges the paragraph texts in the first recall text set and the paragraph texts in the second recall text set together to obtain a target recall text set; constructs a prompt message based on the target recall text set and the target user query text; inputs the prompt message into the second large language model to obtain a target answer text corresponding to the target user query text. Since the first large language model is used to generate a hypothetical answer text based on the target user query text input by the user, the large language model can understand the semantics of the query text. Subsequently, retrieval and recall texts are performed based on the hypothetical answer text and the target user query text to construct rich prompt messages, and the prompt messages are input into the second large language model, enabling the second large language model to accurately generate the corresponding target answer text based on the prompt messages, improving the accuracy of the generated target answer text. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0045] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0046] Figure 1 Schematic flowchart provided for the first embodiment of the text processing method of the present application;
[0047] Figure 2 Schematic flowchart of a refined process provided for step S20 in the first embodiment of the text processing method of the present application;
[0048] Figure 3 Schematic flowchart of a refined process provided for step S22 in the first embodiment of the text processing method of the present application;
[0049] Figure 4 Schematic flowchart of a refined process provided for step S222 in the first embodiment of the text processing method of the present application;
[0050] Figure 5 Schematic flowchart of a refined process provided for step S30 in the first embodiment of the text processing method of the present application;
[0051] Figure 6 Schematic flowchart of a refined process provided for step S21 in the first embodiment of the text processing method of the present application;
[0052] Figure 7 Schematic flowchart provided for the second embodiment of the text processing method of the present application;
[0053] Figure 8 Schematic flowchart provided for the third embodiment of the text processing method of the present application;
[0054] Figure 9 Schematic flowchart of the text processing method of the present application;
[0055] Figure 10 Schematic diagram of the device structure of the hardware operating environment involved in the text processing method in the embodiments of the present application.
[0056] The realization of the purpose, functional features, and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0057] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0058] To better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.
[0059] At present, the large language model can generate corresponding reply text for the input query text with its excellent reasoning ability. In actual application scenarios, the query text can be constructed according to actual needs and input into the large language model to obtain the reply text corresponding to the query text. However, there are often many interference factors in the query text input into the large language model, such as ambiguous expressions in the text, a word or phrase may have multiple meanings, which makes it difficult for the model to accurately judge its accurate semantics in a specific context; or the query text has incomplete expressions and ambiguous references, which makes the model unable to obtain sufficient and clear information. In addition, the non-standard text format of the query text, such as a large number of special characters, non-standard typesetting, etc., will also interfere with the model's extraction of the core semantics of the query text. Due to the above interference factors, it is difficult for the large language model to accurately understand the true semantics of the query text, which ultimately reduces the accuracy of the reply text generated by the large language model and cannot meet the user's demand for accurate and reliable information.
[0060] In view of the above problems, the main solution of the present application is: to perform semantic analysis on the target user query text through the first largest language model to obtain the hypothetical answer text corresponding to the target user query text; to retrieve and recall the target user query text and the hypothetical answer text respectively to obtain the first recall text set corresponding to the target user query text and the second recall text set corresponding to the hypothetical answer text; to rearrange the paragraph text in the first recall text set and the paragraph text in the second recall text set together to obtain the target recall text set; to construct prompt information according to the target recall text set and the target user query text; to input the prompt information into the second largest language model to obtain the target answer text corresponding to the target user query text.
[0061] The solution of the present application uses the reply text generated by the large language model based on the query text input by the user, and retrieves and recalls the query text input by the user in the database to obtain multiple recall texts to construct the prompt words of the large language model, and inputs the prompt words into the large language model to obtain the answer corresponding to the query text. By refining the semantic expression of the query text input into the large language model, the accuracy of the reply text generated by the large language model is improved.
[0062] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a text processing device, etc. The following takes a text processing device as an example to illustrate this embodiment and the following embodiments.
[0063] It should be noted that the first large language model, the second large language model, and the third large language model can be the same large language model, or two of them can be the same, or they can be three different large language models trained from different datasets. The first preset index word library and the second preset index word library can be the same preset index word library, or two different preset index word libraries. The first preset ranking, the second preset ranking, and the third preset ranking can be the same for all three, or two of them can be the same, or all three can be different. The paragraph text can be a paragraph composed of no less than one sentence.
[0064] Based on this, an embodiment of the present application provides a text processing method. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the text processing method of the present application.
[0065] In this embodiment, the text processing method includes steps S10 to S50:
[0066] Step S10, perform semantic analysis on the target user query text through the first large language model to obtain a hypothetical answer text corresponding to the target user query text.
[0067] It should be noted that the first large language model can be a large language model (LLM, Large Language Model). A large language model refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. The model can provide in-depth knowledge and language production on various topics through training on a vast dataset. The first large language model refers to the large language model obtained after training using the training dataset. Among them, the data source of the training dataset is extensive, including but not limited to web pages, books, news articles, academic papers, social media posts, encyclopedias, etc. on the Internet.
[0068] In addition, the target user query text is the text input by the user into the interaction interface when interacting with the large language model, aiming to clearly and precisely convey the user's intention, so that the first large language model can generate a hypothetical answer text corresponding to the target user query text based on this.
[0069] The hypothetical response text is the response text output by the first large language model according to the target user's query text. Since there are many interfering factors in the target user's query text input to the first large language model, the first large language model cannot accurately understand the semantics of the target user's query text. Therefore, in this application, the first large language model performs semantic analysis on the target user's query text to analyze the semantics of the target user's query text and generate a hypothetical response text, which can facilitate the subsequent second large language model to accurately determine the semantics of the target user's query text and make the finally generated response text more accurate.
[0070] In this embodiment, the first large language model receives the target user's query text and performs steps such as encoding, neural network semantic extraction, and probability distribution calculation on the target user's query text to obtain a hypothetical response text corresponding to the target user's query text, so that the Retrieval - Augmented Generation (RAG) technology integrates the hypothetical response text into the inference of the large language model to improve the response accuracy of the large language model.
[0071] Step S20: Perform retrieval and recall on the target user's query text and the hypothetical response text respectively to obtain a first recall text set corresponding to the target user's query text and a second recall text set corresponding to the hypothetical response text.
[0072] It should be noted that retrieval and recall refers to the work of initially screening a batch of documents that may be relevant to the user's query from a large - scale document collection, such as keyword - based recall, vector - based recall, hybrid recall, etc. The first recall text set is the set of all paragraph texts retrieved and related to the target user's query text, and the second recall text set is the set of all paragraph texts retrieved and related to the hypothetical response text.
[0073] In a feasible implementation manner, perform hybrid recall on the target user's query text and the hypothetical response text respectively to obtain a first recall text set corresponding to the target user's query text and a second recall text set corresponding to the hypothetical response text.
[0074] In this embodiment, perform keyword recall and vector recall on the target user's query text to obtain a first recall text set, and perform keyword recall and vector recall on the hypothetical response text to obtain a second recall text set, so as to integrate the advantages of keyword recall and vector recall, utilize both the rapidity of keyword recall and the semantic understanding ability of vector recall, thereby improving the accuracy and efficiency of recall.
[0075] Further, referring to Figure 2 , step S20 includes steps S21 - S22:
[0076] Step S21: Perform vector retrieval and word retrieval on the target user's query text to obtain each first paragraph text similar to the target user's query text, and combine each first paragraph text to obtain a first recall text set.
[0077] It should be noted that vector retrieval converts the target user's query text into a vector representation, calculates the spatial distance between vectors to obtain the similarity between data, and then obtains data with a high similarity to the target user's query text as the query result; word retrieval uses keywords related to the target user's query text to perform retrieval in the database and obtains data containing the keyword as the query result.
[0078] Step S22: Also, perform word retrieval on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text, and combine each second paragraph text to obtain a second recall text set.
[0079] Furthermore, referring to Figure 3 , step S22 includes steps S221 to S222:
[0080] Step S221: Use a word segmentation tool to perform word segmentation on the hypothetical answer text to obtain second entry data.
[0081] It should be noted that word segmentation is the process of splitting a continuous text sequence into individual words or phrases according to certain rules, and a word segmentation tool is software or a library that implements this function. Entry data is a list containing the word segmentation results of the target user's query text. The word segmentation tool can be jieba in Python (a programming language), which is a word segmentation tool applied to Chinese text processing, to split the Chinese text according to word boundaries. The second entry data is the result of the word segmentation of the hypothetical answer text.
[0082] Step S222: Retrieve multiple second preset entry data whose similarity rankings with the second entry data are before the third preset ranking in the second preset index word library, and use the preset paragraph text associated with each second preset entry data as the second paragraph text.
[0083] It should be noted that before retrieval, first parse and split the text corpus into paragraph data, perform word segmentation on these paragraph data and then store them in the second preset index word library. The preset paragraph text associated with the second preset entry data is the paragraph text corresponding to the second preset entry data before word segmentation.
[0084] Furthermore, referring to Figure 4 , step S222 includes steps S2221 to S2224:
[0085] Step S2221: According to the second entry data and each second preset entry data, determine the inverse document frequency calculation result corresponding to each word in the second entry data, and the word frequency calculation result corresponding to each word relative to each second preset entry data.
[0086] It should be noted that the term frequency (TF) measures the number of times a word appears in a document. The more frequently it appears, the more important the word may be in representing the document's theme. The inverse document frequency (IDF) measures the rarity of a word in the entire document collection. If a word appears less frequently in the documents, the higher its corresponding inverse document frequency value, indicating a greater distinction between this word and other words. Hereinafter, TF is used to represent the term frequency, and IDF is used to represent the inverse document frequency.
[0087] The formula for calculating the inverse document frequency is shown in formula (1):
[0088]
[0089] where q i is the i-th word in the entry data, n(q i ) is the number of preset entry data containing the word q i , and N is the total number of preset entry data.
[0090] The word frequency calculation result TF(q i , D) is the number of times the word q i appears in the preset entry data D. For example, if the word "apple" appears 3 times in the preset entry data D, then TF(apple, D) = 3.
[0091] In this embodiment, determine the word q in the second entry data, q i is the i-th word in the second entry data, calculate the word frequency calculation result of the word q i in each second preset entry data, and calculate the inverse document frequency calculation result of the word q i in all preset entry data using formula (1).
[0092] Step S2222: According to the inverse document frequency calculation result corresponding to each word, and the word frequency calculation result corresponding to each word relative to each second preset entry data, obtain the score corresponding to each word relative to each second preset entry data.
[0093] Step S2223: According to the sum of the scores corresponding to each word relative to each second preset entry data, obtain the correlation score of each second preset entry data.
[0094] The formula for calculating the correlation score is shown in formula (2):
[0095]
[0096] Among them, |D| is the length of the preset entry data D, usually measured by the number of words, avgdl is the average length of the preset entry data, k1 and b are adjustable parameters, k1 controls the influence degree of word frequency on the score, the larger k1 is, the slower the marginal contribution of word frequency to the score increases, avoiding a single high word frequency from dominating the score. b is used to control the degree of document length normalization, the larger b is, the more significant the influence of document length on the score. n is the total number of words in the entry data, Q is the entry data code, and the explanations of IDF(q i ) and TF(q i , D) refer to the description in step S2221.
[0097] As shown in formula (2), by calculating the inverse document frequency calculation result IDF(q i ) and the product of the normalized TF(q i , D) of all words q i in the preset entry data D with respect to the entry data Q, the score of each word q i with respect to the preset entry data D is obtained. By calculating the sum of the scores of all words q i in the preset entry data D, the relevance score of the preset entry data D is obtained.
[0098] In this embodiment, by using formula (2), by calculating the inverse document frequency calculation result IDF(q i ) and the product of the normalized TF(q i , D) of all words q i in the second preset entry data D with respect to the second entry data Q, the score of each word q i with respect to the second preset entry data D is obtained. By calculating the sum of the scores of all words q i in the second preset entry data D, the relevance score of the second preset entry data D is obtained. Repeating the above steps, the relevance scores of all second preset entry data are obtained, and this is used as the basis for screening the second preset entry data related to the second entry data.
[0099] Step S2224, according to the relevance scores, sort all the second preset entry data in descending order, and select multiple second preset entry data whose rankings are before the third preset ranking.
[0100] It should be noted that the larger the value of the relevance score, the more similar the second preset entry data is to the second entry data.
[0101] In this embodiment, all the second preset entry data are sorted in descending order according to the correlation scores calculated with the second entry data Q by formula (2), and the third preset ranking is N. The top N second preset entry data are selected from the descending distribution.
[0102] Step S30: Rearrange the paragraph texts in the first recall text set and the paragraph texts in the second recall text set together to obtain the target recall text set.
[0103] It should be noted that in the context of information retrieval and question answering systems, "rearrangement" refers to reordering the results obtained from the preliminary retrieval to improve the relevance and quality of the results to the user's needs. A rearrangement model, i.e., the Rerank model, can be selected to complete the rearrangement operation. The rearrangement model quantifies and evaluates the relevance between each paragraph text and the query text based on the features extracted from the paragraph text, using its own algorithms and the knowledge learned from training, and gives a specific score.
[0104] Further, referring to Figure 5 , step S30 includes steps S31 to S34:
[0105] Step S31: Extract features from each paragraph text in the first recall text set and each paragraph text in the second recall text set respectively to obtain the lexical features, syntactic features, and semantic features of each paragraph text.
[0106] It should be noted that lexical features such as word frequency and keyword occurrence positions, syntactic features such as sentence structure and dependency relationships, and semantic features such as word vectors and text vector representations.
[0107] In this embodiment, when the rearrangement model extracts features from the paragraph texts in the first recall text set and the second recall text set, in terms of lexical feature extraction, a predefined dictionary and regular expressions are used to match words and specific pattern words, or a character-level convolutional neural network (CNN) is used to capture local character patterns, and a recurrent neural network (RNN) and its variants are used to process the character sequence order to identify word boundaries; in terms of syntactic feature extraction, a dependency syntactic analysis tool is called to analyze the dependency relationships between words and represent them as a graph, and a graph neural network (GNN) is used to process them, or part-of-speech tagging is first performed, and syntactic features are extracted based on part-of-speech and grammar rules; in terms of semantic feature extraction, a pre-trained word vector model is used to combine word vectors in an average, weighted, or other ways, or an attention mechanism is used to highlight important words to obtain the lexical, syntactic, and semantic features of the paragraph text.
[0108] Step S32: Perform weighted calculation according to the lexical features, syntactic features, and semantic features of each paragraph text, and the weight coefficients corresponding to each feature, to obtain the feature scores corresponding to each paragraph text.
[0109] In this embodiment, the rearrangement model extracts multiple features of each paragraph text in the recall text set, including lexical features, syntactic features, and semantic features. The rearrangement model then calculates the weights and other information corresponding to each feature according to its own algorithm and training, and performs weighted calculation on the lexical features, syntactic features, and semantic features of each paragraph text according to the corresponding weights to obtain the feature score corresponding to each paragraph text, so as to measure the relevance between the paragraph texts in the first recall text set and the second recall text set and the target user query text.
[0110] Step S33: Select multiple paragraph texts with feature scores greater than the preset feature score among all paragraph texts.
[0111] Step S34: Combine the selected multiple paragraph texts to obtain the target recall text set.
[0112] It should be noted that the setting of the preset feature score can be determined according to the number of paragraph texts to be selected.
[0113] In this embodiment, according to the feature scores calculated in step S32, the paragraph texts in the recall text set are sorted in descending order, and multiple paragraph texts with feature scores greater than the preset feature score are selected and combined into the target recall text set to further refine the recall texts to be input into the large language model for reasoning, and improve the accuracy of the reply texts generated by the large language model.
[0114] Step S40: Construct a prompt message according to the target recall text set and the target user query text.
[0115] Step S50: Input the prompt message into the second large language model to obtain the target answer text corresponding to the target user query text.
[0116] It should be noted that the prompt message is a text instruction input by the user to the large language model, which is used to guide the model to generate a reply of a specific type, style, or content. The second large language model refers to the large language model obtained by training with the training data set.
[0117] In this embodiment, by constructing the target recall text set and the target user query text into a prompt message according to the prompt template and inputting it into the second large language model to obtain the corresponding target answer text, the prompt message input into the second large language model can accurately express the semantic information of the target user query text, and then enable the second large language model to reply to the target user query text semantically, improving the accuracy of the target answer text.
[0118] Further, referring to Figure 6 , step S21 includes steps S211 to S214:
[0119] Step S211, use a preset embedding model to vectorize the target user query text to obtain a first vector representation of the target user query text.
[0120] It should be noted that the embedding model is a type of model widely used in the fields of machine learning and deep learning. Its core function is to map high-dimensional discrete data to a low-dimensional continuous vector space, such as words in text, class labels, etc., so as to capture the internal features and semantic information of the data. The preset embedding model is an embedding model trained using a text corpus. Here, the text corpus is a collection composed of a large number of text corpora, and the text corpora include but are not limited to web pages, books, news articles, academic papers, social media posts, encyclopedias, etc. on the Internet.
[0121] In this embodiment, the target user query text is subjected to text cleaning and word segmentation processing to remove special characters in the text and convert the text into a unified format. Then, the word vectors of relevant words in the target user query text are found through the preset embedding model, and the first vector representation corresponding to the target user query text is obtained by calculating the average value or weighted average value of all word vectors, completing the vectorization process. In this way, the semantic information of complex text is mapped into the vector space to achieve the rapid calculation of text similarity.
[0122] Step S212, retrieve multiple preset vector representations in the preset vector database whose similarity rankings with the first vector representation are before the first preset ranking, and use the preset paragraph text associated with each preset vector representation as the first paragraph text.
[0123] It should be noted that before retrieval, the text corpus is first parsed and split into paragraph data, and these paragraph data are vectorized and then stored in the preset vector database. The preset paragraph text associated with the preset vector representation is the paragraph text corresponding to the preset vector representation before vectorization processing. The first preset ranking is a preset ranking threshold.
[0124] In this embodiment, the Euclidean distance between the preset vector representations in the preset vector database and the first vector representation is calculated, and the Euclidean distance value is used as the similarity index between vectors. The smaller the Euclidean distance value, the higher the similarity between vectors. The preset vector representations in the preset vector database are sorted in ascending order according to the Euclidean distance value. The first preset ranking is M, and the M preset paragraph texts corresponding to the first M preset vector representations are selected as the first paragraph text to find multiple paragraph texts that are similar to the target user query text in the vector space dimension.
[0125] Step S213, and use a word segmentation tool to perform word segmentation processing on the target user query text to obtain the first entry data.
[0126] Step S214, retrieve multiple first preset entry data in the first preset thesaurus whose similarity rankings with the first entry data are before the second preset ranking, and use the preset paragraph text associated with each first preset entry data as the first paragraph text.
[0127] It should be noted that the first entry data is the result of word segmentation processing on the target user's query text. Before retrieval, the text corpus is first parsed and split into paragraph data, and after these paragraph data are subjected to word segmentation processing, they are stored in the first preset thesaurus. The preset paragraph text associated with the first preset entry data is the paragraph text corresponding to the second preset entry data before word segmentation processing.
[0128] In this embodiment, the Jieba word segmentation tool in Python is used to perform word segmentation processing on the target user's query text to obtain the first entry data. Then, using the word indexing method described in steps S2221 to S2224, the correlation score between each first preset entry data in the first preset thesaurus and the first entry data is calculated, and this is used as a measure of the similarity between the first preset entry data and the first entry data. The second preset ranking is Y. All the first preset entry data are sorted in descending order according to the correlation score, and the first Y first preset entry data are selected from the descending distribution, and the preset paragraph text associated with these Y first preset entry data is used as the first paragraph text.
[0129] In some embodiments of the present application, referring to Figure 7 , the text processing method further includes steps S01 to S03:
[0130] Step S01, parse and split the document data samples in the knowledge database into a set of paragraph text data.
[0131] It should be noted that the document data in the knowledge database can be sourced from any accessible data source, such as web pages, newspapers, social media, journals, books, news articles, academic papers, social media posts, encyclopedias, and so on.
[0132] In this embodiment, the paragraph splitting of the document data is achieved by identifying the line break characters or specific delimiters between paragraphs, and a set of paragraph text data is obtained.
[0133] Step S02, perform vectorization processing on the set of paragraph text data to obtain a set of paragraph text data vectors, and use the set of paragraph text data vectors to construct a preset vector database.
[0134] It should be noted that the vector database architecture can choose an open-source Faiss (Facebook AI Similarity Search, a vector similarity search library developed by Facebook) vector library or a Milvus vector library (an open-source vector database).
[0135] In this embodiment, the vectorization processing method described in the first embodiment is used to implement the vectorization processing of the paragraph text data set, obtaining a paragraph text data vector set for constructing a preset vector database to provide a recall basis for subsequent retrieval and recall of text.
[0136] Step S03, and perform word segmentation on the paragraph text data set to obtain a paragraph text data word entry set, and use the paragraph text data word entry set to construct a first preset index word library.
[0137] It should be noted that the word entry index library architecture can choose an index library such as ElasticSearch (an open-source distributed search engine based on Lucene).
[0138] In this embodiment, the word segmentation processing method described in the first embodiment is used to implement the word segmentation processing of the paragraph text data set, obtaining a paragraph text data word entry set for constructing a first preset index word library to provide a recall basis for subsequent retrieval and recall of text. The construction method of the second preset index word library can be the same as that of the first preset index word library.
[0139] This embodiment is implemented based on the first embodiment. In this embodiment, before step S10 of the text processing method, referring to Figure 8 , the following steps A10 to A40 are further included:
[0140] Step A10, obtain the current user's query text.
[0141] It should be noted that the current user's query text is the query text currently input by the user to prompt the large language model to generate a corresponding reply text.
[0142] In this embodiment, by receiving the text information input by the user in the interaction interface, such as a web search box, an input box of an application (APP, Application), etc., the current user's query text is obtained.
[0143] Step A20, determine whether there is a historical Q&A pair related to the current user's query text in the historical record.
[0144] In a feasible implementation manner, select a corresponding query method according to the storage method of the historical record data to determine whether there is a historical Q&A pair related to the current user's query text in the historical record.
[0145] In this embodiment, the historical records are stored using the file system. By reading the file content line by line and using a string matching algorithm, it is determined whether the Q&A pairs recorded in the file are relevant to the current user's query text, and further whether there are historical Q&A pairs in the historical records that are relevant to the current user's query text.
[0146] Step A30, if there are historical Q&A pairs in the historical records that are relevant to the current user's query text, generate a query text with background information based on the third large language model, the historical Q&A pairs, and the current user's query text, and use the query text with background information as the target user query text.
[0147] It should be noted that the third large language model is a large language model obtained by training using a training data set.
[0148] Step A40, if there are no historical Q&A pairs in the historical records that are relevant to the current user's query text, use the current user's query text as the target user query text.
[0149] In this embodiment, if there are historical Q&A pairs in the historical records that are relevant to the current user's query text, then these historical Q&A pairs and the current user's query text are constructed into a prompt text according to the prompt template, and passed into the first large language model to obtain a query text with background information, and use the query text with background information as the target user query text; if there are no historical Q&A pairs in the historical records that are relevant to the current user's query text, then directly use the current user's query text as the target user query text.
[0150] Exemplarily, to help understand the implementation process of the text processing method obtained by combining this embodiment with the above Embodiment 1, please refer to Figure 9 , Figure 9 A brief flow schematic diagram of a text processing method is provided. Specifically:
[0151] After receiving the current user query text input by the user, if there is a historical Q&A pair related to the current user query text in the historical record, the large language model is used to combine the historical Q&A pair and the current user query text to generate a query text with background information as the target user query text. If there is no historical Q&A pair related to the current user query text in the historical record, the current user query text is directly used as the target user query text. The large language model is used to generate a hypothetical answer text corresponding to the target user query text. The first recall text set is obtained by using vector retrieval and word retrieval on the target user query text, and the second recall text set is obtained by using word retrieval on the hypothetical answer text. The first recall text set and the second recall text set are rearranged to obtain the target recall text set. The target recall text set and the target user query text are combined with the prompt template to obtain the prompt information, and the prompt information is input into the large language model to obtain the target answer text.
[0152] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the text processing method of the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.
[0153] The present application provides a text processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the text processing method in the first embodiment above.
[0154] Refer to the following Figure 10 , which shows a schematic structural diagram of a text processing device suitable for implementing the embodiments of the present application. The text processing device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 10 The text processing device shown is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.
[0155] As Figure 10As shown, the text processing device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the text processing device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the text processing device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a text processing device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.
[0156] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0157] The text processing device provided in the present application adopts the text processing method in the above-mentioned embodiments, and can solve the technical problem of relatively low accuracy of the response text generated by the large language model. Compared with the prior art, the beneficial effects of the text processing device provided in the present application are the same as those of the text processing method provided in the above-mentioned embodiments, and other technical features in the text processing device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0158] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0159] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0160] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the text processing method in the above embodiments.
[0161] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROMs), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.
[0162] The above computer-readable storage medium can be included in the text processing device; or it can exist alone without being assembled into the text processing device.
[0163] The above computer-readable storage medium carries one or more programs, which, when executed by a text processing device, cause the text processing device to: perform semantic analysis on a target user query text through a first large language model to obtain a hypothetical answer text corresponding to the target user query text; perform retrieval and recall on the target user query text and the hypothetical answer text respectively to obtain a first recall text set corresponding to the target user query text and a second recall text set corresponding to the hypothetical answer text; rearrange the passage texts in the first recall text set and the passage texts in the second recall text set together to obtain a target recall text set; construct a prompt message according to the target recall text set and the target user query text; input the prompt message into a second large language model to obtain a target answer text corresponding to the target user query text. Since the reply text generated by the large language model based on the query text input by the user is used, combined with the query text input by the user to perform retrieval and recall in the database to obtain multiple recall texts to construct the prompt words of the large language model, and the prompt words are input into the large language model to obtain the answer corresponding to the query text.
[0164] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0165] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0166] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0167] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above text processing method, and can solve the technical problem of low accuracy of the response text generated by the large language model. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the text processing method provided by the above embodiments, and will not be elaborated here.
[0168] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. All equivalent structural transformations made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A text processing method, characterized in that: The text processing method comprises: Get the current user query text; Determine whether there is a historical question-answer pair related to the current user query text in the historical records; If there is a historical question-answer pair related to the current user query text in the historical record, generating a target user query text according to the third language model, the historical question-answer pair and the current user query text; If there is no historical question-answer pair related to the current user query text in the historical records, taking the current user query text as the target user query text; Performing semantic analysis on the target user query text by using the first language model to obtain a hypothetical answer text corresponding to the target user query text; Performing vector retrieval and word retrieval on the target user query text to obtain each first paragraph text similar to the target user query text, and combining each first paragraph text to obtain a first recalled text set; and Performing a word search on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text, and combining each second paragraph text to obtain a second recalled text set; Performing feature extraction on each paragraph text in the first recalled text set and each paragraph text in the second recalled text set to obtain lexical features, syntactic features, and semantic features of each paragraph text; A weighted calculation is performed based on the lexical features, syntactic features and semantic features of each paragraph text, as well as the weight coefficient corresponding to each feature, to obtain the feature score corresponding to each paragraph text; Selecting multiple paragraph texts whose feature scores are greater than a preset feature score from all the paragraph texts; Combine the selected multiple paragraph texts to obtain a target recall text set; Constructing prompt information according to the target recall text set and the target user query text; The prompt information is input into the second largest language model to obtain a target answer text corresponding to the target user query text.
2. The text processing method according to claim 1, characterized in that: The step of performing vector retrieval and word retrieval on the target user query text to obtain each first paragraph text similar to the target user query text comprises: Vectorizing the target user query text using a preset embedding model to obtain a first vector representation of the target user query text; Retrieving a plurality of preset vector representations whose similarity rankings to the first vector representation are before a first preset ranking in a preset vector database, and taking a preset paragraph text associated with each of the preset vector representations as the first paragraph text; and Using a word segmentation tool to perform word segmentation processing on the target user query text to obtain first word entry data; A plurality of first preset term data items whose similarity ranking with the first term data is before a second preset ranking are retrieved in a first preset index word library, and a preset paragraph text associated with each of the first preset term data items is used as the first paragraph text.
3. The text processing method according to claim 2, characterized in that: The text processing method further includes: Split the document data parsing samples in the knowledge database into paragraph text data sets; Vectorizing the paragraph text data set to obtain a paragraph text data vector set, and using the paragraph text data vector set to construct the preset vector database; and The paragraph text data set is segmented to obtain a paragraph text data entry set, and the first preset index word library is constructed using the paragraph text data entry set.
4. The text processing method according to claim 1, characterized in that: The step of performing a word search on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text comprises: Using a word segmentation tool to perform word segmentation processing on the hypothetical answer text to obtain second word entry data; A plurality of second preset term data items whose similarity ranking with the second term data is before a third preset ranking are retrieved in the second preset index word library, and a preset paragraph text associated with each second preset term data item is used as the second paragraph text.
5. The text processing method according to claim 4, characterized in that: The step of retrieving a plurality of second preset term data items whose similarity ranking with the second term data is before a third preset ranking in the second preset index word library comprises: Determine, according to the second term data and each of the second preset term data, a calculation result of the inverse document frequency corresponding to each word in the second term data, and a calculation result of the term frequency corresponding to each word relative to each of the second preset term data; Obtaining a score corresponding to each word relative to each second preset term data according to the calculation result of the inverse document frequency corresponding to each word and the calculation result of the term frequency corresponding to each word relative to each second preset term data; Obtaining a relevance score for each of the second preset term data according to the sum of the scores corresponding to each word relative to each of the second preset term data; According to the relevance score, all the second preset term data are sorted in descending order, and a plurality of second preset term data ranked before the third preset rank are selected.
6. A text processing device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the text processing method according to any one of claims 1 to 5.
7. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the text processing method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Question and answer method and device, equipment and medium
CN118445395A
RAG text processing method and device based on multi-path recall and medium
CN119167921A