Text processing method and device and storage medium

By semantic analysis and retrieval recall of the query text input by the large language model, prompt information is constructed and inputted to the second largest language model, the problem of semantic understanding difficulty when generating reply text is solved, and the accuracy of reply text is improved.

CN119990142AActive Publication Date: 2025-05-13SHENZHEN SMARTCITY TECH DEV GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510457589.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

When generating reply text, the large language model has interfering factors in the input query text, which makes it difficult to accurately understand the semantics of the query text, thereby reducing the accuracy of the reply text.

Method used

Through the first major language model, the target user query text is semantic analysis, hypothetical answer text is generated, and the query text and hypothetical answer text are searched and recalled separately to obtain a collection of recall texts. The recall text is then rearranged, prompt information is constructed and the second largest language model is entered to generate more accurate target answer text.

Benefits of technology

By refining the semantic expression of query text input to the large language model, the accuracy of the reply text generated by the large language model is improved, and the user's needs for accurate information are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990142A_ABST
    Figure CN119990142A_ABST
Patent Text Reader

Abstract

The invention discloses a text processing method and device and a storage medium, and relates to the technical field of natural language process.The text processing method comprises the steps that semantic analysis is conducted on a target user query text through a first large language model, and a hypothetical answer text corresponding to the target user query text is obtained; performing retrieval recall on the target user query text and the hypothetical answer text to obtain a first recall text set corresponding to the target user query text and a second recall text set corresponding to the hypothetical answer text; rearranging the paragraph texts in the first recalled text set and the paragraph texts in the second recalled text set to obtain a target recalled text set; according to the target recall text set and the target user query text, constructing prompt information; and inputting the prompt information into the second large language model to obtain a target answer text corresponding to the target user query text, so that the accuracy of the answer text generated by the large language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a text processing method, device and storage medium. Background Art

[0002] At present, the large language model can generate corresponding reply text for the input query text through its powerful reasoning ability. In the related technology, the query text can be generated according to the actual application needs, and the query text can be input into the large language model to generate the reply text of the query text. However, due to the presence of many interference factors in the query text input into the large language model, the large language model cannot accurately understand the semantics of the query text, which leads to the low accuracy of the reply text finally generated.

[0003] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0004] The main purpose of this application is to provide a text processing method, device and storage medium, aiming to solve the technical problem of lack of accuracy of reply text generated by large language models.

[0005] To achieve the above purpose, the present application proposes a text processing method, which includes: Perform semantic analysis on the target user's query text through the first language model to obtain a hypothetical answer text corresponding to the target user's query text; Retrieve and recall the target user query text and the hypothetical answer text respectively to obtain a first recalled text set corresponding to the target user query text and a second recalled text set corresponding to the hypothetical answer text; Rearranging the paragraph texts in the first recalled text set and the paragraph texts in the second recalled text set together to obtain a target recalled text set; Construct prompt information based on the target recall text set and the target user query text; The prompt information is input into the second largest language model to obtain the target answer text corresponding to the target user query text.

[0006] In one embodiment, the steps of respectively retrieving and recalling the target user query text and the hypothetical answer text to obtain a first recalled text set corresponding to the target user query text and a second recalled text set corresponding to the hypothetical answer text include: Performing vector retrieval and word retrieval on the target user query text to obtain each first paragraph text similar to the target user query text, and combining each first paragraph text to obtain a first recalled text set; and, A word search is performed on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text, and each second paragraph text is combined to obtain a second recalled text set.

[0007] In one embodiment, the step of performing vector search and word search on the target user query text to obtain each first paragraph text similar to the target user query text includes: Vectorize the target user's query text using a preset embedding model to obtain a first vector representation of the target user's query text; Retrieving multiple preset vector representations whose similarity rankings to the first vector representation are before a first preset ranking in the preset vector database, and taking the preset paragraph text associated with each preset vector representation as the first paragraph text; and Using a word segmentation tool to perform word segmentation processing on the target user's query text to obtain first term data; A plurality of first preset term data items whose similarity ranking with the first term data is before a second preset ranking are retrieved in the first preset index word library, and a preset paragraph text associated with each first preset term data item is used as the first paragraph text.

[0008] In one embodiment, the text processing method further includes: Split the document data parsing samples in the knowledge database into paragraph text data sets; Vectorizing the paragraph text data set to obtain a paragraph text data vector set, and using the paragraph text data vector set to construct a preset vector database; and, The paragraph text data set is segmented to obtain a paragraph text data entry set, and the paragraph text data entry set is used to construct a first preset index word library.

[0009] In one embodiment, the step of performing a word search on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text includes: Using a word segmentation tool to perform word segmentation on the hypothetical answer text to obtain second word entry data; A plurality of second preset term data items whose similarity ranking with the second term data is before a third preset ranking are retrieved in the second preset index word library, and a preset paragraph text associated with each second preset term data item is used as the second paragraph text.

[0010] In one embodiment, the step of retrieving a plurality of second preset term data items whose similarity ranking with the second term data items is before a third preset ranking in the second preset index word library comprises: Determine, based on the second entry data and each second preset entry data, a calculation result of the inverse document frequency corresponding to each word in the second entry data, and a calculation result of the word frequency corresponding to each word relative to each second preset entry data; According to the calculation result of the inverse document frequency corresponding to each word and the calculation result of the word frequency corresponding to each word relative to each second preset word data, the score corresponding to each word relative to each second preset word data is obtained; Obtaining a relevance score for each second preset term data according to the sum of the scores corresponding to each word relative to each second preset term data; According to the relevance scores, all second preset term data are sorted in descending order, and a plurality of second preset term data ranked before the third preset rank are selected.

[0011] In one embodiment, the step of rearranging the paragraph texts in the first recalled text set and the paragraph texts in the second recalled text set together to obtain the target recalled text set includes: Performing feature extraction on each paragraph text in the first recalled text set and each paragraph text in the second recalled text set to obtain lexical features, syntactic features, and semantic features of each paragraph text; A weighted calculation is performed based on the lexical features, syntactic features and semantic features of each paragraph text, as well as the weight coefficient corresponding to each feature, to obtain the feature score corresponding to each paragraph text; Select multiple paragraph texts whose feature scores are greater than a preset feature score from all paragraph texts; The selected multiple paragraph texts are combined to obtain a target recall text set.

[0012] In one embodiment, before the step of performing semantic analysis on the target user query text by using the first language model to obtain a hypothetical answer text corresponding to the target user query text, the step further includes: Get the current user query text; Determine whether there are historical question-answer pairs related to the current user's query text in the historical records; If there are historical question-answer pairs related to the current user query text in the historical records, generate the target user query text according to the third language model, the historical question-answer pairs and the current user query text; If there is no historical question-answer pair related to the current user query text in the historical records, the current user query text is used as the target user query text.

[0013] In addition, to achieve the above objectives, the present application also proposes a text processing device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the above text processing method.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the above-mentioned text processing method are implemented.

[0015] The present application performs semantic analysis on the target user query text through the first language model to obtain the hypothetical answer text corresponding to the target user query text; retrieves and recalls the target user query text and the hypothetical answer text respectively to obtain the first recall text set corresponding to the target user query text and the second recall text set corresponding to the hypothetical answer text; rearranges the paragraph text in the first recall text set and the paragraph text in the second recall text set together to obtain the target recall text set; constructs prompt information according to the target recall text set and the target user query text; inputs the prompt information into the second language model to obtain the target answer text corresponding to the target user query text. Since the hypothetical answer text is generated by the first language model based on the target user query text input by the user, the large language model can understand the semantics of the query text, and then retrieves and recalls the text based on the hypothetical answer text and the target user query text to construct rich prompt information, and inputs the prompt information into the second language model, so that the second language model can accurately generate the corresponding target answer text based on the prompt information, thereby improving the accuracy of the generated target answer text. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0018] Figure 1 A flowchart of the first embodiment of the text processing method of the present application is provided; Figure 2 A detailed flow chart of step S20 in the first embodiment of the text processing method of this application; Figure 3A detailed flow chart of step S22 in the first embodiment of the text processing method of this application; Figure 4 A detailed flow chart of step S222 in the first embodiment of the text processing method of this application; Figure 5 A detailed flow chart of step S30 in the first embodiment of the text processing method of this application; Figure 6 A detailed flow chart of step S21 in the first embodiment of the text processing method of this application; Figure 7 A flowchart diagram of the second embodiment of the text processing method of the present application; Figure 8 A flowchart diagram of the third embodiment of the text processing method of the present application; Fig. 9 This is a flowchart of the method for processing the application text; Fig.10 Schematic diagram of the device structure of the hardware operating environment involved in the text processing method in the embodiment of the present application.

[0019] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0021] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0022] At present, the large language model can generate corresponding reply text for the input query text with its excellent reasoning ability. In actual application scenarios, the query text can be constructed according to actual needs and input into the large language model to obtain the reply text corresponding to the query text. However, there are often many interference factors in the query text input into the large language model, such as ambiguous expressions in the text, a word or phrase may have multiple meanings, which makes it difficult for the model to accurately judge its accurate semantics in a specific context; or the query text has incomplete expressions and ambiguous references, which makes the model unable to obtain sufficient and clear information. In addition, the non-standard text format of the query text, such as a large number of special characters, non-standard typesetting, etc., will also interfere with the model's extraction of the core semantics of the query text. Due to the above interference factors, it is difficult for the large language model to accurately understand the true semantics of the query text, which ultimately reduces the accuracy of the reply text generated by the large language model and cannot meet the user's demand for accurate and reliable information.

[0023] In view of the above problems, the main solution of the present application is: to perform semantic analysis on the target user query text through the first largest language model to obtain the hypothetical answer text corresponding to the target user query text; to retrieve and recall the target user query text and the hypothetical answer text respectively to obtain the first recall text set corresponding to the target user query text and the second recall text set corresponding to the hypothetical answer text; to rearrange the paragraph text in the first recall text set and the paragraph text in the second recall text set together to obtain the target recall text set; to construct prompt information according to the target recall text set and the target user query text; to input the prompt information into the second largest language model to obtain the target answer text corresponding to the target user query text.

[0024] The solution of the present application uses the reply text generated by the large language model based on the query text input by the user, and retrieves and recalls the query text input by the user in the database to obtain multiple recall texts to construct the prompt words of the large language model, and inputs the prompt words into the large language model to obtain the answer corresponding to the query text. By refining the semantic expression of the query text input into the large language model, the accuracy of the reply text generated by the large language model is improved.

[0025] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a text processing device, etc. The following takes a text processing device as an example to illustrate this embodiment and the following embodiments.

[0026] It should be noted that the first language model, the second language model and the third language model can be the same large language model, or two identical large language models, or three different large language models trained from different data sets. The first preset index vocabulary and the second preset index vocabulary can be the same preset index vocabulary, or two different preset index vocabulary. The first preset ranking, the second preset ranking and the third preset ranking can be the same, or two identical, or different. The paragraph text can be a paragraph consisting of no less than one sentence.

[0027] Based on this, the present application embodiment provides a text processing method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the text processing method of this application.

[0028] In this embodiment, the text processing method includes steps S10 to S50: Step S10: semantically analyze the target user query text by using the first language model to obtain a hypothetical answer text corresponding to the target user query text.

[0029] It should be noted that the first language model can be a large language model (LLM, Large Language Model), which refers to a deep learning model trained with a large amount of text data, so that the model can generate natural language text or understand the meaning of language text. The model can provide in-depth knowledge and language production on various topics by training on a large data set. The first language model refers to a large language model trained using a training data set, where the data source of the training data set is wide, including but not limited to web pages on the Internet, books, news articles, academic papers, social media posts, encyclopedias, etc.

[0030] In addition, the target user query text is the text input into the interactive interface when the user interacts with the large language model, which is intended to clearly and accurately convey the user's intention, so that the first large language model can generate a hypothetical answer text corresponding to the target user query text.

[0031] The hypothetical answer text is the reply text output by the first language model according to the target user query text. Since there are many interference factors in the target user query text input into the first language model, the first language model cannot accurately understand the semantics of the target user query text. Therefore, the present application performs semantic analysis on the target user query text through the first language model to analyze the semantics of the target user query text and generate a hypothetical answer text. The hypothetical answer text can facilitate the subsequent second language model to accurately determine the semantics of the target user query text, so that the final generated reply text is more accurate.

[0032] In this embodiment, the first large language model receives a target user query text, and performs steps such as encoding, neural network semantic extraction, and probability distribution calculation on the target user query text to obtain a hypothetical answer text corresponding to the target user query text, so that the retrieval-augmented generation (RAG) technology can integrate the hypothetical answer text into the reasoning of the large language model to improve the response accuracy of the large language model.

[0033] Step S20 , searching and recalling the target user query text and the hypothetical answer text respectively, to obtain a first recalled text set corresponding to the target user query text and a second recalled text set corresponding to the hypothetical answer text.

[0034] It should be noted that retrieval recall refers to the work of preliminarily screening out a batch of documents that may be relevant to the user's query from a large-scale document collection, such as keyword-based recall, vector-based recall, hybrid recall, etc. The first recall text set is the set of all paragraph texts retrieved that are relevant to the target user's query text, and the second recall text set is the set of all paragraph texts retrieved that are relevant to the hypothetical answer text.

[0035] In a feasible implementation, the target user query text and the hypothetical answer text are respectively mixed and recalled to obtain a first recalled text set corresponding to the target user query text and a second recalled text set corresponding to the hypothetical answer text.

[0036] In this embodiment, keyword recall and vector recall are performed on the target user query text to obtain a first recalled text set, and keyword recall and vector recall are performed on the hypothetical answer text to obtain a second recalled text set, so as to integrate the advantages of keyword recall and vector recall, utilize the rapidity of keyword recall, and make use of the semantic understanding ability of vector recall, thereby improving the accuracy and efficiency of recall.

[0037] Further, see Figure 2 , step S20 includes steps S21-S22: Step S21 , performing vector search and word search on the target user query text, obtaining each first paragraph text similar to the target user query text, and combining each first paragraph text to obtain a first recalled text set.

[0038] It should be noted that vector retrieval converts the target user's query text into a vector representation, calculates the spatial distance between the vectors to obtain the similarity between the data, and then obtains the data with high similarity to the target user's query text as the query result; word retrieval uses the keywords related to the target user's query text to search in the database and obtains the data containing the keyword as the query result.

[0039] Step S22, and, perform a word search on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text, and combine each second paragraph text to obtain a second recalled text set.

[0040] Further, see Figure 3 , step S22 includes steps S221~S222: Step S221, using a word segmentation tool to perform word segmentation processing on the hypothetical answer text to obtain second term data.

[0041] It should be noted that word segmentation is the process of dividing a continuous text sequence into individual words or phrases according to certain rules, and the word segmentation tool is the software or library that implements this function. The entry data is a list of word segmentation results of the target user's query text. The word segmentation tool can be Jieba (a word segmentation tool used for Chinese text processing) in Python (a programming language) to segment Chinese text according to word boundaries. The second entry data is the result of word segmentation processing of the hypothetical answer text.

[0042] Step S222, searching the second preset index word library for a plurality of second preset term data whose similarity ranking to the second term data is before a third preset ranking, and taking the preset paragraph text associated with each second preset term data as the second paragraph text.

[0043] It should be noted that before retrieval, the text corpus is parsed and split into paragraph data, and these paragraph data are stored in the second preset index vocabulary after word segmentation. The preset paragraph text associated with the second preset term data is the paragraph text corresponding to the second preset term data before word segmentation.

[0044] Further, see Figure 4 , step S222 includes steps S2221~S2224: Step S2221, based on the second entry data and each second preset entry data, determine the inverse document frequency calculation result corresponding to each word in the second entry data, and the word frequency calculation result corresponding to each word relative to each second preset entry data.

[0045] It should be noted that term frequency (TF) measures the number of times a word appears in a document. The more frequently it appears, the more important the word may be in representing the topic of the document. Inverse document frequency (IDF) measures the rarity of a word in the entire document collection. The less a word appears in a document, the higher its corresponding inverse document frequency value, indicating that the word is more distinguishable from other words. In the following, TF is used to represent term frequency and IDF is used to represent inverse document frequency.

[0046] The inverse document frequency calculation formula is shown in formula (1):

[0047] Among them, q i is the i-th word in the entry data, n(q i ) contains the word q i The number of preset entry data, N is the total number of preset entry data.

[0048] The word frequency calculation result TF(q i ,D) is word q iThe number of times the word "apple" appears in the preset entry data D. For example, if the word "apple" appears 3 times in the preset entry data D, then TF(apple, D)=3.

[0049] In this embodiment, determine the word q in the second word data, q i For the i-th word in the second entry data, calculate the word q in the second entry data i The word frequency calculation result in each second preset word data, and the word q in the second word data are calculated using formula (1) i The inverse document frequency calculation result among all preset term data.

[0050] Step S2222, based on the inverse document frequency calculation result corresponding to each word and the word frequency calculation result corresponding to each word relative to each second preset term data, obtain the score corresponding to each word relative to each second preset term data.

[0051] Step S2223, obtaining a relevance score for each second preset term data according to the sum of the scores corresponding to each word with respect to each second preset term data.

[0052] The calculation formula of the correlation score is shown in formula (2):

[0053] Among them, |D| is the length of the preset entry data D, usually measured in terms of the number of words, avgdl is the average length of the preset entry data, k1 and b are adjustable parameters, k1 controls the influence of word frequency on the score, the larger the k1, the slower the marginal contribution of word frequency to the score, to avoid a single word frequency dominating the score. b is used to control the degree of document length normalization, the larger the b, the more significant the influence of document length on the score. n is the total number of words in the entry data, Q is the entry data code, IDF (q i ) and TF(q i ,D) refers to the description of step S2221.

[0054] As shown in formula (2), by calculating the preset entry data D relative to all words q in the entry data Q i The inverse document frequency calculation result IDF(q i ) and the normalized TF(q i ,D) to get each word q i Relative to the score of the preset entry data D, by calculating all the words q i The score of is summed to obtain the relevance score of the preset term data D.

[0055] In this embodiment, by using formula (2), by calculating the second preset word data D relative to all words q in the second word data Q i The inverse document frequency calculation result IDF(q i ) and the normalized TF(q i ,D), we get each word q i With respect to the score of the second preset entry data D, by calculating the scores of all words q i The scores of the second preset term data D are summed to obtain the relevance score of the second preset term data D, and the above steps are repeated to obtain the relevance scores of all second preset term data, which are used as the basis for screening the second preset term data related to the second term data.

[0056] Step S2224: sort all second preset term data in descending order according to the relevance scores, and select a plurality of second preset term data that are ranked before the third preset ranking.

[0057] It should be noted that, the greater the value of the relevance score is, the more similar the second preset term data is to the second term data.

[0058] In this embodiment, according to the correlation scores with the second term data Q calculated by formula (2) for all the second preset term data, all the second preset term data are sorted in descending order, the third preset ranking is N, and the first N second preset term data are selected in the descending order distribution.

[0059] Step S30: rearrange the paragraph texts in the first recalled text set and the paragraph texts in the second recalled text set together to obtain a target recalled text set.

[0060] It should be noted that in the context of information retrieval and question-answering systems, "reranking" refers to reordering the results obtained from the initial search to improve the relevance and quality of the results to user needs. The reranking model can be used to complete the reranking operation. The reranking model is based on the features extracted from the paragraph text, using its own algorithm and the knowledge learned from training to quantitatively evaluate the relevance of each paragraph text to the query text and give a specific score.

[0061] Further, see Figure 5 , step S30 includes steps S31 to S34: Step S31 , performing feature extraction on each paragraph text in the first recalled text set and each paragraph text in the second recalled text set, respectively, to obtain lexical features, syntactic features and semantic features of each paragraph text.

[0062] It should be noted that lexical features include word frequency and keyword occurrence position, syntactic features include sentence structure and dependency relationship, and semantic features include word vector and text vector representation.

[0063] In this embodiment, when the rearrangement model extracts features from the paragraph texts in the first recalled text set and the second recalled text set, in terms of lexical feature extraction, a predefined dictionary and regular expressions are used to match vocabulary and specific pattern words, or a character-level convolutional neural network (CNN) is used to capture local character patterns, and a recurrent neural network (RNN) and its variants are used to process the order of character sequences to identify word boundaries; in terms of syntactic feature extraction, a dependency syntactic analysis tool is called to analyze the dependency relationship between words and represent it in a graph, and a graph neural network (GNN) is used for processing, or part-of-speech tagging is performed first, and syntactic features are extracted based on part-of-speech and grammatical rules; in terms of semantic feature extraction, a pre-trained word vector model is used in combination with averaging, weighting, etc. to combine word vectors, or an attention mechanism is used to highlight important words, so as to obtain the lexical, syntactic and semantic features of the paragraph text.

[0064] Step S32, performing weighted calculation according to the lexical features, syntactic features and semantic features of each paragraph text and the weight coefficient corresponding to each feature, to obtain the feature score corresponding to each paragraph text.

[0065] In this embodiment, the rearrangement model extracts multiple features of each paragraph text in the recall text set, including lexical features, syntactic features, and semantic features. The rearrangement model then learns the weights corresponding to each feature based on its own algorithm and training, and performs weighted calculations on the lexical features, syntactic features, and semantic features of each paragraph text according to the corresponding weights to obtain the feature scores corresponding to each paragraph text, thereby measuring the correlation between the paragraph texts in the first recall text set and the second recall text set and the target user's query text.

[0066] Step S33, selecting a plurality of paragraph texts whose feature scores are greater than a preset feature score from all paragraph texts.

[0067] Step S34, combining the selected multiple paragraph texts to obtain a target recalled text set.

[0068] It should be noted that the setting of the preset feature score can be determined according to the number of paragraph texts that need to be selected.

[0069] In this embodiment, the paragraph texts in the recalled text set are sorted in descending order based on the feature scores calculated in step S32, and multiple paragraph texts with feature scores greater than the preset feature scores are selected and combined into a target recalled text set to further refine the recalled texts that need to be input into the large language model for reasoning, thereby improving the accuracy of the response text generated by the large language model.

[0070] Step S40: construct prompt information according to the target recall text set and the target user query text.

[0071] Step S50: input the prompt information into the second largest language model to obtain a target answer text corresponding to the target user query text.

[0072] It should be noted that the prompt information is the text instruction input by the user to the large language model, which is used to guide the model to generate a response of a specific type, style or content. The second largest language model refers to the large language model obtained by training using the training data set.

[0073] In this embodiment, the target recall text set and the target user query text are constructed into prompt information according to a prompt word (Prompt) template and input into the second largest language model to obtain the corresponding target answer text, so that the prompt information input into the second largest language model can accurately express the semantic information of the target user query text, and then the second largest language model can semantically reply to the target user query text, thereby improving the accuracy of the target answer text.

[0074] Further, see Figure 6 , step S21 includes steps S211 to S214: Step S211, using a preset embedding model to vectorize the target user query text to obtain a first vector representation of the target user query text.

[0075] It should be noted that the embedding model is a type of model widely used in the fields of machine learning and deep learning. Its core function is to map high-dimensional discrete data to low-dimensional continuous vector space, such as words and category labels in text, so as to capture the intrinsic characteristics and semantic information of the data. The preset embedding model is an embedding model trained using a text corpus, where the text corpus is a collection of a large number of text corpora, including but not limited to web pages, books, news articles, academic papers, social media posts, encyclopedias, etc. on the Internet.

[0076] In this embodiment, the target user query text is cleaned and segmented to remove special characters in the text and convert the text into a unified format, and then the word vectors of the relevant words in the target user query text are found through a preset embedding model, and the first vector representation corresponding to the target user query text is obtained by calculating the average or weighted average of all word vectors to complete the vectorization process. The complex text information semantics is mapped into the vector space to achieve rapid calculation of text similarity.

[0077] Step S212: Retrieve multiple preset vector representations whose similarity rankings to the first vector representation are before a first preset ranking in the preset vector database, and use the preset paragraph text associated with each preset vector representation as the first paragraph text.

[0078] It should be noted that before searching, the text corpus is parsed and split into paragraph data, and these paragraph data are vectorized and stored in the preset vector database. The preset paragraph text associated with the preset vector representation is the paragraph text corresponding to the preset vector representation before vectorization. The first preset ranking is a preset ranking threshold.

[0079] In this embodiment, the Euclidean distance between the preset vector representation in the preset vector database and the first vector representation is calculated, and the Euclidean distance value is used as the similarity index between the vectors. The smaller the Euclidean distance value, the higher the similarity between the vectors. The preset vector representations in the preset vector database are arranged in ascending order according to the Euclidean distance value. The first preset ranking is M, and the M preset paragraph texts corresponding to the first M preset vector representations are selected as the first paragraph text to find multiple paragraph texts that are similar to the target user query text in the vector space dimension.

[0080] Step S213, and using a word segmentation tool to perform word segmentation processing on the target user query text to obtain first term data.

[0081] Step S214, searching the first preset index word library for a plurality of first preset term data items whose similarity ranking to the first term data items is before a second preset ranking, and taking the preset paragraph text associated with each first preset term data item as the first paragraph text.

[0082] It should be noted that the first term data is the result of the word segmentation processing of the target user's query text. Before retrieval, the text corpus is parsed and split into paragraph data, and these paragraph data are stored in the first preset index word library after word segmentation processing. The preset paragraph text associated with the first preset term data is the corresponding paragraph text before the second preset term data is processed by word segmentation.

[0083] In this embodiment, the Jieba word segmentation tool in Python is used to perform word segmentation processing on the target user query text to obtain the first term data. Then, the word indexing method described in steps S2221 to S2224 is used to calculate the relevance score between each first preset term data and the first term data in the first preset index word library, and this is used as a similarity measurement indicator between the first preset term data and the first term data. The second preset ranking is Y, and all the first preset term data are sorted in descending order according to the relevance score. The first Y first preset term data are selected in the descending order distribution, and the preset paragraph text associated with these Y first preset term data is used as the first paragraph text.

[0084] In some embodiments of the present application, reference Figure 7 The text processing method further includes steps S01 to S03: Step S01, splitting the document data parsing samples in the knowledge database into paragraph text data sets.

[0085] It should be noted that the document data in the knowledge database can come from any accessible data source, such as web pages, newspapers, social media, journals, books, news articles, academic papers, social media posts, encyclopedias, etc.

[0086] In this embodiment, the paragraph splitting of the document data is achieved by identifying line breaks or specific separators between paragraphs to obtain a paragraph text data set.

[0087] Step S02, vectorizing the paragraph text data set to obtain a paragraph text data vector set, and using the paragraph text data vector set to construct a preset vector database.

[0088] It should be noted that the vector database architecture can use the open source Faiss (Facebook AI SimilaritySearch, a vector similarity search library developed by Facebook) vector library or the Milvus vector library (an open source vector database).

[0089] In this embodiment, the vectorization processing method described in the first embodiment is used to implement vectorization processing of a paragraph text data set to obtain a paragraph text data vector set for constructing a preset vector database to provide a recall basis for subsequent retrieval of recalled text.

[0090] Step S03, and, performing word segmentation processing on the paragraph text data set to obtain a paragraph text data entry set, and using the paragraph text data entry set to construct a first preset index word library.

[0091] It should be noted that the term index library architecture can use index libraries such as ElasticSearch (an open source distributed search engine based on Lucene).

[0092] In this embodiment, the word segmentation processing method described in the first embodiment is used to implement the word segmentation processing of the paragraph text data set, and obtain the paragraph text data entry set, which is used to construct the first preset index word library to provide a recall basis for subsequent retrieval of the recalled text. The construction method of the second preset index word library can be the same as the construction method of the first preset index word library.

[0093] This embodiment is implemented based on the first embodiment. In this embodiment, before step S10 of the text processing method, refer to Figure 8, further comprising steps A10 to A40: Step A10, obtaining the current user query text.

[0094] It should be noted that the current user query text is the query text currently input by the user, which is used to prompt the large language model to generate the corresponding reply text.

[0095] In this embodiment, the current user query text is obtained by receiving text information input by the user in an interactive interface, such as a webpage search box, an application (APP, Application) input box, etc.

[0096] Step A20, determining whether there are historical question-answer pairs related to the current user query text in the historical records.

[0097] In a feasible implementation manner, a corresponding query method is selected according to the storage method of the historical record data to determine whether there is a historical question-answer pair related to the current user query text in the historical record.

[0098] In this embodiment, the historical records are stored using the file system. The file content is read line by line and a string matching algorithm is used to determine whether the question and answer pairs recorded in the file are related to the current user query text, and then to determine whether there are historical question and answer pairs related to the current user query text in the historical records.

[0099] Step A30, if there are historical question-answer pairs related to the current user query text in the historical records, generate a query text with background information based on the third language model, the historical question-answer pairs and the current user query text, and use the query text with background information as the target user query text.

[0100] It should be noted that the third largest language model is a large language model trained using the training data set.

[0101] Step A40: If there is no historical question-answer pair related to the current user query text in the historical records, the current user query text is used as the target user query text.

[0102] In this embodiment, if there are historical question-answer pairs related to the current user query text in the historical records, these historical question-answer pairs and the current user query text are constructed into prompt texts according to the prompt word template, and passed into the first language model to obtain query texts with background information, and the query texts with background information are used as the target user query texts; if there are no historical question-answer pairs related to the current user query text in the historical records, the current user query text is directly used as the target user query text.

[0103] For example, to help understand the implementation process of the text processing method obtained by combining this embodiment with the above-mentioned embodiment 1, please refer to Fig. 9 , Fig. 9 A brief flowchart of a text processing method is provided, specifically: A current user query text input by a user is received. If there are historical question-answer pairs related to the current user query text in the historical records, a query text with background information is generated by combining the historical question-answer pairs and the current user query text with a large language model as a target user query text. If there are no historical question-answer pairs related to the current user query text in the historical records, the current user query text is directly used as the target user query text, and a hypothetical answer text corresponding to the target user query text is generated by a large language model. A first recall text set is obtained by using vector retrieval and word retrieval for the target user query text, and a second recall text set is obtained by using word retrieval for the hypothetical answer text. The first recall text set and the second recall text set are rearranged to obtain a target recall text set. The target recall text set and the target user query text are combined with a prompt word template to obtain prompt information, and the prompt information is input into the large language model to obtain a target answer text.

[0104] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the text processing method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0105] The present application provides a text processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the text processing method in the above-mentioned embodiment 1.

[0106] Reference below Fig.10 , which shows a schematic diagram of the structure of a text processing device suitable for implementing the embodiment of the present application. The text processing device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDA, Personal Digital Assistant), tablet computers (PAD, Portable Application Description), portable multimedia players (PMP, Portable Media Player), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.10 The text processing device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0107] like Fig.10 As shown, the text processing device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the text processing device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the text processing device to communicate with other devices wirelessly or by wire to exchange data. Although the text processing device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0108] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0109] The text processing device provided by the present application adopts the text processing method in the above embodiment, which can solve the technical problem of low accuracy of the reply text generated by the large language model. Compared with the prior art, the beneficial effects of the text processing device provided by the present application are the same as the beneficial effects of the text processing method provided by the above embodiment, and the other technical features in the text processing device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0110] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0111] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0112] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the text processing method in the above-mentioned embodiment.

[0113] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, Erasable Programmable ReadOnly Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequencies (RF, Radio Frequency), etc., or any suitable combination of the above.

[0114] The computer-readable storage medium may be included in the text processing device; or may exist independently without being installed in the text processing device.

[0115] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the text processing device, the text processing device: performs semantic analysis on the target user query text through the first large language model to obtain the hypothetical answer text corresponding to the target user query text; retrieves and recalls the target user query text and the hypothetical answer text respectively to obtain the first recall text set corresponding to the target user query text and the second recall text set corresponding to the hypothetical answer text; rearranges the paragraph text in the first recall text set and the paragraph text in the second recall text set together to obtain the target recall text set; constructs prompt information according to the target recall text set and the target user query text; inputs the prompt information into the second large language model to obtain the target answer text corresponding to the target user query text. Since the reply text generated by the large language model based on the query text input by the user is used, the query text input by the user is retrieved and recalled in the database to obtain multiple recall texts to construct the prompt words of the large language model, and the prompt words are input into the large language model to obtain the answer corresponding to the query text.

[0116] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0117] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0118] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0119] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned text processing method, and can solve the technical problem of low accuracy of the reply text generated by the large language model. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the text processing method provided in the above-mentioned embodiment, and will not be elaborated here.

[0120] The above are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A text processing method, characterized in that: The text processing method comprises: Performing semantic analysis on the target user query text by using the first language model to obtain a hypothetical answer text corresponding to the target user query text; Retrieve and recall the target user query text and the hypothetical answer text respectively to obtain a first recalled text set corresponding to the target user query text and a second recalled text set corresponding to the hypothetical answer text; Rearranging the paragraph texts in the first recalled text set and the paragraph texts in the second recalled text set together to obtain a target recalled text set; Constructing prompt information according to the target recall text set and the target user query text; The prompt information is input into the second largest language model to obtain a target answer text corresponding to the target user query text.

2. The text processing method according to claim 1, characterized in that: The step of respectively retrieving and recalling the target user query text and the hypothetical answer text to obtain a first recalled text set corresponding to the target user query text and a second recalled text set corresponding to the hypothetical answer text comprises: Performing vector retrieval and word retrieval on the target user query text to obtain each first paragraph text similar to the target user query text, and combining each first paragraph text to obtain the first recalled text set; and A word search is performed on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text, and each second paragraph text is combined to obtain the second recalled text set.

3. The text processing method according to claim 2, characterized in that: The step of performing vector retrieval and word retrieval on the target user query text to obtain each first paragraph text similar to the target user query text comprises: Vectorizing the target user query text using a preset embedding model to obtain a first vector representation of the target user query text; Retrieving a plurality of preset vector representations whose similarity rankings to the first vector representation are before a first preset ranking in a preset vector database, and taking a preset paragraph text associated with each of the preset vector representations as the first paragraph text; and Using a word segmentation tool to perform word segmentation processing on the target user query text to obtain first word entry data; A plurality of first preset term data items whose similarity ranking with the first term data is before a second preset ranking are retrieved in a first preset index word library, and a preset paragraph text associated with each of the first preset term data items is used as the first paragraph text.

4. The text processing method according to claim 3, characterized in that: The text processing method further includes: Split the document data parsing samples in the knowledge database into paragraph text data sets; Vectorizing the paragraph text data set to obtain a paragraph text data vector set, and using the paragraph text data vector set to construct the preset vector database; and The paragraph text data set is segmented to obtain a paragraph text data entry set, and the first preset index word library is constructed using the paragraph text data entry set.

5. The text processing method according to claim 2, characterized in that: The step of performing a word search on the hypothetical answer text to obtain each second paragraph text similar to the hypothetical answer text comprises: Using a word segmentation tool to perform word segmentation processing on the hypothetical answer text to obtain second word entry data; A plurality of second preset term data items whose similarity ranking with the second term data is before a third preset ranking are retrieved in the second preset index word library, and a preset paragraph text associated with each second preset term data item is used as the second paragraph text.

6. The text processing method according to claim 5, characterized in that: The step of retrieving a plurality of second preset term data items whose similarity ranking with the second term data is before a third preset ranking in the second preset index word library comprises: Determine, according to the second term data and each of the second preset term data, a calculation result of the inverse document frequency corresponding to each word in the second term data, and a calculation result of the term frequency corresponding to each word relative to each of the second preset term data; Obtaining a score corresponding to each word relative to each second preset term data according to the calculation result of the inverse document frequency corresponding to each word and the calculation result of the term frequency corresponding to each word relative to each second preset term data; Obtaining a relevance score for each of the second preset term data according to the sum of the scores corresponding to each word relative to each of the second preset term data; According to the relevance score, all the second preset term data are sorted in descending order, and a plurality of second preset term data ranked before the third preset rank are selected.

7. The text processing method according to claim 1, characterized in that: The step of rearranging the paragraph texts in the first recalled text set and the paragraph texts in the second recalled text set together to obtain a target recalled text set comprises: Performing feature extraction on each paragraph text in the first recalled text set and each paragraph text in the second recalled text set to obtain lexical features, syntactic features, and semantic features of each paragraph text; A weighted calculation is performed based on the lexical features, syntactic features and semantic features of each paragraph text, as well as the weight coefficient corresponding to each feature, to obtain the feature score corresponding to each paragraph text; Selecting multiple paragraph texts whose feature scores are greater than a preset feature score from all the paragraph texts; The selected multiple paragraph texts are combined to obtain the target recalled text set.

8. The text processing method according to claim 1, characterized in that: Before the step of performing semantic analysis on the target user query text by using the first language model to obtain a hypothetical answer text corresponding to the target user query text, the method further includes: Get the current user query text; Determine whether there is a historical question-answer pair related to the current user query text in the historical records; If there is a historical question-answer pair related to the current user query text in the historical record, generating the target user query text according to the third language model, the historical question-answer pair and the current user query text; If there is no historical question-answer pair related to the current user query text in the historical records, the current user query text is used as the target user query text.

9. A text processing device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the text processing method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the text processing method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Question and answer method and device, equipment and medium

    CN118445395A

  • RAG text processing method and device based on multi-path recall and medium

    CN119167921A

  • Question and answer processing method and device, electronic equipment and storage medium

    CN119202151A