Text retrieval method and device, equipment, storage medium and product

By calculating the semantic similarity and keyword similarity between the question and the candidate texts and dynamically adjusting the weights, the candidate texts with the highest comprehensive similarity are screened out, solving the problem of poor long text retrieval effect in existing technologies and achieving more efficient and accurate long text retrieval.

CN120653762APending Publication Date: 2025-09-16CHINA MOBILE SHANGHAI ICT CO LTD +2
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510785530.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing technology has the problem of poor retrieval effect in long text retrieval, especially in knowledge bases in the form of long texts or documents, where it is difficult for the existing technology to accurately perform knowledge retrieval.

Method used

By obtaining the question and candidate text input by the user, the semantic similarity and keyword similarity between the question and the candidate text are calculated, and based on the sentence features of the question and the document features of the candidate text, the weights of the semantic similarity and keyword similarity are dynamically adjusted to finally determine the comprehensive similarity and select the candidate text with the largest comprehensive similarity as the answer candidate set.

Benefits of technology

It improves the accuracy and efficiency of long text retrieval, can better match user needs, and improves the quality of retrieval results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653762A_ABST
    Figure CN120653762A_ABST
Patent Text Reader

Abstract

The invention discloses a text retrieval method and device, equipment, a storage medium and a product. The method comprises the steps of obtaining a question input by a user and at least one candidate text; for each candidate text, determining semantic similarity and keyword similarity between the question and the candidate text; sentence features of the questions and document features of the candidate texts are extracted respectively, and the weight of semantic similarity and the weight of keyword similarity are determined according to the sentence features and the document features; based on the weight of the semantic similarity and the weight of the keyword similarity, according to the semantic similarity and the keyword similarity of the question and each candidate text, calculating the comprehensive similarity of the question and each candidate text; at least one candidate text with the maximum comprehensive similarity is screened out, and an answer candidate set is obtained; and inputting the answer candidate set and the question into a preset large language model, and outputting a retrieval result. According to the method, the accuracy and efficiency of long text retrieval can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a text retrieval method, apparatus, device, storage medium, and product. Background Art

[0002] To improve the performance of large language models in specific domain tasks and overcome the knowledge lag problem of large language models, existing technologies use external knowledge bases to enable large language models to provide higher quality, more realistic, and more diverse outputs. When using external knowledge bases, accurate knowledge retrieval within the knowledge base is crucial to the effectiveness and accuracy of model output results.

[0003] In practical applications, knowledge is usually stored in the form of long texts or documents. Existing retrieval technologies have the problem of poor retrieval effect on long texts. Summary of the Invention

[0004] To solve the above technical problems, the present application provides a text retrieval method, apparatus, device, storage medium and product, which can improve the accuracy and efficiency of long text retrieval.

[0005] The present invention provides a text search method, including:

[0006] Get the question and at least one candidate text entered by the user;

[0007] For each candidate text, determining the semantic similarity and keyword similarity between the question and the candidate text;

[0008] Extracting sentence features of the question and document features of the candidate text respectively, and determining the weight of the semantic similarity and the weight of the keyword similarity according to the sentence features and the document features;

[0009] Based on the weight of the semantic similarity and the weight of the keyword similarity, determining the comprehensive similarity between the question and each of the candidate texts according to the semantic similarity and keyword similarity between the question and each of the candidate texts;

[0010] Filtering out at least one candidate text with the greatest comprehensive similarity to obtain a candidate answer set;

[0011] The answer candidate set and the question are input into a preset large language model, and a retrieval result is output.

[0012] As an improvement to the above solution, the candidate text is obtained by the following method, including:

[0013] Obtaining at least one first original text from a pre-built knowledge base;

[0014] For each first original text, determining a text distance between the question and each sentence in the first original text;

[0015] Selecting a set number of sentences as candidate sentences in ascending order of the text distance;

[0016] The first original text corresponding to the candidate sentence is used as the candidate text.

[0017] As an improvement to the above solution, the knowledge base is constructed in the following way:

[0018] Obtain at least one second original text, and segment each of the second original text into at least one sentence;

[0019] Input the sentence into a preset inter-sentence relationship discrimination model to obtain the inter-sentence relationship between the previous and next sentences;

[0020] reorganizing each sentence in the second original text according to the inter-sentence relationship to obtain the first original text;

[0021] Perform vector conversion on the first original text to obtain the knowledge base.

[0022] As an improvement to the above solution, determining the semantic similarity and keyword similarity between the question and the candidate text includes:

[0023] Determining the semantic similarity between the question and each sentence in the candidate text;

[0024] Determining the semantic similarity between the question and each of the candidate texts based on the semantic similarity between the question and each of the sentences in the candidate texts;

[0025] Extracting entity words from the candidate text, and filtering the entity words to obtain keywords of the candidate text based on word frequency and inverse document frequency algorithms;

[0026] Perform word segmentation on the question to obtain question words;

[0027] Performing word sense disambiguation and entity alignment on the keywords and the question words to obtain disambiguated keywords and question words;

[0028] The keyword similarity between the question and each of the candidate texts is determined based on the disambiguated keywords and question words.

[0029] As an improvement to the above-mentioned solution, the sentence features include sentence length, complexity and query intent; the document features include text length, keyword density and content type; the weight of the semantic similarity is positively correlated with the sentence length, complexity and text length, and negatively correlated with the keyword density; the weight of the keyword similarity is negatively correlated with the sentence length, complexity and text length, and positively correlated with the keyword density.

[0030] As an improvement to the above solution, the method further includes:

[0031] Obtaining user query data and user behavior data associated with the user query data; wherein the user query data includes at least one of a current question and a previous question, and the user behavior data includes at least one of a click-through rate, dwell time, and bounce rate of the user for each candidate text;

[0032] The weight of the semantic similarity and the weight of the keyword similarity are adjusted according to the user query data and the user behavior data.

[0033] The present application also provides a text search device, including:

[0034] A data acquisition module, configured to acquire a question input by a user and at least one candidate text;

[0035] A similarity calculation module, configured to determine, for each candidate text, the semantic similarity and keyword similarity between the question and the candidate text;

[0036] A weight determination module, configured to extract sentence features of the question and document features of the candidate text, respectively, and determine the weight of the semantic similarity and the weight of the keyword similarity based on the sentence features and the document features;

[0037] A comprehensive similarity module, configured to determine a comprehensive similarity between the question and each of the candidate texts based on the weight of the semantic similarity and the weight of the keyword similarity and according to the semantic similarity and keyword similarity between the question and each of the candidate texts;

[0038] An answer screening module is used to screen out at least one candidate text with the greatest comprehensive similarity to obtain a candidate answer set;

[0039] The retrieval result module is used to input the answer candidate set and the question into a preset large language model and output the retrieval result.

[0040] An embodiment of the present application also provides a computer device, including a processor and a memory, wherein a computer program is stored in the memory, and the computer program is configured to be executed by the processor, and when the processor executes the computer program, any of the above-mentioned text retrieval methods is implemented.

[0041] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned text retrieval methods.

[0042] An embodiment of the present application further provides a computer program product, including a computer program, which implements any of the above-mentioned text retrieval methods when executed by a processor.

[0043] Compared with the prior art, the text retrieval method, apparatus, device, storage medium, and product provided by the embodiments of the present application have the following beneficial effects:

[0044] By calculating the semantic similarity and keyword similarity between the question and the candidate texts, and determining the weight of the semantic similarity and the weight of the keyword similarity based on the sentence features of the question and the document features of the candidate texts, and then performing weighted calculations on the semantic similarity and keyword similarity based on the weights, the comprehensive similarity between the question and each candidate text is obtained, and at least one candidate text with the largest comprehensive similarity is selected as the answer candidate set, and finally the retrieval results are selected from the answer candidate set using a large language model. It can be seen from this that the present application can comprehensively consider keywords and semantics to determine the comprehensive similarity between the question and each candidate text, and then can select an accurate answer candidate set based on accurate comprehensive similarity, thereby effectively improving the accuracy of long text retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flowchart of a text retrieval method provided in an embodiment of the present application;

[0046] Figure 2 This is a schematic diagram of a relationship determination process of an inter-sentence relationship determination model provided in an embodiment of the present application;

[0047] Figure 3 This is a flow chart of entity disambiguation in a text retrieval method provided in an embodiment of the present application;

[0048] Figure 4 This is a structural diagram of a text retrieval device provided in an embodiment of the present application;

[0049] Figure 5 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0051] Large language models (LLMs) are language models trained on large text corpora and containing tens of billions or more parameters. They can perform a wide range of natural language tasks, including text classification, text translation, and entity recognition. However, LLMs also have certain drawbacks. First, as the model size grows, the computing power and data required for model fine-tuning increase exponentially, making it difficult to fine-tune the model for specific tasks once it is trained. Second, compared to other models, LLMs have higher training requirements, requiring large amounts of data and a long training cycle, resulting in a lag in model knowledge. To address this, existing technologies use external knowledge bases to enable LLMs to provide higher-quality output when performing specific tasks. To ensure the quality of LLM output, it is necessary to accurately retrieve relevant knowledge from the external knowledge base. Since most knowledge is stored in the form of long text or documents in real-world situations, existing retrieval technologies are mostly applied to short text, which leads to limited input length and low recall rates for longer texts.

[0052] Therefore, an embodiment of the present application provides a text retrieval method that can perform accurate and efficient long text retrieval based on LLM.

[0053] See also Figure 1 , Figure 1 : is a flowchart of a text search method provided in an embodiment of the present application. The text search method includes:

[0054] S1: Get the question and at least one candidate text input by the user;

[0055] S2: for each candidate text, determining the semantic similarity and keyword similarity between the question and the candidate text;

[0056] S3: extracting sentence features of the question and document features of the candidate text respectively, and determining the weight of the semantic similarity and the weight of the keyword similarity according to the sentence features and the document features;

[0057] S4: Based on the weight of the semantic similarity and the weight of the keyword similarity, and according to the semantic similarity and keyword similarity between the question and each of the candidate texts, determining a comprehensive similarity between the question and each of the candidate texts;

[0058] S5: Filter out at least one candidate text with the greatest comprehensive similarity to obtain a candidate answer set;

[0059] S6: Input the answer candidate set and the question into a preset large language model and output a retrieval result.

[0060] Specifically, before step S1, a knowledge base is pre-built, and a number of candidate texts are stored in the knowledge base. The candidate texts can be long texts and / or short texts, and are specifically stored in the form of vectorized documents. The embodiment of the present application performs text retrieval on the question input by the user and can return corresponding retrieval results as answers to the question based on the query intent of the user's question, where the retrieval results are one or more texts in the knowledge base.

[0061] In step S1, the question input by the user is obtained, and then several candidate texts are obtained from the knowledge base based on the question input by the user. In step S2, the semantic similarity and keyword similarity between the question input by the user and each candidate text in the knowledge base are calculated respectively. Among them, semantic similarity performs text matching by capturing the semantic information between the question and the candidate text. Keyword similarity can quickly and effectively evaluate the degree of matching between the question and the key information in the candidate text through specific word matching. Furthermore, the present application improves the accuracy of long text retrieval by comprehensively utilizing global semantic information and local keyword information.

[0062] Exemplarily, methods for calculating semantic similarity include, but are not limited to, methods for calculating semantic similarity based on deep learning models such as Transformer, Twin Neural Network, BERT (Bidirectional Encoder Representations from Transformers), or RoBERTa (Robustly Optimized BERT Approach), and methods for calculating semantic similarity based on word embedding models such as Word2Vec (Word to Vector), GloVe (Global Vectors for Word Representation), or FastText. Methods for calculating keyword similarity include, but are not limited to, word frequency statistics algorithm, TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, BM25 (Best Match 25) algorithm, Simhash (similarity hashing) algorithm, Jaccard similarity algorithm, and the like.

[0063] In step S3, the sentence features of the question and the document features of the candidate texts are extracted. Sentence features include sentence length, complexity, and query intent, while document features include text length, keyword density, and content type. Furthermore, based on the sentence and document features, the weights for semantic similarity and keyword similarity are determined. Based on these weights, the semantic similarity and keyword similarity between the question and each candidate text are weighted and summed to calculate the overall similarity between the question and each candidate text. Finally, in steps S4 and S5, the candidate texts are reordered based on the overall similarity, and the candidate texts with the highest overall similarity are selected to form a candidate answer set. The large language model is then used to filter the answer to the question from the candidate answer set to obtain the retrieval results.

[0064] As one of the optional embodiments, the candidate text is obtained by the following method, including:

[0065] Obtaining at least one first original text from a pre-built knowledge base;

[0066] For each first original text, determining a text distance between the question and each sentence in the first original text;

[0067] Selecting a set number of sentences as candidate sentences in ascending order of the text distance;

[0068] The first original text corresponding to the candidate sentence is used as the candidate text.

[0069] Specifically, the text stored in the knowledge base is vectorized first-source text. Each first-source text is stored in the form of a document and consists of several sentences. When a user enters a question, the question is first vectorized and encoded, converting it into a corresponding vector q. Then, the text distance between question q and each sentence in the knowledge base is calculated based on a preset distance function, including but not limited to cosine similarity, Jaccard similarity, and Euclidean distance.

[0070] For example, cosine similarity is used as the distance function to calculate the cosine similarity between the question and each sentence in the first original text, that is, the text distance. The specific formula is as follows:

[0071]

[0072] Among them, cos(·) is the cosine similarity, q is the vector corresponding to the question, e ij is the vector corresponding to the jth sentence in the i-th original text.

[0073] After calculating the text distance between the question and each sentence, the sentences in the knowledge base are sorted according to the size of the text distance, and the top k sentences are selected in the order of text distance from small to large as the candidate sentence set S = {s1, s2, ..., s k}, and then take the first original text corresponding to the candidate sentence as the candidate text set D = {D1, D2, ..., D m}, where m is the number of candidate texts, m≤k, and each candidate text includes at least one candidate sentence.

[0074] Furthermore, the candidate texts are texts recalled from the knowledge base based on the question. The embodiment of the present application re-sorts the candidate texts based on the preliminary recalled candidate texts by calculating the comprehensive similarity between the question and the candidate texts, and then filters the candidate texts according to the new sorting to obtain a candidate answer set.

[0075] As an optional embodiment, the knowledge base is constructed in the following manner:

[0076] Obtain at least one second original text, and segment each of the second original text into at least one sentence;

[0077] Input the sentence into a preset inter-sentence relationship discrimination model to obtain the inter-sentence relationship between the previous and next sentences;

[0078] reorganizing each sentence in the second original text according to the inter-sentence relationship to obtain the first original text;

[0079] Perform vector conversion on the first original text to obtain the knowledge base.

[0080] Specifically, the knowledge base is constructed from the original document set (i.e., the second original text), and the original document set is set as D ′ ={D ′ 1,D ′ 2,…,D ′ p}, p is the total number of the second original text, and the text length in the second original text is not limited and can be a long text or a short text.

[0081] First, each second original text is divided into several sentences according to the preset length limit or punctuation marks, where the sentence set corresponding to the i-th second original text is represented as D ′ i ={s i1 ,s i2 ,…,s in}, i=1,2,…,p, where s in represents the nth sentence in the i-th second original text.

[0082] Furthermore, in order to ensure the continuity and relevance of the text, an inter-sentence relationship discrimination model is constructed to identify the logical relationship between adjacent sentences, such as causality, transition, parallelism, progression, explanation, etc., and continuous sentences with strong correlation in each second original text are merged to perform sentence reorganization.

[0083] Among them, the inter-sentence relationship discrimination model can be built based on machine learning methods, such as support vector machines, logistic regression models, random forests, etc.; it can also be built based on deep learning methods, such as RNN (Recurrent Neural Network), LSTM (Long Short-term Memory Networks), GRU (Gated Recurrent Unit), Transformer, BERT, etc.

[0084] Preferably, the inter-sentence relationship discrimination model is a fusion model of an encoder and a bidirectional long short-term memory network.

[0085] See also Figure 2 , is a schematic diagram of a relationship discrimination flow of an inter-sentence relationship discrimination model provided by an embodiment of the present application. Specifically, a multi-head Transformer is used as an encoder, and the encoder is used to decode the input sentence s i Encode the sentence s i Converted into sentence vector e i, then use the bidirectional long short-term memory network Bi-Lstm to classify the continuous sentence vectors into inter-sentence relationship categories, and finally output the inter-sentence relationship pre through the Relu activation function and the Softmax function i .

[0086] The formula of the inter-sentence relationship discrimination model is as follows:

[0087] e i =Encoder(s i )

[0088] p i =BiLstm(e1,e2,…,e j ), i=1,2,…,n-1, j=1,2,…,n

[0089] pre i =softmax(relu(p i ))

[0090] Among them, s i is the i-th sentence input, e i is the i-th sentence vector output by the encoder, n is the number of sentences in the second original text, pre i e is the output of the sentence relationship discrimination model i With e i+1 The relationship between sentences.

[0091] Furthermore, based on the inter-sentence relationships between adjacent sentences in the second original text, sentences with inter-sentence relationships that fall into a preset relationship category are merged, and the n sentences in the second original text are reorganized into m sentences, resulting in a first original text consisting of the reorganized sentences. The preset relationship category is set based on the actual task requirements. For example, sentences with a progressive relationship are merged to ensure text continuity and relevance. Finally, each sentence in the first original text is converted into a vector and stored to obtain a knowledge base.

[0092] The embodiment of the present application uses a multi-head Transformer as an encoder to capture the complex semantic information of a sentence; by using a Bi-Lstm that is good at processing sequence data, it can consider the influence of the context of the sentence; and furthermore, it can accurately identify the logical relationship between sentences, reorganize the sentences in the second original text, and achieve refined segmentation of the original document, which is conducive to improving the efficiency and effectiveness of subsequent text retrieval.

[0093] As one of the optional embodiments, determining the semantic similarity and keyword similarity between the question and the candidate text includes:

[0094] Determining the semantic similarity between the question and each sentence in the candidate text;

[0095] Determining the semantic similarity between the question and each of the candidate texts based on the semantic similarity between the question and each of the sentences in the candidate texts;

[0096] Extracting entity words from the candidate text, and filtering the entity words to obtain keywords of the candidate text based on word frequency and inverse document frequency algorithms;

[0097] Perform word segmentation on the question to obtain question words;

[0098] Performing word sense disambiguation and entity alignment on the keywords and the question words to obtain disambiguated keywords and question words;

[0099] The keyword similarity between the question and each of the candidate texts is determined based on the disambiguated keywords and question words.

[0100] In an embodiment of the present application, for each candidate text, a deep learning model is used to determine the semantic similarity between the question and each sentence in the candidate text, and then the semantic similarity corresponding to all sentences in the candidate text is combined to obtain the semantic similarity between the question and the candidate text; the semantic similarity between the question and each candidate text is calculated separately.

[0101] Specifically, BERT is an improved model of Transformer. Based on BERT, this embodiment of the application trains a model for judging the semantic similarity of sentences, which is denoted as g(·). For the jth candidate text D j , candidate text D j The set of sentences contained in is D j ={s1,s2,…,s n}, s i is the i-th sentence, i=1,2,…,n, n is the candidate text D j The total number of sentences in .

[0102] In one specific example, each sentence in the candidate text represents all sentences in the first original text; in another specific example, each sentence in the candidate text represents the candidate sentence in the first original text.

[0103] Question q and candidate text D j The calculation formula of semantic similarity is as follows:

[0104] s` i =concant(s i ,q)

[0105] qi =g(s` i )

[0106] log it i =Dense(dropout(q i ))

[0107] p i =softmax(log it i )

[0108]

[0109] Among them, concant(s i ,q) represents the candidate text D j Sentences in i The concatenation with question q is judged by using seq characters; q i The encoding vector corresponding to the [CLS] character output by the encoder-decoder of the BERT model; i Perform dropout (regularization) and full connection layer operations to obtain the vector log it i ; Finally, according to the vector log it i , the question q and sentence s are obtained by calculating the activation function softmax i Furthermore, by synthesizing the candidate text D j The semantic similarity between each sentence and question q is calculated to obtain the semantic similarity between question q and candidate text D j The semantic similarity D score1 .

[0110] Furthermore, based on the word frequency and inverse document frequency TF-IDF algorithm, the keyword similarity between the question and the candidate text is calculated. j , first extract the candidate text D j The entity words in question q and the question words in question q, that is, keywords, are specifically segmented using a word segmentation algorithm to obtain the question word set A = {u1, u2, ..., u m}; for candidate text D j Perform entity recognition and obtain candidate text D j The initial keyword set B={v1,v2,...,v n}, and then use the TF-IDF algorithm to calculate the TF-IDF value of each entity word in the initial keyword set B. The formula is as follows:

[0111]

[0112] Among them, TF-IDF k is the TF-IDF value of the k-th entity word in B, n k The kth entity word in B in the candidate text D j The number of times that entity words in B appear in candidate text D is n. j The total number of times it appears in the text, N is the total number of candidate texts, N k is the number of candidate texts containing the k-th entity word in B.

[0113] After calculating the TF-IDF value of each entity word in the initial keyword set B, filter the words whose TF-IDF value is greater than the preset threshold to obtain the candidate text D j The keyword set C={c1,c2,...,c q}.

[0114] Furthermore, based on the problem word set A and the keyword set C, word sense disambiguation and entity alignment are performed to obtain the disambiguated problem word set A ` ={u ` 1,u2 ` ,...,u m `} and keyword set C ` ={c ` 1,c ` 2,...,c ` q Since a word may have multiple different meanings, for example, the word "apple" may mean "rose-like plant", "American technology company" or "song sung by Stephy Tang", etc., the embodiment of the present application uses a pre-trained classification model to distinguish the meaning of the word and determine the exact meaning of the word. Specifically, the sentence in which the word is located and the preset meaning category of the word are spliced, and the splicing result is input into the pre-trained classification model for classification to obtain the actual meaning of the word and achieve word sense disambiguation. Furthermore, according to the preset dictionaries of different meaning categories, the dictionaries of the corresponding categories are used to perform entity alignment on the words. For example, the word "apple" which means "American technology company" is classified as the same entity as the word "iPhone" of the corresponding category.

[0115] For example, see Figure 3, the sentence sentence_a containing the word "apple" is concatenated with the preset meaning category sentence_b of the word and then input into the pre-trained classification model. The classification model is composed of transformers. The input model data is first encoded in transformers to obtain variables corresponding to [CLS] characters and [seq] characters. These variables pass through the fully connected layer Dense and the activation function softmax, and finally the output result of the classification model is obtained. It is determined that the actual meaning of the word "apple" is "American technology company". Then, the entity words "iPhone", "iPhone phone", and "Apple" that represent the same meaning are aligned to achieve word sense disambiguation and entity alignment.

[0116] Furthermore, according to the disambiguated question word set A ` ={u ` 1,u2 ` ,...,u m `} and keyword set C ` ={c ` 1,c ` 2,...,c ` q}, calculate the set A ` With set C ` The similarity between them is used to obtain the candidate text D j Similarity D with the keywords of question q score2 , the specific formula is as follows:

[0117]

[0118] Furthermore, the semantic similarity D score1 Similarity with keywords D score2 Dynamic weighting is performed to extract the sentence features of the question and the document features of the candidate text respectively, so as to dynamically adjust the weights of semantic similarity and keyword similarity according to the sentence features and document features.

[0119] As one of the optional embodiments, the sentence features include sentence length, complexity and query intent; the document features include text length, keyword density and content type; the weight of the semantic similarity is positively correlated with the sentence length, complexity and text length, and negatively correlated with the keyword density; the weight of the keyword similarity is negatively correlated with the sentence length, complexity and text length, and positively correlated with the keyword density.

[0120] Specifically, the sentence features of the question include:

[0121] Sentence length: Questions can be divided into short and long queries based on their sentence length. Short queries are more likely to match keywords, while long queries require better semantic understanding. Therefore, the longer the sentence, the greater the weight of semantic similarity and the smaller the weight of keyword similarity.

[0122] Complexity: Indicates the complexity of the query. This can be analyzed using natural language processing techniques. For example, the more complex the syntactic structure and the deeper the dependency tree in the question, the higher the complexity of the query. Alternatively, through entity recognition analysis, the more entities identified in the question, the more entity matching requirements the query involves, indicating a higher complexity. The higher the complexity, the more dependent it is on semantic understanding. Therefore, the higher the complexity, the greater the weight of semantic similarity and the smaller the weight of keyword similarity.

[0123] Query intent: The query intent of a question can be divided into information query, navigation query, transaction query, etc. according to actual needs. The similarity weight ratio needs to be determined based on the specific query intent. Professionals can pre-set the weight adjustment coefficient corresponding to each query intent. For example, information query tends to semantic understanding, and the corresponding semantic similarity weight needs to be increased; navigation query tends to keyword matching, and the corresponding keyword similarity weight needs to be increased.

[0124] Document features of candidate texts include:

[0125] Text length: Texts can be divided into short documents and long documents according to their length. Short documents can be directly matched by keywords, while long documents need to rely on semantic understanding. Therefore, the longer the text, the greater the weight of semantic similarity and the smaller the weight of keyword similarity.

[0126] Keyword density: The higher the keyword density in the candidate text, the better the keyword matching effect may be. Candidate texts with scattered keywords require better contextual semantic analysis. Therefore, the higher the keyword density, the smaller the weight of semantic similarity and the greater the weight of keyword similarity.

[0127] Content type: refers to the field or type of document. Documents of different content types have different emphases on semantic matching and keyword matching. For example, technical documents or legal documents may require more semantic analysis, while other types of documents may be more suitable for simple keyword matching. The weight corresponding to the specific content type is set according to actual needs.

[0128] Furthermore, the specific calculation formula for dynamically adjusting the weights of semantic similarity and keyword similarity based on sentence features and document features can be set to a linear function, piecewise function, Sigmoid function, etc. according to actual needs. For example, a calculation formula for the similarity weight adjustment coefficient based on sentence length is as follows:

[0129]

[0130]

[0131] in, is the weight adjustment coefficient of semantic similarity based on sentence length, is the weight adjustment coefficient of keyword similarity based on sentence length, L q is the sentence length, and T is an adjustable threshold used to distinguish long and short queries.

[0132] Another calculation formula for the similarity weight adjustment coefficient based on sentence length is as follows:

[0133]

[0134] Where C is the adjustment constant.

[0135] A calculation formula for the similarity weight adjustment coefficient based on complexity is as follows:

[0136]

[0137] in, is the weight adjustment coefficient of semantic similarity based on complexity, is the weight adjustment coefficient of keyword similarity based on complexity, C q For complexity.

[0138] Another calculation formula for the similarity weight adjustment coefficient based on complexity is as follows:

[0139]

[0140] Among them, k is the adjustment parameter.

[0141] A calculation formula for the similarity weight adjustment coefficient based on text length is as follows:

[0142]

[0143]

[0144] in, is the weight adjustment coefficient of semantic similarity based on text length, is the weight adjustment coefficient of keyword similarity based on text length, L d is the text length, T d is an adjustable threshold used to distinguish long and short documents.

[0145] Another calculation formula for the similarity weight adjustment coefficient based on text length is as follows:

[0146]

[0147] Among them, C ′ is the adjustment constant.

[0148] A calculation formula for the similarity weight adjustment coefficient based on keyword density is as follows:

[0149]

[0150] in, is the weight adjustment coefficient of semantic similarity based on keyword density, is the weight adjustment coefficient of keyword similarity based on keyword density, D q is the keyword density, and D0 is the preset keyword density threshold.

[0151] Another calculation formula for the similarity weight adjustment coefficient based on keyword density is as follows:

[0152]

[0153] Finally, the corresponding weight adjustment coefficients are calculated based on the sentence features Q of the question and the document features D of the candidate text, and then the weight α(Q, D) of semantic similarity and the weight β(Q, D) of keyword similarity are obtained by combining these weight adjustment coefficients.

[0154] Furthermore, according to the weight of semantic similarity α(Q, D) and the weight of keyword similarity β(Q, D), the candidate text D j Semantic similarity D with question q score1 Similarity with keywords D score2 Perform weighted addition to obtain the comprehensive similarity D between the question and each candidate text 综合 :

[0155] D 综合 =α(Q,D)·D score1 +β(Q,D)·D score2

[0156] Further, re - sort the candidate texts according to the size of the comprehensive similarity, and select the top p (p < m) candidate texts with the largest comprehensive similarity to construct an answer candidate set. Then, splice the question, the answer candidate set and a preset prompt template and input them into a preset large - language model, and the large - language model outputs the retrieved answer corresponding to the question. Here, the internal understanding ability of the large - language model for knowledge is mainly utilized. This application does not use the large - language model for text generation, but uses the large - language model for text selection, which improves the efficiency of text retrieval. Among them, the prompt template is used to limit the large - language model to perform the selection task and restrict the model from fabricating and generating answers. Exemplarily, the answer candidate set is {doc1, doc2, … doc k}, the question is {query}, and the prompt template can be: "Please answer the question according to the known information {doc1, doc2, … doc k}, and select one or more documents from the known information as the answer according to the user's question {query}, and it is prohibited to fabricate answers", "The known retrieved answers are {doc1, doc2, … doc k}, and the user's question is {query}. Please select the corresponding answer according to the user's question. Your answer can only be one or more documents in the retrieved answers, and it is prohibited to fabricate answers", etc. It should be noted that since the user's question may have multiple intentions, the retrieved answer recalled by the large - language model can be a single document or multiple documents corresponding to multiple intentions.

[0157] Among them, the preset large - language model is obtained through model fine - tuning. The training data used for fine - tuning is historical question - answer pairs, including questions, multiple answer candidate sets and correct answer labels. Standard supervised learning methods are used for fine - tuning to maximize the accuracy of the model in the selection task, and finally the fine - tuned large - language model is obtained.

[0158] As an optional embodiment, the method further includes:

[0159] Obtain user query data and user behavior data associated with the user query data; wherein, the user query data includes at least one of the current question and previous questions, and the user behavior data includes at least one of the click - through rate, dwell time and bounce rate of the user on each of the candidate texts;

[0160] Adjust the weights of the semantic similarity and the keyword similarity according to the user query data and the user behavior data.

[0161] Specifically, in the process of user querying data, the embodiment of the present application automatically adjusts the weight of semantic similarity and the weight of keyword similarity according to the user query data and user behavior data, and specifically fine-tunes the weight through online learning or supervised learning.

[0162] The weights are fine-tuned through online learning, mainly by monitoring user behavior data, such as the click-through rate, dwell time, and bounce rate of users for each candidate text, and then dynamically adjusting the similarity weights based on the characteristics corresponding to the user behavior data. For example, if user behavior data shows that documents with more user clicks are results with higher keyword similarity, the weight of keyword similarity is increased under similar queries; if user behavior data shows that users are interested in documents with higher semantic matching, the weight of semantic similarity is increased under similar queries.

[0163] Weights are fine-tuned through supervised learning. First, based on historical user query data and a collection of candidate texts, sentence features from each user query and document features from the candidate texts are extracted to construct a feature dataset. This is then combined with user behavior data to obtain optimal weights as labels, forming a training dataset. A machine learning model is then constructed and trained using the training dataset. The machine learning model can be a linear regression, random forest, or neural network model. During supervised learning, the model predicts semantic similarity and keyword similarity weights based on the input sentence and document features. Ultimately, the similarity weights for similar queries are fine-tuned based on the model's predictions.

[0164] In this embodiment, by dynamically adjusting the similarity weights during the retrieval process based on the behavioral data of the user during the query, the retrieval results can be made more in line with the user's needs, thereby improving the accuracy of text retrieval and user experience.

[0165] Compared with the prior art, the text retrieval method provided in the embodiment of the present application calculates the semantic similarity and keyword similarity between the question and the candidate texts, and determines the weight of the semantic similarity and the weight of the keyword similarity based on the sentence features of the question and the document features of the candidate texts. Then, the semantic similarity and keyword similarity are weightedly calculated based on the weights to obtain the comprehensive similarity between the question and each candidate text, and several candidate texts with the largest comprehensive similarity are screened out as the answer candidate set. Finally, the retrieval results are selected from the answer candidate set using a large language model. The method can comprehensively consider keywords and semantics to determine the comprehensive similarity between the question and each candidate text, and then can screen out an accurate answer candidate set based on the accurate comprehensive similarity, thereby effectively improving the accuracy of long text retrieval.

[0166] By identifying the relationship between sentences in the second original text and reorganizing the sentences in the second original text, the second original text is finely segmented and a knowledge base is constructed, which is conducive to improving the efficiency of subsequent text retrieval. By first performing a preliminary search in the knowledge base to obtain candidate texts and then re-screening the candidate texts, the quality of the retrieval results is further improved.

[0167] By dynamically adjusting the weights of semantic similarity and keyword similarity based on the sentence features of the question and the document features of the candidate text, it is possible to more flexibly balance semantic matching and keyword matching in different query scenarios, thereby improving the accuracy of text retrieval and user experience.

[0168] See also Figure 4 , Figure 4 Schematic diagram of a text search device provided in an embodiment of the present application. The text search device includes:

[0169] A data acquisition module 401 is used to acquire a question input by a user and at least one candidate text;

[0170] A similarity calculation module 402 is used to determine, for each candidate text, the semantic similarity and keyword similarity between the question and the candidate text;

[0171] A weight determination module 403 is configured to extract sentence features of the question and document features of the candidate text, and determine the weight of the semantic similarity and the weight of the keyword similarity based on the sentence features and the document features;

[0172] A comprehensive similarity module 404 is configured to determine a comprehensive similarity between the question and each of the candidate texts based on the weight of the semantic similarity and the weight of the keyword similarity and the semantic similarity and keyword similarity between the question and each of the candidate texts;

[0173] An answer screening module 405 is configured to screen out at least one candidate text with the greatest comprehensive similarity to obtain a candidate answer set;

[0174] The retrieval result module 406 is used to input the answer candidate set and the question into a preset large language model and output a retrieval result.

[0175] Preferably, the text retrieval device is further used for:

[0176] Obtaining at least one first original text from a pre-built knowledge base;

[0177] For each first original text, determining a text distance between the question and each sentence in the first original text;

[0178] Selecting a set number of sentences as candidate sentences in ascending order of the text distance;

[0179] The first original text corresponding to the candidate sentence is used as the candidate text.

[0180] Preferably, the knowledge base is constructed in the following manner:

[0181] Obtain at least one second original text, and segment each of the second original text into at least one sentence;

[0182] Input the sentence into a preset inter-sentence relationship discrimination model to obtain the inter-sentence relationship between the previous and next sentences;

[0183] reorganizing each sentence in the second original text according to the inter-sentence relationship to obtain the first original text;

[0184] Perform vector conversion on the first original text to obtain the knowledge base.

[0185] Preferably, the similarity calculation module 402 is specifically used to:

[0186] Determining the semantic similarity between the question and each sentence in the candidate text;

[0187] Determining the semantic similarity between the question and each of the candidate texts based on the semantic similarity between the question and each of the sentences in the candidate texts;

[0188] Extracting entity words from the candidate text, and filtering the entity words to obtain keywords of the candidate text based on word frequency and inverse document frequency algorithms;

[0189] Perform word segmentation on the question to obtain question words;

[0190] Performing word sense disambiguation and entity alignment on the keywords and the question words to obtain disambiguated keywords and question words;

[0191] The keyword similarity between the question and each of the candidate texts is determined based on the disambiguated keywords and question words.

[0192] Preferably, the sentence features include sentence length, complexity and query intent; the document features include text length, keyword density and content type; the weight of the semantic similarity is positively correlated with the sentence length, complexity and text length, and negatively correlated with the keyword density; the weight of the keyword similarity is negatively correlated with the sentence length, complexity and text length, and positively correlated with the keyword density.

[0193] Preferably, the text retrieval device is further used for:

[0194] Obtaining user query data and user behavior data associated with the user query data; wherein the user query data includes at least one of a current question and a previous question, and the user behavior data includes at least one of a click-through rate, dwell time, and bounce rate of the user for each candidate text;

[0195] The weight of the semantic similarity and the weight of the keyword similarity are adjusted according to the user query data and the user behavior data.

[0196] In specific implementation, the working principle, control process and technical effects achieved by the text retrieval device provided in the embodiment of the present application are the same as those of the text retrieval method in the above embodiment, and will not be repeated here.

[0197] See also Figure 5 , Figure 5 5 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device includes: a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the processor 501 executes the computer program, the steps of the above-described text retrieval method embodiment are implemented. Alternatively, when the processor 501 executes the computer program, the functions of the modules / units in the above-described apparatus embodiments are implemented.

[0198] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 502 and executed by the processor 501 to implement the present application. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device.

[0199] The computer device may include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art will appreciate that the schematic diagram is merely an example of a computer device and does not limit the computer device. The computer device may include more or fewer components than shown in the diagram, or may combine certain components or different components. For example, the computer device may also include input and output devices, network access devices, buses, and the like.

[0200] The processor 501 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor 501 is the control center of the computer device, connecting various parts of the entire computer device using various interfaces and lines.

[0201] The memory 502 can be used to store the computer programs and / or modules. The processor 501 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 502 and calling the data stored in the memory 502. The memory 502 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the mobile phone, etc. In addition, the memory 502 can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0202] Wherein, if the module / unit integrated in the computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor 501, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc.

[0203] An embodiment of the present application further provides a computer-readable storage medium, which includes a stored computer program. When the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the text retrieval method described in any of the above embodiments.

[0204] An embodiment of the present application further provides a computer program product, comprising a computer program / instruction, which implements the text retrieval method described in any of the above embodiments when the computer program / instruction processor is executed.

[0205] The methods described herein can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described herein are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, a core network device, an OAM, or other programmable device.

[0206] The computer program or instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; an optical medium, such as a digital video disk; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.

[0207] The computer programs / instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0208] The computer program instructions for performing the operation of the present application can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or object code written in any combination of one or more programming languages, wherein the programming language includes object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or executed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (such as by using an Internet service provider to connect to the Internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to personalize electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLAs), the electronic circuits can execute computer-readable program instructions, thereby realizing various aspects of the present application.

[0209] Various aspects of the present application are described herein with reference to flowcharts and / or block diagrams according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0210] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0211] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0212] Each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction comprises one or more executable instructions for realizing the logical function of regulation. In some alternative implementations, the function annotated in the box can also occur in a sequence different from that annotated in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the function involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a dedicated hardware-based system that performs the function or action of regulation, or can be realized by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that realizing by hardware, realizing by software, and realizing by the combination of software and hardware are all equivalent.

[0213] The embodiments of the present application provide a text retrieval method, apparatus, device, storage medium and product, the beneficial effects of which are: by calculating the semantic similarity and keyword similarity between the question and the candidate text, and determining the weight of the semantic similarity and the weight of the keyword similarity based on the sentence features of the question and the document features of the candidate text, and then performing weighted calculation on the semantic similarity and keyword similarity based on the weight, the comprehensive similarity between the question and each candidate text is obtained, and at least one candidate text with the largest comprehensive similarity is screened out as the answer candidate set, and finally the retrieval result is selected from the answer candidate set using a large language model. It can be seen from this that the present embodiment can comprehensively consider keywords and semantics to determine the comprehensive similarity between the question and each candidate text, and then can screen out an accurate answer candidate set based on accurate comprehensive similarity, thereby effectively improving the accuracy of long text retrieval.

[0214] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications are also considered to be within the scope of protection of the present application.

Claims

1. A text retrieval method, characterized in that: include: Get the question and at least one candidate text entered by the user; For each candidate text, determining the semantic similarity and keyword similarity between the question and the candidate text; Extracting sentence features of the question and document features of the candidate text respectively, and determining the weight of the semantic similarity and the weight of the keyword similarity according to the sentence features and the document features; Based on the weight of the semantic similarity and the weight of the keyword similarity, determining the comprehensive similarity between the question and each of the candidate texts according to the semantic similarity and keyword similarity between the question and each of the candidate texts; Filtering out at least one candidate text with the greatest comprehensive similarity to obtain a candidate answer set; The answer candidate set and the question are input into a preset large language model, and a retrieval result is output.

2. The text retrieval method according to claim 1, wherein: The candidate text is obtained by the following methods, including: Obtaining at least one first original text from a pre-built knowledge base; For each first original text, determining a text distance between the question and each sentence in the first original text; Selecting a set number of sentences as candidate sentences in ascending order of the text distance; The first original text corresponding to the candidate sentence is used as the candidate text.

3. The text retrieval method according to claim 2, wherein: The knowledge base is constructed in the following way: Obtain at least one second original text, and segment each of the second original text into at least one sentence; Input the sentence into a preset inter-sentence relationship discrimination model to obtain the inter-sentence relationship between the previous and next sentences; reorganizing each sentence in the second original text according to the inter-sentence relationship to obtain the first original text; Perform vector conversion on the first original text to obtain the knowledge base.

4. The text retrieval method according to claim 1, wherein: Determining the semantic similarity and keyword similarity between the question and the candidate text includes: Determining the semantic similarity between the question and each sentence in the candidate text; Determining the semantic similarity between the question and each of the candidate texts based on the semantic similarity between the question and each of the sentences in the candidate texts; Extracting entity words from the candidate text, and filtering the entity words to obtain keywords of the candidate text based on word frequency and inverse document frequency algorithms; Perform word segmentation on the question to obtain question words; Performing word sense disambiguation and entity alignment on the keywords and the question words to obtain disambiguated keywords and question words; The keyword similarity between the question and each of the candidate texts is determined based on the disambiguated keywords and question words.

5. The text retrieval method according to claim 1, wherein: The sentence features include sentence length, complexity and query intent; the document features include text length, keyword density and content type; the weight of the semantic similarity is positively correlated with the sentence length, complexity and text length, and negatively correlated with the keyword density; the weight of the keyword similarity is negatively correlated with the sentence length, complexity and text length, and positively correlated with the keyword density.

6. The text retrieval method according to claim 1, wherein: The method further comprises: Obtaining user query data and user behavior data associated with the user query data; wherein the user query data includes at least one of a current question and a previous question, and the user behavior data includes at least one of a click-through rate, dwell time, and bounce rate of the user for each candidate text; The weight of the semantic similarity and the weight of the keyword similarity are adjusted according to the user query data and the user behavior data.

7. A text retrieval device, characterized in that: include: A data acquisition module, configured to acquire a question input by a user and at least one candidate text; A similarity calculation module, configured to determine, for each candidate text, the semantic similarity and keyword similarity between the question and the candidate text; A weight determination module, configured to extract sentence features of the question and document features of the candidate text, respectively, and determine the weight of the semantic similarity and the weight of the keyword similarity based on the sentence features and the document features; A comprehensive similarity module, configured to determine a comprehensive similarity between the question and each of the candidate texts based on the weight of the semantic similarity and the weight of the keyword similarity and according to the semantic similarity and keyword similarity between the question and each of the candidate texts; An answer screening module is used to screen out at least one candidate text with the greatest comprehensive similarity to obtain a candidate answer set; The retrieval result module is used to input the answer candidate set and the question into a preset large language model and output the retrieval result.

8. A computer device, characterized in that: The invention comprises a processor and a memory, wherein a computer program is stored in the memory and the computer program is configured to be executed by the processor, and when the processor executes the computer program, the text retrieval method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the device where the computer-readable storage medium is located executes the computer program, the text retrieval method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the text retrieval method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Intelligent customer service question and answer processing method, server and storage medium

    CN121051215A

  • Intelligent customer service question and answer processing method, server and storage medium

    CN121051215B

  • Retrieval method and device for automobile standard document

    CN121188169A

  • Document recall method and device, medium and program product

    CN121434382A

  • Bidirectional sorting-based large model question and answer method and device and storage medium

    CN121579662A