Data retrieval method and device, equipment and storage medium

By calculating the word frequency and inverse document frequency of the keywords in the to-process search statements, combined with the fixed slicing document setting values, the problems of high computational cost and complex storage management in the dynamic corpus are solved, and the accuracy and efficiency of data retrieval is improved.

CN119988579APending Publication Date: 2025-05-13CHINA UNITED NETWORK COMM GRP CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510045458.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When the prior art applies the BM25 algorithm in a dynamic corpus, it faces problems such as high computing costs and complex storage management, resulting in reduced accuracy and efficiency of data retrieval.

Method used

By obtaining the pending search statement input by the user, the word frequency and inverse document frequency of each pending keyword in the document are calculated, and the relevance score of the document is determined. Set a fixed slicing document setting value to avoid recalculating the word frequency vector when the document increases, and improve retrieval efficiency.

Benefits of technology

Improve the accuracy and efficiency of data retrieval, and prioritize the display of documents with high correlation scores by analyzing the keyword contributions in the search statements to be processed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988579A_ABST
    Figure CN119988579A_ABST
Patent Text Reader

Abstract

The invention provides a data retrieval method and device, equipment and a storage medium. The method comprises the following steps: firstly, obtaining a to-be-processed retrieval statement input by a user, and then determining at least one document associated with a to-be-processed keyword in the to-be-processed retrieval statement and a correlation score in a corpus according to the to-be-processed retrieval statement, the correlation score of each document is the sum of first scores corresponding to the to-be-processed keywords, each first score is the product of a first word frequency vector and a first inverse document frequency vector, and the first word frequency vector is determined based on the word frequency value of the to-be-processed keyword in the document and a preset value of a preset segmented document; the first inverse document frequency vector is determined based on a first total number of all the documents in the corpus and a second total number of the documents containing the keywords to be processed, and finally displaying the at least one document according to the at least one document and the correlation score. According to the method provided by the invention, the accuracy and the retrieval efficiency of data retrieval are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information retrieval, and in particular to a data retrieval method, device, equipment and storage medium. Background Art

[0002] Information retrieval technology is one of the key technologies in modern data processing, especially in search engines, recommendation systems and large-scale text analysis. The performance of information retrieval algorithms directly affects the system's response speed and user experience. As a classic document ranking algorithm based on a probability model, the BM25 (Best Matching 25) algorithm has been widely used in the field of information retrieval.

[0003] In the existing technology, the application of BM25 algorithm is relatively mature in the traditional static corpus. BM25 algorithm determines the ranking score of the document by calculating the correlation between the query term and the document, thereby realizing the priority display of the most relevant documents.

[0004] However, with the emergence of dynamic corpora, the BM25 algorithm faces challenges in computational efficiency and storage management. It has problems of high computational cost and complex storage management, which leads to reduced data retrieval accuracy and retrieval efficiency. Summary of the invention

[0005] The present application provides a data retrieval method, device, equipment and storage medium to solve the technical problems of low accuracy and retrieval efficiency of data retrieval.

[0006] In a first aspect, the present application provides a data retrieval method, comprising:

[0007] Get the search statement to be processed input by the user;

[0008] According to the search sentence to be processed, at least one document associated with the keyword to be processed in the search sentence to be processed and a relevance score of the at least one document are determined in the corpus; for the relevance score of each document, the relevance score is the sum of the first scores corresponding to each keyword to be processed; for each first score, the first score is the product of a first term frequency vector and a first inverse document frequency vector, the first term frequency vector is determined based on the term frequency value of the keyword to be processed in the document and a preset setting value of the segmented document, and the first inverse document frequency vector is determined based on a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed;

[0009] The at least one document is displayed according to the at least one document and a relevance score of the at least one document.

[0010] In one or more embodiments, before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the method further includes:

[0011] For each keyword to be processed, obtaining a word frequency value of the keyword to be processed in the document and a setting value of the segmented document;

[0012] A first word frequency vector of the keyword to be processed is determined according to the word frequency value of the keyword to be processed in the document and the set value of the segmented document.

[0013] In one or more embodiments, before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the method further includes:

[0014] For each keyword to be processed, obtaining a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed;

[0015] A first inverse document frequency vector of the keyword to be processed is determined according to a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed.

[0016] In one or more embodiments, before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the method further includes:

[0017] Get the newly added source document file;

[0018] According to the setting value of the segmented document, the source document file is segmented to obtain multiple documents;

[0019] For each document, word segmentation is performed on the document to obtain multiple first words corresponding to the document.

[0020] In one or more embodiments, the method further comprises:

[0021] For each first word, updating the word frequency value of the first word in a preset vocabulary;

[0022] A first total of all documents in the corpus is updated.

[0023] In one or more embodiments, displaying the at least one document according to the at least one document and the relevance score of the at least one document includes:

[0024] The at least one document is displayed in sequence on the user interface according to the order of the relevance scores.

[0025] In one or more embodiments, the method further comprises:

[0026] In response to a user's selection operation of a target document in the at least one document, acquiring source document information corresponding to the target document;

[0027] Display the source document information corresponding to the target document.

[0028] In a second aspect, the present application provides a data retrieval device, comprising:

[0029] An acquisition module is used to acquire the search statement to be processed input by the user;

[0030] A determination module, used for determining, in a corpus, at least one document associated with a keyword to be processed in the search sentence to be processed and a relevance score of the at least one document according to the search sentence to be processed; for the relevance score of each document, the relevance score is the sum of first scores corresponding to each keyword to be processed; for each first score, the first score is the product of a first term frequency vector and a first inverse document frequency vector, the first term frequency vector is determined based on a term frequency value of the keyword to be processed in the document and a preset setting value of the segmented document, and the first inverse document frequency vector is determined based on a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed;

[0031] The display module displays the at least one document according to the at least one document and the relevance score of the at least one document.

[0032] In one or more embodiments, before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the determination module is further configured to:

[0033] For each keyword to be processed, obtaining a word frequency value of the keyword to be processed in the document and a setting value of the segmented document;

[0034] A first word frequency vector of the keyword to be processed is determined according to the word frequency value of the keyword to be processed in the document and the set value of the segmented document.

[0035] In one or more embodiments, before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the determination module is further configured to:

[0036] For each keyword to be processed, obtaining a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed;

[0037] A first inverse document frequency vector of the keyword to be processed is determined according to a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed.

[0038] In one or more embodiments, before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the determination module is further configured to:

[0039] Get the newly added source document file;

[0040] According to the setting value of the segmented document, the source document file is segmented to obtain multiple documents;

[0041] For each document, word segmentation is performed on the document to obtain multiple first words corresponding to the document.

[0042] In one or more embodiments, the determining module is further configured to:

[0043] For each first word, updating the word frequency value of the first word in a preset vocabulary;

[0044] A first total of all documents in the corpus is updated.

[0045] In one or more embodiments, the display module is specifically used for:

[0046] The at least one document is displayed in sequence on the user interface according to the order of the relevance scores.

[0047] In one or more embodiments, the display module is further used for:

[0048] In response to a user's selection operation of a target document in the at least one document, acquiring source document information corresponding to the target document;

[0049] Display the source document information corresponding to the target document.

[0050] In a third aspect, the present application provides an electronic device, including:

[0051] a processor, and a memory communicatively connected to the processor;

[0052] The memory stores computer-executable instructions;

[0053] The processor executes the computer-executable instructions stored in the memory to implement the method described in the first aspect and any one of the embodiments above.

[0054] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect and any one of the embodiments.

[0055] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the data retrieval method as described in the first aspect and various possible implementations of the first aspect.

[0056] The data retrieval method, device, equipment and storage medium provided by the present application first obtain the search sentence to be processed input by the user; then, according to the search sentence to be processed, determine at least one document associated with the keyword to be processed in the search sentence to be processed and the relevance score of at least one document in the corpus; for the relevance score of each document, the relevance score is the sum of the first scores corresponding to each keyword to be processed; for each first score, the first score is the product of the first word frequency vector and the first inverse document frequency vector, the first word frequency vector is determined based on the word frequency value of the keyword to be processed in the document and the preset setting value of the segmented document, and the first inverse document frequency vector is determined based on the first total number of all documents in the corpus and the second total number of documents containing the keyword to be processed; finally, according to the relevance score of at least one document and at least one document, display at least one document. In the above method, according to the keyword to be processed in the search sentence to be processed input by the user, determine at least one document related to it in the corpus, and according to the word frequency and inverse document frequency of each keyword to be processed in at least one document, determine the relevance score of at least one document, which can analyze the contribution of each keyword to be processed in the user's search sentence to be processed and improve the accuracy of data retrieval. The first word frequency vector is determined based on the word frequency value of the keyword to be processed in the document and the preset setting value of the segmented document. By setting a fixed setting value of the segmented document, it is possible to avoid recalculating the first word frequency vector when adding documents, thereby quickly determining the relevance score of at least one document, effectively improving data retrieval efficiency. At the same time, the at least one document retrieved is sorted and displayed according to the relevance score, and at least one document with a higher relevance score can be displayed preferentially. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0058] Figure 1 Schematic diagram of the data retrieval method provided in the embodiment of the present application Figure 1 ;

[0059] Figure 2 Schematic diagram of the data retrieval method provided in the embodiment of the present application Figure 2 ;

[0060] Figure 3 Schematic diagram of the data retrieval method provided in the embodiment of the present application Figure 3 ;

[0061] Figure 4 Schematic diagram of the data retrieval method provided in the embodiment of the present application Figure 4 ;

[0062] Figure 5 A schematic diagram of the structure of a data retrieval device provided in an embodiment of the present application;

[0063] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0064] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0065] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0066] First, the terms used in this application are explained:

[0067] Retrieval-Augmented Generation (RAG): is a model that combines retrieval and generation techniques. It generates answers or content by referencing information from an external knowledge base. It has strong interpretability and customization capabilities and is suitable for multiple natural language processing tasks such as question-answering systems, document generation, and intelligent assistants. The advantages of the RAG model are its strong versatility, instant knowledge updates, and more efficient and accurate information services through end-to-end evaluation methods.

[0068] BM25 (Best Matching 25) algorithm: is an algorithm used for information retrieval and text mining, and is widely used in search engines and related fields. The main principle of this algorithm is to parse the query and generate vocabulary, and then for each document search result, calculate the relevance score of each vocabulary and document, and perform weighted summation to obtain the relevance score between the query and the document.

[0069] Term Frequency (TF): refers to the frequency of a word appearing in a document;

[0070] Inverse Document Frequency (IDF): refers to the frequency with which a word appears in the entire corpus or file set. The inverse document frequency can also be called the anti-document frequency.

[0071] Secondly, the technical background technology involved in this application is described as follows:

[0072] Information retrieval technology is one of the key technologies in modern data processing, especially in search engines, recommendation systems and large-scale text analysis. The performance of information retrieval algorithms directly affects the system's response speed and user experience. As a classic document ranking algorithm based on a probability model, the BM25 (Best Matching 25) algorithm has been widely used in the field of information retrieval. The BM25 algorithm calculates the correlation between the query term and the document, determines the ranking score of the document, and thus prioritizes the display of the most relevant documents.

[0073] In traditional static corpora, the application of BM25 algorithm is relatively mature. The core of BM25 is to calculate the matching degree between query and document by combining term frequency (TF) and inverse document frequency (IDF). IDF is used to measure the rarity of a word in the corpus, while TF is used to measure the importance of a word in a specific document.

[0074] However, with the emergence of dynamic corpora (i.e., environments where document sets are constantly updated or expanded), the BM25 algorithm faces challenges in computational efficiency and storage management. In particular, in the Retrieval-Augmented Generation (RAG) scenario, the application of BM25 needs to efficiently perform document recall operations on a large number of dynamically changing document sets. When new documents are added to the corpus, the two core parameters in the BM25 algorithm, TF and IDF, need to be recalculated, which will lead to an increase in computational costs and increased complexity in storage management, which will further reduce the accuracy and efficiency of data retrieval.

[0075] The data retrieval method provided by the present application is intended to solve the above technical problems of the prior art. The inventive concept of the present application is as follows: when searching according to the pending search statement input by the user, it is necessary to calculate the word frequency and inverse document frequency of each pending keyword in at least one document to determine the relevance score of at least one document, and when a document is added to the corpus, it is necessary to recalculate all word frequencies and inverse document frequencies, and the calculation cost is high. In order to reduce the calculation cost, by setting a fixed setting value for segmenting the document, the word frequency of the pending keyword in at least one document is calculated, and the calculated word frequency value is stored. When adding a document, the word frequency of the corresponding pending keyword that has been stored can be directly obtained, avoiding the need to recalculate the word frequency of the pending keyword in at least one document during retrieval, and then the relevance score of at least one document can be quickly calculated, and at least one document can be displayed according to the relevance score, effectively improving the efficiency of data retrieval.

[0076] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0077] Figure 1 Schematic diagram of the data retrieval method provided in the embodiment of the present application Figure 1 .like Figure 1 As shown, the data retrieval method includes the following steps:

[0078] S110: Obtain the search statement to be processed input by the user.

[0079] In this step, the user can input a search statement to be processed according to actual needs, so as to retrieve documents related to the search statement to be processed and perform other subsequent processing.

[0080] The search sentence to be processed refers to a complete user query sentence containing at least one word.

[0081] In a possible implementation, the search sentence to be processed contains at least one keyword to be processed, and the search sentence to be processed is segmented to obtain at least one keyword to be processed, wherein the keyword to be processed can be a specific word and can contain one character, two characters, three characters, etc.

[0082] For example, when a user needs to inquire about the principle of the Tyndall effect, the user can enter the search sentence to be processed as "Tyndall's principle", segment the search sentence to be processed, and obtain the keywords to be processed as "Tyndall effect" and "principle".

[0083] S120, according to the search statement to be processed, determining in the corpus at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of at least one document;

[0084] Among them, for the relevance score of each document, the relevance score is the sum of the first scores corresponding to each keyword to be processed; for each first score, the first score is the product of the first term frequency vector and the first inverse document frequency vector, the first term frequency vector is determined based on the term frequency value of the keyword to be processed in the document and the preset setting value of the segmented document, and the first inverse document frequency vector is determined based on the first total number of all documents in the corpus and the second total number of documents containing the keyword to be processed.

[0085] In this step, firstly, the keywords to be processed in the search sentence to be processed are determined, and then the keywords to be processed are matched in the corpus to determine at least one document associated with the keywords to be processed in the search sentence to be processed.

[0086] For each document, the frequency of occurrence of the keyword to be processed in the document is counted and recorded as the word frequency value, and each document is obtained by segmenting according to the preset setting value of the segmented document. Based on the word frequency value of the keyword to be processed in the document and the preset setting value of the segmented document, the first word frequency vector is calculated according to the calculation formula of the first word frequency vector. At the same time, the total number of all documents in the corpus is recorded as the first total number, and the number of documents containing the keyword to be processed in the corpus is counted and recorded as the second total number. Based on the first total number of all documents in the corpus and the second total number of documents containing the keyword to be processed, the first inverse document frequency vector is calculated according to the first inverse document frequency vector formula.

[0087] The first word frequency vector obtained is multiplied by the first inverse document frequency vector to obtain a first score, and the first scores corresponding to each keyword to be processed are added together to finally obtain a relevance score for each document.

[0088] In one possible implementation, the first word frequency vector is recorded as , the first inverse document frequency vector is recorded as , then the first score corresponding to each keyword to be processed is , further, the relevance score of each document The calculation formula is as follows:

[0089]

[0090] Where D represents each document, Q represents the search statement to be processed, represents the i-th keyword to be processed, and n represents the number of keywords to be processed.

[0091] S130: Display at least one document according to the at least one document and the relevance score of the at least one document.

[0092] In this step, at least one document is matched with at least one relevance score of the document in a one-to-one manner, and the documents are sorted according to the relevance scores. The at least one document corresponding to the relevance score is sorted and displayed to the user.

[0093] In a possible implementation, step S130 may include the following implementation: displaying at least one document in sequence on the user interface according to the order of relevance scores.

[0094] In one possible implementation, the relevance score of at least one document may be sorted in descending order, and documents corresponding to high relevance scores contribute more to the user's query statement and may be displayed first, while documents corresponding to low relevance scores contribute less to the user's query statement and may be displayed later.

[0095] For example, when a user queries the principle of the Tyndall effect, through the above steps, the relevance score of document A is 0.78, the relevance score of document B is 0.92, and the relevance score of document C is 0.85. The relevance scores of the documents are sorted from high to low as 0.92, 0.85, and 0.78, and are displayed in the user interface in the order of document B, document C, and document A.

[0096] In a possible implementation, the method may further include the following implementations: in response to a user's selection operation on a target document in at least one document, acquiring source document information corresponding to the target document; and displaying the source document information corresponding to the target document.

[0097] Exemplarily, the user may select one of the at least one documents as a target document, and the corresponding source document information may be determined according to the target document, and the source document information may be displayed.

[0098] In a possible implementation, the target document may be the document with the highest relevance score, and the source document information corresponding to the target document may include the source and origin of the source document, and the position of the target document in the corresponding source document (which may be a chapter, paragraph or page number). The source document information may be displayed in a certain format, such as "Source X, Chapter Y, Paragraph Z".

[0099] The data retrieval method provided in the embodiment of the present application first obtains the search sentence to be processed input by the user; then, according to the search sentence to be processed, at least one document associated with the keyword to be processed in the search sentence to be processed and the relevance score of at least one document are determined in the corpus; for the relevance score of each document, the relevance score is the sum of the first scores corresponding to each keyword to be processed; for each first score, the first score is the product of the first word frequency vector and the first inverse document frequency vector, the first word frequency vector is determined based on the word frequency value of the keyword to be processed in the document and the preset setting value of the segmented document, and the first inverse document frequency vector is determined based on the first total number of all documents in the corpus and the second total number of documents containing the keyword to be processed; finally, according to the relevance score of at least one document and at least one document, at least one document is displayed. In this embodiment, according to the keyword to be processed in the search sentence to be processed input by the user, at least one document related to it in the corpus is determined, and according to the word frequency and inverse document frequency of each keyword to be processed in at least one document, the relevance score of at least one document is determined, so that the contribution of each keyword to be processed in the user's search sentence to be processed can be analyzed to improve the accuracy of data retrieval. The first word frequency vector is determined based on the word frequency value of the keyword to be processed in the document and the preset setting value of the segmented document. By setting a fixed setting value of the segmented document, it is possible to avoid recalculating the first word frequency vector when adding documents, thereby quickly determining the relevance score of at least one document, effectively improving data retrieval efficiency. At the same time, the at least one document retrieved is sorted and displayed according to the relevance score, and at least one document with a higher relevance score can be displayed preferentially.

[0100] Based on the above embodiments, Figure 2 Schematic diagram of the data retrieval method provided in the embodiment of the present application Figure 2 .like Figure 2 As shown, before the above step S120, a possible implementation includes the following steps:

[0101] S210: For each keyword to be processed, obtain the word frequency value of the keyword to be processed in the document and the setting value for segmenting the document.

[0102] In this step, the document can be segmented to obtain the total number of words in the document. For each keyword to be processed, the number of times the keyword to be processed appears in the document is counted. The number of times the keyword to be processed appears in the document is compared with the total number of words in the document. The ratio obtained is the word frequency value of the keyword to be processed in the document.

[0103] In a possible implementation, the document is segmented according to a preset fixed value, which is a set value for segmenting the document, wherein the set value for segmenting the document may be a set fixed number of characters for segmenting the document.

[0104] For example, in the large model RAG application scenario, in order to improve the retrieval efficiency, it is necessary to segment the large document into multiple documents to help the model better understand and process relevant information. The setting value of the segmented document can be a fixed number of characters or words, or a number of paragraphs or sentences. The segmented documents can be added to the corpus for user query and subsequent processing.

[0105] S220 . Determine a first word frequency vector of the keyword to be processed according to the word frequency value of the keyword to be processed in the document and a set value for segmenting the document.

[0106] In this step, the setting value of the segmented document is a preset fixed number of characters. The word frequency value of the keyword to be processed in the document and the setting value of the segmented document are substituted into the calculation formula of the first word frequency vector to obtain the first word frequency vector of the keyword to be processed.

[0107] In a possible implementation, the first word frequency vector The calculation formula is as follows:

[0108]

[0109] Where D represents each document, represents the i-th keyword to be processed, Indicates keywords to be processed The word frequency value in document D. k and b are adjustment parameters. Usually, k is between 1.2 and 2.0, b is around 0.75, and chunksize represents the preset value for segmenting documents.

[0110] In one possible implementation, the first word frequency vector of the keyword to be processed can be stored in a vector database. If a new document is added to the corpus, the first word frequency vector of the keyword to be processed can be directly obtained from the vector database, avoiding recalculation of the first word frequency vector of the keyword to be processed that has been stored, thereby reducing the calculation cost.

[0111] For example, in large-model RAG application scenarios, the RAG model needs to frequently retrieve and recall relevant information quickly and accurately from a massive corpus or document library that is constantly updated. When a document in the corpus is updated, the first word frequency vector of the keyword to be processed can be determined directly from the vector database to avoid excessive computational burden. Taking a corpus with N documents as an example, when a new document is added to the corpus, it may be necessary to update and calculate the first word frequency vector N times and calculate the first word frequency vector of the newly added document once, for a total of N+1 times. To determine the first word frequency vector of the keyword to be processed directly from the vector database, it is only necessary to calculate the first word frequency vector of the newly added document once.

[0112] The data retrieval method provided in the embodiment of the present application obtains the word frequency value of the keyword to be processed in the document and the set value of the segmented document for each keyword to be processed, and then determines the first word frequency vector of the keyword to be processed based on the word frequency value of the keyword to be processed in the document and the set value of the segmented document. In this embodiment, by obtaining the word frequency value of each keyword to be processed and the set value of the segmented document, the importance of each keyword to be processed in a single document can be accurately measured, and then the word frequency value of the keyword to be processed and the set value of the segmented document are calculated according to the calculation formula of the first word frequency vector to determine the first word frequency vector of the keyword to be processed. The use of the set value of the segmented document can effectively avoid regenerating all word frequency vectors after the document is updated, reducing calculation time and cost, and the generation of the word frequency vector reflects the distribution of each keyword to be processed in the document, which helps to improve the accuracy of data retrieval.

[0113] Based on the above embodiments, Figure 3 Schematic diagram of the data retrieval method provided in the embodiment of the present application Figure 3 .like Figure 3 As shown, before the above step S120, a possible implementation includes the following steps:

[0114] S310: For each keyword to be processed, obtain a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed.

[0115] In this step, first obtain the first total number of all documents in the corpus, and segment all documents in the corpus. Then, for each keyword to be processed, search in the corpus, match the keyword to be processed with the document segmentation in the corpus, obtain the document containing the keyword to be processed, and record the obtained number of documents as the second total number of documents containing the keyword to be processed.

[0116] In one possible implementation, all documents in the corpus need to be segmented, which can be done by a semantic segmenter. Semantic segmentation is text segmentation based on contextual semantics. It is not just segmentation based on conventional word boundaries, but takes into account context and contextual information, integrates the meaning of words into the segmentation process, avoids polysemy and ambiguity in different contexts, and fully considers the grammatical and semantic relationships between words in the segmentation process to improve the accuracy of segmentation.

[0117] For example, taking a specific sentence as an example, the sentence "the future development of artificial intelligence" is segmented according to the traditional word segmentation method, without considering the sentence context, the word segmentation results are "artificial", "intelligent", "of", "future", and "development". However, using the semantic word segmentation method, considering the sentence context, the word segmentation results are "artificial intelligence", "of", "future", and "development". Segmenting "artificial intelligence" as a whole can improve the accuracy of data retrieval.

[0118] S320: Determine a first inverse document frequency vector of the keyword to be processed according to a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed.

[0119] In this step, the first total number of all documents in the corpus and the second total number of documents containing the keyword to be processed may be substituted into the calculation formula of the first inverse document frequency vector to obtain the first inverse document frequency vector of the keyword to be processed.

[0120] In one possible implementation, the first inverse document frequency vector The calculation formula is as follows:

[0121]

[0122] Where N is the total number of all documents in the corpus. Indicates that it contains keywords to be processed The second total number of documents.

[0123] In a possible implementation, when a document is added to the corpus, the first total number of all documents in the corpus is increased by one. The second total number of documents containing the keyword to be processed may change or remain unchanged. It is necessary to segment the added document, and then match the keyword to be processed with the added document after segmentation according to the vocabulary. If the keyword to be processed is matched in the added document, the second total number of documents of the keyword to be processed is increased by one, otherwise it remains unchanged. Afterwards, the first inverse document frequency vector of the keyword to be processed is recalculated according to the calculation formula of the first inverse document frequency vector.

[0124] For example, taking the principle of the aforementioned user querying the Tyndall effect as an example, for the keyword to be processed "Tyndall effect", its first inverse document frequency vector is calculated, and the first total number of all documents in the corpus is recorded as W. The documents after word segmentation in the corpus can be matched according to the "Tyndall effect" to obtain all documents containing the "Tyndall effect", and the total number of documents is recorded as the second total number V. Substituting W and V into the calculation formula of the first inverse document frequency vector, the first inverse document frequency vector of the keyword to be processed "Tyndall effect" can be obtained. If a document is added to the corpus, the first total number of all documents is recorded as W+1. If the added document contains the keyword to be processed "Tyndall effect", the second total number V+1 is recalculated according to the calculation formula of the first inverse document frequency vector.

[0125] The data retrieval method provided in the embodiment of the present application obtains, for each keyword to be processed, a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed, and then determines a first inverse document frequency vector of the keyword to be processed based on the first total number of all documents in the corpus and the second total number of documents containing the keyword to be processed. In this embodiment, by obtaining the first total number of all documents in the corpus and the second total number of documents containing the keyword to be processed, the first inverse document frequency vector of the keyword to be processed can be calculated and determined according to the calculation formula of the first inverse document frequency vector. Through the first inverse document frequency vector, the importance of the keyword to be processed in the entire corpus can be accurately evaluated, thereby improving the accuracy of data retrieval.

[0126] Based on the above embodiments, Figure 4 Schematic diagram of the data retrieval method provided in the embodiment of the present application Figure 4 .like Figure 4 As shown, before the above step S120, a possible implementation includes the following steps:

[0127] S410: Obtain a newly added source document file.

[0128] In this step, source document files can be obtained according to different information channels, and newly added source document files can be recorded.

[0129] In a possible implementation, different information channels may include open databases, knowledge forums, learning websites, etc. The updated source document files can be obtained by regularly accessing the information channels, or scripts or automated tools can be used to regularly scan API interfaces to automatically record the information of newly added source document files. When adding new source document files, the system or manual records the incremental information to ensure that existing source document files are not recorded repeatedly.

[0130] S420: Split the source document file according to the setting value of splitting the document to obtain multiple documents.

[0131] In this step, the set value for segmenting the document refers to a fixed number of characters. The fixed set value for segmenting the document can be set according to the document size that is easy for subsequent processing. The source document file is segmented based on a unified segmentation standard to obtain multiple documents with the same number of characters.

[0132] In a possible implementation, the document segmentation method may be: segmentation by page: segmentation according to the number of pages of each document; segmentation by the number of words or characters: segmentation according to the specified number of words or characters, which is suitable for text files; segmentation by chapter or paragraph: segmentation according to the structure of the document content, such as chapters, titles, paragraphs, etc.; segmentation by specific identifiers: for example, segmentation according to certain keywords or specific formatting tags.

[0133] S430: For each document, perform word segmentation processing on the document to obtain multiple first words corresponding to the document.

[0134] In a possible implementation, for each document, after word segmentation, multiple word segments can be obtained. The multiple word segments obtained need to be further deduplicated, that is, repeated words in the multiple word segments are removed to ensure the uniqueness of each word in the document. The unique words after deduplication are recorded as multiple first words corresponding to the document.

[0135] In a possible implementation, the method may further include the following implementations: for each first word, updating the frequency value of the first word in a preset vocabulary; and updating the first total number of all documents in the corpus.

[0136] Exemplarily, each document in the corpus is segmented, and multiple words after segmentation and the word frequency value of each word in the document are stored in a preset vocabulary table. For each first word, it can be matched with the words in the preset vocabulary table, and the word frequency value of the first word in the preset vocabulary table is updated according to the matching result. At the same time, the first total number of all documents in the corpus is updated.

[0137] In a possible implementation, updating the frequency value of the first word may be based on whether the first word exists in a preset vocabulary. If the first word already exists in the preset vocabulary, its corresponding frequency value is increased by one; if the first word does not exist, the preset vocabulary is expanded, that is, the first word is added to the preset vocabulary, and the initial frequency value of the first word is set to one.

[0138] The data retrieval method provided in the embodiment of the present application first obtains a newly added source document file, then segments the source document file according to the set value of the segmented document to obtain multiple documents, and finally performs word segmentation processing on each document to obtain multiple first words corresponding to the document. In this embodiment, by segmenting the newly added source document file according to the set value of the segmented document, a document is split into multiple documents for processing, which can effectively reduce the processing complexity of a single newly added source document file, and then performs word segmentation processing on each of the split documents to obtain multiple first words corresponding to the document, so that in each split document, the semantic relationship of the text can be more accurately identified, the word segmentation accuracy can be improved, and the accuracy and efficiency of data retrieval can be improved.

[0139] Based on the above embodiments, the following are device embodiments involved in this application:

[0140] Figure 5 This is a schematic diagram of the structure of the data retrieval device provided in the embodiment of the present application. Figure 5 As shown, the data retrieval device 500 comprises:

[0141] The acquisition module 510 is used to acquire the search statement to be processed input by the user;

[0142] A determination module 520 is used to determine, in the corpus, at least one document associated with the keyword to be processed in the search sentence to be processed and a relevance score of at least one document according to the search sentence to be processed; for the relevance score of each document, the relevance score is the sum of the first scores corresponding to each keyword to be processed; for each first score, the first score is the product of a first term frequency vector and a first inverse document frequency vector, the first term frequency vector is determined based on the term frequency value of the keyword to be processed in the document and a preset setting value for segmenting the document, and the first inverse document frequency vector is determined based on a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed;

[0143] The display module 530 displays at least one document according to the at least one document and the relevance score of the at least one document.

[0144] In an optional embodiment, before determining in the corpus according to the search statement to be processed at least one document associated with the keyword to be processed in the search statement to be processed and the relevance score of at least one document, the determination module 520 is further configured to:

[0145] For each keyword to be processed, obtain the word frequency value of the keyword to be processed in the document and the setting value for segmenting the document;

[0146] According to the word frequency value of the keyword to be processed in the document and the set value of segmenting the document, a first word frequency vector of the keyword to be processed is determined.

[0147] In an optional embodiment, before determining in the corpus according to the search statement to be processed at least one document associated with the keyword to be processed in the search statement to be processed and the relevance score of at least one document, the determination module 520 is further configured to:

[0148] For each keyword to be processed, obtaining a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed;

[0149] A first inverse document frequency vector of the keyword to be processed is determined according to a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed.

[0150] In an optional embodiment, before determining in the corpus according to the search statement to be processed at least one document associated with the keyword to be processed in the search statement to be processed and the relevance score of at least one document, the determination module 520 is further configured to:

[0151] Get the newly added source document file;

[0152] According to the setting value of the split document, the source document file is split to obtain multiple documents;

[0153] For each document, the document is segmented to obtain multiple first words corresponding to the document.

[0154] In an optional embodiment, the determination module 520 is further configured to:

[0155] For each first word, updating the word frequency value of the first word in a preset vocabulary;

[0156] Updates the first total of all documents in the corpus.

[0157] In an optional embodiment, the display module 530 is specifically used for:

[0158] At least one document is displayed in sequence on the user interface according to the order of relevance scores.

[0159] In an optional embodiment, the display module 530 is further used for:

[0160] In response to a user's selection operation of a target document in at least one document, acquiring source document information corresponding to the target document;

[0161] Displays the source document information corresponding to the target document.

[0162] Based on the above embodiments, Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 6As shown, the electronic device 600 includes: a processor 610, a memory 620 and a bus 630;

[0163] The memory 620 is used to store computer-executable instructions of the processor 610;

[0164] The processor 610 is configured to execute the technical solution of any of the aforementioned method embodiments by executing computer execution instructions.

[0165] Optionally, the memory 620 may be independent or integrated with the processor 610 .

[0166] Optionally, the memory 620 may include a random access memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0167] The bus 630 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the drawings of the present application, but it does not mean that there is only one bus or one type of bus.

[0168] The above-mentioned processor can be a general-purpose processor, including a central processing unit CPU, a network processor (NP), etc.; it can also be a digital signal processor DSP, an application-specific integrated circuit ASIC, a field programmable gate array FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0169] The electronic device is used to execute the technical solution of any of the aforementioned method embodiments, and its implementation principle and technical effect are similar and will not be repeated here.

[0170] An embodiment of the present application also provides a computer-readable storage medium on which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the technical solution provided by any of the above method embodiments.

[0171] An embodiment of the present application also provides a computer program product, including a computer program, which includes computer instructions stored in a computer-readable storage medium. When the computer program is executed by a processor, it is used to implement the technical solution provided by any of the above method embodiments.

[0172] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.

[0173] It should be further noted that, although the various steps in the flowchart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0174] It should be understood that the above-mentioned device embodiments are only illustrative, and the device of the present application can also be implemented in other ways. For example, the division of units / modules in the above-mentioned embodiments is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.

[0175] In addition, unless otherwise specified, each functional unit / module in each embodiment of the present application may be integrated into one unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The above-mentioned integrated unit / module may be implemented in the form of hardware or in the form of a software program module.

[0176] If the integrated unit / module is implemented in the form of hardware, the hardware may be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the processor may be any appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, an ASIC, etc. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high-bandwidth memory (HBM), a hybrid memory cube (HMC), etc.

[0177] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0178] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0179] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0180] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A data retrieval method, characterized in that: include: Get the search statement to be processed input by the user; According to the search sentence to be processed, determining in a corpus at least one document associated with a keyword to be processed in the search sentence to be processed and a relevance score of the at least one document; For each document's relevance score, the relevance score is the sum of first scores corresponding to each keyword to be processed; for each first score, the first score is the product of a first term frequency vector and a first inverse document frequency vector, the first term frequency vector is determined based on a term frequency value of the keyword to be processed in the document and a preset setting value of the segmented document, and the first inverse document frequency vector is determined based on a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed; The at least one document is displayed according to the at least one document and a relevance score of the at least one document.

2. The method according to claim 1, characterized in that Before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the method further comprises: For each keyword to be processed, obtaining a word frequency value of the keyword to be processed in the document and a setting value of the segmented document; A first word frequency vector of the keyword to be processed is determined according to the word frequency value of the keyword to be processed in the document and the set value of the segmented document.

3. The method according to claim 1, characterized in that Before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the method further comprises: For each keyword to be processed, obtaining a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed; A first inverse document frequency vector of the keyword to be processed is determined according to a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed.

4. The method according to claims 1-3, characterized in that: Before determining, in a corpus according to the search statement to be processed, at least one document associated with the keyword to be processed in the search statement to be processed and a relevance score of the at least one document, the method further comprises: Get the newly added source document file; According to the setting value of the segmented document, the source document file is segmented to obtain multiple documents; For each document, word segmentation is performed on the document to obtain multiple first words corresponding to the document.

5. The method according to claim 4, further comprising: For each first word, updating the word frequency value of the first word in a preset vocabulary; A first total of all documents in the corpus is updated.

6. The method according to claims 1-3, characterized in that: The step of displaying the at least one document according to the at least one document and the relevance score of the at least one document comprises: The at least one document is displayed in sequence on the user interface according to the order of the relevance scores.

7. The method according to claims 1-3, characterized in that: The method further comprises: In response to a user's selection operation of a target document in the at least one document, acquiring source document information corresponding to the target document; Display the source document information corresponding to the target document.

8. A data retrieval device, characterized in that: include: An acquisition module is used to acquire the search statement to be processed input by the user; A determination module, configured to determine, in a corpus, at least one document associated with a keyword to be processed in the search statement to be processed and a relevance score of the at least one document according to the search statement to be processed; For each document's relevance score, the relevance score is the sum of first scores corresponding to each keyword to be processed; for each first score, the first score is the product of a first term frequency vector and a first inverse document frequency vector, the first term frequency vector is determined based on a term frequency value of the keyword to be processed in the document and a preset setting value of the segmented document, and the first inverse document frequency vector is determined based on a first total number of all documents in the corpus and a second total number of documents containing the keyword to be processed; The display module displays the at least one document according to the at least one document and the relevance score of the at least one document.

9. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.