Document content retrieval method and device, computer equipment and readable storage medium
Through the combination of the statement search judgment model and the document search database, the problem of inefficient document search in traditional APIs is solved, and fast and accurate document content retrieval is achieved, which improves the work efficiency and accuracy of developers.
Patent Information
- Application Number
- CN202510103803.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional API documents are static documents, and it is difficult for developers to quickly locate and accurately understand the functional interfaces of specific needs, resulting in inefficient development and increased risk of errors.
By obtaining the query statement input by the user, using the trained statement search to determine whether the model needs to be searched, perform statement analysis, and search content with the pre-constructed document search database, and use the intelligent question and answer model to generate query results.
It improves the efficiency and accuracy of document content retrieval, reduces unnecessary search overhead, and enhances the response efficiency of simple query statements and the accuracy of query results.
Smart Images

Figure CN120045665A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a document content retrieval method, device, computer equipment and readable storage medium. Background Art
[0002] As the size of enterprise code bases and enterprise API (Application Programming Interface) documents gradually expands, developers need to quickly locate and accurately understand function interfaces for specific requirements during the software and hardware development process.
[0003] However, traditional API documents are static documents, and developers often have difficulty quickly locating the required interface functions during document content retrieval. Moreover, the located interface functions are usually not quickly and accurately understood, which reduces development efficiency and increases the risk of errors. Summary of the invention
[0004] Based on this, it is necessary to provide a document content retrieval method, apparatus, computer device and readable storage medium to address the above technical problems, which can improve the efficiency and accuracy of document content retrieval.
[0005] In a first aspect, the present application provides a document content retrieval method, comprising:
[0006] Obtain the query statement entered by the user based on the front-end interface;
[0007] When the trained sentence retrieval judgment model is used to predict that a query sentence needs to be retrieved, the query sentence is parsed to obtain target retrieval information;
[0008] Based on the pre-built document retrieval database, content retrieval is performed on the target retrieval information to obtain content retrieval results;
[0009] According to the query statement and the content retrieval result, the statement query result corresponding to the query statement is determined.
[0010] In a second aspect, the present application provides a document content retrieval device, comprising:
[0011] The acquisition module is used to obtain the query statement input by the user based on the front-end interface;
[0012] A parsing module is used to parse the query statement to obtain target retrieval information when the trained sentence retrieval judgment model is used to predict that the query statement needs to be retrieved;
[0013] A retrieval module is used to perform content retrieval on target retrieval information based on a pre-built document retrieval database to obtain content retrieval results;
[0014] The determination module is used to determine the statement query result corresponding to the query statement according to the query statement and the content retrieval result.
[0015] In a third aspect, the present application provides a computer device, the computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above method when executing the computer program.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above method when executed by a processor.
[0017] In a fifth aspect, the present application provides a computer program product, which includes a computer program, and the computer program implements the steps in the above method when executed by a processor.
[0018] The document content retrieval method, device, computer equipment and readable storage medium above parse the query statement when determining whether the query statement input by the user needs to be retrieved, perform content retrieval based on the document retrieval database, and determine the statement query result based on the content retrieval result and the query statement. Compared with the prior art, the technical solution of the present application reduces unnecessary retrieval overhead, thereby improving the response efficiency to simple query statements; improves the query efficiency and query accuracy of the statement query results, thereby improving the retrieval efficiency and retrieval accuracy of the document content. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 An application environment diagram of a document content retrieval method provided in an embodiment of the present application;
[0020] Figure 2 A flowchart of a document content retrieval method provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of the processing flow structure of a document content retrieval method provided in an embodiment of the present application;
[0022] Figure 4 A structural block diagram of a document content retrieval device provided in an embodiment of the present application;
[0023] Figure 5 An internal structure diagram of a computer device provided in an embodiment of the present application;
[0024] Figure 6 An internal structure diagram of another computer device provided in an embodiment of the present application;
[0025] Figure 7An internal structure diagram of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0027] The document content retrieval method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through a communication network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be but is not limited to various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented with an independent server or a server cluster consisting of multiple servers.
[0028] like Figure 2 As shown, the embodiment of the present application provides a document content retrieval method, which is applied to Figure 1 The terminal 102 or the server 104 in the example is used for explanation. It is understandable that the computer device may include at least one of the terminal and the server. The method comprises the following steps:
[0029] S102: Obtain the query statement input by the user based on the front-end interface.
[0030] S104: When the trained sentence retrieval judgment model is used to predict that a query sentence needs to be retrieved, the query sentence is parsed to obtain target retrieval information.
[0031] S106: Based on the pre-built document retrieval database, perform content retrieval on the target retrieval information to obtain content retrieval results.
[0032] S108: Determine a statement query result corresponding to the query statement according to the query statement and the content retrieval result.
[0033] The user may be a developer or tester who has a statement query requirement, or other non-technical personnel with query authority. The front-end interface may be a page with a statement query entry, for example, the query entry may be a statement query input box, a Boolean query button, or a drop-down query list, etc., which is not limited in this embodiment.
[0034] The query statement may be a string or text entered by the user in the front-end interface based on the user's query requirements. For example, the query statement may be "How to use the calculate_sum function to calculate the sum of the two numbers 234 and 345?".
[0035] The sentence retrieval judgment model is a pre-trained binary classification model, which is specifically used to determine whether the query sentence input by the user needs to be retrieved. This application provides a model training method for the sentence retrieval judgment model, and the specific implementation process is as follows:
[0036] Obtain historical query statements in a historical time period and use the historical query statements as query statement samples; or, artificially or randomly generate query statements as query statement samples. Perform statement labeling on each query statement in the query statement sample, specifically, label whether the query statement is a query statement that needs to be retrieved; if so, the query statement can be labeled as "1", if not, the query statement can be labeled as "0", and obtain the sample label value corresponding to each query statement in the query statement sample.
[0037] A binary classification model is pre-built, and each query statement in the query statement sample and its corresponding sample label value are input into the binary classification model for processing to obtain the prediction retrieval judgment result output by the model; the loss value is determined according to the prediction retrieval judgment result and sample label value of the corresponding query statement, and the binary classification model is trained based on the loss value until the preset model training end condition is met to obtain a sentence retrieval judgment model that has completed training. Among them, the model training end condition can be that the loss value reaches a set threshold or the loss value tends to be stable, or the number of model iterations reaches a preset iteration number threshold, etc.
[0038] It is understandable that the sentence retrieval judgment model obtained by the above model training method needs to collect a large number of training samples in advance and label a large number of training samples, so as to train the sentence judgment model based on the labeled training samples. In this process, a large amount of human resources are consumed to build training samples, resulting in low efficiency in the training sample construction process, which further reduces the training efficiency of the sentence retrieval judgment model and increases the time it takes to put the sentence retrieval judgment model into use.
[0039] In summary, the present application also provides another model training method for the sentence retrieval judgment model, which can realize the simultaneous model training and model optimization of the sentence retrieval judgment model, without the need to pre-build a large number of training samples, thereby improving the model training efficiency and further shortening the time it takes for the sentence retrieval judgment model to be put into use.
[0040] In some embodiments, after obtaining the query statement input by the user based on the front-end interface, the method further includes:
[0041] Obtaining a simulated query statement input by a user based on a front-end interface and a judgment result on whether to perform a search on the simulated query statement;
[0042] Generate a statement label of a simulated query statement according to the judgment result;
[0043] Input the simulated query statement and its corresponding statement label into the preset statement retrieval judgment model for processing, and output the retrieval prediction result corresponding to the simulated query statement;
[0044] According to the sentence labels and retrieval prediction results of the simulated query sentences, the preset sentence retrieval judgment model is trained until the first model training end condition is met to obtain a trained sentence retrieval judgment model.
[0045] The simulated query statement may be a query statement input by the user based on the statement query entry of the front-end interface, and the user may determine whether the input simulated query statement requires statement retrieval, and select the judgment result of whether the simulated query statement needs to be retrieved through the retrieval selection entry of the front-end interface. The retrieval selection entry may be a Boolean type selection button, which may include "Yes (need retrieval)" and "No (no retrieval)".
[0046] Specifically, obtain a simulated query statement and its corresponding judgment result. If the judgment result is that a search is required, the sentence label of the simulated query statement is automatically marked as "1". If the judgment result is that a search is not required, the sentence label of the simulated query statement is automatically marked as "0". Input the simulated query statement and its corresponding sentence label into the current sentence retrieval judgment model to obtain the retrieval prediction result of the simulated query statement output by the model. Determine the loss value based on the sentence label and retrieval prediction result of the simulated query statement, and perform model training on the sentence retrieval judgment model based on the loss value until the first model training end condition is met to obtain a trained sentence retrieval judgment model.
[0047] It should be noted that the current sentence retrieval judgment model can be an untrained model in an initial state, or a trained model that has undergone at least one round of iterative training. If the current sentence retrieval judgment model is a trained model, the current round of iteration process can be regarded as a model optimization process for the sentence retrieval judgment model.
[0048] It can be seen that in this embodiment, by allowing the user to input a simulated query statement and input the judgment result of the content retrieval of the simulated query statement in the front-end interface, training samples for model training or model optimization are automatically generated, thereby reducing human resource costs, realizing the simultaneous training process and optimization process of the semantic retrieval judgment model, improving the model training efficiency of the sentence retrieval judgment model, and shortening the time it takes for the sentence retrieval judgment model to be put into use.
[0049] If it is necessary to perform content retrieval on the query statement, the query statement is parsed to obtain target retrieval information. The query statement may be parsed by extracting keywords from the query statement based on a vocabulary extraction algorithm. For example, the vocabulary extraction algorithm may be an IDF (Inverse Document Frequency) algorithm. Specifically, the query statement is parsed using a pre-selected vocabulary extraction algorithm, i.e., keyword extraction, to obtain target retrieval information.
[0050] The document retrieval database is a pre-built database for storing relevant function information of at least one API document.
[0051] Specifically, the target search information is searched for content with the relevant function information of the API document in the document search database, and matching function information with a high degree of matching with the target search information is obtained, and the matching function information is used as the content search result. The query statement and its corresponding content search result can be directly used as the statement query result.
[0052] In order to further improve the query accuracy of the statement query results, statement query can also be performed based on the intelligent question-answering model. In some embodiments, according to the query statement and the content retrieval results, the statement query result corresponding to the query statement is determined, including:
[0053] Determine the target search statement according to the query statement and content search results;
[0054] The target search sentence is input into the trained search question answering model for processing, and the sentence query result corresponding to the query sentence is output.
[0055] Specifically, the content search result can be used as a prompt word of the query sentence to obtain the target search sentence.
[0056] The retrieval question-answering model may be an existing intelligent question-answering model, such as an LLM (Large Language Model), or a question-answering model pre-trained by relevant technicians based on the retrieval scenario of this embodiment.
[0057] Specifically, a sample data set is prepared, including retrieval sample sentences and standard question and answer results corresponding to the retrieval sample sentences; the retrieval sample sentences and their corresponding standard question and answer results are input into a pre-selected question and answer model (such as LLM) to obtain the predicted question and answer results output by the model; the question and answer model is trained according to the standard question and answer results and the predicted question and answer results of the retrieval sample sentences to obtain a trained retrieval question and answer model.
[0058] It can be seen that in this embodiment, by adopting a method of generating a target search sentence including a prompt word and using an intelligent question-answering model to query the target search sentence that needs to be queried, accurate determination of the sentence query result of the query sentence is achieved.
[0059] If there is no need to perform content retrieval on the query statement, the query statement can be directly input into an existing intelligent question-answering model, or input into the retrieval question-answering model that is combined with the above-mentioned scenario and trained by the model to obtain the statement query result output by the model.
[0060] It can be seen that in the embodiment of the present application, by judging whether the query statement input by the user needs to be retrieved, unnecessary retrieval overhead is reduced, thereby improving the response efficiency of simple query statements. By parsing the query statement when it is determined that a retrieval is required, performing content retrieval based on the document retrieval database, and determining the statement query result based on the content retrieval result and the query statement, the efficiency and accuracy of querying the statement query result are improved, thereby improving the efficiency and accuracy of document content retrieval.
[0061] It should be noted that, in the above-mentioned process of parsing the query statement, in order to improve the accuracy of the parsing of the statement, thereby improving the accuracy of determining the target retrieval information, the present application also provides another method of parsing the query statement.
[0062] In some embodiments, the query statement is parsed to obtain target retrieval information, including:
[0063] The query statement is input into the trained named entity recognition model for processing and the target retrieval information is output.
[0064] The named entity recognition model is pre-trained by relevant technical personnel based on the retrieval scenario of API documents. The named entity recognition model is used to identify and predict the keywords in the query sentence that are related to the API document retrieval scenario.
[0065] By using the named entity recognition model to parse the query statement to obtain the target retrieval information, the retrieval scenario of the API document is introduced into the model parsing process, so that the sentence parsing process is more accurate. And due to the limitation of the model scenario of the named entity recognition model, the prediction efficiency is higher than other text recognition models.
[0066] The present application also provides a model training method for a named entity recognition model. In some embodiments, after obtaining a query statement input by a user based on a front-end interface, the method further includes:
[0067] Performing named entity category annotation on the acquired training text sample data to obtain standard named entity categories of the training text sample data;
[0068] Input the training text sample data and its corresponding standard named entity category into a preset named entity recognition model for processing, and output the predicted named entity category;
[0069] According to the standard named entity categories and predicted named entity categories of the training text sample data, the preset named entity recognition model is trained until the second model training end condition is met to obtain a trained named entity recognition model.
[0070] Among them, the training text sample data can be randomly generated based on the API documentation in the current scenario, or manually generated. Standard named entity categories can include named entities and entity categories. For example, if the training text sample data is "How to use the calculate_sum function to calculate the sum of the two numbers num1 and num2?", then the named entities include "calculate_sum function", "num1", "num2" and "Use the calculate_sum function to calculate the sum of the two numbers num1 and num2". Among them, the entity category of the named entity "calculate_sum function" can be "function name"; the entity category of the named entities "num1" and "num2" can be "parameter name"; the entity category of the named entity "Use the calculate_sum function to calculate the sum of the two numbers num1 and num2" can be "requirement description".
[0071] The training text sample data and its corresponding standard named entity category are input into a preset named entity recognition model (Named Entity Recognition, NER) to obtain the predicted named entity category output by the model. The named entity recognition model can be a NER model based on deep learning, such as BERT (Bidirectional Encoder Representations from Transformers).
[0072] According to the standard named entity categories and predicted named entity categories of the training text sample data, the loss value is determined, and the named entity recognition model is trained based on the loss value until the second model training end condition is met, thereby obtaining a trained named entity recognition model. The second model training end condition may be that the loss value tends to be stable, or that the loss value reaches a set loss threshold, or that the number of model iterations reaches a set iteration number threshold, which is not limited in this embodiment.
[0073] It can be seen that the above technical solution pre-trains the named entity recognition model by combining the named entity prediction scenario of the API document, so that the trained named entity recognition model can predict the target retrieval information of the query statement based on the API document in this embodiment with higher accuracy, thereby further improving the accuracy of the statement query results of subsequent query statements.
[0074] It should be noted that since the present application performs content retrieval on target retrieval information based on data stored in a document retrieval database, the present application also provides a method for constructing a document retrieval database. In some embodiments, before performing content retrieval on target retrieval information based on a pre-constructed document retrieval database and obtaining content retrieval results, the method further includes:
[0075] Input the acquired document to be processed into the trained named entity recognition model for processing, and output document retrieval information; the document retrieval information includes function keyword information and function description information corresponding to at least one function block;
[0076] Generate a keyword index of function keyword information of each function block, and generate a vector index of function description information of each function block;
[0077] The function keyword information of each function block is stored in a pre-built document retrieval database based on the keyword index, and the function description information of each function block is stored in the pre-built document retrieval database based on the vector index.
[0078] Among them, since the processing logic and processing subject of the document content processing of the document to be processed and the statement content processing of the query statement are the same, therefore, in the construction stage of the document retrieval database, the named entity recognition model used when processing the document to be processed and the named entity recognition model used when processing the query statement can be the same model.
[0079] Specifically, the trained named entity recognition model is used to parse the content of the document to be processed to obtain at least one structured function block, as well as function keyword information and function description information corresponding to each function block. The function keyword information may include function name and function parameters; the function description information may include function requirement description and function usage example description. Exemplarily, the specific content of the document retrieval information may be as follows:
[0080] Function name: used for quick positioning function;
[0081] Function parameters: important content as search keywords;
[0082] Function requirement description: describes the function's function and usage scenarios, suitable for semantic retrieval;
[0083] Function usage example description: Code examples to help users understand the specific usage of the function.
[0084] Specifically, in order to facilitate the subsequent information processing of the document retrieval information, regular expressions or NLP (Natural Language Processing) tools can be used to automatically extract the content of the document retrieval information and convert it into a structured data format, such as JOSN (JavaScript Object Notation) format or XML (eXtensible Markup Language) format.
[0085] Generate a keyword index of the function keyword information of each function block. Specifically, an inverted index can be used to build a keyword index based on the function name and function parameters to achieve fast query capabilities. Specifically, the keyword index construction method can be as follows:
[0086] "calculate_sum"→[Block_ID_1]
[0087] "num1" → [Block_ID_1, Block_ID_2]
[0088] Among them, the above "calculate_sum" is the function name, "num1" is the function parameter, Block_ID_1 represents the function block with ID (Identity) 1, and Block_ID_2 represents the function block with ID 2.
[0089] Generate a vector index of the function description information of each function block. Specifically, you can use an open source natural language processing model, such as BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language representation model based on Transformer) to encode the function requirement description and function usage example description in the function description information into fixed-length vectors, and establish a vector index based on the vectors corresponding to the function requirement description and the function usage example description, so as to achieve fast query capabilities. Specifically, the vector index can be constructed as follows:
[0090] "block_id" → [Block_ID_1]
[0091] "vector_1" → [0.12, 0.45, …, 0.78]
[0092] "vector_2" → [0.41, 0.57, …, 0.69]
[0093] Among them, the above-mentioned "block_id" is the function block ID, Block_ID_1 represents the function block with ID 1, vector_1 represents the fixed-length vector corresponding to the function requirement description, and vector_2 represents the fixed-length vector corresponding to the function usage example description.
[0094] It can be seen that in this embodiment, by respectively generating a keyword index of the function keyword information of each function block and a vector index of the function description information, and storing the function keyword information and function description information of the function block in the document retrieval database in an index-based manner, it is possible to quickly locate the corresponding retrieval information in the subsequent retrieval process, thereby improving the retrieval efficiency.
[0095] It is understandable that, since the index storage methods of the function keyword information and the function description information of the function block in the document retrieval database are different, information matching needs to be performed separately in the target retrieval information matching process based on different index storage methods.
[0096] In some embodiments, the target retrieval information includes sentence keyword information and sentence description information; the document retrieval database stores at least one function block and function keyword information and / or function description information corresponding to each function block;
[0097] Accordingly, based on the pre-built document retrieval database, content retrieval is performed on the target retrieval information to obtain content retrieval results, including:
[0098] Using the sentence keyword information and the function keyword information corresponding to each function block in the document retrieval database to perform content matching, to obtain a first function block that meets a preset first matching condition and a first matching similarity corresponding to the first function block;
[0099] Using the sentence description information and the function description information corresponding to each function block in the document retrieval database to perform content matching, to obtain a second function block that meets a preset second matching condition and a second matching similarity corresponding to the second function block;
[0100] A content retrieval result is determined according to the first function block and its corresponding first matching similarity, the second function block and its corresponding second matching similarity.
[0101] For the sentence keyword information in the target retrieval information, the sentence keyword information also includes function name and function parameters; for the function keyword information corresponding to the function block in the document retrieval database, the function keyword information also includes function name and function parameters.
[0102] Specifically, the function name in the sentence keyword information is matched with the function name of each function block in the document retrieval database to obtain a function name matching result. The function name matching result includes a first sub-function block that meets the first matching condition, and a first sub-matching similarity corresponding to the first sub-function block. The first matching condition may be that the similarity matching result reaches a preset first similarity threshold, and the first similarity threshold may be pre-set by relevant technical personnel according to actual needs. For example, the first similarity threshold may be set to 90%.
[0103] The function parameters in the sentence keyword information are matched with the function parameters of each function block in the document retrieval database to obtain a function parameter matching result, wherein the function parameter matching result includes a second sub-function block that meets the first matching condition and a second sub-matching similarity corresponding to the second sub-function block.
[0104] The first sub-function block and the second sub-function block are combined to obtain a first function block. The first sub-matching similarity of the same first function sub-block and the second sub-matching similarity of the second sub-function block are averaged to obtain the first matching similarity of the corresponding first function block.
[0105] Exemplarily, if the first sub-function block includes sub-block A, sub-block B and sub-block C; wherein the first sub-matching similarity corresponding to sub-block A is 90%, the first sub-matching similarity corresponding to sub-block B is 92%, and the first sub-matching similarity corresponding to sub-block C is 96%. The second sub-function block includes sub-block A, sub-block B and sub-block D; wherein the second sub-matching similarity corresponding to sub-block A is 94%, the second sub-matching similarity corresponding to sub-block B is 90%, and the second sub-matching similarity corresponding to sub-block D is 98%. Then the first function block includes the first function block A (sub-block A), the first function block B (sub-block B), the first function block C (sub-block C) and the first function block D (sub-block D). The first matching similarity corresponding to the first function block A is the average of the first sub-matching similarity of block A (90%) and the second sub-matching similarity of block A (94%), that is, the first matching similarity corresponding to the first function block A is 92%. Similarly, the first matching similarity corresponding to the first function block B is 91%. The first matching similarity corresponding to the first function block C is 96% (i.e., the first sub-matching similarity corresponding to block C of the first sub-function block). The first matching similarity corresponding to the first function block D is 98% (i.e., the second sub-matching similarity corresponding to block D of the second sub-function block).
[0106] For the statement description information in the target retrieval information, the statement description information also includes a function requirement description and a function usage example description; for the function description information of the function block in the document retrieval database, the function description information also includes a function requirement description and a function usage example description.
[0107] Specifically, the function requirement description in the statement description information is matched with the function requirement description of each function block in the document retrieval database for similarity, and a function requirement description matching result is obtained. The function requirement description matching result includes a third sub-function block that satisfies the second matching condition, and a third sub-matching similarity corresponding to the third sub-function block. The second matching condition may be the same as or different from the first matching condition. Specifically, the second matching condition may be that the similarity matching result reaches a preset second similarity threshold. The second similarity threshold may be pre-set by relevant technical personnel according to actual needs. For example, the second similarity threshold may be set to 90%.
[0108] The function usage example description in the statement description information is matched with the function usage example description of each function block in the document retrieval database to obtain a usage example matching result, wherein the usage example matching result includes a fourth sub-function block that meets the second matching condition and a fourth sub-matching similarity corresponding to the fourth sub-function block.
[0109] The third sub-function block and the fourth sub-function block are combined to obtain a second function block. The third sub-matching similarity of the same third function sub-block and the fourth sub-matching similarity of the fourth sub-function block are averaged to obtain the second matching similarity of the corresponding second function block. The specific implementation method is the same as the implementation method of the first matching similarity of the first function block mentioned above, and this embodiment will not be repeated here.
[0110] Specifically, the first function block and the second function block are block-merged to obtain a merged function block. The first matching similarity and the second matching similarity between the same function blocks are averaged to obtain a merged matching similarity of the merged function block. The merged function block with the highest merged matching similarity is determined as the target function block. The function keyword information and function description information corresponding to the target function block are determined as the content retrieval result.
[0111] It can be seen that in this embodiment, by performing similarity matching on the sentence keyword information and the sentence description information in the sentence keyword information respectively, the accuracy of determining the target block function that meets the similarity conditions is improved, thereby improving the accuracy of determining the content retrieval results, and further improving the accuracy of determining the subsequent sentence query results.
[0112] This embodiment further provides a preferred embodiment based on the above embodiment. The method of this embodiment includes two stages, namely, the stage of constructing a document retrieval database and the stage of document content retrieval.
[0113] The construction phase of the document retrieval database includes the following steps:
[0114] Step a1: input the acquired document to be processed into the trained named entity recognition model for processing, and output document retrieval information.
[0115] The document retrieval information includes function keyword information and function description information corresponding to at least one function block.
[0116] Specifically, the trained named entity recognition model is used to parse the content of the document to be processed to obtain at least one structured function block, as well as function keyword information and function description information corresponding to each function block. The function keyword information may include function name and function parameters; the function description information may include function requirement description and function usage example description. Exemplarily, the specific content of the document retrieval information may be as follows:
[0117] Function name: used for quick positioning function;
[0118] Function parameters: important content as search keywords;
[0119] Function requirement description: describes the function's function and usage scenarios, suitable for semantic retrieval;
[0120] Function usage example description: Code examples to help users understand the specific usage of the function.
[0121] Specifically, in order to facilitate the subsequent information processing of the document retrieval information, regular expressions or NLP (Natural Language Processing) tools can be used to automatically extract the content of the document retrieval information and convert it into a structured data format, such as JOSN (JavaScript Object Notation) format or XML (eXtensible Markup Language) format.
[0122] For example, the above document retrieval information is automatically extracted using regular expressions or NLP tools and stored in JSON format. Each function block can be expressed as follows:
[0123]
[0124]
[0125] Step a2: Generate a keyword index of the function keyword information of each function block, and generate a vector index of the function description information of each function block.
[0126] Generate a keyword index of the function keyword information of each function block. Specifically, an inverted index can be used to establish a keyword index based on the function name and function parameters to achieve fast query capabilities.
[0127] Specifically, the keyword index construction method can be as follows:
[0128] "calculate_sum"→[Block_ID_1]
[0129] "num1" → [Block_ID_1, Block_ID_2]
[0130] Among them, the above "calculate_sum" is the function name, "num1" is the function parameter, Block_ID_1 represents the function block with ID (Identity) 1, and Block_ID_2 represents the function block with ID 2.
[0131] Generate a vector index of the function description information of each function block. Specifically, you can use an open source natural language processing model, such as BERT, to encode the function requirement description and the function usage example description in the function description information into fixed-length vectors, and build a vector index based on the vectors corresponding to the function requirement description and the function usage example description, to achieve fast query capabilities. Specifically, the vector index can be constructed as follows:
[0132] "block_id" → [Block_ID_1]
[0133] "vector_1" → [0.12, 0.45, …, 0.78]
[0134] "vector_2" → [0.41, 0.57, …, 0.69]
[0135] Among them, the above-mentioned "block_id" is the function block ID, Block_ID_1 represents the function block with ID 1, vector_1 represents the fixed-length vector corresponding to the function requirement description, and vector_2 represents the fixed-length vector corresponding to the function usage example description.
[0136] Step a3: storing the function keyword information of each function block in a pre-built document retrieval database based on the keyword index, and storing the function description information of each function block in a pre-built document retrieval database based on the vector index.
[0137] For the document content retrieval stage, such as Figure 3 The schematic diagram of the processing flow structure of a document content retrieval method shown includes the following steps:
[0138] Step b1: The user inputs a query statement based on the front-end interface.
[0139] A selection input box is provided in the front-end interface, and the user enters a query statement in the selection input box of the front-end interface. For example, the query statement is "How to use the calculate_sum function to calculate the sum of the two numbers 234 and 345?".
[0140] Step b2: Use the trained sentence retrieval judgment model to predict whether the query sentence performs content retrieval; if so, execute steps b3 to b6; if not, execute step b7.
[0141] Step b3: Use the trained named entity recognition model to parse the query statement to obtain target retrieval information.
[0142] For the query statement input by the user above, the target retrieval information obtained by parsing with the named entity recognition model is as follows:
[0143] Function name: "calculate_sum";
[0144] Function parameters: "num1", "num2";
[0145] Function requirement description: "Use the calculate_sum function to calculate the sum of the two numbers 234 and 345."
[0146] Step b4: Based on the document retrieval database, sentence keyword information retrieval and sentence description information retrieval are performed on the target retrieval information to obtain content retrieval results.
[0147] Use the inverted index to quickly match the relevant blocks of function names and function parameters, return a first set of function blocks with a higher matching degree, and vectorize the function requirement description, calculate the similarity (such as cosine similarity) with the vector corresponding to the function requirement description in the vector index in the document retrieval database, and return a second set of function blocks with a higher similarity. Remove and merge the function blocks between the first set of function blocks and the second set of function blocks, and sort them from high to low by similarity to obtain the target function block; use the function keyword information and function description information of the target function block as the content retrieval result.
[0148] Step b5: Generate a target search statement based on the content search results and the query statement.
[0149] Exemplarily, a template for generating a target search statement is pre-constructed, and the template content may be as follows:
[0150]
[0151]
[0152] Specifically, the generated template may further include Context: {function category}. The function category may be used to store the function category of the function block in the document retrieval database in advance during the process of parsing the document to be processed.
[0153] Step b6: Input the target search statement into LLM for question-answering processing, and obtain the statement query result corresponding to the query statement output by LLM.
[0154] Step b7: Input the query statement into LLM for question and answer processing, and obtain the statement query result corresponding to the query statement output by LLM.
[0155] It should be understood that, although the steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indications of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0156] Based on the same inventive concept, the embodiment of the present application also provides a document content retrieval device. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more document content retrieval device embodiments provided below can refer to the limitations on the document content retrieval method above, and will not be repeated here.
[0157] like Figure 4 As shown, the embodiment of the present application provides a document content retrieval device 400, including:
[0158] The acquisition module 402 is used to acquire the query statement input by the user based on the front-end interface;
[0159] The parsing module 404 is used to parse the query sentence to obtain target retrieval information when the trained sentence retrieval judgment model is used to predict that the query sentence needs to be searched for content;
[0160] The retrieval module 406 is used to perform content retrieval on the target retrieval information based on the pre-built document retrieval database to obtain content retrieval results;
[0161] The determination module 408 is used to determine the statement query result corresponding to the query statement according to the query statement and the content retrieval result.
[0162] In some embodiments, in parsing the query statement to obtain the target search information, the parsing module 404 is specifically used to:
[0163] The query statement is input into the trained named entity recognition model for processing and the target retrieval information is output.
[0164] In some embodiments, the target retrieval information includes sentence keyword information and sentence description information; the document retrieval database stores at least one function block and function keyword information and / or function description information corresponding to each function block; accordingly, in performing content retrieval on the target retrieval information based on the pre-built document retrieval database to obtain content retrieval results, the retrieval module 406 is specifically used to:
[0165] Using the sentence keyword information and the function keyword information corresponding to each function block in the document retrieval database to perform content matching, to obtain a first function block that meets a preset first matching condition and a first matching similarity corresponding to the first function block;
[0166] Using the sentence description information and the function description information corresponding to each function block in the document retrieval database to perform content matching, to obtain a second function block that meets a preset second matching condition and a second matching similarity corresponding to the second function block;
[0167] A content retrieval result is determined according to the first function block and its corresponding first matching similarity, the second function block and its corresponding second matching similarity.
[0168] In some embodiments, in determining the statement query result corresponding to the query statement according to the query statement and the content search result, the determination module 408 is specifically used to:
[0169] Determine the target search statement according to the query statement and content search results;
[0170] The target search sentence is input into the trained search question answering model for processing, and the sentence query result corresponding to the query sentence is output.
[0171] In some embodiments, the apparatus 400 further includes a model training module, which is used to:
[0172] Acquire a simulated query statement input by a user based on a front-end interface and a judgment result on whether to search for the simulated query statement; generate a sentence label of the simulated query statement based on the judgment result; input the simulated query statement and its corresponding sentence label into a preset sentence retrieval judgment model for processing, and output a retrieval prediction result corresponding to the simulated query statement; train the preset sentence retrieval judgment model based on the sentence label and retrieval prediction result of the simulated query statement until the first model training end condition is met, thereby obtaining a trained sentence retrieval judgment model.
[0173] In some embodiments, the device 400 also includes a data storage module, which is used to: input the acquired document to be processed into a trained named entity recognition model for processing, and output document retrieval information; the document retrieval information includes function keyword information and function description information corresponding to at least one function block; generate a keyword index of the function keyword information of each function block, and generate a vector index of the function description information of each function block; store the function keyword information of each function block in a pre-built document retrieval database based on the keyword index, and store the function description information of each function block in a pre-built document retrieval database based on the vector index.
[0174] In some embodiments, the model training module is further used to:
[0175] The acquired training text sample data are labeled with named entity categories to obtain standard named entity categories of the training text sample data; the training text sample data and its corresponding standard named entity categories are input into a preset named entity recognition model for processing, and the predicted named entity categories are output; the preset named entity recognition model is trained according to the standard named entity categories and the predicted named entity categories of the training text sample data until the second model training end condition is met, thereby obtaining a trained named entity recognition model.
[0176] Each module in the document content retrieval device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0177] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to document content retrieval. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in the above-mentioned document content retrieval method are implemented.
[0178] In some embodiments, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, the steps in the document content retrieval method described above are implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen; the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0179] Those skilled in the art will understand that Figure 5 or Figure 6The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0180] In some embodiments, a computer device is provided. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0181] In some embodiments, Figure 7 The figure shows an internal structure diagram of a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0182] In some embodiments, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0183] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0184] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0185] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0186] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A document content retrieval method, characterized in that: include: Obtain the query statement entered by the user based on the front-end interface; When the trained sentence retrieval judgment model is used to predict that the query sentence needs to be retrieved, the query sentence is parsed to obtain target retrieval information; Based on a pre-built document retrieval database, content retrieval is performed on the target retrieval information to obtain content retrieval results; According to the query statement and the content retrieval result, a statement query result corresponding to the query statement is determined.
2. The method according to claim 1, characterized in that The query statement is parsed to obtain target search information, including: The query statement is input into a trained named entity recognition model for processing, and target retrieval information is output.
3. The method according to claim 1, characterized in that The target retrieval information includes sentence keyword information and sentence description information; the document retrieval database stores at least one function block and function keyword information and / or function description information corresponding to each of the function blocks; Accordingly, the content retrieval of the target retrieval information is performed based on the pre-built document retrieval database to obtain the content retrieval results, including: Using the sentence keyword information and the function keyword information corresponding to each of the function blocks in the document retrieval database to perform content matching, to obtain a first function block that meets a preset first matching condition and a first matching similarity corresponding to the first function block; Using the statement description information and the function description information corresponding to each of the function blocks in the document retrieval database to perform content matching, to obtain a second function block that meets a preset second matching condition and a second matching similarity corresponding to the second function block; A content retrieval result is determined according to the first function block and its corresponding first matching similarity, the second function block and its corresponding second matching similarity.
4. The method according to claim 1, characterized in that: The step of determining, according to the query statement and the content search result, a statement query result corresponding to the query statement includes: Determining a target search statement according to the query statement and the content search result; The target search sentence is input into the trained search question answering model for processing, and the sentence query result corresponding to the query sentence is output.
5. The method according to any one of claims 1 to 4, characterized in that: After obtaining the query statement input by the user based on the front-end interface, the method further includes: Acquire a simulated query statement input by a user based on the front-end interface and a judgment result of whether to perform a search on the simulated query statement; According to the judgment result, generating a statement label of the simulated query statement; Inputting the simulated query statement and the corresponding statement label into a preset statement retrieval judgment model for processing, and outputting the retrieval prediction result corresponding to the simulated query statement; The preset sentence retrieval judgment model is trained according to the sentence label and the retrieval prediction result of the simulated query sentence until the first model training end condition is met to obtain the trained sentence retrieval judgment model.
6. The method according to claim 1, characterized in that Before performing content retrieval on the target retrieval information based on the pre-built document retrieval database to obtain content retrieval results, the method further includes: Input the acquired document to be processed into the trained named entity recognition model for processing, and output document retrieval information; the document retrieval information includes function keyword information and function description information corresponding to at least one function block; Generate a keyword index of the function keyword information of each of the function blocks, and generate a vector index of the function description information of each of the function blocks; The function keyword information of each of the function blocks is stored in a pre-built document retrieval database based on the keyword index, and the function description information of each of the function blocks is stored in the pre-built document retrieval database based on the vector index.
7. The method according to claim 2 or 6, characterized in that: After obtaining the query statement input by the user based on the front-end interface, the method further includes: Performing named entity category annotation on the acquired training text sample data to obtain standard named entity categories of the training text sample data; Inputting the training text sample data and its corresponding standard named entity category into a preset named entity recognition model for processing, and outputting a predicted named entity category; The preset named entity recognition model is trained according to the standard named entity categories and predicted named entity categories of the training text sample data until a second model training end condition is met to obtain the trained named entity recognition model.
8. A document content retrieval device, characterized in that: include: The acquisition module is used to obtain the query statement input by the user based on the front-end interface; A parsing module, for parsing the query statement to obtain target retrieval information when the trained sentence retrieval judgment model is used to predict that the query statement needs to be retrieved; A retrieval module, used to perform content retrieval on the target retrieval information based on a pre-built document retrieval database to obtain content retrieval results; The determination module is used to determine the statement query result corresponding to the query statement according to the query statement and the content retrieval result.
9. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.