Document search method and device, electronic device and storage medium
By performing performance evaluation of the document search model and classification processing of knowledge fields, filtering and optimizing model parameters, the problem of insufficient search performance in specific knowledge fields of traditional document search models is solved, and the search accuracy is improved.
Patent Information
- Application Number
- CN202410975185.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-07-19
AI Technical Summary
The search performance of the traditional document search model in specific knowledge fields has not been optimized, resulting in low search accuracy.
By obtaining the running log data of the preset document search model, performing performance evaluation, and classifying the performance evaluation data based on multiple preset knowledge fields, filtering out the model training data, optimizing the model parameters, and obtaining the target document search model.
It improves the search accuracy of the document search model in specific knowledge fields, breaks through performance bottlenecks, and achieves more accurate document acquisition.
Smart Images

Figure CN118964593B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and in particular to a document search method and device, an electronic device, and a storage medium. Background Art
[0002] The document search model is a natural language processing model based on deep learning, which can understand human language and find target documents. However, in the traditional document search model optimization, the search performance of the document search model in a specific knowledge field is not considered. Instead, the content of all knowledge fields is directly input into the document search model for updating. The model is not optimized for the performance bottleneck of the document search model, which leads to low search accuracy of the document search model. Therefore, how to improve the search accuracy of the document search model has become an urgent problem to be solved. Summary of the invention
[0003] The main purpose of the embodiments of the present application is to propose a document search method and device, an electronic device and a storage medium, aiming to improve the search accuracy of the document search model.
[0004] To achieve the above object, a first aspect of an embodiment of the present application proposes a document search method, the method comprising:
[0005] Obtain the log data of the preset document search model during the running process to obtain the running log data;
[0006] Performing a performance evaluation on the preset document search model according to the operation log data to obtain performance evaluation data;
[0007] Classify the performance evaluation data according to multiple preset knowledge fields to obtain target evaluation data for each preset knowledge field; wherein the target evaluation data is used to indicate the search performance of the preset document search model in the preset knowledge field;
[0008] Filtering model training data from a preset database according to the target evaluation data;
[0009] Optimizing the parameters of the preset document search model according to the model training data to obtain a target document search model;
[0010] In response to the document search request, a document search is performed using the target document search model and the document search request to obtain a target document.
[0011] In some embodiments, the operation log data includes: a model error log, a model response log, and a user feedback log; and the performance evaluation of the preset document search model according to the operation log data to obtain the performance evaluation data includes:
[0012] Perform response efficiency evaluation on the preset document search model according to the model response log to obtain efficiency evaluation data;
[0013] Performing a model output satisfaction evaluation on the preset document search model according to the user feedback log to obtain satisfaction evaluation data;
[0014] Performing an operation status evaluation on the preset document search model according to the model error log to obtain operation status evaluation data;
[0015] The preset document search model is subjected to a performance evaluation according to the efficiency evaluation data, the satisfaction evaluation data and the operation status evaluation data to obtain the performance evaluation data.
[0016] In some embodiments, the model response log includes: model response time and model response resource occupancy rate; the response efficiency evaluation of the preset document search model according to the model response log to obtain efficiency evaluation data includes:
[0017] Comparing a preset response time threshold with the model response time to obtain response time comparison data;
[0018] Compare a preset resource occupancy rate threshold with the model response resource occupancy rate to obtain resource occupancy rate comparison data;
[0019] The efficiency of the preset document search model is evaluated according to the response time comparison data and the resource occupancy rate comparison data to obtain the efficiency evaluation data.
[0020] In some embodiments, the user feedback log includes direct feedback information, the number of repetitions of the same question, and the number of modifications to the same question; and performing a model output satisfaction evaluation on the preset document search model according to the user feedback log to obtain satisfaction evaluation data includes:
[0021] Performing feedback evaluation on the preset document search model according to the direct feedback information to obtain feedback evaluation data;
[0022] Compare the preset repetition threshold with the number of repetitions of the same question to obtain repetition number analysis data;
[0023] Compare the preset modification threshold with the modification times of the same problem to obtain modification times analysis data;
[0024] The output satisfaction evaluation of the preset document search model is performed according to the feedback evaluation data, the repetition number analysis data and the modification number analysis data to obtain the satisfaction evaluation data.
[0025] In some embodiments, the step of filtering out model training data from a preset database according to the target evaluation data includes:
[0026] Taking the preset knowledge domain corresponding to the smallest target evaluation data as the target knowledge domain;
[0027] The model training data is screened out from the preset database according to the target knowledge domain.
[0028] In some embodiments, the document search request includes model text input information, text publication time, and text category; performing document search through the target document search model and the document search request to obtain the target document includes:
[0029] Performing document search on the model text input information through the target document search model to obtain candidate documents;
[0030] The candidate documents are screened according to the text publication time and the text category to obtain the target document.
[0031] In some embodiments, the model training data includes: document training data and selected documents; the step of optimizing parameters of the preset document search model according to the model training data to obtain a target document search model includes:
[0032] Inputting the document training data into the preset document search model to perform document search to obtain a predicted document;
[0033] Calculate the loss value according to the selected document and the predicted document to obtain the model loss value;
[0034] The parameters of the preset document search model are optimized according to the model loss value to obtain the target document search model.
[0035] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a document search device, the device comprising:
[0036] An acquisition module is used to acquire log data of a preset document search model during operation to obtain operation log data;
[0037] An evaluation module, used to perform a performance evaluation on the preset document search model according to the operation log data to obtain performance evaluation data;
[0038] A classification module, used for classifying the performance evaluation data according to a plurality of preset knowledge fields to obtain target evaluation data for each of the preset knowledge fields; wherein the target evaluation data is used to indicate the search performance of the preset document search model in the preset knowledge field;
[0039] A screening module, used to screen out model training data from a preset database according to the target evaluation data;
[0040] An optimization module, used to optimize the parameters of the preset document search model according to the model training data to obtain a target document search model;
[0041] The search module is used to respond to a document search request, perform a document search through the target document search model and the document search request, and obtain the target document.
[0042] To achieve the above objectives, a third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0043] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0044] The present application proposes a document search method and device, an electronic device and a storage medium, which obtains operation log data of a preset document search model during operation, and then performs a performance evaluation on the preset document search model according to the operation log data to obtain performance evaluation data, and then classifies the performance evaluation data according to multiple preset knowledge fields to obtain the performance search performance of the preset document search model in each knowledge field to obtain target evaluation data, and then screens out model training data from a preset database according to the target evaluation data, and then optimizes the parameters of the preset document search model according to the model training data to obtain a target search model, so as to achieve targeted updating of the preset document model according to the search performance of the preset document search model in each knowledge field, thereby breaking through the performance bottleneck of the preset document search model and obtaining a target document search model; further, in response to a document search request, a document search is performed through the target document search model and the document search request to achieve accurate acquisition of the target document, thereby improving the accuracy of the search. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of a document search method provided by an embodiment of the present application;
[0046] Figure 2 yes Figure 1 Flow chart of step S102 in FIG.
[0047] Figure 3 yes Figure 2 Flow chart of step S201 in FIG.
[0048] Figure 4 yes Figure 2 Flow chart of step S202 in FIG.
[0049] Figure 5 yes Figure 1 Flow chart of step S104 in FIG.
[0050] Figure 6 yes Figure 1 Flow chart of step S105 in FIG.
[0051] Figure 7 yes Figure 1 Flow chart of step S106 in FIG.
[0052] Figure 8 is a schematic diagram of the structure of a document search device provided in an embodiment of the present application;
[0053] Fig. 9 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0055] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0057] First, some nouns involved in this application are analyzed:
[0058] Document search model
[0059] Domain of Knowledge: refers to specific knowledge and skills concentrated in a specific subject or professional scope. In a domain of knowledge, it usually contains a set of related concepts, theories, methods, technologies and their application systems. These knowledge and skills are used to solve specific types of problems or complete specific tasks. Common domains of knowledge include but are not limited to medicine, engineering, law, education, economics, computer science, etc.
[0060] Document Retrieval Model: Document retrieval model is a technology used to retrieve the most relevant documents from a large collection of documents. The main task of the document retrieval model is to match the documents in the document library with the query terms provided by the user, and return the results in order of relevance. Common document retrieval models include Boolean model, vector space model, probability model and language model.
[0061] The document search model is a natural language processing model based on deep learning, which can understand human language and find target documents. However, in the traditional document search model optimization, the search performance of the document search model in a specific knowledge field is not considered. Instead, the content of all knowledge fields is directly input into the document search model for updating. The model is not optimized for the performance bottleneck of the document search model, which leads to low search accuracy of the document search model. Therefore, how to improve the search accuracy of the document search model has become an urgent problem to be solved.
[0062] Based on this, the embodiments of the present application provide a document search method and device, an electronic device and a storage medium, aiming to improve the search accuracy of the document search model.
[0063] A document search method and device, electronic device and storage medium provided in the embodiments of the present application are specifically described through the following embodiments. First, the document search method in the embodiments of the present application is described.
[0064] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0065] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0066] The document search method provided in the embodiment of the present application relates to the field of information technology. The document search method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the document search method, etc., but is not limited to the above forms.
[0067] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0068] Figure 1 is an optional flowchart of the document search method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S106.
[0069] Step S101, obtaining log data of a preset document search model during operation to obtain operation log data;
[0070] Step S102, performing a performance evaluation on a preset document search model according to the operation log data to obtain performance evaluation data;
[0071] Step S103, classifying and processing the performance evaluation data according to a plurality of preset knowledge fields to obtain target evaluation data for each preset knowledge field; wherein the target evaluation data is used to indicate the search performance of the preset document search model in the preset knowledge field;
[0072] Step S104, filtering out model training data from a preset database according to the target evaluation data;
[0073] Step S105, optimizing parameters of a preset document search model according to the model training data to obtain a target document search model;
[0074] Step S106, in response to the document search request, a document search is performed through the target document search model and the document search request to obtain the target document.
[0075] Steps S101 to S106 shown in the embodiment of the present application are obtained by obtaining the operation log data of the preset document search model during operation, and then evaluating the performance of the preset document search model according to the operation log data to obtain performance evaluation data, and then classifying the performance evaluation data according to multiple preset knowledge fields to obtain the performance search performance of the preset document search model in each knowledge field to obtain target evaluation data, and then filtering out model training data from the preset database according to the target evaluation data, and then optimizing the parameters of the preset document search model according to the model training data to obtain the target search model, so as to achieve targeted updating of the preset document model according to the search performance of the preset document search model in each knowledge field, thereby breaking through the performance bottleneck of the preset document search model and obtaining the target document search model; further, in response to the document search request, document search is performed through the target document search model and the document search request to achieve accurate acquisition of the target document, thereby improving the accuracy of the search.
[0076] In step S101 of some embodiments, the preset document search model is a model based on natural language processing technology. After obtaining the user's request, the user's request is parsed to obtain the user's input text, and then the user input text is subjected to stop word removal, word segmentation, vectorization and other operations to obtain the target input tensor, and then a document search is performed based on the target input tensor to obtain the document the user is looking for. Operation log data refers to the log generated by the preset document search model during operation, and the logic of log generation is specified by the developer. For example: The user enters the text at 18:00: Please help me find books related to Lagrange interpolation. The preset document search model outputs advanced mathematics. The log can be: (user name, 18:00, please help me find books related to Lagrange interpolation, mathematics, advanced mathematics). It should be noted that this application does not impose specific restrictions on the specific content and specific format of the log.
[0077] See also Figure 2 In some embodiments, the operation log data includes: model error log, model response log and user feedback log; step S102 may include but is not limited to steps S201 to S204:
[0078] Step S201, evaluating the response efficiency of a preset document search model according to the model response log to obtain efficiency evaluation data;
[0079] Step S202, performing a model output satisfaction evaluation on a preset document search model according to the user feedback log to obtain satisfaction evaluation data;
[0080] Step S203, evaluating the running status of the preset document search model according to the model error log to obtain running status evaluation data;
[0081] Step S204, performing performance evaluation on the preset document search model according to the efficiency evaluation data, the satisfaction evaluation data and the operation status evaluation data to obtain performance evaluation data.
[0082] In steps S201 to S204 shown in the embodiment of the present application, the response efficiency of the preset document search model is evaluated according to the model response log to obtain efficiency evaluation data, and then the model output satisfaction of the preset document search model is evaluated according to the user feedback log to obtain satisfaction evaluation data, and the operation status of the preset document search model is evaluated according to the model error log to obtain operation status evaluation data. Finally, the performance of the preset document search model is evaluated according to the efficiency evaluation data, the satisfaction evaluation data and the operation status evaluation data to obtain performance evaluation data, so as to obtain the performance evaluation data of the preset document search model, thereby providing a data basis for subsequent performance evaluation data to judge the bottleneck of the model, thereby improving the accuracy of subsequent model optimization.
[0083] See also Figure 3 In some embodiments, the model response log includes: model response time and model response resource occupancy rate; step S201 may include but is not limited to steps S301 to S303:
[0084] Step S301, comparing the preset response time threshold with the model response time to obtain response time comparison data;
[0085] Step S302, comparing a preset resource occupancy rate threshold with a model response resource occupancy rate to obtain resource occupancy rate comparison data;
[0086] Step S303: Efficiency evaluation is performed on the preset document search model according to the response time comparison data and the resource occupancy rate comparison data to obtain efficiency evaluation data.
[0087] In steps S301 to S303 shown in the embodiment of the present application, a preset response time threshold is compared with the model response time to obtain response time comparison data, and then a preset resource occupancy rate threshold is compared with the model response resource occupancy rate to obtain resource occupancy rate comparison data. Finally, the efficiency of the preset document search model is evaluated based on the response time comparison data and the resource occupancy rate comparison data to obtain efficiency evaluation data, thereby achieving the model efficiency of the preset document search model in each knowledge field, thereby providing a data basis for subsequent judgment of the bottleneck of the model based on the model efficiency, thereby improving the accuracy of subsequent model optimization.
[0088] In step S301 of some embodiments, the model response time is the time taken for the preset document search model to accept the input text and search after the user inputs the text, and finally output the document. When the preset response time threshold is greater than the model response time, the response time comparison data is determined to be the difference between the preset response time threshold and the model response time.
[0089] In step S302 of some embodiments, the model response resource occupancy includes CPU occupancy, memory occupancy and video memory occupancy, the preset resource occupancy threshold includes CUP threshold, memory threshold and video memory threshold, which correspond one-to-one to the model response resource occupancy, and the difference between the preset resource occupancy threshold and the model response resource occupancy is obtained to obtain resource occupancy comparison data.
[0090] In step S303 of some embodiments, for each knowledge field, the sum of the corresponding response time comparison data and the resource occupancy rate comparison data is obtained to obtain efficiency evaluation data, that is, the smaller the efficiency evaluation data, the higher the efficiency of the preset document search model, and the larger the efficiency evaluation data, the lower the efficiency of the preset document search model.
[0091] It should be noted that the present application does not impose any specific restrictions on the specific forms of the response comparison data, resource occupancy comparison data, and efficiency evaluation data, and they may be other values in other scenarios.
[0092] See also Figure 4 In some embodiments, the user feedback log includes direct feedback information, the number of times the same question is repeated, and the number of times the same question is modified. Step S202 may include, but is not limited to, steps S401 to S404:
[0093] Step S401, performing feedback evaluation on a preset document search model according to direct feedback information to obtain feedback evaluation data;
[0094] Step S402, comparing the preset repetition threshold with the number of repetitions of the same question to obtain repetition number analysis data;
[0095] Step S403, comparing the preset modification threshold with the number of times the same question has been modified to obtain modification number analysis data;
[0096] Step S404, performing output satisfaction evaluation on the preset document search model according to the feedback evaluation data, the repetition number analysis data and the modification number analysis data to obtain satisfaction evaluation data.
[0097] In steps S401 to S404 shown in the embodiment of the present application, feedback evaluation is performed on the preset document search model based on direct feedback information to obtain feedback evaluation data, and then the preset repetition threshold is compared with the number of repetitions of the same question to obtain repetition analysis data, and the preset modification threshold is compared with the number of modifications of the same question to obtain modification analysis data, and finally, the output satisfaction evaluation is performed on the preset document search model based on the feedback evaluation data, the repetition analysis data and the modification analysis data to obtain satisfaction evaluation data, thereby achieving the acquisition of satisfaction evaluation data of the preset document search model in each knowledge field, thereby providing a data basis for subsequent judgment of the bottleneck of the model based on the satisfaction evaluation data, thereby improving the accuracy of subsequent model optimization.
[0098] In step S401 of some embodiments, the direct feedback information is the user's feedback evaluation on the preset document search model. In one embodiment, the user feedback evaluation is judged by a preset sentiment judgment model, and the direct feedback information is input into the sentiment judgment model to obtain feedback judgment data.
[0099] For example, the direct feedback information is "I think this document search is very poor." The direct feedback information is input into the preset sentiment judgment model, and the feedback evaluation data of -0.95 is output.
[0100] For example, if the direct feedback information is “I think this document search is OK”, the direct feedback information is input into the preset sentiment judgment model, and the feedback evaluation data of 0.46 is output.
[0101] In step S402 of some embodiments, the number of times the same question is repeated is the number of times the user repeats a document search question. When the model output of the preset document search model does not meet the user's expectations, the user can choose to repeat the query based on the same prompt word to obtain different search results of the preset document search model. The preset repetition threshold and the number of times the same question is repeated are compared. When the preset repetition threshold is greater than the number of times the same question is repeated, the repetition number analysis data is 0. When the preset repetition threshold is greater than the number of times the same question is repeated, the repetition number analysis data is the difference between the number of times the same question is repeated and the preset repetition threshold, that is, the repetition number analysis data is obtained.
[0102] In step S403 of some embodiments, the number of same question modifications is the number of times the user modifies the prompt word input in a document search request. When the model output of the preset document search model does not meet the user's expectations, the user can repeat the query by modifying the prompt word to obtain different search results of the preset document search model. The preset modification threshold is compared with the number of same question modifications. When the preset modification threshold is greater than the number of same question modifications, the modification number analysis data is 0. When the preset modification threshold is less than the number of same question modifications, the difference between the number of same question modifications and the preset modification threshold is obtained to obtain the modification number analysis data.
[0103] In step S404 of some embodiments, satisfaction evaluation is performed based on the feedback evaluation data, the repetition number analysis data, and the modification number analysis data to obtain output satisfaction evaluation data. The satisfaction evaluation method can be weighted calculation, average value calculation, percentage evaluation, etc., which is not specifically limited in this application.
[0104] In step S203 of some embodiments, the model error log is a log of a program error occurring during the operation of the preset document search model. The model error log is classified according to the preset knowledge field, and the error status of a program error occurring when the preset document search model searches in each knowledge field is obtained. The number of errors is counted to obtain the operation status evaluation data.
[0105] In step S204 of some embodiments, a performance evaluation is performed on the preset document search model according to the efficiency evaluation data, the satisfaction evaluation data and the operation status evaluation data to obtain performance evaluation data. The performance evaluation method can be weighted calculation, average value calculation and percentage evaluation, etc., which is not specifically limited in this application.
[0106] In step S103 of some embodiments, the performance evaluation data is classified according to the preset knowledge domain to obtain target evaluation data for each knowledge domain. The target evaluation data indicates the search performance of the preset document search model in the corresponding preset knowledge domain.
[0107] See also Figure 5 In some embodiments, step S104 includes but is not limited to steps S501 to S502:
[0108] Step S501, taking the preset knowledge domain corresponding to the minimum target evaluation data as the target knowledge domain;
[0109] Step S502: filter out model training data from a preset database according to the target knowledge domain.
[0110] In steps S501 to S502 shown in the embodiment of the present application, the preset knowledge field corresponding to the smallest target evaluation data is used as the target knowledge field, and model training data is filtered out from a preset database according to the target knowledge field, so as to obtain the model training data corresponding to the knowledge field with the worst performance of the preset document search model, thereby training the preset document search model for the knowledge field with the worst search performance, thereby improving the search accuracy of the preset document search model.
[0111] In step S501 of some embodiments, the target evaluation data includes evaluation data corresponding to multiple preset knowledge fields, characterizing the search performance of the preset document search model in the preset knowledge field, and obtaining the minimum target evaluation data, that is, obtaining the preset knowledge field with the worst search performance of the preset document search model, to obtain the target knowledge field.
[0112] In step S502 of some embodiments, the preset database is classified according to the preset knowledge fields, training data of multiple knowledge fields are obtained, and then training data corresponding to the target knowledge field is obtained to obtain model training data.
[0113] See also Figure 6 In some embodiments, the model training data includes: document training data and selected documents, and step S105 includes but is not limited to steps S601 to S603:
[0114] Step S601, inputting document training data into a preset document search model to perform document search and obtain a predicted document;
[0115] Step S602, calculating the loss value according to the selected document and the predicted document to obtain the model loss value;
[0116] Step S603, optimizing the parameters of the preset document search model according to the model loss value to obtain the target document search model.
[0117] In steps S601 to S603 shown in the embodiment of the present application, document training data is input into a preset document search model to perform document search to obtain a predicted document, and loss values are calculated based on the selected document and the predicted document to obtain a model loss value, and then parameters of the preset document search model are optimized based on the model loss value to obtain a target document search model, thereby improving the search accuracy of the preset document search model.
[0118] In step S601 of some embodiments, after obtaining the document training data, the preset document search model performs document search according to the document training data to obtain a predicted document. For example, input "please help me find books related to Lagrange interpolation method" to the preset document search model, and the model outputs: "Advanced mathematics".
[0119] In step S602 of some embodiments, the selected document is the document corresponding to the document training data, and the loss value is calculated based on the selected document and the predicted document. The calculation method can be mean square error, cross entropy error, etc., and this application does not make specific restrictions.
[0120] In step S603 of some embodiments, the parameters of the preset document search model are optimized according to the loss value, so as to improve the search accuracy of the preset document search model and obtain the target document search model.
[0121] See also Figure 7 In some embodiments, the document search request includes model text input information, text publishing time, and text category. Step S106 may include, but is not limited to, steps S701 to S702:
[0122] Step S701, performing document search on model text input information through a target document search model to obtain candidate documents;
[0123] Step S702, screening candidate documents according to text publication time and text category to obtain target documents.
[0124] In steps S701 to S702 shown in the embodiment of the present application, a document search is performed on the model text input information through a target document search model to obtain candidate documents, and the candidate documents are screened according to the text publication time and text category to obtain the target document, thereby improving the search accuracy of the target document search model.
[0125] In step S701 of some embodiments, the model input text information is input into the target document search model, and the target document search model performs a document search to obtain a candidate document.
[0126] In step S702 of some embodiments, the text publication time is the publication time of the document, and the document category is the type of document, including but not limited to journals, books, newspapers, etc. This application does not make specific restrictions. The candidate documents are screened by text publication time and text category to obtain the target document, thereby improving the search accuracy of the target document search model.
[0127] See also Figure 8 The embodiment of the present application also provides a document search device, which can implement the above document search method, and the device includes:
[0128] The acquisition module 801 is used to acquire the log data of the preset document search model during the operation process to obtain the operation log data;
[0129] An evaluation module 802 is used to perform a performance evaluation on a preset document search model according to the operation log data to obtain performance evaluation data;
[0130] The classification module 803 is used to classify the performance evaluation data according to a plurality of preset knowledge fields to obtain target evaluation data for each preset knowledge field; wherein the target evaluation data is used to indicate the search performance of the preset document search model in the preset knowledge field;
[0131] A screening module 804 is used to screen out model training data from a preset database according to target evaluation data;
[0132] The optimization module 805 is used to optimize the parameters of the preset document search model according to the model training data to obtain the target document search model;
[0133] The search module 806 is used to respond to the document search request, perform document search through the target document search model and the document search request, and obtain the target document.
[0134] The specific implementation of the document search device is substantially the same as the specific implementation of the document search method described above, and will not be described in detail herein.
[0135] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above document search method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0136] See also Fig. 9 , Fig. 9 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0137] The processor 901 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0138] The memory 902 may be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 may store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the document search method of the embodiment of this application;
[0139] Input / output interface 903, used to implement information input and output;
[0140] Communication interface 904, used to realize communication interaction between the device and other devices, which can be realized through wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WI FI, Bluetooth, etc.);
[0141] A bus 905 that transmits information between various components of the device (e.g., the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);
[0142] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0143] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned document search method is implemented.
[0144] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0145] The document search method, document search device, electronic device and storage medium provided in the embodiments of the present application obtain operation log data of a preset document search model during operation, and then evaluate the performance of the preset document search model according to the operation log data to obtain performance evaluation data, and then classify and process the performance evaluation data according to multiple preset knowledge fields to obtain the performance search performance of the preset document search model in each knowledge field to obtain target evaluation data, and then filter out model training data from a preset database according to the target evaluation data, and then optimize the parameters of the preset document search model according to the model training data to obtain a target search model, so as to achieve targeted updating of the preset document model according to the search performance of the preset document search model in each knowledge field, thereby breaking through the performance bottleneck of the preset document search model and obtaining a target document search model; further, in response to a document search request, a document search is performed through the target document search model and the document search request to achieve accurate acquisition of the target document, thereby improving the accuracy of the search.
[0146] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0147] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0148] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0149] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0150] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0151] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0152] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0153] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0154] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0155] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0156] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A document search method, characterized in that: The method comprises: Obtaining log data of the preset document search model during operation to obtain operation log data; the operation log data includes: model error log, model response log and user feedback log; Perform response efficiency evaluation on the preset document search model according to the model response log to obtain efficiency evaluation data; Performing a model output satisfaction evaluation on the preset document search model according to the user feedback log to obtain satisfaction evaluation data; Obtaining the error status of the preset document search model when searching in each knowledge field according to the model error log, and counting the number of errors of the preset document search model when searching in each knowledge field according to the model error log, and using the error status and the number of errors as operation status evaluation data; Performing a performance evaluation on the preset document search model according to the efficiency evaluation data, the satisfaction evaluation data and the operation status evaluation data to obtain performance evaluation data; Classify the performance evaluation data according to multiple preset knowledge fields to obtain target evaluation data for each preset knowledge field; wherein the target evaluation data is used to indicate the search performance of the preset document search model in the preset knowledge field; Filtering model training data from a preset database according to the target evaluation data; Optimizing the parameters of the preset document search model according to the model training data to obtain a target document search model; In response to the document search request, a document search is performed using the target document search model and the document search request to obtain a target document.
2. The method according to claim 1, characterized in that: The model response log includes: model response time and model response resource occupancy rate; the response efficiency evaluation of the preset document search model is performed according to the model response log to obtain efficiency evaluation data, including: Comparing a preset response time threshold with the model response time to obtain response time comparison data; Compare a preset resource occupancy rate threshold with the model response resource occupancy rate to obtain resource occupancy rate comparison data; The efficiency of the preset document search model is evaluated according to the response time comparison data and the resource occupancy rate comparison data to obtain the efficiency evaluation data.
3. The method according to claim 1, characterized in that The user feedback log includes direct feedback information, the number of repetitions of the same question, and the number of modifications to the same question; the satisfaction evaluation data obtained by performing a model output satisfaction evaluation on the preset document search model according to the user feedback log includes: Performing feedback evaluation on the preset document search model according to the direct feedback information to obtain feedback evaluation data; Compare the preset repetition threshold with the number of repetitions of the same question to obtain repetition number analysis data; Compare the preset modification threshold with the modification times of the same problem to obtain modification times analysis data; The output satisfaction evaluation of the preset document search model is performed according to the feedback evaluation data, the repetition number analysis data and the modification number analysis data to obtain the satisfaction evaluation data.
4. The method according to claim 1, characterized in that: The step of selecting model training data from a preset database according to the target evaluation data includes: Taking the preset knowledge domain corresponding to the smallest target evaluation data as the target knowledge domain; The model training data is screened out from the preset database according to the target knowledge domain.
5. The method according to claim 1, characterized in that The document search request includes model text input information, text publishing time and text category; the document search is performed through the target document search model and the document search request to obtain the target document, including: Performing document search on the model text input information through the target document search model to obtain candidate documents; The candidate documents are screened according to the text publication time and the text category to obtain the target document.
6. The method according to claim 1, characterized in that The model training data includes: document training data and selected documents; the parameter optimization of the preset document search model according to the model training data to obtain the target document search model includes: Inputting the document training data into the preset document search model to perform document search to obtain a predicted document; Calculate the loss value according to the selected document and the predicted document to obtain the model loss value; The parameters of the preset document search model are optimized according to the model loss value to obtain the target document search model.
7. A document search device, characterized in that: The device comprises: An acquisition module is used to acquire log data of a preset document search model during operation to obtain operation log data; the operation log data includes: model error log, model response log and user feedback log; An evaluation module, used to evaluate the response efficiency of the preset document search model according to the model response log to obtain efficiency evaluation data; Performing a model output satisfaction evaluation on the preset document search model according to the user feedback log to obtain satisfaction evaluation data; Obtaining the error status of the preset document search model when searching in each knowledge field according to the model error log, and counting the number of errors of the preset document search model when searching in each knowledge field according to the model error log, and using the error status and the number of errors as operation status evaluation data; Performing a performance evaluation on the preset document search model according to the efficiency evaluation data, the satisfaction evaluation data and the operation status evaluation data to obtain performance evaluation data; A classification module, used for classifying the performance evaluation data according to a plurality of preset knowledge fields to obtain target evaluation data for each of the preset knowledge fields; wherein the target evaluation data is used to indicate the search performance of the preset document search model in the preset knowledge field; A screening module, used to screen out model training data from a preset database according to the target evaluation data; An optimization module, used to optimize the parameters of the preset document search model according to the model training data to obtain a target document search model; The search module is used to respond to a document search request, perform document search through the target document search model and the document search request, and obtain the target document.
8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the document search method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the document search method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Document searching method, device and equipment and computer readable storage medium
CN114969287A
Classification retrieval method and system based on GPT large model
CN117891898A
Dynamic adaptation question answering system and method based on hierarchical structure and retrieval enhancement
CN118193714A