Document retrieval method, automatic question and answer method and legal provision recommendation method

By using document index filtering and model training during document search, the accuracy problem caused by semantic correlation in document search is solved, and a more efficient and accurate document search effect is achieved.

CN120523924APending Publication Date: 2025-08-22ALIBABA (CHINA) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410186807.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

In the prior art, the correlation between the query conditions entered by users and the document during document searching is mainly semantic, resulting in poor retrieval accuracy. Especially in the field of legal document searching, due to the abstraction and diversity of professional terms, it becomes difficult to accurately match user queries and related legal provisions.

Method used

By obtaining the data to be retrieved for target documents, using the document index for filtering, determining the target document index related to the data to be retrieved, and inputting it into the document search model, searching based on multiple sample search data and the model trained in the results, narrowing the search scope and avoiding misjudgments caused by pure semantic text matching.

Benefits of technology

It improves the efficiency and accuracy of document retrieval, reduces the risk of misjudgment, and ensures the relevance and accuracy of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523924A_ABST
    Figure CN120523924A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a document retrieval method, an automatic question and answer method and a legal provision recommendation method. The document retrieval method comprises the steps of obtaining to-be-retrieved data for a target document; at least one target document index corresponding to the to-be-retrieved data is screened out from the multiple document indexes of the target document, and the target document indexes are used for positioning document content related to the to-be-retrieved data in the target document; the to-be-retrieved data and the at least one target document index are input into a document retrieval model, a document retrieval result of the to-be-retrieved data is obtained, and the document retrieval model is obtained based on training of the multiple pieces of sample retrieval data and sample retrieval results corresponding to the multiple pieces of sample retrieval data. According to the document retrieval method and device, the retrieval range of the document retrieval model is narrowed through the mutual mapping relation between the target document and the document index and the mutual mapping relation between the document index and the document content, the misjudgment risk caused by pure semantic text matching is avoided, and the document retrieval efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a document retrieval method, an automatic question-answering method, and a legal provision recommendation method. Background Art

[0002] With the development of computer technology, automated document retrieval has gradually become a research focus. Document retrieval refers to the process of finding and returning relevant information in a large collection of documents based on the query conditions entered by the user.

[0003] Currently, document retrieval is typically performed through text similarity calculations. However, during the document retrieval process, the correlation between the user's query and the document is mostly semantic, rather than simply textually similar. Calculating text similarity alone can result in poor document retrieval accuracy. Therefore, a highly accurate document retrieval solution is urgently needed. Summary of the Invention

[0004] In light of this, embodiments of this specification provide a document retrieval method. One or more embodiments of this specification also include an automatic question-answering method, a legal text recommendation method, a document retrieval device, an automatic question-answering device, a legal text recommendation device, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a document retrieval method is provided, comprising:

[0006] Obtain the data to be retrieved for the target document;

[0007] Filtering at least one target document index corresponding to the data to be retrieved from multiple document indexes of the target document, wherein the target document index is used to locate document content related to the data to be retrieved in the target document;

[0008] The data to be retrieved and at least one target document index are input into a document retrieval model to obtain a document retrieval result of the data to be retrieved, wherein the document retrieval model is trained based on a plurality of sample retrieval data and the sample retrieval results corresponding to the plurality of sample retrieval data.

[0009] According to a second aspect of the embodiments of this specification, there is provided an automatic question-answering method, comprising:

[0010] Get the questions to be answered for the target document;

[0011] Filtering at least one target document index corresponding to the question to be answered from multiple document indexes of the target document, wherein the target document index is used to locate document content in the target document that is relevant to the question to be answered;

[0012] The question to be answered and at least one target document index are input into a document retrieval model to obtain an answer result for the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and sample retrieval results corresponding to the multiple sample retrieval data.

[0013] According to a third aspect of the embodiments of this specification, a method for recommending legal provisions is provided, including:

[0014] Get answers to questions about the target legal document;

[0015] Filtering at least one target document index corresponding to the question to be answered from multiple document indexes of the target legal document, wherein the target document index is used to locate legal provisions in the target legal document that are relevant to the question to be answered;

[0016] The question to be answered and at least one target document index are input into a document retrieval model to obtain a legal provision recommendation result for the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and sample retrieval results corresponding to the multiple sample retrieval data.

[0017] According to a fourth aspect of the embodiments of this specification, a document retrieval device is provided, comprising:

[0018] A first acquisition module is configured to acquire data to be retrieved for a target document;

[0019] A first screening module is configured to screen out at least one target document index corresponding to the data to be retrieved from a plurality of document indexes of the target document, wherein the target document index is used to locate document content related to the data to be retrieved in the target document;

[0020] The first input module is configured to input the data to be retrieved and at least one target document index into a document retrieval model to obtain a document retrieval result of the data to be retrieved, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

[0021] According to a fifth aspect of the embodiments of this specification, an automatic question-answering device is provided, comprising:

[0022] A second acquisition module is configured to acquire questions to be answered for the target document;

[0023] A second screening module is configured to screen out at least one target document index corresponding to the question to be answered from the multiple document indexes of the target document, wherein the target document index is used to locate document content in the target document that is relevant to the question to be answered;

[0024] The second input module is configured to input the question to be answered and at least one target document index into the document retrieval model to obtain an answer result for the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

[0025] According to a sixth aspect of the embodiments of this specification, a device for recommending legal provisions is provided, including:

[0026] a third acquisition module, configured to acquire questions to be answered regarding the target legal document;

[0027] a third screening module configured to screen out at least one target document index corresponding to the question to be answered from the plurality of document indexes of the target legal document, wherein the target document index is used to locate legal provisions in the target legal document that are relevant to the question to be answered;

[0028] The third input module is configured to input the question to be answered and at least one target document index into the document retrieval model to obtain the legal text recommendation results of the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

[0029] According to a seventh aspect of the embodiments of this specification, a computing device is provided, including:

[0030] memory and processor;

[0031] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.

[0032] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the method provided in the first aspect, the second aspect, or the third aspect.

[0033] According to the ninth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method provided in the first aspect, the second aspect, or the third aspect.

[0034] One embodiment of the present specification provides a document retrieval method, comprising: obtaining data to be retrieved for a target document; screening at least one target document index corresponding to the data to be retrieved from multiple document indexes of the target document, wherein the target document index is used to locate document content in the target document related to the data to be retrieved; and inputting the data to be retrieved and the at least one target document index into a document retrieval model to obtain a document retrieval result for the data to be retrieved, wherein the document retrieval model is trained based on multiple sample retrieval data and sample retrieval results corresponding to the multiple sample retrieval data. By utilizing the mutual mapping relationship between the target document and the document index, and between the document index and the document content, the retrieval scope of the document retrieval model is narrowed. Furthermore, by accurately locating the target document index corresponding to the data to be retrieved and performing document retrieval based on the target document index, the accurate mutual mapping relationship between the document index and the document content is utilized, thereby avoiding the risk of misjudgment caused by pure semantic text matching and improving the efficiency and accuracy of document retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is an architectural diagram of a document retrieval system provided by one embodiment of this specification;

[0036] Figure 2 This is an architectural diagram of another document retrieval system provided by one embodiment of this specification;

[0037] Figure 3 This is a flowchart of a document retrieval method provided by one embodiment of this specification;

[0038] Figure 4 This is a flow chart of an automatic question-answering method provided by one embodiment of this specification;

[0039] Figure 5 This is a flowchart of a method for recommending legal provisions provided in one embodiment of this specification;

[0040] Figure 6 This is a flowchart of a process for recommending legal provisions provided by one embodiment of this specification;

[0041] Figure 7 This is a schematic diagram of a document retrieval interface provided by an embodiment of this specification;

[0042] Figure 8 This is a schematic diagram of the structure of a document retrieval device provided by one embodiment of this specification;

[0043] Figure 9 This is a schematic diagram of the structure of an automatic question-answering device provided by one embodiment of this specification;

[0044] Figure 10This is a schematic diagram of the structure of a legal provision recommendation device provided in one embodiment of this specification;

[0045] Figure 11 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0046] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0047] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0048] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0049] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0050] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. By pre-training a large model with large-scale unlabeled corpus, a pre-trained model with more than 100 million parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large-scale language model (LLM) and a multi-modal pre-training model.

[0051] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image description (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0052] First, the terms involved in one or more embodiments of this specification are explained.

[0053] Text recall: Recall documents based on the similarity between the query content and the text content of the document.

[0054] Semantic recall: There may not be textual similarity between the query content and the document, but there is a certain semantic connection. For example, recalling the answer in the knowledge base based on the question is often achieved through semantic recall.

[0055] Large-scale language models: AI systems trained using deep learning techniques that can understand and generate natural language text. These models can include more than 10 trillion model parameters, enabling them to capture the complexity and subtleties of language and flexibly handle text generation, translation, summarization, reasoning, question answering, sentiment analysis, and more complex tasks.

[0056] Single-tower model: concatenates the query content and the document content, and then uses a single neural network to encode them, so that the query and document are in the same vector space, and finally outputs the relevance score between them.

[0057] Twin-tower model: Use two independent neural network encoders to process queries and documents respectively. Generally, the document vectors can be stored locally in advance, and then the query encoder is used to convert the query into a vector. The query vector is then compared with the vectors in the document vector library through a similarity calculation method.

[0058] Cause of Action Category: A cause of action category is a name for a case formed by the People's Court's summary of the nature of the legal relationship involved in the litigation case. In the examples of this specification, 10 first-level cause of action categories are organized according to the relevant standards for cause of action categories, including "Personal Rights Disputes," "Marriage, Family, and Inheritance Disputes," "Property Rights Disputes," "Contract, Gratuitous Management, and Unjust Enrichment Disputes," "Intellectual Property and Competition Disputes," "Labor Disputes and Personnel Disputes," "Maritime and Maritime Disputes," "Civil Disputes Related to Companies, Securities, Insurance, and Bills of Exchange," "Tort Liability Disputes," and "Causes of Action in Cases Applicable to Special Procedures."

[0059] Code catalog category: Most of the current codes are written according to the hierarchical structure of "section-chapter-section", such as "Civil Code-Marriage and Family Section-Family Relations Chapter-Spousal Relations Section", "Interpretation of the Marriage and Family Section of the Civil Code-Spousal Relations Section", etc.

[0060] With the advancement of computer technology, automated document retrieval has gradually become a research focus. Document retrieval includes, but is not limited to, legal document retrieval, academic document retrieval, business document retrieval, and technical document retrieval. Taking legal document retrieval as an example, this field often faces a challenge: the correlation between user queries and relevant legal provisions is mostly semantic, rather than simply textually similar. This challenge is further complicated by the abstract nature of legal texts, making it difficult to accurately match user queries with relevant legal provisions.

[0061] Currently, legal document retrieval is typically performed using a combination of text recall (coarse ranking based on text similarity) and model-based fine ranking. First, a smaller set of candidate documents (e.g., 10-100) is selected from a large set of candidate documents using text similarity. Then, a single-tower or dual-tower model is used to calculate the relevance score between the query and each candidate document in the candidate set. However, in the field of legal document retrieval, the query and legal text often have only a low textual relevance, making this initial coarse ranking based on text similarity difficult to perform. Another approach is document vector clustering combined with model-based ranking. After saving the document vector data, the vector data is clustered so that each document belongs to a cluster center. Then, for the query vector, the relevance score is calculated only for the candidate documents within the cluster center closest to the query vector. However, in the legal industry, there are numerous terms with similar semantics but completely different definitions and usage scenarios, such as "name rights" and "title rights," "reputation rights" and "honor rights." Simply using semantic clustering to narrow the search scope can easily lead to omissions and misidentifications.

[0062] To address the above-mentioned issues, the embodiments of this specification, through in-depth analysis of officially published legal guides and codes, discovered that each code chapter is typically associated with a specific case category. This finding is also applicable to document retrieval tasks in other fields. Based on this finding, the embodiments of this specification propose a document retrieval method based on document index enhancement, which obtains data to be retrieved for a target document; selects at least one target document index corresponding to the data to be retrieved from multiple document indexes of the target document, wherein the target document index is used to locate document content in the target document that is related to the data to be retrieved; and inputs the data to be retrieved and the at least one target document index into a document retrieval model to obtain document retrieval results for the data to be retrieved, wherein the document retrieval model is trained based on multiple sample retrieval data and sample retrieval results corresponding to the multiple sample retrieval data. By utilizing the mutual mapping relationship between the target document and the document index, and between the document index and the document content, the retrieval scope of the document retrieval model is narrowed. Furthermore, by accurately locating the target document index corresponding to the data to be retrieved and performing document retrieval based on the target document index, the risk of misjudgment caused by pure semantic text matching is avoided, thereby improving the efficiency and accuracy of document retrieval.

[0063] In this specification, a document retrieval method is provided. This specification also involves an automatic question-answering method, a legal provision recommendation method, a document retrieval device, an automatic question-answering device, a legal provision recommendation device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0064] See also Figure 1 , Figure 1 1 shows an architecture diagram of a document retrieval system provided by an embodiment of this specification. The document retrieval system may include a client 100 and a server 200.

[0065] The client 100 is used to send the data to be retrieved for the target document to the server 200;

[0066] The server 200 is configured to filter out at least one target document index corresponding to the data to be retrieved from multiple document indexes of the target document, wherein the target document index is used to locate document content in the target document that is related to the data to be retrieved; input the data to be retrieved and the at least one target document index into a document retrieval model to obtain a document retrieval result for the data to be retrieved, wherein the document retrieval model is trained based on multiple sample retrieval data and sample retrieval results corresponding to the multiple sample retrieval data; and send the document retrieval result to the client 100.

[0067] The client 100 is also used to receive the document retrieval results sent by the server 200.

[0068] The solution of the embodiments of this specification narrows the retrieval scope of the document retrieval model by utilizing the mutual mapping relationship between the target document and the document index, and between the document index and the document content. In addition, by accurately locating the target document index corresponding to the data to be retrieved, document retrieval is performed based on the target document index, avoiding the risk of misjudgment caused by pure semantic text matching and improving the efficiency and accuracy of document retrieval.

[0069] See also Figure 2 , Figure 2 The following diagram illustrates the architecture of another document retrieval system provided by one embodiment of this specification. The document retrieval system may include multiple clients 100 and a server 200. The clients 100 may include end-side devices, and the server 200 may include cloud-side devices. Multiple clients 100 can establish communication connections through the server 200. In a document retrieval scenario, the server 200 provides document retrieval services between the multiple clients 100. The multiple clients 100 can each act as a sender or receiver, communicating through the server 200.

[0070] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100. In the document retrieval scenario, users can publish data streams to the server 200 through the client 100. The server 200 generates document retrieval results based on the data stream and pushes the document retrieval results to other clients with which communication has been established.

[0071] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.

[0072] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application. The client 100 can be based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0073] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that support background training for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server that integrates a blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0074] It is worth noting that the document retrieval methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server and thus execute the document retrieval methods provided in the embodiments of this specification. In other embodiments, the document retrieval methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0075] See also Figure 3 , Figure 3 A flowchart of a document retrieval method provided by an embodiment of this specification is shown, which specifically includes the following steps:

[0076] Step 302: Obtain the data to be retrieved for the target document.

[0077] In one or more embodiments of the present specification, data to be retrieved for a target document may be obtained, and the target document may be retrieved based on the data to be retrieved to obtain a document retrieval result for the data to be retrieved.

[0078] Specifically, the target document is used to provide the retrieval content required for the document retrieval process. The number of target documents can be one or more. The target documents can be documents from different fields, such as code documents in the legal field, research documents in the academic field, financial documents in the business field, operational documents in the technical field, and so on. The target document can be a user-specified document, or it can be all documents in the user-specified field. Of course, the user can choose not to participate in the specification of the target document and use all knowledge base documents as target documents. The data to be retrieved is used to describe the document retrieval intention. The data to be retrieved can be understood as a query. The data to be retrieved can be data in different formats, such as voice data, text data, video data, and so on. The data to be retrieved can also be data in different languages, such as English data, Chinese data, and so on. The data to be retrieved can also be data from different scenarios, such as financial retrieval data in business scenarios, legal text retrieval data in legal scenarios, and so on.

[0079] In practical applications, there are multiple ways to obtain the data to be retrieved for a target document, and the method to be used depends on the actual situation. This specification does not impose any restrictions on this method in the embodiments. In one possible implementation of this specification, the data to be retrieved for a target document can be received from a user. In another possible implementation of this specification, the data to be retrieved for a target document can be read from other data acquisition devices or databases.

[0080] Step 304: Filter out at least one target document index corresponding to the data to be retrieved from the multiple document indexes of the target document, wherein the target document index is used to locate document content related to the data to be retrieved in the target document.

[0081] In one or more embodiments of the present specification, after obtaining the data to be retrieved for the target document, since multiple document indexes correspond to different document contents in the target document, in order to improve the document retrieval efficiency, at least one target document index corresponding to the data to be retrieved in the multiple document indexes can be first determined, thereby narrowing the retrieval scope from the entire target document to the document content corresponding to at least one target document index, and further, document retrieval is performed using the data to be retrieved and at least one target document index.

[0082] Specifically, a document index refers to a data structure for organizing and searching target documents, such as case categories in the legal field and document directories in the academic field. A document index can be a single-level index or a multi-level index. A multi-level index refers to a document index written in a hierarchical structure, such as a code catalog category written in a hierarchical structure of "volume-chapter-section", where the "chapter" level is subordinate to the "volume" level, and the "section" level is subordinate to the "chapter" level and the "volume" level. It can also be understood that the "volume" level includes the "chapter" level and the "section" level, and the "chapter" level includes the "section" level.

[0083] It should be noted that if the multiple document indexes of the target document are known, they can be directly determined based on the target document. If the multiple document indexes of the target document are unknown, the target document can be parsed to determine them. Specifically, the target document can be input into an index generation model to obtain the multiple document indexes of the target document. The index generation model is trained based on multiple sample documents and the index labels of each sample document. Alternatively, the directory generation function of a document editing tool can be used to generate the multiple document indexes of the target document.

[0084] For example, by analyzing the categories of legal cases and the catalog of legal codes, we analyzed the relevant regulations for these categories and identified 10 first-level categories of cases, including "Personal Rights Disputes," "Marriage, Family, and Inheritance Disputes," "Property Rights Disputes," "Contract, Gratuitous Management, and Unjust Enrichment Disputes," "Intellectual Property and Competition Disputes," "Labor Disputes and Personnel Disputes," "Maritime and Maritime Disputes," "Civil Disputes Related to Companies, Securities, Insurance, and Bills of Exchange," "Tort Liability Disputes," and "Causes of Action in Cases Applicable to Special Procedures." Each category of cases has a corresponding code of law referenced in the relevant regulations. Furthermore, taking the first-level case category "marriage, family, and inheritance disputes" as an example, the code catalog categories of "marriage, family, and inheritance disputes" related to it are sorted out. According to the classification of the code catalog categories, a total of 16 second-level case categories are obtained, including "general provisions on marriage and family", "marriage", "husband-wife relationship", "parent-child relationship and other close relatives relationship", "divorce", "adoption relationship", "general provisions on inheritance", "statutory inheritance", "testamentary inheritance and bequest", "handling of estate", "husband-wife debt dispute", "anti-domestic violence", "family protection of minors", "family support and maintenance of the elderly", "foreign-related marriage", and "foreign-related inheritance". For other first-level case categories, a similar sorting method as the first-level case category "marriage, family, and inheritance disputes" can also be adopted, and the embodiments of this specification will not be repeated. Preferably, each first-level case category can contain 10-30 second-level case categories; each second-level case category can contain 10-100 legal provisions involved in its corresponding chapter.

[0085] In actual applications, there are multiple ways to filter out at least one target document index corresponding to the data to be retrieved from multiple document indexes of the target document. The specific method to be selected depends on the actual situation, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, keywords in the data to be retrieved and keywords in multiple document indexes can be identified, and at least one target document index corresponding to the data to be retrieved can be filtered out from multiple document indexes of the target document by using keyword matching. In another possible implementation of this specification, the index screening model's understanding and reasoning ability of document knowledge can be used to input the data to be retrieved and multiple document indexes into the index screening model to obtain at least one target document index corresponding to the data to be retrieved.

[0086] In an optional embodiment of the present specification, taking the multiple document indexes of the target document as a two-level index as an example, that is, the document index includes a first document index and a second document index, and the second document index is subordinate to the first document index; the above-mentioned screening of at least one target document index corresponding to the data to be retrieved from the multiple document indexes of the target document may include the following steps:

[0087] Filtering at least one candidate document index corresponding to the data to be retrieved from the multiple first document indexes of the target document;

[0088] At least one target document index corresponding to the data to be retrieved is screened out from the second document index included in the at least one candidate document index.

[0089] Specifically, the first document index can be understood as a first-level document index, such as a first-level case category in the legal field. The second document index can be understood as a second-level document index, such as a second-level case category in the legal field. The first level includes the second level, and the second document index can be understood as a sub-document index of the first document index.

[0090] It should be noted that before screening out at least one target document index corresponding to the data to be retrieved from the second document index included in at least one candidate document index, the second document index included in at least one candidate document index can be determined. When determining the second document index included in at least one candidate document index, if the first document index includes multiple second document indexes, the second document index included in at least one candidate document index can be directly determined. If the first document index does not include multiple second document indexes, the document content corresponding to the first document index can be parsed to generate multiple second document indexes included in the first document index. Among them, the implementation method of "parsing the document content corresponding to the first document index and generating multiple second document indexes included in the first document index" is the same as the implementation method of "parsing the target document and determining multiple document indexes of the target document" mentioned above, and will not be repeated in this embodiment of the specification.

[0091] In actual applications, the implementation method of "screening out at least one candidate document index corresponding to the data to be retrieved from multiple first document indexes of the target document; and screening out at least one target document index corresponding to the data to be retrieved from the second document indexes included in at least one candidate document index" can refer to the above-mentioned implementation method of "screening out at least one target document index corresponding to the data to be retrieved from multiple document indexes of the target document", and the embodiments of this specification will not be repeated here.

[0092] It is worth noting that if there are multiple candidate document indexes corresponding to the data to be retrieved, a circular logic can be used to filter out at least one target document index corresponding to the data to be retrieved from the second document index included in each candidate document index.

[0093] Using the solutions of the embodiments of this specification, at least one candidate document index corresponding to the data to be retrieved is screened from multiple first document indexes of a target document; and at least one target document index corresponding to the data to be retrieved is screened from the second document indexes included in the at least one candidate document index. By screening the target document index from the second document indexes included in the at least one candidate document index, the document retrieval scope can be narrowed, thereby improving document retrieval efficiency.

[0094] In an optional embodiment of the present specification, a method for selecting candidate document indexes using an index selection model is described. That is, selecting at least one candidate document index corresponding to the data to be retrieved from the multiple first document indexes of the target document may include the following steps:

[0095] Get screening prompt information;

[0096] The screening prompt information, the data to be retrieved and the plurality of first document indexes are input into the index screening model to obtain at least one candidate document index corresponding to the data to be retrieved.

[0097] Specifically, the filtering prompt information (Prompt) is used to guide the index filtering model to perform index filtering, such as "You are a professional. Which of the following document indexes {document index} do you think the problem {data to be retrieved} belongs to? Only return the category name, do not return other content, and use ";" to separate multiple categories. There are many ways to obtain filtering prompt information, and you can choose according to the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the filtering prompt information sent by the user can be received. In another possible implementation of this specification, the filtering prompt information can be read from other data acquisition devices or databases.

[0098] The index screening model can be a large model, or it can be a model obtained by continuing to train on data in a specified field based on the large model. That is, the index screening model is trained based on multiple sample retrieval data and the sample document index and index retrieval label of each sample retrieval data. By training the large model on data in a specified field, the large model can master the knowledge in the specified field and can make inferences and judgments based on the knowledge in the specified field. Index screening models include but are not limited to BERT (Bidirectional Encoder Representations from Transformers) models and T5 (Text-to-Text Transfer Transformer) models. The candidate document index refers to the first document index that matches the data to be retrieved, and the target document index refers to the second document index that matches the data to be retrieved.

[0099] Using the solutions of the embodiments of this specification, filtering prompt information is obtained; the filtering prompt information, the data to be retrieved, and multiple first document indexes are input into an index filtering model to obtain at least one candidate document index corresponding to the data to be retrieved. By utilizing the index filtering model for processing, the index filtering model's understanding and reasoning capabilities of document knowledge and the hierarchical information of document indexes are effectively integrated, thereby improving document retrieval accuracy.

[0100] In an optional embodiment of the present specification, the above-mentioned screening of at least one target document index corresponding to the data to be retrieved from the second document index included in the at least one candidate document index may include the following steps:

[0101] Get screening prompt information;

[0102] The screening prompt information, the data to be retrieved, and the second document index are input into the index screening model to obtain at least one target document index corresponding to the data to be retrieved.

[0103] It should be noted that when using the index screening model to perform multiple rounds of document index screening, the screening prompt information for each round can be updated according to the different document indexes. Taking the legal text recommendation scenario as an example, when using the index screening model to screen at least one candidate document index in the first round, the updated screening prompt information can be "You are a professional in the legal field. Which of the following case categories {first-level case categories} do you think the legal consulting issue {data to be retrieved} belongs to? Only return the category name, do not return other content, and use ";" to separate multiple categories." When using the index screening model to screen at least one target document index in the second round, the updated screening prompt information can be "Which of the following second-level case categories {second-level case categories} do you think the legal consulting issue {data to be retrieved} belongs to? Again, only return the category name, do not return other content, and use ";" to separate multiple categories.

[0104] In practical applications, if the data to be retrieved is non-text modal data (such as voice data), the data to be retrieved can be modally converted to obtain text modal data to be retrieved. The screening prompt information, the text modal data to be retrieved, and multiple first document indexes can be input into an index screening model to obtain at least one candidate document index corresponding to the data to be retrieved. The screening prompt information, the text modal data to be retrieved, and the second document index can be input into the index screening model to obtain at least one target document index corresponding to the data to be retrieved. Mode conversion methods include, but are not limited to, modal conversion tools (such as audio-to-text tools) and modal conversion models (such as speech-to-text recognition models).

[0105] Applying the solutions of the embodiments of this specification, filtering prompt information is obtained; the filtering prompt information, the data to be retrieved, and the second document index are input into an index filtering model to obtain at least one target document index corresponding to the data to be retrieved. By utilizing the index filtering model for multiple rounds of processing, the index filtering model's understanding and reasoning capabilities of document knowledge and the hierarchical information of the document index are effectively integrated, improving document retrieval accuracy. Furthermore, by designing specific multi-round filtering prompt information for the index filtering model, the user's data to be retrieved can be accurately mapped to the relevant target document content.

[0106] In an embodiment of the present specification, since at least one target document index output by the index screening model may include redundant content or be inconsistent with the text index of the target document, in order to avoid document retrieval errors caused by the above-mentioned circumstances, the at least one target document index may be normalized so that the at least one target document index is aligned with the document index of the target document. Specifically, after inputting the screening prompt information, the data to be retrieved, and multiple first document indexes into the index screening model and obtaining at least one candidate document index corresponding to the data to be retrieved, candidate association indicators between the at least one candidate document index and the first document index may be determined respectively; based on the first document index and the candidate association indicators, the at least one candidate document index may be updated to obtain at least one updated candidate document index. After inputting the screening prompt information, the data to be retrieved, and the second document index into the index screening model and obtaining at least one target document index corresponding to the data to be retrieved, association indicators between the at least one target document index and the second document index may be determined respectively; based on the second document index and the association indicators, the at least one target document index may be updated to obtain at least one updated target document index.

[0107] In an optional embodiment of the present specification, after inputting the screening prompt information, the data to be retrieved, and the second document index into the index screening model and obtaining at least one target document index corresponding to the data to be retrieved, the following steps may also be included:

[0108] determining a correlation index between at least one target document index and a second document index respectively;

[0109] At least one target document index is updated according to the second document index and the correlation index to obtain at least one updated target document index.

[0110] Specifically, the correlation index represents the correlation degree between the second document index and the target document index. The correlation index can be a correlation degree index, such as strong correlation, correlation, weak correlation, no correlation, etc. The correlation index can also be a correlation degree value, such as 98%.

[0111] It should be noted that there are multiple ways to determine the correlation index between at least one target document index and the second document index, and the specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the correlation index between the target document index and the second document index can be determined using similarity algorithms such as cosine similarity, Euclidean distance, and Natural Language Toolkit (NLTK). In another possible implementation of this specification, the target document index and the second document index can be input into the matching model, and the correlation index between the target document index and the second document index can be output. When updating at least one target document index based on the second document index and the correlation index, the correlation index can be sorted from large to small or from strong to weak, and the second document index with the highest ranking can be determined as the updated target document index.

[0112] For example, assuming that the target document index is "marriage, family and inheritance disputes", the second document index "marriage, family, inheritance disputes" has the largest correlation index with the target document index, then the updated target document index is "marriage, family, inheritance disputes".

[0113] In actual applications, the implementation method of "respectively determining the candidate association index between at least one candidate document index and the first document index; updating the at least one candidate document index according to the first document index and the candidate association index to obtain at least one updated candidate document index" is the same as the above-mentioned implementation method of "respectively determining the association index between at least one target document index and the second document index; updating the at least one target document index according to the second document index and the association index to obtain at least one updated target document index", and this embodiment of the present specification will not be repeated in detail.

[0114] By applying the solution of the embodiment of this specification, at least one target document index is updated according to the second document index and the associated index to obtain at least one updated target document index, so that the at least one target document index is aligned with the document index of the target document, thereby ensuring the accuracy of document retrieval.

[0115] Step 306: Input the data to be retrieved and at least one target document index into a document retrieval model to obtain a document retrieval result of the data to be retrieved, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

[0116] In one or more embodiments of the present specification, data to be retrieved for a target document is obtained; after filtering out at least one target document index corresponding to the data to be retrieved from multiple document indexes of the target document, the data to be retrieved and the at least one target document index can be further input into a document retrieval model to obtain a document retrieval result for the data to be retrieved.

[0117] Specifically, the document retrieval model can be understood as a semantic retrieval model or a retrieval recall model. The document retrieval model can be a single-tower model or a dual-tower model, and the specific selection is based on the actual situation. This specification does not impose any restrictions on this. The document retrieval result is related to the data to be retrieved. If the data to be retrieved is financial retrieval data for a business scenario, the document retrieval result is financial knowledge; if the data to be retrieved is legal text retrieval data for a legal scenario, the document retrieval result is legal text.

[0118] Applying the solutions of the embodiments of this specification, the document retrieval method described above effectively integrates the understanding and reasoning capabilities of the index screening model with the hierarchical information of the document index. By designing multiple rounds of screening prompts for the index screening model, it is possible to accurately map the data to be retrieved to the relevant document content. Subsequently, the document retrieval model is activated only for the document content corresponding to the target document index, significantly reducing the search scope and significantly improving the accuracy of document retrieval results.

[0119] In an optional embodiment of the present specification, the document retrieval model includes a retrieval encoding unit and a retrieval unit; the inputting of the data to be retrieved and at least one target document index into the document retrieval model to obtain the document retrieval result of the data to be retrieved may include the following steps:

[0120] The retrieval coding unit encodes the data to be retrieved to obtain a retrieval coding feature;

[0121] The retrieval unit retrieves the document retrieval result of the data to be retrieved based on the retrieval coding feature and the target document feature corresponding to at least one target document index.

[0122] Specifically, the retrieval encoding unit can be understood as a query encoding unit, which is used to encode the user's data to be retrieved and generate retrieval encoding features. Encoding refers to the process of converting the data to be retrieved into a continuous vector representation. This vector is called an embedding, which is also known as a retrieval encoding feature. The retrieval encoding feature can capture the semantic and grammatical relationships between features. The target document feature corresponding to at least one target document index refers to the encoding feature of the document content corresponding to at least one target document index in the target document.

[0123] It should be noted that there are multiple ways to obtain the target document features corresponding to at least one target document index, and the specific selection should be made according to the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the document content corresponding to at least one target document index can be screened out from the target document, and the document content can be encoded to obtain the target document features corresponding to at least one target document index. In another possible implementation of this specification, the complete document features of the target document can be obtained, a feature index can be written for the document features, and the feature index of the document features can be further matched with at least one target document index, and the document features corresponding to the feature index that is the same as the at least one target document index can be determined as the target document features corresponding to the at least one target document index. Among them, the document features of the target document can be obtained by encoding the target document in real time, or by obtaining pre-encoded document features.

[0124] Furthermore, obtaining a document retrieval result for the data to be retrieved based on the retrieval coding feature and the target document feature corresponding to at least one target document index by the retrieval unit may include the following steps: the retrieval unit calculates a feature matching index between the retrieval coding feature and the target document feature corresponding to at least one target document index, and obtaining a document retrieval result for the data to be retrieved from the target document based on the feature matching index. The feature matching index represents the degree of association between the retrieval coding feature and the target document feature. The feature matching index may be a feature matching degree index, such as strong match, match, weak match, mismatch, etc. The feature matching index may also be a feature matching degree numerical value, such as 98%.

[0125] In practical applications, the implementation method for "respectively calculating the feature matching index between the retrieval encoding feature and the target document feature corresponding to at least one target document index" is similar to the implementation method for "respectively determining the association index between at least one target document index and a second document index" described above, and this specification does not impose any limitations on this. When a document retrieval result for the data to be retrieved is obtained from the target document based on the feature matching index, the target document features corresponding to the at least one target document index can be ranked from largest to smallest or from strongest to weakest based on the feature matching index, and the document content corresponding to the target document feature with the highest ranking can be determined as the document retrieval result.

[0126] For example, when a user asks a legal question, the retrieval encoding unit processes the question and converts it into a vector-formatted retrieval encoding feature. The retrieval unit then calculates similarity with the target document features corresponding to at least one target document index, generating a relevance score ranking. Finally, the 5-10 legal provisions with the highest relevance scores are selected based on the relevance score ranking and presented to the user, providing immediate and accurate legal guidance.

[0127] Using the solution of the embodiments of this specification, the retrieval encoding unit encodes the data to be retrieved to obtain a retrieval encoding feature. The retrieval unit then retrieves the document retrieval results for the data to be retrieved based on the retrieval encoding feature and the target document features corresponding to at least one target document index. By only applying the document retrieval model to the document content corresponding to the target document index, the search scope is significantly reduced and the accuracy of the document retrieval results is significantly improved.

[0128] In an optional embodiment of the present specification, taking the document retrieval model as a dual-tower model as an example, the document retrieval model further includes a document encoding unit; before the retrieval unit retrieves the document retrieval results of the to-be-retrieved data based on the retrieval encoding features and the target document features corresponding to at least one target document index, the following steps may be further included:

[0129] The document encoding unit encodes the document content corresponding to at least one target document index in the target document to obtain a target document feature corresponding to the at least one target document index.

[0130] Specifically, the document encoding unit is an encoding unit independent of the retrieval encoding unit and is used to encode the target document and generate target document features.

[0131] It should be noted that before encoding the document content corresponding to at least one target document index in the target document via the document encoding unit, the document content corresponding to at least one target document index in the target document may also be determined. Specifically, a search engine may be used to search the target document for the document content corresponding to at least one target document index. Alternatively, the document content corresponding to at least one target document index in the target document may be directly read using a file operation function in a programming language. The method for determining the document content corresponding to at least one target document index in the target document is selected based on actual circumstances and is not limited in any way by the embodiments of this specification.

[0132] Using the solutions of the embodiments of this specification, the document encoding unit encodes the document content corresponding to at least one target document index in the target document to obtain target document features corresponding to the at least one target document index. By encoding only the document content corresponding to the at least one target document index, the amount of encoded data is reduced, the search scope is significantly narrowed, and the accuracy of document retrieval results is significantly improved.

[0133] In an optional embodiment of the present specification, the document retrieval model further includes a document encoding unit; before the retrieval unit retrieves the document retrieval result of the to-be-retrieved data based on the retrieval encoding feature and the target document feature corresponding to at least one target document index, the following steps may be further included:

[0134] The target document is encoded by the document encoding unit to obtain document features;

[0135] Determining a correspondence between document features and document indexes according to multiple document indexes of a target document;

[0136] According to the corresponding relationship, a target document feature corresponding to at least one target document index is determined.

[0137] It should be noted that after the target document is encoded by the document encoding unit and document features are obtained, the document features can be stored locally. At this point, the document retrieval model can be viewed as an efficient vector representation system, which includes a vector representation of all document content, a retrieval encoding unit, and a retrieval unit. Subsequent retrieval tasks for the target document can directly obtain the pre-encoded document features, eliminating the need to re-encode the target document, thus saving document retrieval time.

[0138] In practical applications, when determining the correspondence between document features and document indexes based on multiple document indexes of a target document, since the document features correspond to the document content of the target document and the document index is an index of the document content of the target document, the document index corresponding to the document content can be determined as the feature index of the document feature of the document content, thereby obtaining the correspondence between the document features and the document indexes. Finally, based on the correspondence, the target document feature corresponding to at least one target document index is selected from the document features.

[0139] Using the solutions of the embodiments of this specification, a document encoding unit encodes a target document to obtain document features. Based on multiple document indexes of the target document, a correspondence between the document features and the document indexes is determined. Based on the correspondence, a target document feature corresponding to at least one target document index is determined. By indexing the document features of all document content in the target document, once the index screening model completes its index screening, the document retrieval model can quickly and directly match the data to be retrieved with the document content associated with the target document index.

[0140] In an optional embodiment of the present specification, after inputting the data to be retrieved and at least one target document index into the document retrieval model and obtaining the document retrieval results for the data to be retrieved, the following steps may be further included:

[0141] Receive adjustment information sent by the user based on the document retrieval result, and adjust the model parameters of the document retrieval model according to the adjustment information.

[0142] It should be noted that after obtaining the document retrieval results of the data to be retrieved, the document retrieval results can be sent to the client so that the client can display the document retrieval results to the user. There are many ways for the client to display the document retrieval results to the user, and the specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the document retrieval results can be directly displayed to the user. In another possible implementation of this specification, the document retrieval results can be displayed to the user based on the user's display demand information. Among them, the display demand information represents the user's demand for viewing the document retrieval results, and the display demand information includes but is not limited to displaying only the document retrieval results and displaying a link to the document retrieval results.

[0143] In actual applications, users may be dissatisfied with the document retrieval results. In this case, the system can receive adjustment information sent by the user based on the document retrieval results and adjust the model parameters of the document retrieval model based on the adjustment information. The adjustment information includes, but is not limited to, updated document retrieval results and updated sample documents. Furthermore, the model parameters of the model can be filtered based on the type of adjustment information. The implementation of "adjusting the model parameters of the document retrieval model based on the adjustment information" can be found in the document retrieval model training method and will not be further described in this embodiment of the specification.

[0144] By applying the solution of the embodiments of this specification, adjustment information sent by the user based on the document retrieval results is received, and the model parameters of the document retrieval model are adjusted according to the adjustment information, the accuracy of the document retrieval results is improved, and at the same time, the user experience is improved.

[0145] In an optional embodiment of the present specification, a method for training a document retrieval model is described. Before inputting the data to be retrieved and at least one target document index into the document retrieval model and obtaining the document retrieval results for the data to be retrieved, the following steps may be further included:

[0146] Acquire multiple sample documents, wherein the sample documents include multiple sample retrieval data, sample retrieval results corresponding to the multiple sample retrieval data, and document index labels;

[0147] Input multiple sample retrieval data and document index labels into the document retrieval model to obtain predicted retrieval results corresponding to the multiple sample retrieval data;

[0148] According to the sample retrieval results and the predicted retrieval results, the model parameters of the document retrieval model are adjusted to obtain a trained document retrieval model.

[0149] Specifically, sample retrieval data refers to the processing object of the model training process. Sample retrieval results refer to the actual retrieval results corresponding to the sample retrieval data, and the sample retrieval results are the retrieval targets of the document retrieval model. Document index tags are used to accurately locate the sample document content related to the sample retrieval data in the sample document.

[0150] In practical applications, there are various ways to obtain multiple sample documents, and the method to be used depends on the actual situation. This specification does not impose any restrictions on this method. In a first possible implementation of this specification, multiple sample documents can be read from other data acquisition devices or databases. In a second possible implementation of this specification, multiple sample documents can be received from a user.

[0151] Furthermore, when adjusting the model parameters of the document retrieval model based on the sample retrieval results and the predicted retrieval results, a retrieval loss value can be calculated based on the sample retrieval results and the predicted retrieval results, and the model parameters of the document retrieval model can be adjusted based on the retrieval loss value until a preset stopping condition is reached, thereby obtaining a trained document retrieval model. There are many functions for calculating the retrieval loss value, such as the cross entropy loss function, the L1 norm loss function, the maximum loss function, the mean square error loss function, the logarithmic loss function, etc. The specific selection should be based on the actual situation, and the embodiments of this specification do not impose any restrictions on this.

[0152] In actual applications, the preset stopping conditions include but are not limited to the retrieval loss value being less than or equal to a preset threshold and the number of iterations reaching a preset number of iterations, wherein the preset threshold and the preset number of iterations are selected according to actual conditions, and the embodiments of this specification do not impose any restrictions on this.

[0153] In one possible implementation of this specification, after calculating the retrieval loss value, the retrieval loss value is compared with a preset threshold. Specifically, if the retrieval loss value is greater than the preset threshold, it indicates that the difference between the predicted retrieval result and the sample retrieval result is large, and the document retrieval model has poor retrieval capabilities for the sample retrieval data. In this case, the model parameters of the document retrieval model can be adjusted, and the process of inputting multiple sample retrieval data and document index labels into the document retrieval model to obtain the predicted retrieval results corresponding to the multiple sample retrieval data can be returned to. The document retrieval model is then trained until the retrieval loss value is less than or equal to the preset threshold, indicating that the difference between the predicted retrieval result and the sample retrieval result is small, and the preset stopping condition is met, thereby obtaining a fully trained document retrieval model.

[0154] In another possible implementation of this specification, in addition to comparing the retrieval loss value with a preset threshold, the number of iterations may also be used to determine whether the current document retrieval model is fully trained. Specifically, if the retrieval loss value is greater than the preset threshold, the model parameters of the document retrieval model are adjusted, and the process returns to the step of inputting multiple sample retrieval data and document index labels into the document retrieval model to obtain predicted retrieval results corresponding to the multiple sample retrieval data. Training of the document retrieval model continues until the preset number of iterations is reached, at which point iterations are terminated to obtain a fully trained document retrieval model.

[0155] Using the solution of the embodiments of this specification, a retrieval loss value is calculated based on the predicted retrieval results and the sample retrieval results. This retrieval loss value is then compared with a preset stopping condition. If the preset stopping condition is not met, the document retrieval model is trained continuously until the preset stopping condition is met, completing the training and obtaining the document retrieval model. By continuously adjusting the model parameters of the document retrieval model, the resulting document retrieval model can be made more accurate.

[0156] The following combined Figure 4 , taking the application of the document retrieval method provided in this specification in the intelligent question-answering scenario as an example, the document retrieval method is further explained. Figure 4 A flowchart of an automatic question-answering method provided by an embodiment of this specification is shown, which specifically includes the following steps:

[0157] Step 402: Obtain questions to be answered for the target document.

[0158] Step 404: Filter out at least one target document index corresponding to the question to be answered from the multiple document indexes of the target document, wherein the target document index is used to locate document content related to the question to be answered in the target document.

[0159] Step 406: Input the question to be answered and at least one target document index into a document retrieval model to obtain an answer result for the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and sample retrieval results corresponding to the multiple sample retrieval data.

[0160] It should be noted that the implementation of steps 402 to 406 can refer to the implementation of steps 302 to 306 above, and will not be described in detail in this embodiment of the specification.

[0161] The solution of the embodiments of this specification narrows the retrieval scope of the document retrieval model by utilizing the mutual mapping relationship between the target document and the document index, and between the document index and the document content. In addition, by accurately locating the target document index corresponding to the question to be answered, intelligent question answering is performed based on the target document index, avoiding the risk of misjudgment caused by pure semantic text matching and improving the efficiency and accuracy of automatic question answering.

[0162] The following combined Figure 5 , taking the application of the document retrieval method provided in this specification in a legal scenario as an example, the document retrieval method is further explained. Figure 5 A flowchart of a method for recommending legal provisions provided in one embodiment of this specification is shown, which specifically includes the following steps:

[0163] Step 502: Obtain questions to be answered for the target legal document.

[0164] Step 504: Filter out at least one target document index corresponding to the question to be answered from the multiple document indexes of the target legal document, wherein the target document index is used to locate the legal provisions related to the question to be answered in the target legal document.

[0165] Step 506: Input the question to be answered and at least one target document index into the document retrieval model to obtain the legal provisions recommendation results for the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

[0166] It should be noted that the implementation of steps 502 to 506 can refer to the implementation of steps 302 to 306 above, and will not be described in detail in this embodiment of the specification.

[0167] The solution of the embodiments of this specification narrows the retrieval scope of the document retrieval model by utilizing the mutual mapping relationship between the target legal document and the document index, and between the document index and the legal text. In addition, by accurately locating the target document index corresponding to the question to be answered, legal text recommendations are made based on the target document index, avoiding the risk of misjudgment caused by pure semantic text matching and improving the efficiency and accuracy of legal text recommendations.

[0168] In an optional embodiment of the present specification, after inputting the question to be answered and at least one target document index into the document retrieval model and obtaining the legal provision recommendation result for the question to be answered, the following steps may also be included:

[0169] generating optimization prompt information in response to the model feedback information sent by the client, wherein the optimization prompt information is used to guide the client to send model optimization data for optimizing the document retrieval model;

[0170] Send optimization prompt information to the client, and receive model optimization data sent by the client based on the optimization prompt information;

[0171] Adjust model parameters of the document retrieval model based on model tuning data.

[0172] Specifically, model feedback information can include feedback on the accuracy of the document retrieval model, such as "the model is inaccurate," or feedback on the applicable domain of the document retrieval model, such as "can it handle retrieval tasks in the XXX domain?" Model optimization data can include either an accurate optimization sample set or the optimization domain of the model. The choice is based on actual circumstances and is not limited in this specification.

[0173] It should be noted that there are multiple ways to generate optimization prompt information in response to the model feedback information sent by the client, and the specific selection depends on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation method of this specification, you can directly obtain pre-set optimization prompt information, such as "I am very sorry to have brought you inaccurate information. Please point out the specific inaccuracies or provide the correct answers to related questions. I will correct and optimize my answers as soon as possible to better serve you." In another possible implementation method of this specification, you can perform type identification on the model feedback information to determine the information type of the model feedback information, and further match the information type with the prompt type of each prompt information in the prompt information library, and determine the prompt information with the same prompt type as the information type as the optimization prompt information.

[0174] Furthermore, after the optimization prompt information is generated, the optimization prompt information can be sent to the client, the client can display the optimization prompt information to the user, and the user can send model optimization data to the client based on the optimization prompt information. After receiving the model optimization data sent by the client based on the optimization prompt information, if the model optimization data is an accurate optimization sample set, the optimization sample set can be directly used to adjust the model parameters of the document retrieval model to obtain an updated document retrieval model. If the model optimization data is the optimization field of the model, such as the XXX field, a sample set of the XXX field can be obtained, and the model parameters of the document retrieval model can be adjusted using the sample set of the XXX field to obtain an updated document retrieval model. Among them, the process of adjusting the model parameters of the document retrieval model is the same as the training process of the above-mentioned document retrieval model, and will not be repeated in the embodiment of this specification.

[0175] Using the solution of the embodiments of this specification, optimization prompt information is generated in response to model feedback information sent by the client, wherein the optimization prompt information is used to guide the client to send model optimization data for optimizing the document retrieval model; the optimization prompt information is sent to the client, and the model optimization data sent by the client based on the optimization prompt information is received; and the model parameters of the document retrieval model are adjusted according to the model optimization data. By obtaining the model optimization data through interactive guidance and optimizing the parameters of the document retrieval model, the document retrieval model is made more accurate, while improving interactivity with the user and increasing user satisfaction.

[0176] See also Figure 6 , Figure 6 The following is a flowchart showing a process of a method for recommending legal provisions according to an embodiment of the present specification, which specifically includes:

[0177] Obtain the unanswered question "Who should pay child support after divorce" for the target legal document; input the filtering prompt information "You are a professional in the legal field. Which of the following case categories {first document index} do you think the legal consultation question {unanswered question} belongs to. Only return the category name, do not return other content. If multiple categories are involved, use "; to separate", the data to be retrieved and multiple first document indexes into the index filtering model to obtain the candidate document index "marriage, family and inheritance disputes" corresponding to the data to be retrieved; perform index standardization on the candidate document index "marriage, family and inheritance disputes" to obtain the updated candidate document index "marriage, family, inheritance disputes"; input the filtering prompt information "Which of the following case categories {marriage, family, inheritance disputes} do you think {unanswered question} belongs to {marriage, family, inheritance disputes} {second document index Similarly, only the category name is returned, and other content is not returned. For multiple categories, use "; to separate". Input the data to be retrieved and the second document index into the index screening model to obtain the target document index "parent-child relationship and other close relatives relationship; divorce" corresponding to the data to be retrieved; perform index standardization on the target document index to obtain the updated target document index "parent-child relationship and other close relatives relationship; divorce"; input the question to be answered and the target document index into the document retrieval model, and use the document retrieval model to sort out a total of 55 legal provisions under the categories of "marriage and family, inheritance disputes-divorce" and "marriage and family-parent-child relationship and other close relatives relationship", and determine that the recommended legal provision for the question to be answered is "Article 1,085 of the "XXX Code": After divorce, children..."

[0178] The solution of the embodiment of this specification cleverly utilizes the mutual mapping relationship between case cause categories and code chapters, and between code chapters and legal provisions, thereby avoiding the reliance on text similarity in traditional strategies, and also avoiding the risk of misjudgment and omission based on the purely semantic "document vector clustering" method, and also achieving the purpose of narrowing the search scope. In addition, the solution also fully integrates the index screening model's understanding of legal knowledge and powerful reasoning and judgment capabilities. It can accurately locate the case cause category whose definition and usage scenario are consistent with the problem based on the description of the legal consultation problem, ensuring the high reliability of the case cause classification. At the same time, it supports the output of multiple candidate case cause categories to achieve feedback on complex problems, effectively avoiding the risk of misjudgment or omission in professional fields based on the purely semantic "document vector clustering" method. In addition, considering that the output of the index screening model may not be completely aligned with the document index, a lightweight index normalization processing method is adopted to ensure the consistency of the model output and the document index without basically losing model performance.

[0179] See also Figure 7 , Figure 7This diagram illustrates a document retrieval interface provided by one embodiment of this specification. The document retrieval interface is divided into a request input interface and a result display interface. The request input interface includes a request input box, an "OK" control, and a "Cancel" control. The result display interface includes a result display box.

[0180] The user enters a document retrieval request into the request input box displayed on the client, where the document retrieval request includes the data to be retrieved. The user clicks the "OK" control. The server receives the data to be retrieved from the client and selects at least one target document index corresponding to the data to be retrieved from multiple document indexes of the target document. The target document index is used to locate document content in the target document that is related to the data to be retrieved. The server then inputs the data to be retrieved and the at least one target document index into a document retrieval model to obtain a document retrieval result for the data to be retrieved. The document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data. The server then sends the document retrieval result to the client. The client displays the document retrieval result in a result display box.

[0181] In actual applications, users can operate controls by clicking, double-clicking, touching, hovering the mouse, sliding, long pressing, voice control, or shaking, etc. The specific selection is based on the actual situation, and the embodiments of this specification do not impose any restrictions on this.

[0182] Corresponding to the above document retrieval method embodiment, this specification also provides a document retrieval device embodiment, Figure 8 FIG. 1 shows a schematic diagram of the structure of a document retrieval device provided by an embodiment of this specification. Figure 8 As shown, the device includes:

[0183] A first acquisition module 802 is configured to acquire data to be retrieved for a target document;

[0184] A first screening module 804 is configured to screen out at least one target document index corresponding to the data to be retrieved from the multiple document indexes of the target document, wherein the target document index is used to locate document content related to the data to be retrieved in the target document;

[0185] The first input module 806 is configured to input the data to be retrieved and at least one target document index into a document retrieval model to obtain a document retrieval result of the data to be retrieved, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

[0186] Optionally, the document retrieval model includes a retrieval encoding unit and a retrieval unit; the first input module 806 is further configured to encode the data to be retrieved through the retrieval encoding unit to obtain retrieval encoding features; through the retrieval unit, the document retrieval results of the data to be retrieved are retrieved based on the retrieval encoding features and the target document features corresponding to at least one target document index.

[0187] Optionally, the document retrieval model also includes a document encoding unit; the device also includes: a first encoding module, configured to encode the document content corresponding to at least one target document index in the target document through the document encoding unit, and obtain the target document feature corresponding to the at least one target document index.

[0188] Optionally, the document retrieval model also includes a document encoding unit; the device also includes: a second encoding module, configured to encode the target document through the document encoding unit to obtain document features; determine the correspondence between document features and document indexes based on multiple document indexes of the target document; and determine the target document features corresponding to at least one target document index based on the correspondence.

[0189] Optionally, the device further includes: a first adjustment module configured to receive adjustment information sent by a user based on the document retrieval result, and adjust model parameters of the document retrieval model according to the adjustment information.

[0190] Optionally, the document index includes a first document index and a second document index, and the second document index is subordinate to the first document index; the first screening module 804 is further configured to screen out at least one candidate document index corresponding to the data to be retrieved from multiple first document indexes of the target document; and to screen out at least one target document index corresponding to the data to be retrieved from the second document index included in at least one candidate document index.

[0191] Optionally, the first screening module 804 is further configured to obtain screening prompt information; input the screening prompt information, the data to be retrieved and the plurality of first document indexes into the index screening model to obtain at least one candidate document index corresponding to the data to be retrieved.

[0192] Optionally, the first screening module 804 is further configured to obtain screening prompt information; input the screening prompt information, the data to be retrieved and the second document index into the index screening model to obtain at least one target document index corresponding to the data to be retrieved.

[0193] Optionally, the device also includes: an updating module, configured to respectively determine the association index between at least one target document index and the second document index; update the at least one target document index according to the second document index and the association index to obtain the at least one updated target document index.

[0194] Optionally, the device also includes: a document retrieval model training module, configured to obtain multiple sample documents, wherein the sample documents include multiple sample retrieval data, sample retrieval results corresponding to the multiple sample retrieval data, and document index labels; inputting the multiple sample retrieval data and document index labels into the document retrieval model to obtain predicted retrieval results corresponding to the multiple sample retrieval data; adjusting the model parameters of the document retrieval model according to the sample retrieval results and the predicted retrieval results to obtain a trained document retrieval model.

[0195] The solution of the embodiments of this specification narrows the retrieval scope of the document retrieval model by utilizing the mutual mapping relationship between the target document and the document index, and between the document index and the document content. In addition, by accurately locating the target document index corresponding to the data to be retrieved, document retrieval is performed based on the target document index, avoiding the risk of misjudgment caused by pure semantic text matching and improving the efficiency and accuracy of document retrieval.

[0196] The above is a schematic diagram of a document retrieval device according to this embodiment. It should be noted that the technical solution of the document retrieval device and the technical solution of the document retrieval method described above are based on the same concept. For details not described in detail in the technical solution of the document retrieval device, please refer to the description of the technical solution of the document retrieval method described above.

[0197] Corresponding to the above-mentioned automatic question-answering method embodiment, this specification also provides an automatic question-answering device embodiment, Figure 9 FIG1 shows a schematic diagram of the structure of an automatic question-answering device provided by an embodiment of this specification. Figure 9 As shown, the device includes:

[0198] The second acquisition module 902 is configured to acquire questions to be answered for the target document;

[0199] The second screening module 904 is configured to screen out at least one target document index corresponding to the question to be answered from the multiple document indexes of the target document, wherein the target document index is used to locate document content in the target document that is relevant to the question to be answered;

[0200] The second input module 906 is configured to input the question to be answered and at least one target document index into the document retrieval model to obtain an answer result for the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

[0201] The solution of the embodiments of this specification narrows the retrieval scope of the document retrieval model by utilizing the mutual mapping relationship between the target document and the document index, and between the document index and the document content. In addition, by accurately locating the target document index corresponding to the question to be answered, intelligent question answering is performed based on the target document index, avoiding the risk of misjudgment caused by pure semantic text matching and improving the efficiency and accuracy of automatic question answering.

[0202] The above is a schematic diagram of an automatic question-answering device according to this embodiment. It should be noted that the technical solution of this automatic question-answering device and the technical solution of the automatic question-answering method described above are based on the same concept. For details not described in detail in the technical solution of the automatic question-answering device, please refer to the description of the technical solution of the automatic question-answering method described above.

[0203] Corresponding to the above-mentioned legal provisions recommendation method embodiment, this specification also provides a legal provisions recommendation device embodiment. Figure 10 FIG1 shows a schematic diagram of the structure of a legal provision recommendation device provided by an embodiment of this specification. Figure 10 As shown, the device includes:

[0204] The third acquisition module 1002 is configured to acquire questions to be answered for the target legal document;

[0205] The third screening module 1004 is configured to screen out at least one target document index corresponding to the question to be answered from the multiple document indexes of the target legal document, wherein the target document index is used to locate the legal provisions in the target legal document that are relevant to the question to be answered;

[0206] The third input module 1006 is configured to input the question to be answered and at least one target document index into the document retrieval model to obtain the legal text recommendation results of the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

[0207] Optionally, the device also includes: a second adjustment module, configured to generate optimization prompt information in response to model feedback information sent by the client, wherein the optimization prompt information is used to guide the client to send model optimization data for optimizing the document retrieval model; send the optimization prompt information to the client, and receive the model optimization data sent by the client based on the optimization prompt information; adjust the model parameters of the document retrieval model according to the model optimization data.

[0208] The solution of the embodiments of this specification narrows the retrieval scope of the document retrieval model by utilizing the mutual mapping relationship between the target legal document and the document index, and between the document index and the legal text. In addition, by accurately locating the target document index corresponding to the question to be answered, legal text recommendations are made based on the target document index, avoiding the risk of misjudgment caused by pure semantic text matching and improving the efficiency and accuracy of legal text recommendations.

[0209] The above is a schematic diagram of a legal provision recommendation device according to this embodiment. It should be noted that the technical solution of this legal provision recommendation device and the technical solution of the aforementioned legal provision recommendation method are based on the same concept. For details not described in detail in the technical solution of the legal provision recommendation device, please refer to the description of the technical solution of the aforementioned legal provision recommendation method.

[0210] Figure 11 11 is a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0211] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a World Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0212] In one embodiment of the present specification, the above components of the computing device 1100 and Figure 11 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 11 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0213] Computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1100 may also be a mobile or stationary server.

[0214] The processor 1120 is used to execute a computer program / instruction, which, when executed by the processor, implements the steps of the above-mentioned document retrieval method, automatic question-answering method, or legal provision recommendation method.

[0215] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the document retrieval method, automatic question-answering method, and legal provision recommendation method described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the document retrieval method, automatic question-answering method, or legal provision recommendation method described above.

[0216] An embodiment of the present specification also provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned document retrieval method, automatic question-answering method, or legal provision recommendation method.

[0217] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the document retrieval method, automatic question-answering method, and legal provision recommendation method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the document retrieval method, automatic question-answering method, or legal provision recommendation method described above.

[0218] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned document retrieval method, automatic question-answering method, or legal provision recommendation method.

[0219] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product is based on the same concept as the technical solutions of the document retrieval method, automatic question-answering method, and legal provision recommendation method described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the document retrieval method, automatic question-answering method, or legal provision recommendation method described above.

[0220] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0221] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0222] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0223] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0224] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A document retrieval method, comprising: Obtain the data to be retrieved for the target document; Filtering at least one target document index corresponding to the data to be retrieved from the multiple document indexes of the target document, wherein the target document index is used to locate document content related to the data to be retrieved in the target document; The data to be retrieved and the at least one target document index are input into a document retrieval model to obtain a document retrieval result of the data to be retrieved, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

2. The method according to claim 1, wherein the document retrieval model comprises a retrieval encoding unit and a retrieval unit; The step of inputting the data to be retrieved and the at least one target document index into a document retrieval model to obtain a document retrieval result for the data to be retrieved includes: The retrieval encoding unit encodes the data to be retrieved to obtain a retrieval encoding feature; The retrieval unit retrieves a document retrieval result of the data to be retrieved based on the retrieval coding feature and the target document feature corresponding to the at least one target document index.

3. The method according to claim 2, wherein the document retrieval model further comprises a document encoding unit; Before obtaining the document retrieval result of the data to be retrieved by the retrieval unit according to the retrieval coding feature and the target document feature corresponding to the at least one target document index, the method further includes: The document encoding unit encodes the document content corresponding to the at least one target document index in the target document to obtain a target document feature corresponding to the at least one target document index.

4. The method according to claim 2, wherein the document retrieval model further comprises a document encoding unit; Before obtaining the document retrieval result of the data to be retrieved by the retrieval unit according to the retrieval coding feature and the target document feature corresponding to the at least one target document index, the method further includes: The target document is encoded by the document encoding unit to obtain document features; Determining a correspondence between the document features and the document indexes according to the plurality of document indexes of the target document; According to the corresponding relationship, a target document feature corresponding to the at least one target document index is determined.

5. The method according to claim 1, further comprising: inputting the data to be retrieved and the at least one target document index into a document retrieval model to obtain a document retrieval result for the data to be retrieved; Receive adjustment information sent by a user based on the document retrieval result, and adjust model parameters of the document retrieval model according to the adjustment information.

6. The method according to claim 1, wherein the document index comprises a first document index and a second document index, wherein the second document index is subordinate to the first document index; The step of selecting at least one target document index corresponding to the data to be retrieved from the plurality of document indexes of the target document includes: Filtering at least one candidate document index corresponding to the data to be retrieved from the multiple first document indexes of the target document; At least one target document index corresponding to the data to be retrieved is screened out from the second document index included in the at least one candidate document index.

7. The method according to claim 6, wherein the step of selecting at least one candidate document index corresponding to the data to be retrieved from the plurality of first document indexes of the target document comprises: Get screening prompt information; The screening prompt information, the data to be retrieved and the multiple first document indexes are input into an index screening model to obtain at least one candidate document index corresponding to the data to be retrieved.

8. The method according to claim 6, wherein the step of selecting at least one target document index corresponding to the data to be retrieved from the second document index included in the at least one candidate document index comprises: Get screening prompt information; The screening prompt information, the data to be retrieved and the second document index are input into an index screening model to obtain at least one target document index corresponding to the data to be retrieved.

9. The method according to claim 8, further comprising: inputting the screening prompt information, the data to be retrieved, and the second document index into an index screening model to obtain at least one target document index corresponding to the data to be retrieved; respectively determining association indicators between the at least one target document index and the second document index; The at least one target document index is updated according to the second document index and the association index to obtain the at least one updated target document index.

10. The method according to claim 1, before inputting the data to be retrieved and the at least one target document index into a document retrieval model and obtaining a document retrieval result for the data to be retrieved, further comprising: Acquire a plurality of sample documents, wherein the sample documents include a plurality of sample retrieval data, sample retrieval results corresponding to the plurality of sample retrieval data, and document index tags; Inputting the plurality of sample retrieval data and the document index tags into a document retrieval model to obtain predicted retrieval results corresponding to the plurality of sample retrieval data respectively; According to the sample retrieval results and the predicted retrieval results, the model parameters of the document retrieval model are adjusted to obtain a trained document retrieval model.

11. An automatic question-answering method, comprising: Get the questions to be answered for the target document; Filtering at least one target document index corresponding to the question to be answered from the multiple document indexes of the target document, wherein the target document index is used to locate document content in the target document that is relevant to the question to be answered; The question to be answered and the at least one target document index are input into a document retrieval model to obtain an answer result for the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and sample retrieval results corresponding to the multiple sample retrieval data.

12. A method for recommending legal provisions, comprising: Get answers to questions about the target legal document; Filtering at least one target document index corresponding to the question to be answered from a plurality of document indexes of the target legal document, wherein the target document index is used to locate legal provisions in the target legal document that are relevant to the question to be answered; The question to be answered and the at least one target document index are input into a document retrieval model to obtain a legal provision recommendation result for the question to be answered, wherein the document retrieval model is trained based on multiple sample retrieval data and the sample retrieval results corresponding to the multiple sample retrieval data.

13. The method according to claim 12, wherein after inputting the question to be answered and the at least one target document index into a document retrieval model and obtaining a legal provision recommendation result for the question to be answered, the method further comprises: generating optimization prompt information in response to model feedback information sent by the client, wherein the optimization prompt information is used to guide the client to send model optimization data for optimizing the document retrieval model; Sending the optimization prompt information to the client, and receiving the model optimization data sent by the client based on the optimization prompt information; The model parameters of the document retrieval model are adjusted according to the model optimization data.

14. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 10 or claim 11 or any one of claims 12 to 13 are implemented.

15. A computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10 or claim 11 or any one of claims 12 to 13.

16. A computer program product comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10 or claim 11 or any one of claims 12 to 13.

Citation Information

Patent Citations

  • Intelligent retrieval method, device and equipment and storage medium

    CN114153946A

  • Online technical document searching method and device and medium

    CN116775830A

  • Document-based search with facet information

    US20150046443A1