A clause retrieval method, apparatus, device, medium and program product
By constructing a correspondence between a set of clauses and preset questions, and using a large language model to determine the similarity of the target questions, the problem of low clause retrieval efficiency is solved, and efficient clause information location and answering are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2025-08-14
- Publication Date
- 2026-06-02
Smart Images

Figure CN122132513A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, the application of large models in retrieval scenarios, and the field of information security, specifically to a clause retrieval method, apparatus, device, medium, and program product. Background Technology
[0002] With the widespread adoption of digitalization, users can easily and conveniently access the clauses of regulations or standards online, such as internal management regulations within an organization. However, current methods for searching these clauses are relatively inefficient. Summary of the Invention
[0003] In view of the above problems, this application provides a method, apparatus, device, medium and program product for improving the efficiency of clause retrieval.
[0004] According to a first aspect of this application, a clause retrieval method is provided, comprising: determining a target question; determining the correspondence between a set of clauses and a preset question; wherein any set of clauses contains clause information for answering the corresponding preset question; determining a preset question whose question similarity to the target question is greater than a preset question similarity threshold; and determining the set of clauses corresponding to the determined preset question as the answer basis information for answering the target question.
[0005] Optionally, determining the correspondence between the set of clauses and the preset questions includes: determining the correspondence between the set of clauses and the preset original questions, and the correspondence between the set of clauses and the preset derived questions; the preset original questions are determined based on the corresponding set of clauses; the preset derived questions are determined based on the preset questions corresponding to the corresponding set of clauses.
[0006] Optionally, each set of clauses has a corresponding scope of application; the preset derivative question is determined based on the preset question corresponding to the corresponding set of clauses and the scope of application of the corresponding set of clauses.
[0007] Optionally, the preset questions include: questions that can be answered by the corresponding set of terms, determined based on a large language model.
[0008] Optionally, the method for determining the preset question includes at least one of the following: for any set of clauses, based on a large language model, generating a question that can be answered by the set of clauses, and determining the generated question as the corresponding preset question; for any set of clauses, based on a large language model, generating a question that can be answered by the set of clauses and conforms to the scope of application of the corresponding clauses, and determining the generated question as the corresponding preset question; for any set of clauses, based on a large language model, generating a question that can be answered by the set of clauses and conforms to the scope of application of the corresponding clauses, and determining the generated question as the corresponding preset question; for any set of clauses, based on a large language model, rewriting the preset question corresponding to the set of clauses according to the scope of application of the corresponding clauses, and determining the rewritten question as the corresponding preset question.
[0009] Optionally, determining the preset questions whose question similarity to the target question is greater than a preset question similarity threshold includes at least one of the following: determining preset original questions whose question similarity to the target question is greater than a first preset question similarity threshold; if no preset original questions whose question similarity to the target question is greater than the first preset question similarity threshold are determined, determining preset derived questions whose question similarity to the target question is greater than a second preset question similarity threshold are determined; if no preset original questions whose question similarity to the target question is greater than the first preset question similarity threshold are determined, and no preset derived questions whose question similarity to the target question is greater than the second preset question similarity threshold are determined, for the set of clauses corresponding to the determined preset derived questions, determining preset original questions whose question similarity to the target question meets a preset similarity condition from the preset original questions corresponding to the set of clauses; wherein the second preset question similarity threshold is greater than the third preset question similarity threshold.
[0010] Optionally, the method further includes: determining the answer to the target question based on a large language model, according to the target question and the determined answer basis information.
[0011] Optionally, any set of terms may contain: a single term or multiple terms that are related.
[0012] Optionally, any set of clauses may have a scope of application; the scope of application of any set of clauses may include: the intersection of the scopes of application of the clauses corresponding to the clauses in any set of clauses, and / or the union of the scopes of application of the clauses corresponding to the clauses in any set of clauses.
[0013] Optionally, the scope of application of the terms includes at least one of the following: the objects to which the terms apply, the time period in which the terms apply, the business in which the terms apply, and the subject to which the terms apply.
[0014] Optionally, any set of clauses corresponds to a scope of application of the clauses; determining the set of clauses corresponding to the predetermined question as the basis for answering the target question includes: in the set of clauses corresponding to the predetermined question, determining the set of clauses that the target question conforms to the scope of application of the corresponding clauses as the basis for answering the target question.
[0015] Optionally, any set of clauses corresponds to a scope of application of the clauses; determining the preset question whose question similarity to the target question is greater than a preset question similarity threshold includes: determining the set of clauses that the target question conforms to the scope of application of the corresponding clauses; and among the preset questions corresponding to the determined set of clauses, determining the preset question whose question similarity to the target question is greater than a preset question similarity threshold.
[0016] A second aspect of this application provides a clause retrieval device, comprising: a question unit for determining a target question; a correspondence unit for determining the correspondence between a set of clauses and a preset question; wherein any set of clauses contains clause information for answering the corresponding preset question; and a retrieval unit for determining preset questions whose question similarity to the target question is greater than a preset question similarity threshold, and determining the set of clauses corresponding to the determined preset question as the answer basis information for answering the target question.
[0017] Optionally, the correspondence unit is used to: determine the correspondence between the clause set and the preset original question, and the correspondence between the clause set and the preset derived question; the preset original question is determined based on the corresponding clause set; the preset derived question is determined based on the preset question corresponding to the corresponding clause set.
[0018] Optionally, each set of clauses has a corresponding scope of application; the preset derivative question is determined based on the preset question corresponding to the corresponding set of clauses and the scope of application of the corresponding set of clauses.
[0019] Optionally, the preset questions include: questions that can be answered by the corresponding set of terms, determined based on a large language model.
[0020] Optionally, the method for determining the preset question includes at least one of the following: for any set of clauses, based on a large language model, generating a question that can be answered by the set of clauses, and determining the generated question as the corresponding preset question; for any set of clauses, based on a large language model, generating a question that can be answered by the set of clauses and conforms to the scope of application of the corresponding clauses, and determining the generated question as the corresponding preset question; for any set of clauses, based on a large language model, generating a question that can be answered by the set of clauses and conforms to the scope of application of the corresponding clauses, and determining the generated question as the corresponding preset question; for any set of clauses, based on a large language model, rewriting the preset question corresponding to the set of clauses according to the scope of application of the corresponding clauses, and determining the rewritten question as the corresponding preset question.
[0021] Optionally, the retrieval unit is configured to perform at least one of the following: determine a preset original question whose question similarity to the target question is greater than a first preset question similarity threshold; if no preset original question with a question similarity to the target question greater than the first preset question similarity threshold is determined, determine a preset derived question whose question similarity to the target question is greater than a second preset question similarity threshold; if no preset original question with a question similarity to the target question greater than the first preset question similarity threshold is determined, and no preset derived question with a question similarity to the target question greater than the second preset question similarity threshold is determined, for the set of clauses corresponding to the determined preset derived questions, determine preset original questions whose question similarity to the target question satisfies a preset similarity condition from the preset original questions corresponding to the set of clauses; wherein the second preset question similarity threshold is greater than the third preset question similarity threshold.
[0022] Optionally, the retrieval unit is further configured to: determine the answer to the target question based on a large language model, according to the target question and the determined answer basis information.
[0023] Optionally, any set of terms may contain: a single term or multiple terms that are related.
[0024] Optionally, any set of clauses may have a scope of application; the scope of application of any set of clauses may include: the intersection of the scopes of application of the clauses corresponding to the clauses in any set of clauses, and / or the union of the scopes of application of the clauses corresponding to the clauses in any set of clauses.
[0025] Optionally, the scope of application of the terms includes at least one of the following: the objects to which the terms apply, the time period in which the terms apply, the business in which the terms apply, and the subject to which the terms apply.
[0026] Optionally, any set of clauses corresponds to a scope of application of the clauses; the retrieval unit is used to: in the set of clauses corresponding to the predetermined question, determine the set of clauses that the target question conforms to the scope of application of the corresponding clauses as the answer basis information for answering the target question.
[0027] Optionally, any set of clauses corresponds to a scope of application of the clauses; the retrieval unit is used to: determine the set of clauses that the target question conforms to the scope of application of the corresponding clauses; and among the preset questions corresponding to the determined set of clauses, determine the preset questions whose question similarity to the target question is greater than a preset question similarity threshold.
[0028] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0029] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0030] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description
[0031] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0032] Figure 1 This illustration schematically depicts an application scenario of a clause retrieval method according to an embodiment of this application.
[0033] Figure 2 A flowchart illustrating a clause retrieval method according to an embodiment of this application is shown schematically;
[0034] Figure 3 This schematic diagram illustrates a structural block diagram of a clause retrieval device according to an embodiment of the present application;
[0035] Figure 4 A block diagram schematically illustrates an electronic device suitable for implementing a clause retrieval method according to an embodiment of this application. Detailed Implementation
[0036] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0037] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0038] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0039] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0040] With the popularization of digitalization, users can easily and conveniently access clauses in regulations or standards, such as internal management regulations, through the internet. However, current methods for retrieving clauses are relatively inefficient. This application provides a clause retrieval method. In this method, considering that users usually retrieve clauses because they have questions that need answers, and that users often need to retrieve multiple clauses with potential relationships (e.g., a clause mentions that, according to other clauses), the set of clauses can be associated with the questions the user needs to answer through the clauses. Specifically, this can be done by associating any question with a set of clauses that can answer it, or by associating any set of clauses with questions that can be answered. This allows for the determination of a corresponding question index for the set of clauses, enabling clause retrieval based on the questions, thus improving the efficiency of clause retrieval. Specifically, this can be divided into two stages: a preparation stage and a retrieval stage. The preparation stage involves constructing a question index corresponding to the set of clauses. For ease of description, the questions corresponding to the set of clauses are referred to as preset questions. The set of clauses can be used to answer the corresponding preset questions, or the set of clauses can contain information for answering the corresponding preset questions. Through the preparation phase, a correspondence between multiple sets of clauses and preset questions can be established. It's understood that the same set of clauses can correspond to multiple different preset questions, and the same preset question can correspond to multiple different sets of clauses. In the retrieval phase, based on the correspondence between the multiple sets of clauses and preset questions, for the question asked by the user (referred to as the user question for convenience), similar or even identical preset questions can be searched among the multiple preset questions. If similar or identical preset questions are found, the corresponding set of clauses can be further determined, thus enabling clause retrieval and improving retrieval efficiency.
[0041] Subsequently, based on the retrieved terms, answers can be generated to address user questions. Specifically, this can be achieved using a large language model to enhance the search process. It's understandable that, during the preparation phase, pre-defined questions corresponding to the term set can be determined based on the large language model.
[0042] In a specific example, during the preparation phase, for any given set of clauses, a large language model can be used to determine the questions that the clause set can answer, and these questions are designated as preset questions for that clause set. Specifically, this can be achieved by inputting prompts and the clause set, which are then used by the large language model to generate questions that the clause set can answer, serving as the corresponding preset questions; or by inputting prompts, the clause set, and a pre-prepared set of questions, which are then used by the large language model to determine questions that the clause set can answer from the question set, also serving as the corresponding preset questions. Correspondingly, during the retrieval phase, for each user's input question, the question similarity can be calculated between it and each preset question. For preset questions with a similarity greater than a threshold, the corresponding clause set can be designated as the answering information for the user's question. The user's question and answering information can then be further input into the large language model to generate an answer to the user's question, which can then be returned to the customer, thus enhancing the search. Since the preset questions corresponding to the clause set are determined in advance, and the calculation of question similarity is highly efficient, the retrieval efficiency of the clauses can be improved.
[0043] In a specific example, the set of clauses can include internal clauses related to the reimbursement process. The corresponding pre-defined questions can be identified as "How to claim reimbursement," "How to submit a reimbursement request," and "How are reimbursements processed?" Furthermore, for the user's question "How do I claim reimbursement?", pre-defined questions with high similarity can be identified, allowing for the retrieval of internal clauses related to the reimbursement process. The retrieved clause set can then be displayed to the user for easy viewing, or, based on the retrieved clauses and the user's question, corresponding answers can be generated and provided to the user.
[0044] It should be noted that the methods and apparatus disclosed in the embodiments of this application can be used in the field of artificial intelligence technology, the field of information security, and also in the field of fintech. For example, the above methods can be used to retrieve terms related to the financial sector or internal terms of financial institutions; they can also be applied to any field other than fintech, for example, to retrieve internal terms of institutions in other fields. The application fields of the methods and apparatus disclosed in the embodiments of this application are not limited.
[0045] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse. In scenarios where personal information is used for automated decision-making, the methods, devices, and systems provided in the embodiments of this application all provide users with corresponding operation entry points for users to choose to agree to or refuse the automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making process. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and making decisions accordingly. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0046] Figure 1 The illustration depicts an application scenario of a clause retrieval method according to an embodiment of this application. For example... Figure 1As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc. Users can use the first terminal device 101, the second terminal device 102, or the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (only examples). The first terminal device 101, the second terminal device 102, or the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers. The server 105 can be a server providing various services, such as a backend management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, or the third terminal device 103. The backend management server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0047] It is understood that users can perform clause searches locally on their terminal devices, searching for clauses stored locally on the terminal devices, thereby executing the clause search method provided in this application embodiment through the terminal devices. Users can also interact with the server 105 through their terminal devices, searching for clauses in the server 105, and executing the clause search method provided in this application embodiment through the server 105. It is understood that users can send user questions to the server 105 through their terminal devices for retrieval, and the server 105 can return the retrieved clauses, or the answers to the user questions, to the terminal devices for display to the user.
[0048] It should be noted that the clause retrieval method provided in this application embodiment can generally be executed by server 105, first terminal device 101, second terminal device 102, or third terminal device 103. Correspondingly, the clause retrieval device provided in this application embodiment can generally be located in server 105, first terminal device 101, second terminal device 102, or third terminal device 103. The clause retrieval method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with first terminal device 101, second terminal device 102, third terminal device 103, and / or server 105. Correspondingly, the clause retrieval device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with first terminal device 101, second terminal device 102, third terminal device 103, and / or server 105. It should be understood that... Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. It should be understood that... Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0049] Figure 2 A flowchart illustrating a clause retrieval method according to an embodiment of this application is shown schematically.
[0050] like Figure 2 As shown, the method flow of this embodiment may include operations S210 to S230. This application embodiment does not limit the specific executing entity; it can be any electronic device or any software application, such as a user terminal, server, or so on.
[0051] In operation S210, the target problem is identified.
[0052] In operation S220, the correspondence between the clause set and the preset question is determined; wherein, any clause set contains clause information used to answer the corresponding preset question.
[0053] In operation S230, a preset question with a similarity greater than a preset question similarity threshold to the target question is identified, and the set of clauses corresponding to the identified preset question is determined as the answer basis information for answering the target question.
[0054] This method can improve the efficiency of searching for clauses by matching the similarity between the target question and the preset question.
[0055] The embodiments of this application are not limited to a target question. A target question can be any question. For ease of understanding, any question needed for retrieving terms is referred to as a target question. Optionally, a target question can be a question entered by a user, a question received from an external device or application, or a question composed of further relevant information added to a user-entered question. For example, for a user-entered question, relevant prompts and information can be further added using a large language model to form a more detailed and clear question, which serves as the target question. In a specific example, a target question might be, for instance, "What is the xx process?" or "Who should I report to in xx situation?" Adding information to obtain a target question might result in, for instance, "I am a user of xx business, who should I communicate with in xx situation?"
[0056] The embodiments of this application do not limit the terms, nor do they limit the content or form of the terms. Optionally, the terms may specifically be internal regulations of an organization or terms contained in a standard text. For example, an internal regulation of an organization, "When applying for reimbursement, corresponding evidence must be provided, and the format of the evidence is xx," may be a term, and by searching for terms, it is convenient to determine the reimbursement-related procedures or information.
[0057] The embodiments of this application do not limit the set of clauses. Optionally, the set of clauses may contain one clause; the set of clauses may also contain multiple clauses. The set of clauses may contain multiple clauses that are related. For ease of understanding, in a specific example, a process may be related to multiple clauses; therefore, these clauses can be considered as a set of clauses for clause retrieval. The embodiments of this application do not limit the relationship between multiple clauses in the set of clauses. Optionally, multiple clauses in the set that are related may belong to the same business or be used to determine the same process. Each clause in the set may also be related to at least one other clause in the same set; the specific relationship may be a reference relationship or a collaborative relationship, etc. For example, multiple clauses belonging to the same business "reimbursement business" can be identified as a set of clauses. A clause "Based on the provisions of clause xx, further provisions xxxx" and the "provisions of clause xx" referenced by that clause can be identified as a set of clauses. The embodiments of this application do not limit the number of clauses in the set of clauses, nor do they limit the method of determining the set of clauses. Optionally, a large language model can be used to analyze the related clauses among multiple clauses to determine the clause set, while for clauses that are not related, each individual clause can be determined as a clause set.
[0058] Therefore, optionally, any set of clauses may contain a single clause or multiple clauses that are related. This embodiment can facilitate clause retrieval by using clause sets with different numbers of clauses, thereby improving the accuracy and comprehensiveness of clause retrieval. It is understood that different clause sets may contain the same clauses, different clauses, or no identical clauses. For example, for the intersection of business A and business B, there may be a clause that specifies this, and this clause may be added to the clause set corresponding to business A or the clause set corresponding to business B.
[0059] The embodiments of this application do not limit the preset questions corresponding to the clause set. Optionally, the clause set may contain information for answering the corresponding preset questions. For example, the clause set may contain specific clauses or clause details for answering the corresponding preset questions. Optionally, the preset question may be a question that the corresponding clause set can answer. In a specific example, the clause set may be a set of regulations for process A, and the corresponding preset question may be "How to do process A?" or "How to do step xx in process A?" etc. It is understood that, optionally, the same clause set may correspond to different preset questions, and the same preset question may correspond to different clause sets. Specifically, the same preset question may correspond to different clause sets. For example, the same preset question "How to do step B in process A?" may correspond to a clause set related to process A (which may include clauses related to step B), or it may correspond to a clause set related to step B (which may contain a single clause, i.e., a clause related to step B).
[0060] Optionally, the preset questions include: questions that can be answered by the corresponding set of clauses, determined based on a large language model. In this embodiment, the large language model can be used to determine the questions that can be answered by the set of clauses as the corresponding preset questions, which can improve the efficiency of determining preset questions. It can be understood that, specifically, questions that can be answered by the set of clauses can be generated based on the large language model, and these are used as the preset questions corresponding to the set of clauses; or, based on the large language model, questions that can be answered by the set of clauses can be selected from multiple questions for the set of clauses, and these are used as the preset questions corresponding to the set of clauses. The embodiments of this application do not limit the method of determining the correspondence between the set of clauses and the preset questions. Optionally, it can be to determine a pre-constructed correspondence, or to determine the correspondence in real time, or to select the corresponding correspondence from the set of correspondences between the set of clauses and the preset questions based on the target question. Specifically, it can be to select based on information such as the domain or keywords of the target question. For example, the target question can be a question belonging to the financial business domain, so the correspondences in the financial business domain can be selected from the set of correspondences between the set of clauses and the preset questions. Specifically, N sets of correspondences between the set of clauses and the preset questions can be determined, where N is a positive integer.
[0061] It should be noted that the large language model mentioned in the embodiments of this application can be the same large language model, different large language models, or large language models with the same structure. The embodiments of this application do not limit the specific form of the large language model. Optionally, it can be an intelligent agent or other application or device capable of calling the large language model.
[0062] The embodiments of this application do not limit the method of determining the preset questions. Optionally, the preset questions can be generated based on a large language model for the set of clauses, or the preset questions can be selected from multiple questions for the set of clauses. Specifically, the questions that can be answered by the set of clauses can be selected.
[0063] The embodiments of this application do not limit the specific type of the preset question. Optionally, the preset question may include a preset original question and preset derived questions. The preset original question may be a question determined based on a set of clauses. For example, by analyzing and organizing the set of clauses using a large language model, questions that the set of clauses can answer are generated. Preset derived questions may be questions determined based on the preset questions, specifically based on the preset original question or other preset derived questions. For example, a preset derived question may be obtained by rewriting or expanding the preset original question; similarly, a preset derived question may be obtained by rewriting or expanding other preset derived questions. It is understood that preset derived questions are also questions that can be answered by the corresponding set of clauses.
[0064] In a specific example, for a set of clauses related to the reimbursement process, a pre-defined original question, "How to make an expense report," can be generated using a large language model. Furthermore, the pre-defined original question, "How to make an expense report," can be rewritten or expanded to obtain "For xx business personnel, how to make an expense report" or "If xx situation occurs, how to make an expense report," and so on.
[0065] Optionally, determining the correspondence between the clause set and the preset questions can specifically involve determining the correspondence between the clause set and the preset original questions, as well as the correspondence between the clause set and the preset derived questions. The preset original questions can be determined based on the corresponding clause set; the preset derived questions can be determined based on the preset questions corresponding to the corresponding clause set. This embodiment can determine the preset questions in multiple ways, increasing the number and flexibility of the determined preset questions, which can facilitate improving the comprehensiveness and accuracy of clause retrieval based on questions.
[0066] The embodiments of this application do not limit the specific method for determining the preset original questions corresponding to the clause set, nor do they limit the specific method for determining the preset derived questions corresponding to the clause set. Optionally, based on a large language model, questions that can be answered by the clause set can be generated as the preset original questions corresponding to the clause set. Optionally, based on a large language model, for any clause set and any corresponding preset question, preset derived questions can be obtained by rewriting or expanding the preset question by combining other information. Specifically, the preset original questions or existing preset derived questions corresponding to the clause set can be rewritten or expanded to obtain new preset derived questions. For example, the specific application scenario or scope of application of the clause set can be considered to specifically rewrite or expand the questions, improving the efficiency of determining preset derived questions, facilitating an increase in the number and comprehensiveness of preset derived questions and preset questions, and improving the comprehensiveness and accuracy of clause retrieval based on questions.
[0067] Optionally, any set of clauses may correspond to a scope of application for the clauses; the pre-defined derivative questions may be determined based on the pre-defined questions corresponding to the corresponding set of clauses and the scope of application for the corresponding set of clauses. This embodiment can determine the pre-defined derivative questions by combining the scope of application of the clauses, which can improve the efficiency and comprehensiveness of determining the pre-defined derivative questions, facilitate increasing the number and comprehensiveness of the pre-defined derivative questions and pre-defined questions, and improve the comprehensiveness and accuracy of clause retrieval based on questions.
[0068] The embodiments of this application do not limit the scope of application of the corresponding clauses in the clause set, nor do they limit the scope of application of the individual clauses. For ease of understanding, in a specific example, a clause may state, "For business A and business B, operation xx is not allowed," thus determining that the scope of application of this clause is "business A and business B." A clause may state, "For situation xx, xx processing shall be performed," thus determining that the scope of application of this clause is "situation xx." A clause may state, "For user xx, attention to xxx is required," thus determining that the scope of application of this clause is "user xx."
[0069] Optionally, the scope of application of the clauses corresponding to the clause set can specifically refer to the scope of application of the clauses in the clause set, such as the target object, subject, or business, etc. Since the clause set can contain one or more clauses, the embodiments of this application do not limit the specific method for determining the scope of application of the clauses corresponding to the clause set. Optionally, for a clause set containing a single clause, the scope of application of the clause set can be the scope of application of the single clause. For a clause set containing multiple clauses, the scope of application of the clause set can be the intersection of the scopes of application of the various clauses, or the union of the scopes of application of the various clauses.
[0070] Therefore, optionally, any set of clauses may correspond to a scope of application of the clauses; the scope of application of any set of clauses includes: the intersection of the scopes of application of the clauses corresponding to the clauses in the set, and / or the union of the scopes of application of the clauses corresponding to the clauses in the set. This embodiment can conveniently determine the preset derivative issues based on the scope of application of clauses in multiple situations, which can improve the efficiency and accuracy of determining the preset derivative issues.
[0071] The embodiments of this application do not limit the specific content of the scope of application of the terms. Optionally, the scope of application of the terms may include at least one of the following: the objects to which the terms apply, the time period to which the terms apply, the business to which the terms apply, and the subject to which the terms apply. This embodiment can facilitate the determination of pre-defined derivative issues through various types of scope of application of the terms, thereby improving the efficiency and accuracy of determining pre-defined derivative issues.
[0072] Optionally, the scope of application of the clauses can be characterized by the features of the scope of application. For example, the scope of application of the clauses can be "xx business", so that the features of "xx business" can be used to determine the preset derivative questions. Specifically, the preset derivative questions can be determined based on the feature of "xx business" that "it includes xx process", so that the preset derivative questions "how the xx process in xx business is executed" can be generated based on the preset original question "how xx business is executed".
[0073] In a specific example, given a clause set such as "For users of the xx business, the reimbursement process follows the xx process," the scope of application can be determined to be "users of the xx business." Further, for the pre-set original question "What is the reimbursement process?" corresponding to this clause set, corresponding pre-set derivative questions can be determined by combining the characteristics of the clause's scope of application. These characteristics include, for example, the business type, speaking habits, tone, and common terminology of "users of the xx business." Therefore, based on at least one of these characteristics, the pre-set original question can be rewritten to obtain pre-set derivative questions such as "For users of the xx business, how to get reimbursed?", "How should the reimbursement process be conducted?", and "What is the reimbursement channel?". It is understandable that by rewriting the clauses in conjunction with the characteristics of their scope of application, specifically by combining the tone and other characteristics of the applicable objects / subjects, the resulting pre-set derivative questions are more aligned with user habits, improving the efficiency and accuracy of question matching and clause retrieval.
[0074] Therefore, optionally, any set of clauses may correspond to a scope of application; the pre-defined derivative question can be determined based on the pre-defined question corresponding to the corresponding set of clauses and the characteristics of the scope of application of the corresponding set of clauses. Specifically, the pre-defined derivative question can be obtained by rewriting or expanding the pre-defined question corresponding to the corresponding set of clauses based on the characteristics of the scope of application of the corresponding set of clauses. This application does not limit the characteristics of the scope of application of the clauses. The characteristics of the scope of application of the clauses include, for example, the tone, common terminology, speaking habits, attribute information, etc., of the subject / object of application of the clauses. In a specific example, the pre-defined derivative question can be obtained by rewriting or expanding the pre-defined question corresponding to the corresponding set of clauses based on the characteristics (e.g., tone, common terminology, attributes, etc.) of the subject / object of application of the clauses corresponding to the corresponding set of clauses.
[0075] The embodiments of this application do not limit the method for determining the preset question. Optionally, the corresponding preset question can be determined based on at least one of the following: the clause set itself, the preset question currently corresponding to the clause set, and the scope of application of the clauses corresponding to the clause set. Optionally, the preset question can also be determined from the historical clause retrieval process. Specifically, the historical target question in the historical clause retrieval process, where the answer basis information includes any clause set, can be determined as the preset question corresponding to that clause set. Of course, it is also possible to combine user feedback information to determine whether the answer basis information determined in the historical clause retrieval process is accurate, thereby determining the historical target question as the preset question corresponding to the accurate answer basis information (clause set) provided by the user feedback. By accumulating information from the historical clause retrieval process, the accuracy of the preset questions can be easily improved, the number of preset questions can be increased, and the accuracy and efficiency of clause retrieval can be improved.
[0076] Optionally, the method for determining the preset questions may include at least one of the following: (1) For any set of clauses, based on a large language model, generate questions that can be answered by the set of clauses, and determine the generated questions as the corresponding preset questions. (2) For any set of clauses, based on a large language model, generate questions that can be answered by the set of clauses and conform to the scope of application of the corresponding clauses, and determine the generated questions as the corresponding preset questions. (3) For any set of clauses, based on a large language model, generate questions that can be answered by the set of clauses and conform to the scope of application of the corresponding clauses, and determine the generated questions as the corresponding preset questions. (4) For any set of clauses, based on a large language model, rewrite the preset questions corresponding to the set of clauses, and determine the rewritten questions as the corresponding preset questions. This embodiment can determine the preset questions in multiple ways, which can improve the efficiency and flexibility of determining the preset questions, increase the number of preset questions, and improve the efficiency and accuracy of determining the preset questions.
[0077] The embodiments of this application do not limit the issues that fall within the scope of the terms. For example, the scope of the terms may be "xx business," and if the issue is related to "xx business," then the issue may be considered to fall within the scope of the terms. Conversely, if the issue does not relate to "xx business" at all, then the issue may be considered not to fall within the scope of the terms. Therefore, optionally, an issue that falls within the scope of the terms may be an issue related to the scope of the terms.
[0078] The embodiments of this application do not limit the specific method of determining the target problem. Optionally, determining the target problem may specifically involve determining the target problem input by the user, determining the target problem sent by an external device or application, or determining the target problem generated locally. For example, a problem that is expanded and improved based on a large language model in response to a user-input question can be determined as the target problem. The embodiments of this application do not limit the specific number of target problems determined; one or more target problems may be determined. It is understood that, for any target problem, the embodiments of this application can be executed to perform a corresponding clause search. For different target problems, the embodiments of this application can be executed separately to perform a corresponding clause search.
[0079] The embodiments of this application do not limit the specific method for determining the correspondence between the clause set and the preset question. Optionally, the correspondence between the clause set and the preset question can be pre-built or constructed in real time. Specifically, determining the correspondence between the clause set and the preset question can be done by obtaining a pre-built correspondence between the clause set and the preset question, or by constructing the correspondence between the clause set and the preset question in real time. The embodiments of this application do not limit the number of correspondences between the clause set and the preset question that are specifically determined. Optionally, N sets of correspondences between clause sets and preset questions can be determined, where N is a positive integer.
[0080] The embodiments of this application do not limit the specific method of clause retrieval based on the target question. Optionally, based on the determined correspondence, preset questions with a question similarity greater than a preset question similarity threshold with the target question can be determined. The embodiments of this application do not limit the specific method of determining preset questions similar to the target question, nor do they limit the specific method of calculating question similarity. The embodiments of this application do not limit the method of calculating the question similarity between the target question and the preset questions. Optionally, it can specifically calculate the text similarity between the target question and the preset questions; it can also be done by extracting question features, vectorizing the target question and the preset questions, and then calculating the similarity of the question features; or it can be done by using a large language model to determine the similarity between the target question and the preset questions.
[0081] In one optional embodiment, the corresponding retrieval method can be further determined by combining the preset original question and the preset derived question. Specifically, the clause retrieval can be performed separately based on the characteristics of the preset original question and the preset derived question.
[0082] Optionally, a preset question with a question similarity greater than a preset question similarity threshold is identified, including at least one of the following:
[0083] (1) Determine the preset original questions that have a question similarity greater than the first preset question similarity threshold with the target question.
[0084] (2) If it is determined that there is no preset original question with a question similarity greater than the first preset question similarity threshold to the target question, a preset derived question with a question similarity greater than the second preset question similarity threshold to the target question is determined.
[0085] (3) If it is determined that there is no preset original question with a question similarity greater than the first preset question similarity threshold and no preset derived question with a question similarity greater than the second preset question similarity threshold, then a preset derived question with a question similarity greater than the third preset question similarity threshold is determined. For the set of clauses corresponding to the determined preset derived question, from the preset original questions corresponding to the set of clauses, a preset original question with a question similarity that satisfies the preset similarity condition of the target question is determined; the second preset question similarity threshold is greater than the third preset question similarity threshold.
[0086] (4) Identify a preset derivative question whose question similarity to the target question is greater than the second preset question similarity threshold.
[0087] (5) Determine the preset derivative questions that have a question similarity greater than the third preset question similarity threshold with the target question. For the set of clauses corresponding to the determined preset derivative questions, determine the preset original questions that have a question similarity with the target question that meet the preset similarity conditions from the preset original questions corresponding to the set of clauses.
[0088] This embodiment allows for clause retrieval through multiple search methods, which can improve the comprehensiveness and accuracy of clause retrieval.
[0089] It is understandable that, optionally, if it is determined that there is no preset original question with a question similarity greater than a first preset question similarity threshold to the target question, it can be directly determined that no corresponding clause was found. Alternatively, if it is determined that there is no preset original question with a question similarity greater than a first preset question similarity threshold to the target question, and also no preset derived question with a question similarity greater than a second preset question similarity threshold to the target question, it can be directly determined that no corresponding clause was found.
[0090] The embodiments of this application do not limit the number of preset questions determined, nor do they limit the number of corresponding answer basis information. Optionally, one or more answer basis information can be determined, specifically, one or more sets of clauses can be determined as answer basis information.
[0091] The embodiments of this application are not limited to the subsequent processing methods for the answer basis information. Optionally, an answer to the target question can be generated based on the answer basis information, or the answer basis information can be directly displayed to the user who asked the target question, making it convenient for the user to view the clauses related to the target question and complete the clause search. Specifically, when displaying the answer basis information, the source of the answer basis information can also be displayed, which may be information such as the title of the standard text to which the answer basis information belongs, or information such as a link to the standard text to which the answer basis information belongs, so that users can easily view the full text of the standard text.
[0092] Optionally, the answer to the target question can be determined based on a large language model, according to the target question and the determined answer basis information. This embodiment can improve the accuracy and comprehensiveness of the answer to the target question by combining the target question and the retrieved set of terms and generating the answer based on a large language model.
[0093] In an optional embodiment, the clause search can be further performed in conjunction with the scope of application of the clauses. The embodiments of this application do not limit the specific method of performing clause search in conjunction with the scope of application of the clauses.
[0094] Optionally, after searching based on the similarity between the target question and the preset question, a further search can be conducted based on the scope of application of the clauses to determine whether the target question falls within the scope of application. For example, if the target question is "What is process B of business A?", after the first screening based on the similarity between the target question and the preset question, a second screening can be conducted based on the scope of application of the clauses. Specifically, clause sets that do not involve "business A" or "process B" can be excluded, or clause sets whose scope of application does not include "business A" or "process B". Specifically, any clause set may have a corresponding scope of application; the clause set corresponding to the determined preset question can be identified as the basis for answering the target question. Specifically, within the clause set corresponding to the determined preset question, the clause set whose scope of application matches the target question can be identified as the basis for answering the target question. This embodiment can perform clause retrieval through two screenings, combining the similarity between the target question and the preset question, and whether the target question falls within the scope of application of the clauses. This reduces the number of clause sets that need to be screened in the second screening, improving the accuracy and efficiency of clause retrieval.
[0095] Optionally, a matching set of clauses can be selected first based on whether the target question falls within the scope of the clauses. For example, if the target question is "What is process B of business A?", clauses that do not involve "business A" or "process B" can be excluded. Specifically, clauses whose scope of application does not include "business A" and "process B" can be excluded, thus performing the first screening. Further, for the clause set selected in the first screening, a question similarity match can be performed on the target question from the corresponding preset questions. Optionally, any clause set can correspond to a clause scope of application. Preset questions with a question similarity greater than a preset question similarity threshold are determined. Specifically, this can be done by: determining the clause set where the target question falls within the scope of application of the corresponding clauses; and then, among the preset questions corresponding to the determined clause set, determining the preset questions with a question similarity greater than the preset question similarity threshold. This embodiment can perform clause retrieval through two screenings, combining the similarity between the target question and preset questions, and whether the target question falls within the scope of application of the clauses. This reduces the number of preset questions that need to be screened in the second screening, improving the accuracy and efficiency of clause retrieval.
[0096] In another optional embodiment, the determined answer basis information can be filtered, and the answer to the target question can be determined by combining the target question and the filtered answer basis information. For example, the filtering can be based on the number of clauses in the clause set. Specifically, answer basis information containing more than a preset clause number threshold can be selected for subsequent generation of the answer to the target question. Alternatively, the filtering can be based on the relationship between clause sets. Considering that different clause sets may have proper subset relationships, answer basis information containing clauses that are also present in other answer basis information can be deleted, and the answer to the target question can be generated based on the remaining answer basis information.
[0097] Optionally, the determined information supporting the solution can be synthesized, and the answer to the target question can be generated by combining the target question and the synthesized results. This can involve merging the determined information supporting the solution into a comprehensive set of clauses, thereby removing duplicate clauses between different sets of information supporting the solution, and generating the answer to the target question based on the comprehensive set of clauses and the target question.
[0098] In one optional embodiment, any set of clauses may be associated with a set of candidate questions. Specifically, preset questions corresponding to the set of clauses can be selected from the set of candidate questions. It is understood that the candidate questions in the set of candidate questions may be questions that the set of clauses can answer. The embodiments of this application do not limit the specific method of determining the associated set of candidate questions for a set of clauses.
[0099] Optionally, the candidate question set associated with the clause set may include candidate questions determined based on the associated clause set; candidate questions determined based on other candidate questions; and candidate questions whose answer basis information includes the associated clause set. Specifically, candidate questions determined based on the associated clause set may be related questions generated for the associated clause set based on a large language model; candidate questions determined based on other candidate questions may be questions obtained by expanding or rewriting other candidate questions based on a large language model; and candidate questions whose answer basis information includes the associated clause set may be questions for which the associated clause set is determined as the answer basis information. Specifically, clause retrieval can be performed for any question to determine the answer basis information. This can be done using the clause retrieval method provided in the embodiments of this application.
[0100] Optionally, the candidate question set associated with the clause set may include candidate questions determined based on the set of similar clauses of the associated clause set. The similarity between the associated clause set and the set of similar clauses may be greater than a preset clause similarity threshold. The embodiments of this application do not limit the specific calculation method of the clause set similarity; it may be calculated by measuring the text similarity between clause sets, or by combining the scope of application of the clauses in the clause set, or by determining it based on a large language model.
[0101] The embodiments of this application are not limited to the specific method of determining candidate questions based on a set of similar clauses. For details, please refer to the embodiments for determining candidate questions based on a set of clauses. Optionally, the set of similar clauses, as a set of clauses, may also be associated with a set of candidate questions. Therefore, the candidate questions determined based on the set of similar clauses of the associated set of clauses may include candidate questions from other sets of candidate questions associated with the set of similar clauses.
[0102] The embodiments of this application do not limit the specific method for determining the set of similar clauses. Optionally, the similarity of information can be calculated for any two sets of clauses, thereby making a judgment based on a preset clause similarity threshold; alternatively, a large language model can be used to directly determine whether any two sets of clauses are similar clause sets; or the scope of application of clauses can be determined for any two sets of clauses, and then matched to determine whether they are similar clause sets. Optionally, the method for determining the set of similar clauses includes: clustering multiple sets of clauses to obtain clustering results; for any set of clauses in the clustering results, determining the other sets of clauses in the cluster to which the targeted set of clauses belongs as the set of similar clauses for the targeted set of clauses. This embodiment can determine the set of similar clauses by clustering the sets of clauses, which can improve the accuracy of the set of similar clauses.
[0103] In one optional embodiment, the correspondence between the clause set and the preset questions can be updated. Specifically, new correspondences can be added to improve the comprehensiveness and accuracy of clause retrieval. For example, more corresponding preset questions can be added to a clause set. It is understood that clauses may require updates, such as deletion, modification, or addition, which will also update the correspondence accordingly. For example, the corresponding preset questions can be redefined for a modified clause set, and the corresponding preset questions can be defined for a newly added clause set.
[0104] The embodiments of this application are not limited to the method of determining the corresponding preset question for an updated set of clauses; please refer to the explanations of other embodiments. Optionally, a set of similar clauses can be determined for the modified or added set of clauses, thereby facilitating the acquisition of candidate questions, determining the corresponding preset question, and improving the efficiency of preset question determination.
[0105] It is understood that the steps of updating the correspondence and retrieving the terms can be performed in parallel. This application does not limit the execution order of updating the correspondence and retrieving the terms. In an optional embodiment, the correspondence between the term set and preset questions can be stored in a storage space, such as a database. By matching the similarity between the target question and the preset questions in the storage space, the corresponding term set is retrieved to answer the target question. Optionally, this application does not limit the specific execution order of the steps of updating the correspondence and retrieving the term set; they can be performed in parallel or sequentially. For example, a new correspondence can be added to the storage space (database) while retrieving the term set; see other embodiments for examples of determining the corresponding preset questions for adding a new term set. Alternatively, the correspondence can be updated in the storage space before retrieving the term set; specifically, the term set can be retrieved based on the updated correspondence.
[0106] The embodiments of this application also provide an application embodiment. This embodiment relates to the field of institutional knowledge management and intelligent question answering technology, specifically, it relates to a multi-level retrieval enhancement generation method for institutional internal normative texts. By combining a large language model and a structured retrieval mechanism, it achieves efficient and accurate querying and answering of complex institutional norms. This embodiment provides a normative text question answering method based on multi-level retrieval enhancement. By constructing a normative scenario-based knowledge base and designing a multi-level retrieval mechanism, it achieves the following technical objectives: (1) It solves the semantic gap between users' natural language questions and the formal expressions of normative texts, ensuring that the retrieval results are highly relevant to users' actual needs; (2) It reduces the destruction of the complete context of normative texts and maintains the accuracy of clause interpretation through a full-text association mechanism; (3) It establishes a multi-level screening system based on business category-title-scenario summary to accurately locate specific clauses that can answer users' questions.
[0107] This embodiment significantly improves the accuracy and practicality of standardized text queries through a scenario generalization and hierarchical decision-making mechanism driven by a large language model, providing users with efficient and reliable standardized text retrieval services.
[0108] This embodiment provides a standardized text question answering method based on multi-level retrieval enhancement. Its technical solution mainly includes two parts: a preprocessing stage and a retrieval stage. Through the construction of a scenario-based knowledge base driven by a large language model and a multi-level decision-making mechanism, it achieves efficient and accurate retrieval and answering of standardized text.
[0109] 1. Preprocessing stage. The preprocessing stage is used to build a structured, canonical text knowledge base, and specifically includes the following steps.
[0110] Step 1: Summarize the applicable scenarios for the standard texts. Input the titles and full texts of multiple standard texts into the large language model. The model will automatically analyze the text content and generate a summary of the applicable scenarios for each standard text. This summary must accurately reflect the core scope of application of the clauses, such as "applicable to special cases in the financial reimbursement approval process".
[0111] Step 2: Structured Storage. Store the specification text in the database in a structured format: "Business Category - Title - Scenario Scope Summary - Original Specification Text". The business category is a predefined classification system (e.g., "Human Resources", "Financial Management"), the title is the official name of the specification document, the scenario scope summary is the output from Step 1, and the original specification text is the complete clause text.
[0112] 2. Retrieval Phase. The retrieval phase enables multi-level precise positioning of user questions and generation of answers, specifically including the following steps.
[0113] Step 3: User Input Processing. Users input their query through the interactive interface and can optionally select a relevant business category (multiple selections are supported). If no category is selected, proceed to Step 4; otherwise, proceed to Step 5.
[0114] Step 4: Category Recommendation. Send the user's question and a list of business categories to the language model. The model analyzes the question content and returns a list of possible category names. For example, a user's question "travel allowance standards" might be associated with the categories "financial management" and "administrative management".
[0115] Step 5: Title Retrieval. Based on the category name selected by the user or recommended in Step 4, retrieve all titles and their scene scope summaries under the corresponding category from the database.
[0116] Step 6: Initial Screening. Send the list of titles obtained in Step 5 along with the user's question to the large language model. The model is asked to determine which titles' canonical text might contain the answer and return a list of candidate titles. This step achieves coarse-grained screening through title semantic matching.
[0117] Step 7: Scene Summary Retrieval. Based on the standardized text names returned in Step 6, retrieve the corresponding scene scope summary text from the database.
[0118] Step 8: Fine-tuning. Send the scenario summary, titles, and user questions obtained in Step 7 to the large language model. The model is asked to evaluate the actual relevance of each normative text to the question and return the 1-3 normative text titles that are most likely to answer the question. This step achieves fine-grained filtering through scenario matching.
[0119] Step 9: Question and Answer Generation. Send the complete original text corresponding to the title determined in Step 8 and the user's question to the large language model. The model will generate the final answer based on the full text context, ensuring that the answer strictly conforms to the specifications.
[0120] The beneficial effects of this embodiment are mainly reflected in the following aspects: (1) Improved accuracy of semantic association. By establishing a semantic bridge through the scenario scope generalization generated by the large language model, the semantic gap between the user's natural language question and the formal expression of the standard text is effectively bridged. This embodiment uses explicit scenario description as the retrieval basis to ensure that the retrieval results are highly matched with the user's actual needs. For example, when a user asks "how to handle project delays", the system can accurately associate with the relevant clauses of "approval process and responsibility determination for project progress delays" in the "Procedures for Handling Abnormal Situations in Project Management", rather than relying solely on shallow matching of keywords such as "delay" and "handling". (2) Guarantee of clause integrity. The full-text association mechanism is used to replace the traditional block retrieval. The retrieval stage is always based on the complete standard text to generate answers, ensuring the accuracy and authority of the clause interpretation. Especially for content containing exception clauses or additional conditions (such as "except for special approval, all reimbursements must be submitted within 30 days"), the system can completely retain the limiting conditions of the original clauses, reducing the risk of misinterpretation. (3) Optimization of retrieval path. The multi-level retrieval mechanism (category → title → scenario → full text) forms a progressive filtering funnel, which significantly reduces noise interference compared to traditional one-time full-text retrieval. By narrowing the retrieval scope layer by layer, the system can quickly eliminate irrelevant content and accurately locate the target clause. For example, when querying "travel reimbursement requirements", the system first locks the "administrative management" category, then excludes irrelevant normative texts such as "Company Performance Appraisal Regulations", and finally focuses on the specific clauses of "Travel Reimbursement Regulations". (4) Intelligent knowledge base construction. The preprocessing stage uses a large language model to automatically generate a scenario scope summary, which can reduce manual intervention. The generated scenario description can accurately reflect the core applicable scope of the normative text and maintain semantic alignment with the common questioning methods of users, thus establishing an efficient index foundation for subsequent retrieval. (5) Enhanced system adaptability. It provides targeted solutions for the special characteristics of normative texts (complete but wide scope, formal expression but colloquial query, etc.). Unlike general retrieval enhancement systems, this embodiment effectively addresses the unique query demand dispersion and clause association complexity in institutional management scenarios through the dual structure design of business category classification system and scenario summary. The aforementioned technical effects directly stem from the unique multi-level retrieval architecture and scenario-based knowledge base design of this embodiment. Through structured retrieval path planning and semantically enhanced indexing mechanisms, it achieves a fundamental improvement in standardized text question-and-answer queries from "similar retrieval" to "related retrieval".
[0121] This embodiment also provides another process, primarily addressing the semantic gap between user questions and terms and conditions. Specifically, it can pre-generate possible questions from the corresponding roles (corresponding to "scope of application of the terms") using a large language model.
[0122] 1. Preprocessing stage.
[0123] Step 1: Segmenting Standard Clauses. For each standard document, segment it by item to obtain a set of clauses.
[0124] Step 2: Clause Association. For each specification document and its corresponding set of clauses, submit a large language model to analyze dependencies. Clauses with dependencies are aggregated into a clause group, resulting in a clause set.
[0125] Step 3: Question Extraction. For each clause or clause group in the clause set and clause group set (corresponding to the "clause set"), input the large language model to generate possible user questions, resulting in multiple questions (corresponding to the "preset questions"). Cover as many different angles as possible, such as process consultation.
[0126] Step 4: Subject Extraction. For each clause or clause group in the clause set and clause group set, input the large language model requirements to generate the applicable subject group for the clause or clause group (e.g., frontline / middle-level employees, sales / IT staff, etc.). Simulate the language habits and expressions of different roles, consider differences in professional level, and generate multiple variant expressions.
[0127] Input: Individual clause / clause group text. Output: Structured body tagging system. {"Organizational Hierarchy": ["Senior Management", "Middle Management", "Junior Management"] and "Business Line": ["Research and Development", "Marketing", "Finance"]}.
[0128] (1) Subject Identification: Example prompt: "Identify the applicable objects of the following clauses and output them in three categories: organizational level, business line, and attribute. If the subject is not clearly stated in the clause, it will be inferred from the context. Example clause: 'Employees below level A need departmental approval for business trips' → Output: {'Organizational level': [Level A], 'Business line': All business lines}"
[0129] (2) Subject expansion: Generate semantic variants (synonyms, near-synonyms, colloquial expressions) for the identified subject tags: "Middle level" → ["Department head", "Second-level manager", "Team leader"] "Technology research and development" → ["Development position", "Programmer", "Engineer"].
[0130] Step 5: Rewrite the question tone based on the role-playing of the large model (corresponding to the "preset derived questions"). For each question obtained in Step 3, the large model is required to rewrite it according to the tone of the corresponding multiple subject objects, resulting in a new set of rewritten questions. This creates a multi-level mapping relationship: role-rewritten question --> original question --> corresponding clause / clause group. For example, if the original question is "What is the reimbursement process?", for the role label "new employee", the large model is used to rewrite one of the questions: "How do I get reimbursed when I first arrived?"
[0131] Step 6: Vectorization. Use the vectorization model to vectorize all the problems from Steps 3 and 5.
[0132] 2. Search phase.
[0133] Step 7: Vectorize User Questions. Use the vectorization model to vectorize the user questions.
[0134] Step 8: Original question retrieval. Retrieve the original question vector that is closest to the user's question vector. (1) If there is a question vector with a similarity greater than 0.95, obtain the corresponding clause or clause group and proceed to step 11; (2) If not, proceed to step 9.
[0135] Step 9: Initial screening of rewritten questions. Retrieve the rewritten question vector that is closest to the user's vector. (1) If there is a rewritten question vector with a similarity greater than 0.9, the corresponding clause or clause group is obtained directly, and the process proceeds to step 11; (2) If there is a rewritten question vector with a similarity greater than 0.5, the process proceeds to step 10. (3) If the similarity between all rewritten question vectors and the user's question vector is less than 0.5, the system returns "Unable to answer".
[0136] Step 10: Rewrite and refine the question. For rewritten question vectors with a similarity greater than 0.5 retrieved in the initial screening stage, find the corresponding original questions. Several refinement methods are provided: Method 1: Show the user the original question text and ask them to choose which one they want to ask, or neither. Method 2: Input the user's question and the original question set into a large language model, and determine which original question best answers the current user's question. Method 3: Use a text re-ranking model to directly calculate the similarity between the user's question text and the original question text, obtaining the original question texts that exceed a threshold. Through these three implementation methods, the most similar original question texts are obtained, thus yielding the corresponding clauses or clause sets.
[0137] Step 11: Answer Generation. Input the retrieved terms or terms groups along with the user's question into the large language model to obtain the final answer.
[0138] Feedback Mechanism Step 12: Oral Expression Feedback. The system collects user questions and uses these questions as templates, requiring large models to role-play and imitate the tone. This generates more rewritten questions for each original question in the system and adds them to the system.
[0139] Based on the above method embodiments, this application also provides a clause retrieval device. Figure 3 The diagram illustrates a structural block diagram of a clause retrieval device according to an embodiment of this application.
[0140] like Figure 3As shown, the clause retrieval device 300 of this embodiment includes: a question unit 310, a correspondence unit 320, and a retrieval unit 330.
[0141] Problem unit 310 is used to determine the target problem. In one embodiment, problem unit 310 can be used to perform the operation S210 described above, which will not be repeated here.
[0142] The correspondence unit 320 is used to determine the correspondence between a set of clauses and a preset question; wherein any set of clauses contains clause information for answering the corresponding preset question. In one embodiment, the correspondence unit 320 can be used to perform the operation S220 described above, which will not be repeated here.
[0143] The retrieval unit 330 is used to determine preset questions whose similarity to the target question is greater than a preset question similarity threshold, and to determine the set of clauses corresponding to the determined preset questions as the answer basis information for answering the target question. In one embodiment, the retrieval unit 330 can be used to perform the operation S230 described above, which will not be repeated here.
[0144] Optionally, the correspondence unit 320 is used to: determine the correspondence between the clause set and the preset original question, and the correspondence between the clause set and the preset derived question; the preset original question is determined based on the corresponding clause set; the preset derived question is determined based on the preset question corresponding to the corresponding clause set.
[0145] Optionally, each set of clauses has a corresponding scope of application; the pre-defined derivative questions are determined based on the pre-defined questions corresponding to the corresponding set of clauses and the scope of application of the corresponding set of clauses.
[0146] Optionally, the preset questions include: questions that can be answered by the corresponding set of terms, determined based on the large language model.
[0147] Optionally, the method for determining the preset question includes at least one of the following: (1) For any set of clauses, based on the large language model, generate a question that can be answered by the set of clauses, and determine the generated question as the corresponding preset question; (2) For any set of clauses, based on the large language model, generate a question that can be answered by the set of clauses and conforms to the scope of application of the corresponding clauses, and determine the generated question as the corresponding preset question; (3) For any set of clauses, based on the large language model, generate a question that can be answered by the set of clauses and conforms to the scope of application of the corresponding clauses, and determine the generated question as the corresponding preset question; (4) For any set of clauses, based on the large language model, rewrite the preset question corresponding to the set of clauses, and determine the rewritten question as the corresponding preset question.
[0148] Optionally, the retrieval unit 330 is configured to perform at least one of the following: (1) determine a preset original question whose question similarity to the target question is greater than a first preset question similarity threshold; (2) if it is determined that there is no preset original question whose question similarity to the target question is greater than the first preset question similarity threshold, determine a preset derived question whose question similarity to the target question is greater than a second preset question similarity threshold; (3) if it is determined that there is no preset original question whose question similarity to the target question is greater than the first preset question similarity threshold, and there is no preset derived question whose question similarity to the target question is greater than the second preset question similarity threshold, determine a preset derived question whose question similarity to the target question is greater than a third preset question similarity threshold, and for the clause set corresponding to the determined preset derived question, determine a preset original question whose question similarity to the target question satisfies a preset similarity condition from the preset original questions corresponding to the clause set; the second preset question similarity threshold is greater than the third preset question similarity threshold.
[0149] Optionally, the retrieval unit 330 is also used to: determine the answer to the target question based on the large language model, according to the target question and the determined answer basis information.
[0150] Optionally, any set of terms may contain: a single term or multiple terms that are related.
[0151] Optionally, any set of terms has a scope of application; the scope of application of any set of terms includes: the intersection of the scopes of application of the terms corresponding to the terms in any set of terms, and / or the union of the scopes of application of the terms corresponding to the terms in any set of terms.
[0152] Optionally, the scope of application of these terms includes at least one of the following: the objects to which the terms apply, the time period in which the terms apply, the business in which the terms apply, and the subject to which the terms apply.
[0153] Optionally, any set of clauses corresponds to a scope of application of the clauses; the retrieval unit 330 is used to: in the set of clauses corresponding to the predetermined question, determine the set of clauses that the target question conforms to the scope of application of the corresponding clauses as the answer basis information for answering the target question.
[0154] Optionally, any set of clauses corresponds to a scope of application of the clauses; the retrieval unit 330 is used to: determine the set of clauses that conform to the scope of application of the corresponding clauses for the target question; and among the preset questions corresponding to the determined set of clauses, determine the preset questions whose question similarity to the target question is greater than the preset question similarity threshold.
[0155] According to embodiments of this application, any plurality of modules in the problem unit 310, the correspondence unit 320, and the retrieval unit 330 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules can be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the problem unit 310, the correspondence unit 320, and the retrieval unit 330 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the problem unit 310, the correspondence unit 320, and the retrieval unit 330 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0156] Figure 4 A block diagram schematically illustrates an electronic device suitable for implementing a clause retrieval method according to an embodiment of this application. For example... Figure 4As shown, an electronic device 900 according to an embodiment of this application includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0157] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 902 and / or RAM 903. It should be noted that the programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.
[0158] According to embodiments of this application, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the input / output (I / O) interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.
[0159] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0160] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903 described above.
[0161] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement any of the method embodiments provided in the embodiments of this application.
[0162] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0163] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0164] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0165] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0167] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
Claims
1. A method for searching terms, characterized in that, include: Define the target problem; Determine the correspondence between the clause sets and the preset questions; wherein each clause set contains clause information used to answer the corresponding preset questions; Identify preset questions whose similarity to the target question is greater than a preset question similarity threshold, and determine the set of clauses corresponding to the identified preset questions as the answer basis information for answering the target question.
2. The method according to claim 1, characterized in that, The correspondence between the determined set of terms and the preset questions includes: Determine the correspondence between the set of clauses and the pre-set original questions, as well as the correspondence between the set of clauses and the pre-set derived questions; The preset original question is determined based on the corresponding set of clauses; The preset derivative questions are determined based on the preset questions corresponding to the corresponding set of clauses.
3. The method according to claim 2, characterized in that, Each set of clauses has a corresponding scope of application; the pre-defined derivative question is determined based on the pre-defined question corresponding to the corresponding set of clauses and the scope of application of the corresponding set of clauses.
4. The method according to claim 1, characterized in that, The preset questions include: questions that can be answered by the corresponding set of clauses, determined based on a large language model.
5. The method according to claim 1, characterized in that, The method for determining the preset problem includes at least one of the following: For any set of clauses, based on a large language model, generate questions that can be answered for that set of clauses, and define the generated questions as the corresponding preset questions; For any set of clauses, based on the large language model, according to the scope of application of the clauses corresponding to the set of clauses, generate questions that can be answered by the set of clauses and that conform to the scope of application of the corresponding clauses, and determine the generated questions as the corresponding preset questions. For any set of clauses, based on the large language model, according to the preset question corresponding to the set of clauses and the scope of application of the clauses corresponding to the set of clauses, a question that can be answered by the set of clauses and conforms to the scope of application of the corresponding clauses is generated, and the generated question is determined as the corresponding preset question. For any set of clauses, based on the large language model, according to the scope of application of the clauses corresponding to the set of clauses, the preset questions corresponding to the set of clauses are rewritten, and the rewritten questions are determined as the corresponding preset questions.
6. The method according to claim 2, characterized in that, The predetermined questions that have a similarity greater than a predetermined question similarity threshold with the target question include at least one of the following: Identify preset original questions whose question similarity to the target question is greater than a first preset question similarity threshold; If it is determined that there is no preset original question with a question similarity greater than a first preset question similarity threshold to the target question, then a preset derived question with a question similarity greater than a second preset question similarity threshold to the target question is determined. If it is determined that there are no preset original questions with a question similarity greater than a first preset question similarity threshold to the target question, and no preset derived questions with a question similarity greater than a second preset question similarity threshold to the target question, then preset derived questions with a question similarity greater than a third preset question similarity threshold to the target question are determined. For the set of clauses corresponding to the determined preset derived questions, preset original questions with a question similarity that meets preset similarity conditions to the target question are determined from the preset original questions corresponding to the set of clauses. The second preset question similarity threshold is greater than the third preset question similarity threshold.
7. The method according to claim 1, characterized in that, The method further includes: Based on the large language model, the answer to the target question is determined according to the target question and the determined answer basis information.
8. The method according to claim 1, characterized in that, Any set of terms may contain: a single term or multiple terms that are related.
9. The method according to claim 1 or 3, characterized in that, Each set of clauses has a corresponding scope of application; The scope of application of any set of terms includes: the intersection of the scopes of application of the terms corresponding to the terms in any set of terms, and / or the union of the scopes of application of the terms corresponding to the terms in any set of terms.
10. The method according to claim 3 or 9, characterized in that, The scope of application of these terms includes at least one of the following: the objects to which the terms apply, the time period in which the terms apply, the business activities to which the terms apply, and the subject to which the terms apply.
11. The method according to claim 1, characterized in that, Each set of clauses has a corresponding scope of application; The step of determining the set of clauses corresponding to the predetermined question as the basis for answering the target question includes: In the set of clauses corresponding to the predetermined question, the set of clauses that fall within the scope of application of the corresponding clauses are identified as the basis for answering the question.
12. The method according to claim 1, characterized in that, Each set of clauses has a corresponding scope of application; The determination of preset questions whose similarity to the target question is greater than a preset question similarity threshold includes: Determine the set of clauses that fall within the scope of application of the corresponding clauses for the target question; among the preset questions corresponding to the determined set of clauses, determine the preset questions whose similarity to the target question is greater than a preset question similarity threshold.
13. A clause retrieval device, characterized in that, include: Problem unit, used to define the target problem; The correspondence unit is used to determine the correspondence between a set of clauses and a preset question; wherein, any set of clauses contains clause information for answering the corresponding preset question; The retrieval unit is used to determine preset questions whose similarity to the target question is greater than a preset question similarity threshold, and to determine the set of clauses corresponding to the determined preset questions as the answer basis information for answering the target question.
14. An electronic device comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 12.
16. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 12.