Construction method and device of question-answering system, equipment, storage medium and program product

By acquiring the textual feature attributes of user questions, filtering out target text, and generating question-and-answer results, the problem of inaccurate answers in existing legal question-and-answer systems is solved, improving the accuracy and reliability of legal question-and-answer systems.

CN121029973APending Publication Date: 2025-11-28INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511218599.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing legal question-and-answer systems struggle to accurately locate the relevant legal provisions when dealing with complex or ambiguous legal questions, leading to inaccurate answers.

Method used

By acquiring the textual feature attributes of user questions, target text directly related to the questions is filtered out, and the user questions and target texts are input into a general model to generate question-and-answer results.

Benefits of technology

It improves the accuracy and reliability of legal Q&A systems in complex legal scenarios, and reduces information redundancy and regulatory positioning errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029973A_ABST
    Figure CN121029973A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a question answering system construction method and device, equipment, a storage medium and a program product, and relates to the field of artificial intelligence or the field of financial science and technology. Firstly, user questions are obtained. Next, determining text feature attributes corresponding to the user question, and selecting a target text conforming to the user question from the plurality of texts to be processed according to the feature attributes; and finally, inputting the user question and the target text into the first general model to obtain a question and answer result output by the model. According to the method, by introducing the text feature attributes, the problem that the answer result is inaccurate when an existing legal question-answering system adopts vector retrieval and retrieval enhancement generation technologies is effectively solved, so that the accuracy and correlation of legal questions and answers are improved, and more accurate legal answers are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence or the field of financial technology, in particular to a construction method and device of a question and answer system, equipment, a storage medium and a program product. BACKGROUND

[0002] With the development of knowledge question and answer systems, legal question and answer has gradually become an important online service method. Knowledge question and answer systems can help users quickly obtain legal information and provide convenient legal answers for users.

[0003] Existing legal question and answer systems mainly rely on vector retrieval and retrieval enhancement generation technology. That is, the system converts the user's question into a vector form, retrieves relevant legal documents or cases, and generates an answer in combination with a generative model.

[0004] However, this approach may retrieve redundant or weakly related information when dealing with complex or ambiguously expressed legal questions, making it difficult to accurately locate the directly related legal provisions, thereby resulting in inaccurate answers. SUMMARY

[0005] The present application provides a construction method, device, equipment, storage medium and program product of a question and answer system to solve the technical problem of inaccurate answers of existing legal question and answer systems when using vector retrieval and retrieval enhancement generation technology.

[0006] In a first aspect, the present application provides a construction method of a question and answer system applied to a legal knowledge question and answer scenario, comprising:

[0007] obtaining a user question, wherein the user question is proposed based on a plurality of to-be-processed texts;

[0008] determining a text feature attribute corresponding to the user question;

[0009] determining a target text that the user question conforms to from a plurality of to-be-processed texts according to the text feature attribute;

[0010] inputting the user question and the target text into a first general model to obtain a question and answer result output by the first general model.

[0011] In a possible implementation manner, the determination of the text feature attribute corresponding to the user question comprises:

[0012] obtaining a text label database, wherein the text label database comprises a plurality of candidate text categories and at least one candidate keyword corresponding to each candidate text category;

[0013] determine a target text category and a target keyword, which are consistent with the user question, from the plurality of candidate text categories and at least one candidate keyword corresponding to each of the candidate text categories, as the text feature attribute.

[0014] In a possible implementation, the determining the target text category and the target keyword, which are consistent with the user question, from the plurality of candidate text categories and at least one candidate keyword corresponding to each of the candidate text categories, as the text feature attribute, includes:

[0015] inputting the user question and the plurality of candidate text categories into a second general model to obtain the target text category consistent with the user question;

[0016] determining the target keyword according to the user question and at least one candidate keyword corresponding to the target text category.

[0017] In a possible implementation, the text label database further includes a one-to-one mapping relationship between a plurality of candidate keywords and a plurality of to-be-processed texts, and the determining the target keyword according to the user question and at least one candidate keyword corresponding to the target text category includes:

[0018] inputting the user question and at least one candidate keyword corresponding to the target text category into a third general model to obtain the target keyword output by the third general model;

[0019] The determining the target text, which is consistent with the user question, from a plurality of to-be-processed texts according to the text feature attribute includes:

[0020] determining, based on the mapping relationship, a to-be-processed text corresponding to the target keyword as the target text.

[0021] In a possible implementation, the step of constructing the text label database includes:

[0022] obtaining a plurality of to-be-processed texts;

[0023] performing classification and analysis processing on the plurality of to-be-processed texts to obtain a text category of each of the to-be-processed texts;

[0024] performing extraction processing on the plurality of to-be-processed texts to obtain at least one keyword of each of the to-be-processed texts;

[0025] constructing the text label database according to the plurality of to-be-processed texts, the text category of each of the to-be-processed texts, and the at least one keyword.

[0026] In a possible implementation, the extracting processing on the plurality of texts to be processed to obtain at least one keyword of each of the texts to be processed comprises:

[0027] For any one of the texts to be processed, the text to be processed and a preset non-restricted prompt word are input into a fourth general model to obtain the keyword.

[0028] and / or;

[0029] The text to be processed and a preset restricted prompt word are input into a fifth general model to obtain the keyword.

[0030] In a second aspect, the application provides a construction device of a question and answer system, applied to a legal knowledge question and answer scene, comprising:

[0031] An acquisition module is configured to acquire a user question, wherein the user question is proposed based on a plurality of texts to be processed.

[0032] A determination module is configured to determine a text feature attribute corresponding to the user question.

[0033] The determination module is further configured to determine, according to the text feature attribute, a target text to which the user question conforms from the plurality of texts to be processed.

[0034] An input module is configured to input the user question and the target text into a first general model to obtain a question and answer result output by the first general model.

[0035] In a possible implementation, the acquisition module is further configured to acquire a text label database, wherein the text label database comprises a plurality of candidate text categories and at least one candidate keyword corresponding to each of the candidate text categories.

[0036] The determination module is further configured to determine, from the plurality of candidate text categories and at least one candidate keyword corresponding to each of the candidate text categories, a target text category and a target keyword to which the user question conforms as the text feature attribute.

[0037] In a possible implementation, the input module is further configured to input the user question and the plurality of candidate text categories into a second general model to obtain the target text category to which the user question conforms.

[0038] The determination module is specifically configured to determine the target keyword according to the user question and at least one candidate keyword corresponding to the target text category.

[0039] In a possible implementation, the input module is specifically configured to input the user question and at least one candidate keyword corresponding to the target text category into a third general model, to obtain the target keyword output by the third general model.

[0040] The determination module is specifically configured to determine, based on the mapping relationship, the target keyword corresponding to the target text to be processed.

[0041] In a possible implementation, the obtaining module is further configured to obtain a plurality of texts to be processed.

[0042] The device further includes a processing module.

[0043] The processing module is configured to perform classification and analysis processing on the plurality of texts to be processed, to obtain a text category of each of the texts to be processed.

[0044] The processing module is further configured to perform extraction processing on the plurality of texts to be processed, to obtain at least one keyword of each of the texts to be processed.

[0045] The device further includes a construction module.

[0046] The construction module is configured to construct the text label database according to the plurality of texts to be processed, the text category of each of the texts to be processed, and the at least one keyword.

[0047] In a possible implementation, the input module is further configured to, for any one of the texts to be processed, input the text to be processed and a preset non-limited prompt word into a fourth general model, to obtain the keyword.

[0048] The input module is further configured to input the text to be processed and a preset limited prompt word into a fifth general model, to obtain the keyword.

[0049] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor.

[0050] The memory stores computer execution instructions.

[0051] The processor executes the computer execution instructions stored in the memory, so that the processor executes the first aspect and / or various possible implementation manners of the first aspect.

[0052] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by a processor to implement the first aspect and / or various possible implementation manners of the first aspect.

[0053] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program which, when executed by a processor, implements the first aspect and / or various possible implementation manners of the first aspect.

[0054] The method, device, equipment, storage medium and program product for constructing the question and answer system provided by the present application first acquire a question raised by a user based on a plurality of to-be-processed texts, then determine text feature attributes corresponding to the question of the user, and then filter a target text conforming to the question of the user from the plurality of to-be-processed texts according to the text feature attributes, and finally input the question of the user and the target text into a first general model to obtain a question and answer result output by the model. The method solves the problem of inaccurate answer results of an existing legal question and answer system when a vector retrieval and retrieval enhancement generation technology is used, reduces answer errors caused by information redundancy or law provision deviation, and improves the accuracy and reliability of answers of the legal question and answer system. BRIEF DESCRIPTION OF DRAWINGS

[0055] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the present application.

[0056] Figure 1 Flowchart of the method for constructing the question and answer system provided by the present application Figure 1

[0057] Figure 2 Flowchart of the method for constructing the question and answer system provided by the present application Figure 2

[0058] Figure 3 Flowchart of the method for constructing the question and answer system provided by the present application Figure 3

[0059] Figure 4 Structure diagram of the device for constructing the question and answer system provided by the present application

[0060] Figure 5 Structure diagram of the equipment for constructing the question and answer system provided by the present application

[0061] The above drawings have shown the specific embodiments of the present application, and the following will have a more detailed description. The drawings and the text description are not intended to limit the scope of the concept of the present application by any means, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0062] ​​​The exemplary embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals represent like elements or similar elements, unless otherwise indicated. The following exemplary embodiments described herein represent implementations consistent with the present disclosure. They are presented by way of example only, and are not intended to represent the only implementations consistent with the present disclosure. Rather, they are intended to represent just a few of the many implementations consistent with the present disclosure as detailed in the appended claims.

[0063] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards of relevant countries and regions, necessary security measures are taken, public order and good customs are not violated, and appropriate operation portals are provided for users to choose authorization or refusal.

[0064] And the present application involves big data analysis of user information (including but not limited to personal biological characteristics, identity data, consumption data, asset data, electronic terminal operation data, etc.), and uses artificial intelligence technology for automatic decision-making, and makes technical solutions based on automatic decision-making results that have a significant impact on personal rights and interests, provides appropriate operation portals for users to choose to agree or refuse automatic decision-making results; if the user chooses to refuse, the expert decision-making process is entered.

[0065] It should be noted that the method, device, equipment, storage medium and product for constructing the question and answer system provided by the present application can be used in the field of artificial intelligence or the field of financial technology, and can also be used in any field other than the field of artificial intelligence or the field of financial technology. The application field of the method, device, equipment, storage medium and product for constructing the question and answer system in the present application is not limited.

[0066] With the development of knowledge Q&A system, legal Q&A has gradually become an important online service method. Knowledge Q&A system can help users quickly obtain legal information and provide convenient legal solutions for the public.

[0067] The existing legal Q&A scene mainly relies on vector retrieval and retrieval enhancement generation, that is, the user question is converted into a vector form, the information fragments related to the question are filtered from a large amount of legal knowledge base through vector retrieval, and then these information fragments are input into a generation model to generate an answer content for the user question.

[0068] However, this way may retrieve redundant or weakly relevant information when dealing with complex or ambiguously expressed legal issues, and it is difficult to accurately locate the directly relevant legal provisions, resulting in inaccurate answer results.

[0069] To solve the above problems, the construction method of the question and answer system provided by the present application first acquires a user question, then determines the text feature attribute of the user question, then selects a target text directly related to the question from a plurality of to-be-processed texts, and finally inputs the user question and the target text into a first general model to obtain the question and answer result output by the first general model. This scheme effectively solves the technical problem of inaccurate answer results of existing legal question and answer systems by introducing text feature attributes, and improves the accuracy of the legal question and answer system in complex legal scenarios.

[0070] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments. The embodiments of the present application will be described below with reference to the drawings.

[0071] Figure 1 Flowchart of the virtual resource management method provided by the present application Figure 1 As shown in Figure 1 , the method comprises.

[0072] S101, acquire a user question, the user question is based on a plurality of to-be-processed texts.

[0073] Among them, the purpose of this step of acquiring a user question is to clarify the legal question raised by the user.

[0074] It can be understood that since the user question is the core orientation of the legal question and answer process, only when the legal question of the user is clear can the legal question and answer system carry out subsequent operations in a targeted manner. Therefore, the user question needs to be acquired.

[0075] The way of acquiring the user question in this step may be, for example, manually input by the user through the application software corresponding to the terminal device, or acquired after voice input by the user through the application software corresponding to the terminal device. The present application does not make special limitation on this.

[0076] For example, the user question is "Is it illegal for a company to not pay value-added tax on time?"

[0077] S102, determine the text feature attribute corresponding to the user question.

[0078] Among them, the text feature attribute includes but is not limited to: text category, keyword.

[0079] The purpose of this step is to determine the legal knowledge involved from the user's question.

[0080] It can be understood that, since the user's question is usually presented in natural language, the expression is arbitrary and ambiguous, and direct matching of the text is prone to deviation. After determining the text feature attributes, the legal question and answer system can more accurately grasp the core of the question, thereby improving the accuracy of subsequent screening of target text and laying a foundation for generating accurate question and answer results.

[0081] For example, the user's question: "Is it illegal for a company to not pay value-added tax on time?", based on the above information, it can be determined that the text feature attributes corresponding to the question include: value-added tax, payment period and tax law category.

[0082] S103, according to the text feature attributes, determine the target text that meets the user's question from a plurality of to-be-processed texts.

[0083] The purpose of this step is to determine the target text that best meets the question from the pre-set plurality of to-be-processed texts according to the text feature attributes of the user's question.

[0084] It can be understood that, since the number of to-be-processed texts is large and covers a wide range, if not targeted screening, directly let the legal question and answer system find the answer in all to-be-processed texts, it is easy to appear information redundancy, low correlation, and it is difficult to accurately locate the effective legal basis. Therefore, by relying on the text feature attributes, the target text can be more accurately determined from a plurality of to-be-processed texts.

[0085] For example, suppose the to-be-processed texts include:

[0086] (1) According to Article 12 of the Tax Law, value-added tax payers shall pay tax within the prescribed period and may not delay;

[0087] (2) According to Article 8 of the Value-Added Tax Law, value-added tax payers shall report and pay value-added tax to the tax authorities within the prescribed period;

[0088] (3) According to Article 23 of the Labor Law, the labor contract shall clearly specify the work content, work location, work time, remuneration and other matters.

[0089] Given that the text feature attributes of the user's question include: value-added tax, payment period and tax law category. Then based on the above information, it can be determined that the target text includes: (1) According to Article 12 of the Tax Law, value-added tax payers shall pay tax within the prescribed period and may not delay; (2) According to Article 8 of the Value-Added Tax Law, value-added tax payers shall report and pay value-added tax to the tax authorities within the prescribed period.

[0090] S104, input the user question and the target text into the first general model to obtain a question and answer result output by the first general model.

[0091] The first general model may be a natural language processing model, for example. The input of the first general model is the user question and the target text. The output of the first general model is the question and answer result of the user question.

[0092] The purpose of this step is to generate a question and answer result that meets the user's needs and complies with legal basis.

[0093] As can be understood, after the user question is determined and matched to the corresponding target text, the user question and the target text are input into the first general model at the same time, the user question and the content of the text to be processed can be associated and interpreted by means of the analysis and generation capability of the first general model, and a question and answer result that can accurately answer the user's question and comply with legal basis can be output.

[0094] For example, the user question is "Is it illegal for a company to not pay value-added tax on time?"

[0095] The target text includes: (1) According to Article 12 of the Tax Law, value-added tax payers shall pay tax on time and may not delay; (2) According to Article 8 of the Value-Added Tax Law, value-added tax payers shall report and pay value-added tax to the tax authorities within the prescribed period.

[0096] Therefore, based on the above information, it can be determined that the question and answer result is that the behavior of a company not paying value-added tax on time is illegal. According to Article 12 of the Tax Law, value-added tax payers shall pay tax on time and may not delay; at the same time, according to Article 8 of the Value-Added Tax Law, value-added tax payers shall report and pay value-added tax to the tax authorities within the prescribed period.

[0097] The construction method of the question and answer system provided in this embodiment first acquires a user question, then determines the text feature attributes corresponding to the user question, then selects a target text that meets the user question from a plurality of texts to be processed according to the feature attributes, and finally inputs the user question and the target text into a first general model to obtain a question and answer result output by the model. This method effectively solves the problem of inaccurate answer results when using vector retrieval and retrieval enhancement generation technology in existing legal question and answer systems, thereby improving the accuracy and relevance of legal question and answer, and achieving more accurate legal answers.

[0098] Figure 2 Flowchart of the construction method of the question and answer system provided in this application Figure 2 As shown in the flowchart, the embodiment of the present application first acquires a user question, then determines the text feature attributes corresponding to the user question, then selects a target text that meets the user question from a plurality of texts to be processed according to the feature attributes, and finally inputs the user question and the target text into a first general model to obtain a question and answer result output by the model. Figure 1 Figure 3 ​Based on the embodiment, the construction method of the question and answer system is described in detail, which comprises:

[0099] S201, obtaining a user question, the user question is proposed based on a plurality of preset to-be-processed texts.

[0100] Among them, the explanation of step S201 is similar to the explanation of step S101 described above, and will not be repeated here.

[0101] S202, obtaining a text label database, wherein the text label database comprises a plurality of candidate text categories and at least one candidate keyword corresponding to each candidate text category.

[0102] Among them, the candidate text category includes but is not limited to tax law, labor law and contract law. Among them, the tax law is further divided into value-added tax, enterprise income tax and personal income tax.

[0103] The candidate keywords corresponding to the value-added tax include but are not limited to value-added tax, taxpayer, tax obligation, tax management, tax payment period.

[0104] The candidate keywords corresponding to the enterprise income tax include but are not limited to enterprise income tax, taxable income, tax-deductible, tax year, tax rate, tax-free income.

[0105] The candidate keywords corresponding to the personal income tax include but are not limited to personal income tax, comprehensive income, special additional deduction, tax rate, taxable amount.

[0106] The candidate keywords corresponding to the labor law include but are not limited to labor contract, work content, salary, working hours system, social insurance, labor arbitration, dismissal compensation.

[0107] The candidate keywords corresponding to the contract law include but are not limited to liquidated damages, compensation measures, contract amount, contract terms, breach of contract, legal liability, offer, commitment, contract termination.

[0108] The purpose of this step is to obtain a text label database containing a plurality of candidate text categories and corresponding candidate keywords, to provide standardized reference for subsequent accurate extraction of text feature attributes from user questions.

[0109] It can be understood that, since the expression of user question is often diverse and arbitrary, and the legal terminology is professional and standardized. Therefore, by obtaining the text label database, the legal question and answer system can more accurately locate the text category and keyword involved in the user question according to the standardized text category and keyword, avoid the deviation of subsequent matching caused by the fuzzy identification of question characteristics, and lay the foundation for improving the accuracy of the whole question and answer process.

[0110] S203. Determine, from the multiple candidate text categories and the at least one candidate keyword corresponding to each candidate text category, a target text category and a target keyword that the user question conforms to, as a text feature attribute.

[0111] The target text category is used to represent the text category to which the user question belongs. The target keyword is used to represent the keyword involved in the user question.

[0112] It can be understood that, according to the obtained text label database, the user question is compared and analyzed with the candidate text categories and keywords in the text label database, so as to determine the text category to which the user question belongs, and determine the keyword involved in the user question.

[0113] For example, the user question is whether a company failing to pay value-added tax on time is illegal. Then based on the above information, it can be determined that the target text category is value-added tax, and the target keyword includes value-added tax, taxpayer, tax obligation, and tax payment period.

[0114] Optionally, the application provides a possible implementation manner, which includes:

[0115] Firstly, input the user question and the multiple candidate text categories into a second general model to obtain a target text category that the user question conforms to.

[0116] The second general model may be, for example, a natural language processing model. The input of the second general model is the user question and the multiple candidate text categories. The output of the second general model is the target text category.

[0117] The purpose of this step is to identify the target text category that conforms to the user question.

[0118] It can be understood that, since the expression of the user question may be relatively broad, and the second general model has the ability to understand semantics, it can capture the core demand of the user question. Therefore, by inputting the user question and the multiple candidate text categories into the second general model, the second general model can use its semantic understanding and correlation analysis capabilities to compare the matching degree between the user question and each candidate text category, and then accurately determine the text category to which the user question belongs.

[0119] For example, the user question is whether a company failing to pay value-added tax on time is illegal, and the multiple candidate text categories include value-added tax, corporate income tax, personal income tax, labor law, and contract law. Then based on the above information, it can be determined that the target text category that the user question conforms to is “value-added tax.”

[0120] Secondly, determine the target keyword according to the user question and the at least one candidate keyword corresponding to the target text category.

[0121] The step aims to determine the target keyword that can embody the user question.

[0122] It can be understood that, since the target text category only delimits the text category to which the question belongs, and the candidate keyword is the core term in the category, the candidate keyword can embody the characteristics of the text category. Therefore, by comparing the user question with a large number of candidate keywords, the candidate keyword that can embody the user question can be determined from the candidate keywords, and the candidate keyword is taken as the target keyword.

[0123] For example, the user question is whether a company that does not pay value-added tax on time is illegal, and at least one candidate keyword corresponding to the value-added tax includes value-added tax, taxpayer, tax obligation, tax management, and tax payment period. Then, based on the above information, it can be determined that the target keyword includes value-added tax, taxpayer, tax obligation, and tax payment period.

[0124] Optionally, the present application provides a possible implementation manner for determining the target keyword based on the user question and at least one candidate keyword corresponding to the target text category, which includes: inputting the user question and at least one candidate keyword corresponding to the target text category into a third general model to obtain the target keyword output by the third general model.

[0125] The third general model may be, for example, a natural language processing model. The input of the third general model is the user question and at least one candidate keyword corresponding to the target text category. The output of the third general model is the target keyword.

[0126] It can be understood that, by using the third general model, the semantic understanding ability of the third general model can be used to analyze and compare the user question and the multiple candidate keywords under the target text category, so as to obtain the candidate keyword that can embody the user question output by the third general model.

[0127] S204, according to the text feature attribute, determining the target text that the user question conforms to from the multiple to-be-processed texts.

[0128] Optionally, based on the fact that the text label database includes the one-to-one mapping relationship between the multiple candidate keywords and the multiple to-be-processed texts, the present application provides a possible implementation manner, which includes: based on the mapping relationship, taking the to-be-processed text corresponding to the target keyword as the target text.

[0129] The step aims to determine the to-be-processed text corresponding to the target keyword by using the one-to-one mapping relationship between the candidate keywords and the to-be-processed texts in the text label database, and determine the to-be-processed text as the target text.

[0130] It can be understood that the mapping relationship can not only ensure the association between the keyword and the to-be-processed text is clear and fixed, but also can avoid blind screening in a large number of to-be-processed texts. Therefore, by directly calling the mapping relationship, the legal question and answer system can quickly lock the to-be-processed text related to the core of the user question, and ensure the high matching between the target text and the keyword.

[0131] S205, inputting the user question and the target text into the first general model to obtain the question and answer result output by the first general model.

[0132] The explanation of step S205 is similar to that of step S104, and will not be repeated here.

[0133] The construction method of the question and answer system provided in the embodiment first obtains a question raised by a user based on a plurality of to-be-processed texts, then obtains a text label database containing a plurality of candidate text categories and at least one candidate keyword corresponding to each category, determines the target text category and the target keyword corresponding to the user question from the database, and takes them as text feature attributes, and then screens the corresponding target text from the plurality of to-be-processed texts according to the text feature attributes, and finally inputs the user question and the target text into the first general model to obtain the question and answer result output by the model. The method extracts the target text category and the keyword corresponding to the user question as feature attributes by introducing the text label database, and then efficiently matches the target text, and generates an answer in combination with the general model, which effectively improves the understanding degree of the legal question and answer system to the user question and the accuracy of the to-be-processed text matching, thereby improving the reliability of the legal question and answer.

[0134] Figure 3 Flowchart of the construction method of the question and answer system provided in the present application Figure 4 As shown in the flowchart, the embodiment is based on the above-mentioned embodiment, and the construction process of the text label database is described in detail. The method comprises the following steps: Figure 4 S301, obtaining a plurality of to-be-processed texts.

[0135] The purpose of this step is to provide basic materials for constructing the text label database.

[0136] It can be understood that a single or small amount of to-be-processed texts cannot cover the diversified legal questions that the user may raise. Therefore, only by collecting a sufficient amount of to-be-processed texts can the text label database constructed subsequently be ensured to be comprehensive and representative, thereby providing sufficient resource support for matching the user question to a suitable text, and avoiding the matching limitations caused by insufficient to-be-processed texts.

[0137]

[0138] ​S302, performing classification and analysis processing on the plurality of to-be-processed texts to obtain a text category of each to-be-processed text.

[0139] The purpose of this step is to determine the text category to which each to-be-processed text belongs.

[0140] It can be understood that, due to the large number of original to-be-processed texts and the wide range of fields covered, if they are not classified, the text label database may be disorganized due to too many texts when the text label database is constructed.

[0141] Therefore, by classifying and analyzing the to-be-processed texts to determine the text category of each to-be-processed text, the organization structure of the to-be-processed texts can be made more clear, thereby improving the usability of the text label database.

[0142] S303, performing extraction processing on the plurality of to-be-processed texts to obtain at least one keyword of each to-be-processed text.

[0143] The purpose of this step is to extract keywords that can reflect the core content from each to-be-processed text.

[0144] It can be understood that, due to the complex content and numerous clauses of the original to-be-processed texts, directly using the full text for matching is not only inefficient but also difficult to highlight the core content. By extracting keywords, the key information of the to-be-processed texts can be determined, and a basis for matching user questions with to-be-processed texts can be provided.

[0145] Optionally, for any one to-be-processed text, the following possible implementation modes are provided:

[0146] First, for any one to-be-processed text, the to-be-processed text and a preset non-limited prompt word are input into a fourth general model to obtain a keyword.

[0147] The non-limited prompt word can be understood as not presetting a specific extraction range or direction when extracting the keyword, but only starting from the core content of the to-be-processed text itself to extract the prompt sentence of the keyword. For example, "Please analyze the following to-be-processed clauses and extract one or more keywords therefrom."

[0148] The fourth general model may be, for example, a natural language processing model. The input of the fourth general model is the to-be-processed text and the preset non-limited prompt word, and the output of the fourth general model is the keyword.

[0149] It can be understood that, since the preset non-limiting word does not set a specific extraction range or direction, the fourth general model can fully exert the autonomous analysis capability, so that it mines the keywords from the to-be-processed text itself. In this way, not only can the extracted keywords fully cover the core points of the to-be-processed text, but also can flexibly adapt to to-be-processed texts of different structures.

[0150] Secondly, the to-be-processed text and the preset limited prompt word are input into the fifth general model to obtain the keywords.

[0151] Among them, the limited prompt word can be understood as a prompt sentence that specifies the extraction range or direction of the keywords in advance when extracting the keywords, and screens the keywords from the to-be-processed text. For example, "please analyze the following to-be-processed provisions, and extract keywords according to legal responsibility, legal obligation, and compensation.

[0152] The fifth general model may be, for example, a natural language processing model. The input of the fifth general model is the to-be-processed text and the preset limited prompt word. The output of the fifth general model is the keywords.

[0153] It can be understood that, since the limited prompt word can provide the fifth general model with clear extraction standards, it can avoid the extracted keywords being too broad, so that it extracts keywords that meet a specific range or direction from the to-be-processed text. In this way, the extracted keywords can be focused on the preset range or direction, meeting the specific information needs.

[0154] S304, constructing a text label database according to the plurality of to-be-processed texts, the text category of each to-be-processed text, and the at least one keyword.

[0155] Among them, the data format of the text label database is {text category-keyword-to-be-processed text}. For example, {value-added tax-taxpayer-according to article 23 of the provisional regulations on value-added tax and article 62 of the tax collection and management law, the tax authority can order correction and impose a fine on taxpayers who do not report value-added tax on time}.

[0156] The purpose of this step is to integrate and associate the plurality of to-be-processed texts that have been obtained, the text category corresponding to each text, and the extracted keywords, to construct a structured text label database.

[0157] It can be understood that due to the lack of structured association of single to-be-processed text, it is difficult to be efficiently utilized by the legal question and answer system. By establishing the corresponding relationship of "text category-keyword-to-be-processed text", the category of each to-be-processed text can be determined, and the core content can be accurately positioned through the keyword, so that the legal question and answer system can quickly match the related to-be-processed text when processing the user's question subsequently, avoid the inefficient retrieval or matching deviation caused by scattered information, and provide core support for improving the accuracy and efficiency of legal question and answer.

[0158] The construction method of the question and answer system provided by the embodiment first acquires a plurality of to-be-processed texts, then classifies and analyzes the to-be-processed texts to determine the text category of each text, extracts at least one keyword corresponding to each to-be-processed text, and finally constructs a text label database according to the plurality of to-be-processed texts, the respective text categories and the extracted keywords.

[0159] The method constructs the text label database by classifying and analyzing the to-be-processed texts and extracting the keywords, which can clearly present the association between the to-be-processed texts, the text categories and the keywords, and help to provide a structured and standardized basis for subsequent quick and accurate matching of the corresponding target text from the user's question, thereby effectively improving the accuracy of the legal question and answer system in processing information.

[0160] Figure 5 The structural diagram of the construction device of the question and answer system provided by the embodiment is shown in Figure 5 The construction device 400 of the question and answer system provided by the embodiment includes:

[0161] The acquisition module 401 is configured to acquire a user question, wherein the user question is proposed based on a plurality of to-be-processed texts.

[0162] The determination module 402 is configured to determine a text feature attribute corresponding to the user question.

[0163] The determination module 402 is further configured to determine a target text that meets the user question from a plurality of to-be-processed texts according to the text feature attribute.

[0164] The input module 403 is configured to input the user question and the target text into a first general model to obtain a question and answer result output by the first general model.

[0165] In a possible implementation manner, the acquisition module 401 is further configured to acquire a text label database, wherein the text label database includes a plurality of candidate text categories and at least one candidate keyword corresponding to each candidate text category.

[0166] The determining module 402 is further configured to determine, from the plurality of candidate text categories and the at least one candidate keyword corresponding to each of the candidate text categories, a target text category and a target keyword to which the user question conforms, as the text feature attribute.

[0167] In a possible implementation, the input module 403 is further configured to input the user question and the plurality of candidate text categories into a second general model to obtain the target text category to which the user question conforms.

[0168] The determining module 402 is specifically configured to determine the target keyword according to the user question and the at least one candidate keyword corresponding to the target text category.

[0169] In a possible implementation, the input module 403 is specifically configured to input the user question and the at least one candidate keyword corresponding to the target text category into a third general model to obtain the target keyword output by the third general model.

[0170] The determining module 402 is specifically configured to determine, based on the mapping relationship, a to-be-processed text corresponding to the target keyword as the target text.

[0171] In a possible implementation, the obtaining module 401 is further configured to obtain a plurality of to-be-processed texts.

[0172] The apparatus further includes a processing module 404.

[0173] The processing module 404 is configured to perform classification and analysis processing on the plurality of to-be-processed texts to obtain a text category of each of the to-be-processed texts.

[0174] The processing module 404 is further configured to perform extraction processing on the plurality of to-be-processed texts to obtain at least one keyword of each of the to-be-processed texts.

[0175] The apparatus further includes a constructing module 405.

[0176] The constructing module is configured to construct the text label database according to the plurality of to-be-processed texts, the text category, and the at least one keyword of each of the to-be-processed texts.

[0177] In a possible implementation, the input module 403 is further configured to, for any one of the to-be-processed texts, input the to-be-processed text and a preset non-limited prompt word into a fourth general model to obtain the keyword.

[0178] The input module 403 is further configured to input the to-be-processed text and a preset limited prompt word into a fifth general model to obtain the keyword.

[0179] The question and answer system provided by the embodiment can execute the method provided by the method embodiment, and has similar implementation principles and technical effects. Details are not described here.

[0180] ​ A structural diagram of the question and answer system provided by the embodiment is shown in FIG. 5. As shown in FIG. 5, the electronic device 500 provided by the embodiment includes at least one processor 501 and a memory 502. Optionally, the device 500 further includes a communication component 503. The processor 501, the memory 502, and the communication component 503 are connected through a bus 504. ​

[0181] In the specific implementation process, the at least one processor 501 executes the computer execution instructions stored in the memory 502, so that the at least one processor 501 executes the method described above.

[0182] The specific implementation process of the processor 501 can refer to the method embodiment described above, and has similar implementation principles and technical effects. Details are not described here.

[0183] In the above embodiment, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly embodied as execution completed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0184] The memory can include a random access memory (RAM), and can also include a non-volatile memory (NVM), for example, at least one disk memory.

[0185] ​The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.

[0186] The present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described above.

[0187] The present application also provides a computer readable storage medium, which stores computer execution instructions, and when a processor executes the computer execution instructions, the method described above is implemented.

[0188] The readable storage medium described above can be realized by any type of volatile or non-volatile storage device or their combination, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special purpose computer.

[0189] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium, and can write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.

[0190] The division of units is only a logical functional division, and in actual implementation, there can be another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0191] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units may be selected according to actual needs to achieve the purpose of the embodiment.

[0192] In addition, each functional unit in each embodiment of the application can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.

[0193] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the application essentially or the part of the prior art that contributes to the technical solutions or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of each embodiment of the application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0194] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware. The aforementioned program can be stored in a computer readable storage medium. The program executes the steps of the above-mentioned method embodiments when executed; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, and various media that can store program codes.

[0195] It should be noted that for the above-mentioned method embodiments, in order to simply describe, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the application is not limited by the order of the described actions, because according to the application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily necessary for the application.

[0196] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0197] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0198] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0199] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.

[0200] If the integrated units / modules are implemented in the form of software program modules and sold or used as independent products, they can be stored in a computer readable memory. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a number of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the method of the present application. The aforementioned memory includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0201] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.

[0202] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The application is intended to cover any variations, uses or adaptations of the application following, in general, the principles of the application and including such departures from the present disclosure as come within known or customary practice in the art to which the application pertains or can relate. The specification and examples are to be regarded as exemplary only, and the true scope and spirit of the application are indicated by the following claims.

[0203] It should be understood that the application is not limited to the precise construction that has been described above and shown in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application. The scope of the application is limited only by the appended claims.

Claims

1. A method for constructing a question-answering system, characterized in that, Applied to legal knowledge question-and-answer scenarios, the method includes: Obtain user questions, which are proposed based on a set of pre-defined texts to be processed; Determine the text feature attributes corresponding to the user question; Based on the text feature attributes, the target text that matches the user question is determined from a plurality of texts to be processed; The user's question and the target text are input into the first general model to obtain the question-and-answer results output by the first general model.

2. The method according to claim 1, characterized in that, Determining the text feature attributes corresponding to the user question includes: Obtain a text tag database, wherein the text tag database includes: multiple candidate text categories, and at least one candidate keyword corresponding to each candidate text category; From the plurality of candidate text categories and at least one candidate keyword corresponding to each candidate text category, the target text category and target keyword that the user question matches are determined as the text feature attributes.

3. The method according to claim 2, characterized in that, The step of determining the target text category and target keyword that the user question matches from the plurality of candidate text categories and at least one candidate keyword corresponding to each candidate text category, as the text feature attribute, includes: The user question and the multiple candidate text categories are input into the second general model to obtain the target text category that the user question matches; The target keyword is determined based on the user question and at least one candidate keyword corresponding to the target text category.

4. The method according to claim 3, characterized in that, The text tag database also includes a one-to-one mapping relationship between multiple candidate keywords and multiple texts to be processed. The step of determining the target keyword based on the user question and at least one candidate keyword corresponding to the target text category includes: Input the user question and at least one candidate keyword corresponding to the target text category into the third general model to obtain the target keyword output by the third general model; The step of determining the target text matching the user question from a plurality of texts to be processed based on the text feature attributes includes: Based on the mapping relationship, the text to be processed corresponding to the target keyword is taken as the target text.

5. The method according to claim 4, characterized in that, The steps for constructing the text tag database include: Get multiple texts to be processed; The multiple texts to be processed are classified and parsed to obtain the text category of each text to be processed; Extraction processing is performed on the plurality of texts to be processed to obtain at least one keyword for each text to be processed; The text tag database is constructed based on the plurality of texts to be processed, the text category of each text to be processed, and at least one keyword.

6. The method according to claim 5, characterized in that, The step of extracting and processing the plurality of texts to be processed to obtain at least one keyword for each text to be processed includes: For any text to be processed, the text to be processed and the preset non-limited prompt words are input into the fourth general model to obtain the keywords; and / or; The text to be processed and the preset limiting prompt words are input into the fifth general model to obtain the keywords.

7. A construction device for a question-and-answer system, characterized in that, Applied to legal knowledge Q&A scenarios, including: The acquisition module is used to acquire user questions, which are proposed based on a set of preset texts to be processed. The determination module is used to determine the text feature attributes corresponding to the user question; The determining module is used to determine the target text that matches the user question from a plurality of texts to be processed based on the text feature attributes. The input module is used to input the user question and the target text into the first general model to obtain the question-and-answer results output by the first general model.

8. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.