Legitimacy dialogue answer generation method and device, electronic equipment and storage medium
By conducting supervised fine-tuning training on large language models and constructing training data on illegal issues, the problem of the dialogue model being unable to identify illegal intentions was solved, and illegal issue identification and legal suggestions in complex scenarios were achieved, thereby improving legal security.
Patent Information
- Application Number
- CN202411386336.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-09-30
AI Technical Summary
When faced with complex and hidden illegal issues, existing dialogue models are unable to effectively identify illegal intentions and provide legal warnings or suggestions, increasing legal security risks.
A large language model is used for supervised fine-tuning training, and legal precedents are used to construct portraits of illegal entities and scenarios in which illegal conversations occur. This generates training data on illegal issues, including explicit and implicit illegal issues, and provides identification, warnings, and legal advice on illegal issues.
It has achieved effective identification of illegal issues in complex illegal scenarios and provided detailed legal explanations and suggestions, improving legal security and alignment.
Smart Images

Figure CN119397348B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and in particular to a method, device, electronic device and storage medium for generating legitimate dialogue answers. Background Art
[0002] With the rapid development of generative artificial intelligence, dialogue models based on natural language processing have been widely used in various fields.
[0003] However, when faced with complex and hidden illegal issues, dialogue models often fail to effectively identify illegal intentions, and are unable to effectively warn users or provide legal assistance, increasing potential legal security risks.
[0004] Therefore, finding a method for generating legal dialogue answers that can effectively identify illegal issues and provide effective warnings and legal suggestions for illegal issues has become an urgent problem to be solved. Summary of the Invention
[0005] The present invention provides a method, device, electronic device and storage medium for generating legal dialogue answers, which realize dialogue answers that can effectively identify illegal issues and provide effective warnings and legal suggestions for illegal issues.
[0006] The present invention provides a method for generating legal dialogue answers, which includes: obtaining a question to be answered; inputting the question to be answered into a pre-trained large language model, and obtaining a dialogue answer corresponding to the question to be answered output by the large language model, wherein, in the case that the question to be answered is an illegal question, the dialogue answer includes an identification result that the question to be answered is an illegal question, as well as a warning and legal suggestions corresponding to the question to be answered.
[0007] According to a method for generating dialogue answers about legality provided by the present invention, the large language model is obtained through supervised fine-tuning training in the following manner: obtaining a training data set, wherein the training data set includes multiple groups of training data, the training data consisting of illegal question training data and dialogue answer training data corresponding to the illegal question training data, the illegal question training data and the dialogue answer training data being extracted from legal precedents; training the large language model based on the training data set to obtain a trained large language model.
[0008] According to a method for generating dialogue answers to legality provided by the present invention, the training data is obtained in the following manner: obtaining legal precedents, and profiling the illegal subject in the legal precedent based on the legal precedent to obtain a portrait of the illegal subject; based on the portrait of the illegal subject, constructing an illegal dialogue scene corresponding to the portrait of the illegal subject, and an illegal intention corresponding to the portrait of the illegal subject; constructing illegal problem training data based on the portrait of the illegal subject, the illegal dialogue scene and the illegal intention; calling the legal judgment of the legal precedent to generate dialogue answer training data corresponding to the illegal problem training data; obtaining the training data based on the illegal problem training data and the dialogue answer training data corresponding to the illegal problem training data.
[0009] According to a method for generating a legality dialogue answer provided by the present invention, before constructing the illegal dialogue occurrence scene corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait based on the illegal subject portrait, the method also includes: restoring the illegal event process at different illegal stages based on the legal precedent; constructing the illegal dialogue occurrence scene corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait based on the illegal subject portrait specifically includes: constructing the illegal dialogue occurrence scene corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait based on the illegal event process and the illegal subject portrait.
[0010] According to a method for generating a legality dialogue answer provided by the present invention, based on the illegal event process and the illegal subject portrait, an illegal dialogue scenario corresponding to the illegal subject portrait and an illegal intention corresponding to the illegal subject portrait are constructed, which specifically includes: calling a pre-constructed first large language model; inputting the illegal event process and the illegal subject portrait as first text prompt information into the first large language model, and obtaining the illegal dialogue scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait output by the first large language model.
[0011] According to a method for generating answers to legality dialogues provided by the present invention, the illegal question training data is constructed based on the illegal subject portrait, the illegal dialogue scenario, and the illegal intention, specifically including: calling a pre-constructed second language model; inputting the illegal subject portrait, the illegal dialogue scenario, and the illegal intention as second text prompt information into the second language model, and obtaining the illegal question training data output by the second language model, wherein the illegal question training data includes explicit illegal question training data and implicit illegal question training data, the explicit illegal question training data includes illegal keywords; the implicit illegal question training data does not include illegal keywords.
[0012] The present invention also provides a device for generating legal dialogue answers, which includes: an acquisition module for acquiring questions to be answered; a generation module for inputting the questions to be answered into a pre-trained large language model, and obtaining a dialogue answer corresponding to the questions to be answered output by the large language model, wherein, in the case that the questions to be answered are illegal questions, the dialogue answer includes an identification that the questions to be answered are illegal questions, as well as a warning and legal suggestions corresponding to the questions to be answered.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for generating a legitimacy dialogue answer as described above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for generating a legality dialogue answer.
[0015] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described methods for generating a legitimacy dialogue answer.
[0016] The method, device, electronic device, and storage medium for generating legal dialogue responses provided by the present invention obtain a question to be answered; input the question to be answered into a pre-trained large language model, and obtain a dialogue response corresponding to the question to be answered, output by the large language model. If the question to be answered is illegal, the dialogue response includes an identification result indicating that the question to be answered is illegal, as well as a warning and legal advice corresponding to the question to be answered. This achieves a dialogue response that can effectively identify illegal issues and provide effective warnings and legal advice for illegal issues, thereby effectively achieving legal security and legal alignment. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 It is a flowchart of the method for generating a legality dialogue answer provided by the present invention.
[0019] Figure 2 It is a schematic diagram of the process of constructing training data provided by the present invention.
[0020] Figure 3 The present invention provides a flow chart of constructing an illegal conversation scenario corresponding to the illegal subject portrait and an illegal intention corresponding to the illegal subject portrait based on the illegal subject portrait.
[0021] Figure 4 The present invention provides a flowchart of constructing an illegal conversation scenario corresponding to the illegal subject portrait and an illegal intention corresponding to the illegal subject portrait based on the illegal event process and the illegal subject portrait.
[0022] Figure 5 It is a flowchart provided by the present invention for constructing illegal problem training data based on the illegal subject portrait, the illegal conversation scenario and the illegal intention.
[0023] Figure 6 It is a structural diagram of the legality dialogue answer generation device provided by the present invention.
[0024] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0026] The method for generating legal dialogue answers provided by the present invention can identify and reject illegal requests in complex illegal scenarios, provide detailed legal explanations and suggestions, and thus achieve legal security and legal alignment.
[0027] Figure 1 It is a flowchart of the method for generating a legality dialogue answer provided by the present invention.
[0028] The following will be combined Figure 1 The process of the method for generating a legality dialogue answer provided by the present invention is described.
[0029] In an exemplary embodiment of the present invention, Figure 1 It can be seen that the method for generating a legality dialogue answer may include step 110 and step 120, and each step will be introduced below.
[0030] In step 110, questions to be answered are obtained.
[0031] In step 120, the question to be answered is input into a pre-trained large language model, and a dialogue answer corresponding to the question to be answered is output by the large language model. In the case that the question to be answered is an illegal question, the dialogue answer includes a recognition result that the question to be answered is an illegal question, as well as a warning and legal suggestion corresponding to the question to be answered.
[0032] In one embodiment, the question to be answered may be a question that needs to be identified as to whether it complies with legal rules, wherein the question to be answered may be a legal question or an illegal question. When the question to be answered is a legal question, the dialogue response corresponding to the question to be answered output by the large language model can be used to identify that the question to be answered is a question in a legal scenario. When the question to be answered is an illegal question, the dialogue response corresponding to the question to be answered output by the large language model can be used to identify that the question to be answered is a question in an illegal scenario.
[0033] In another embodiment, the question to be answered can be input into a pre-trained large language model, so that the large language model can output a dialogue answer corresponding to the question to be answered. In order to achieve the technical effect of identifying and rejecting illegal requests in complex illegal scenarios and providing detailed legal explanations and suggestions, where the question to be answered is an illegal question, the dialogue answer includes the result of identifying the question to be answered as an illegal question, as well as a warning and legal suggestions corresponding to the question to be answered. In other words, where the question to be answered is an illegal question, the illegal request can be identified and rejected, the user can be clearly warned of the potential risks of illegal behavior, and detailed legal explanations and suggestions can be provided.
[0034] The method for generating legal dialogue responses provided by the present invention obtains a question to be answered; inputs the question to be answered into a pre-trained large language model, and obtains a dialogue response corresponding to the question to be answered, output by the large language model. If the question to be answered is illegal, the dialogue response includes an identification result indicating that the question to be answered is illegal, as well as a warning and legal advice corresponding to the question to be answered. This method achieves a dialogue response that can effectively identify illegal issues and provide effective warnings and legal advice for illegal issues, thereby effectively achieving legal security and legal alignment.
[0035] In another exemplary embodiment of the present invention, the large language model can be obtained through supervised fine-tuning training in the following manner:
[0036] Obtaining a training data set, wherein the training data set includes multiple sets of training data, the training data consisting of illegal question training data and dialogue answer training data corresponding to the illegal question training data, wherein the illegal question training data and the dialogue answer training data are extracted from legal precedents;
[0037] The large language model is trained based on the training data set to obtain a trained large language model.
[0038] In one embodiment, to ensure that the trained large language model has the ability to identify and reject illegal requests in complex illegal scenarios and provide detailed legal explanations and recommendations, a training dataset can be obtained and the large language model can be trained based on the training dataset to obtain a trained large language model.
[0039] In another embodiment, the training data set may include multiple sets of training data. The training data consists of illegal question training data and corresponding dialogue response training data. It should be noted that the illegal question training data and dialogue response training data are extracted from legal precedents. Because the illegal question training data and dialogue response training data are extracted from legal precedents, the constructed training data is ensured to be referenceable and reasonable, thereby laying the foundation for the trained large language model to be able to effectively identify illegal issues and provide dialogue responses that provide effective warnings and legal advice on illegal issues.
[0040] Figure 2 It is a schematic diagram of the process of constructing training data provided by the present invention.
[0041] In order to further introduce the method for generating legitimate dialogue answers provided by the present invention, Figure 2 Provide explanation.
[0042] In an exemplary embodiment of the present invention, Figure 2 It can be seen that constructing training data may include steps 210 to 250, and each step will be introduced below.
[0043] In step 210, legal precedents are obtained, and based on the legal precedents, the offending subject in the legal precedents is profiled to obtain a portrait of the offending subject.
[0044] In one embodiment, the large language model GPT-4 can be used to construct a character profile from the facts of the violation. Specifically, a profile of the offender in the legal precedent can be generated based on legal precedents. During application, a profile of the offender can be constructed based on the facts of the violation. This profile can include the offender's background, upbringing, personality, language style, and motivation for the violation. The aforementioned facts of the violation can be identified through legal precedents.
[0045] In step 220, based on the illegal subject portrait, the illegal conversation scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait are constructed.
[0046] In another embodiment, based on the portrait of the illegal subject, the output of the large language model GPT-4 can be used to construct the illegal conversation scenario corresponding to the portrait of the illegal subject, as well as the illegal intention corresponding to the portrait of the illegal subject.
[0047] In step 230, based on the illegal subject portrait, the illegal conversation scenario, and the illegal intention, illegal problem training data is constructed.
[0048] In another embodiment, illegal problem training data can be constructed based on the profile of the illegal subject, the illegal conversation scene, and the illegal intention. The illegal problem training data can be embodied in the form of synthesized specific conversation content.
[0049] In step 240, the legal judgment of the legal case is called to generate dialogue answer training data corresponding to the illegal question training data.
[0050] In step 250, the training data is obtained based on the illegal question training data and the dialogue answer training data corresponding to the illegal question training data.
[0051] In another embodiment, legal decisions from legal precedents can be used to generate dialogue response training data corresponding to the illegal question training data. A pre-trained model can also be used to extract keywords from legal decisions from legal precedents, thereby generating dialogue response training data corresponding to the illegal question training data based on these keywords. Because the dialogue response training data corresponding to the illegal question training data is extracted from legal decisions from legal precedents, safe and legally responsible dialogue response training data can be obtained. Regarding safety, responses must include explicit rejections and warnings. A rejection means that when a user makes an illegal request, the response must directly refuse to provide assistance to prevent any potential illegal behavior. A warning means that when a user makes an illegal request, the response should clearly indicate the potential violation of the law and the penalties that should or will be imposed. This not only protects user safety but also mitigates potential legal risks. This also reflects the fact that dialogue responses include the identification result of the question being answered as illegal. Regarding responsibility, responses must include constructive suggestions. Suggestions should point to legal avenues for assistance and provide legal actions that the user can take. In addition to legal advice, responses should also include psychological counseling, helping users identify their underlying needs and offering possible solutions. This sense of responsibility is reflected not only in compliance with the law but also in comprehensive care for users, ensuring they find appropriate solutions within the legal framework while receiving psychological support. This can also be reflected in the fact that dialogue responses include legal advice corresponding to the question being answered.
[0052] Through the above two considerations, when providing legal consultation and advice, Dafa can not only protect the legal safety of users, but also demonstrate a sense of responsibility to users and help them find reliable solutions to complex legal issues.
[0053] Figure 3 The present invention provides a flow chart of constructing an illegal conversation scenario corresponding to the illegal subject portrait and an illegal intention corresponding to the illegal subject portrait based on the illegal subject portrait.
[0054] The following will be combined Figure 3 The process of constructing the illegal conversation scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait based on the illegal subject portrait is explained.
[0055] In an exemplary embodiment of the present invention, Figure 3 It can be seen that based on the portrait of the illegal subject, constructing the illegal conversation scenario corresponding to the portrait of the illegal subject and the illegal intention corresponding to the portrait of the illegal subject can include steps 310 and 320, and each step will be introduced below.
[0056] In step 310, based on the legal precedent, the illegal event process at different illegal stages is restored.
[0057] In step 320, based on the illegal event process and the illegal subject portrait, the illegal conversation scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait are constructed.
[0058] In one embodiment, a large language model (such as GPT-4), which excels in text editing and rewriting, can be used to recount the illegal incident in three stages. GPT-4 is required to describe the three stages of the illegal incident: preparation, execution, and completion. GPT-4 must reconstruct the scene of the illegal incident in plain language. This involves reconstructing the illegal incident at different stages based on legal precedents. These stages can include preparation, execution, and completion.
[0059] Furthermore, at each stage, the GPT-4 model can be used to construct scenarios (corresponding to the illegal conversation scenario) and the illegal intent of the offender (corresponding to the illegal intent) based on the offender's personal profile (corresponding to the illegal subject portrait) and the illegal process (corresponding to the illegal event process). Based on this, illegal dialogue questions can be generated in complex illegal scenarios, laying the foundation for identifying and rejecting illegal requests, warning of potential risks of illegal behavior, and providing detailed legal explanations and recommendations.
[0060] Figure 4 The present invention provides a flowchart of constructing an illegal conversation scenario corresponding to the illegal subject portrait and an illegal intention corresponding to the illegal subject portrait based on the illegal event process and the illegal subject portrait.
[0061] The following will be combined Figure 4 The process of constructing the illegal conversation scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait based on the illegal event process and the illegal subject portrait is explained.
[0062] In an exemplary embodiment of the present invention, Figure 4 It can be seen that based on the illegal event process and the portrait of the illegal subject, the illegal conversation scene corresponding to the portrait of the illegal subject and the illegal intention corresponding to the portrait of the illegal subject are constructed, which can include steps 410 and 420. Each step will be introduced below.
[0063] In step 410 , a pre-built first language model is called.
[0064] In step 420, the illegal event process and the portrait of the illegal subject are input into the first language model as the first text prompt information, and the first language model outputs the illegal dialogue scene corresponding to the portrait of the illegal subject and the illegal intention corresponding to the portrait of the illegal subject.
[0065] In one embodiment, a pre-built first language model, namely the GPT-4 model, can also be invoked. The illegal event process and the offending subject's profile are then input into the first language model as the first text prompt information. This output from the first language model can then produce the illegal conversation scenario corresponding to the offending subject's profile, as well as the illegal intent corresponding to the offending subject's profile. During the application process, at each stage, the GPT-4 model can be used to construct the possible conversation scenario (corresponding to the illegal conversation scenario) and the illegal intent of the offending subject's conversation (corresponding to the illegal intent) based on the offender's personal profile (corresponding to the illegal subject's profile) and the illegal process (corresponding to the illegal event process). Based on this, the purpose of the illegal conversation can be inferred, laying the foundation for identifying and rejecting illegal requests in complex illegal scenarios, and providing detailed legal interpretations and recommendations.
[0066] Figure 5 It is a flowchart provided by the present invention for constructing illegal problem training data based on the illegal subject portrait, the illegal conversation scenario and the illegal intention.
[0067] The following will be combined Figure 5 The process of constructing illegal problem training data based on the illegal subject portrait, the illegal conversation scenario and the illegal intention is explained.
[0068] In an exemplary embodiment of the present invention, Figure 5 It can be seen that constructing illegal problem training data based on the illegal subject portrait, the illegal conversation scene and the illegal intention can include steps 510 and 520, and each step will be introduced below.
[0069] In step 510 , the pre-built second language model is called.
[0070] In step 520, the portrait of the illegal subject, the scene of the illegal conversation, and the illegal intention are input into the second language model as second text prompt information to obtain illegal problem training data output by the second language model, wherein the illegal problem training data includes explicit illegal problem training data and implicit illegal problem training data, the explicit illegal problem training data includes illegal keywords; the implicit illegal problem training data does not include illegal keywords.
[0071] In one embodiment, during the illegal question synthesis phase (i.e., the process of obtaining illegal question training data), the GPT-4 model can be used to simulate offenders. The GPT-4 model then synthesizes specific conversation content (corresponding to the illegal question training data) based on character profiles (corresponding to the offending subject's portrait), illegal scenarios (corresponding to the scenarios in which illegal conversations occurred), and the intentions of the conversations (corresponding to the illegal intentions). Both explicit and implicit illegal question types can be considered. Explicit illegal questions (corresponding to explicit illegal question training data) contain obvious illegal signals, such as illegal vocabulary or illegal behavior. Implicit illegal questions (corresponding to implicit illegal question training data) do not rely on explicit illegal keywords but instead express the intention to participate in or promote illegal behavior.
[0072] Among them, in the implicit violation of the law, two forms of expression can be considered: at the language level, language techniques use modifiers, such as hints, metaphors, obscurity, and roundabout methods; at the intention level, concealing intentions, such as avoiding or falsifying specific criminal acts and motives, and borrowing legal appearances to cover up illegal intentions.
[0073] During the application process, a stage, scene, explicit and implicit form and its manifestation will be randomly selected in each illegal incident to synthesize specific dialogue questions.
[0074] In another embodiment, a pre-built second language model can also be called, and the portrait of the illegal subject, the process of the illegal event, the scene of the illegal conversation, and the illegal intention can be input into the second language model as the second text prompt information, so as to obtain the illegal problem training data output by the second language model.
[0075] Through the aforementioned embodiments, the comprehensiveness of the constructed illegal issues (corresponding to illegal issue training data) can be ensured; by simulating illegal behaviors and extracting complex and hidden illegal intentions from legal documents, various explicit and implicit illegal issues are covered, thereby improving the recognition ability of the large language model trained based on the training data set in complex illegal scenarios.
[0076] In addition, the generated dialogue answer training data corresponding to the illegal question training data can not only directly reject illegal requests, but also explain the potential legal risks in detail, cite relevant laws and regulations, and provide legal advice and psychological counseling.
[0077] As described above, the legitimacy dialogue answer generation method provided by the present invention significantly improves the security and legal alignment of the legitimacy dialogue answer generation method adopted by the present invention in illegal scenarios through comprehensive illegal question construction, information-rich security answer synthesis and model alignment training.
[0078] The following describes the legitimacy dialogue answer generation device provided by the present invention. The legitimacy dialogue answer generation device described below and the legitimacy dialogue answer generation method described above can be referenced to each other.
[0079] Figure 6 It is a structural diagram of the legality dialogue answer generation device provided by the present invention.
[0080] The following will be combined Figure 6 The legitimacy dialogue answer generation device provided by the present invention is described.
[0081] In an exemplary embodiment of the present invention, Figure 6 It can be seen that the legality dialogue answer generation device can include an acquisition module 610 and a generation module 620. Each module will be introduced below.
[0082] The acquisition module 610 may be configured to acquire questions to be answered;
[0083] The generation module 620 can be configured to input the question to be answered into a pre-trained large language model, and obtain a dialogue answer corresponding to the question to be answered output by the large language model, wherein, in the case that the question to be answered is an illegal question, the dialogue answer includes a recognition result that the question to be answered is an illegal question, as well as a warning and legal suggestion corresponding to the question to be answered.
[0084] In an exemplary embodiment of the present invention, the generation module 620 may obtain the large language model through supervised fine-tuning training in the following manner:
[0085] Obtaining a training data set, wherein the training data set includes multiple sets of training data, the training data consisting of illegal question training data and dialogue answer training data corresponding to the illegal question training data, wherein the illegal question training data and the dialogue answer training data are extracted from legal precedents;
[0086] The large language model is trained based on the training data set to obtain a trained large language model.
[0087] In an exemplary embodiment of the present invention, the generation module 620 may obtain the training data in the following manner:
[0088] Obtaining legal precedents, and profiling the offending subject in the legal precedent based on the legal precedents to obtain a portrait of the offending subject;
[0089] Based on the illegal subject portrait, construct the illegal conversation scene corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait;
[0090] Based on the illegal subject portrait, the illegal conversation scene and the illegal intention, illegal problem training data is constructed;
[0091] Invoking the legal judgment of the legal precedent to generate dialogue answer training data corresponding to the illegal question training data;
[0092] The training data is obtained based on the illegal question training data and the dialogue answer training data corresponding to the illegal question training data.
[0093] In an exemplary embodiment of the present invention, the generation module 620 may construct, based on the illegal subject portrait, an illegal conversation scenario corresponding to the illegal subject portrait and an illegal intention corresponding to the illegal subject portrait in the following manner:
[0094] Based on the legal precedents, the illegal incident process at different stages of the violation is restored;
[0095] Based on the illegal event process and the illegal subject portrait, the illegal dialogue scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait are constructed.
[0096] In an exemplary embodiment of the present invention, the generation module 620 may construct, based on the illegal event process and the illegal subject portrait, an illegal conversation scenario corresponding to the illegal subject portrait and an illegal intention corresponding to the illegal subject portrait in the following manner:
[0097] Call the pre-built first language model;
[0098] The illegal event process and the portrait of the illegal subject are input into the first language model as the first text prompt information, and the first language model outputs the illegal dialogue scene corresponding to the portrait of the illegal subject and the illegal intention corresponding to the portrait of the illegal subject.
[0099] In an exemplary embodiment of the present invention, the generation module 620 may construct the illegal problem training data based on the illegal subject profile, the illegal conversation scenario, and the illegal intention in the following manner:
[0100] Call the pre-built second language model;
[0101] The portrait of the illegal subject, the scene in which the illegal conversation occurred, and the illegal intention are input into the second language model as second text prompt information to obtain illegal problem training data output by the second language model, wherein the illegal problem training data includes explicit illegal problem training data and implicit illegal problem training data, the explicit illegal problem training data includes illegal keywords; the implicit illegal problem training data does not include illegal keywords.
[0102] Figure 7 An example of a physical structure diagram of an electronic device is shown below. Figure 7 As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communications bus 740. The processor 710 may invoke logic instructions in the memory 730 to execute a method for generating a legal dialogue answer, which includes: obtaining a question to be answered; inputting the question to be answered into a pre-trained large language model, and obtaining a dialogue answer corresponding to the question to be answered, output by the large language model. If the question to be answered is illegal, the dialogue answer includes a recognition result that the question to be answered is illegal, as well as a warning and legal advice corresponding to the question to be answered.
[0103] Furthermore, the logic instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0104] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the legality dialogue answer generation method provided by the above methods, which method includes: obtaining a question to be answered; inputting the question to be answered into a pre-trained large language model, and obtaining a dialogue answer corresponding to the question to be answered output by the large language model, wherein, in the case that the question to be answered is an illegal question, the dialogue answer includes an identification result that the question to be answered is an illegal question, and includes warnings and legal suggestions corresponding to the question to be answered.
[0105] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the method for generating legal dialogue answers provided by the above-mentioned methods, the method comprising: obtaining a question to be answered; inputting the question to be answered into a pre-trained large language model, and obtaining a dialogue answer corresponding to the question to be answered output by the large language model, wherein, in the case that the question to be answered is an illegal question, the dialogue answer includes an identification result that the question to be answered is an illegal question, as well as a warning and legal suggestion corresponding to the question to be answered.
[0106] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0107] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for generating a legitimate dialogue answer, characterized in that: The method comprises: Get questions to be answered; The question to be answered is input into a pre-trained large language model, and the large language model outputs a dialogue answer corresponding to the question to be answered. If the question to be answered is illegal, the dialogue answer includes a recognition result that the question to be answered is illegal, as well as a warning and legal advice corresponding to the question to be answered. The training data of the large language model is obtained in the following manner: Obtaining legal precedents, and profiling the offending subject in the legal precedent based on the legal precedents to obtain a portrait of the offending subject; Based on the illegal subject portrait, construct the illegal conversation scene corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait; Based on the illegal subject portrait, the illegal conversation scene and the illegal intention, illegal problem training data is constructed; Invoking the legal judgment of the legal precedent to generate dialogue answer training data corresponding to the illegal question training data; The training data is obtained based on the illegal question training data and the dialogue answer training data corresponding to the illegal question training data, wherein: The illegal problem training data is constructed based on the illegal subject portrait, the illegal conversation scene, and the illegal intention, specifically including: Call the pre-built second language model; The portrait of the illegal subject, the scene in which the illegal conversation occurred, and the illegal intention are input into the second language model as second text prompt information to obtain illegal problem training data output by the second language model, wherein the illegal problem training data includes explicit illegal problem training data and implicit illegal problem training data, the explicit illegal problem training data includes illegal keywords; the implicit illegal problem training data does not include illegal keywords.
2. The method for generating a legitimate dialogue answer according to claim 1, wherein: The large language model is trained by supervised fine-tuning in the following way: Obtaining a training data set, wherein the training data set includes multiple sets of training data, the training data consisting of illegal question training data and dialogue answer training data corresponding to the illegal question training data, wherein the illegal question training data and the dialogue answer training data are extracted from legal precedents; The large language model is trained based on the training data set to obtain a trained large language model.
3. The method for generating a legitimate dialogue answer according to claim 1, wherein: Before constructing, based on the illegal subject portrait, an illegal conversation scenario corresponding to the illegal subject portrait and an illegal intention corresponding to the illegal subject portrait, the method further includes: Based on the legal precedents, the illegal incident process at different stages of the violation is restored; The constructing of the illegal conversation scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait based on the illegal subject portrait specifically includes: Based on the illegal event process and the illegal subject portrait, the illegal dialogue scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait are constructed.
4. The method for generating a legitimate dialogue answer according to claim 3, wherein: The illegal conversation scenario corresponding to the illegal subject portrait and the illegal intention corresponding to the illegal subject portrait are constructed based on the illegal event process and the illegal subject portrait, specifically including: Call the pre-built first language model; The illegal event process and the portrait of the illegal subject are input into the first language model as the first text prompt information, and the first language model outputs the illegal dialogue scene corresponding to the portrait of the illegal subject and the illegal intention corresponding to the portrait of the illegal subject.
5. A device for generating a legitimate dialogue answer, characterized in that: The device is used to implement the method for generating a legal dialogue answer according to any one of claims 1 to 4, and the device includes: The acquisition module is used to obtain questions to be answered; A generation module is used to input the question to be answered into a pre-trained large language model, and obtain a dialogue answer corresponding to the question to be answered output by the large language model, wherein, in the case that the question to be answered is an illegal question, the dialogue answer includes the recognition result that the question to be answered is an illegal question, as well as the warning and legal suggestions corresponding to the question to be answered.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for generating a legality dialogue answer as described in any one of claims 1 to 4 is implemented.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a legality dialogue answer as described in any one of claims 1 to 4 is implemented.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a legality dialogue answer as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Interpretable law automatic decision prediction method and device
CN111325387A
Method and system for solving illusion problem of large legal language model
CN117744802A