A method for evaluating robustness of large language model in legal field based on knowledge injection attack
By employing a knowledge injection-based attack method, interference is inflicted on large language models from three levels: major premise, minor premise, and conclusion. This addresses the limitations of existing technologies, such as the simplistic evaluation methods, lack of Chinese language support, and inadequacy in legal expertise. It enables a comprehensive robustness assessment of large language models in the legal field, thereby promoting their application in this area.
Patent Information
- Application Number
- CN202411636101.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Existing robustness assessment methods for large language models in the legal field suffer from problems such as limited assessment tools, lack of Chinese language support, lack of legal expertise, and inability to meet the assessment needs of the legal field. In particular, it is difficult to effectively assess their anti-interference ability in complex legal reasoning scenarios.
We employ a knowledge injection-based attack method to interfere with the large language model from three levels: major premise, minor premise, and conclusion. This is achieved through retrieval enhancement generation attack, similar crime attack, lexical attack, element attack, narrative attack, prior behavior attack, and expert opinion attack. Combined with a fine-grained dataset annotated by a team of legal experts, we evaluate its anti-interference capability in legal reasoning.
It provides a wider range of evaluation methods, enabling a true assessment of the robustness of large language models in the legal field, helping model developers identify and improve model defects, and promoting the development of large language models in the legal field.
Smart Images

Figure CN119494405B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of robustness evaluation of large language models in the legal field, and in particular to a large language model legal field robustness evaluation method based on knowledge injection attacks. BACKGROUND
[0002] With the rapid development of natural language processing (NLP) and artificial intelligence (AI) technologies, intelligent systems based on large language models (LLM) have been widely applied in various fields. However, existing large language models are mainly applied in general fields and use English datasets as the main training and testing language, which cannot meet the high requirements of language understanding in specific professional fields (such as law). The application of large language models in the legal field requires accurate understanding and reasoning of legal concepts, legal provisions, and case facts, so it is particularly important to evaluate their robustness and anti-interference ability. However, existing evaluation techniques have the following significant shortcomings in professional field applications:
[0003] 1. Simple evaluation method: In existing techniques, attacks are usually performed on input samples or prompts. This method is limited to word-level perturbation of data text, and the evaluation method is single, cannot cover various real error scenarios, and cannot effectively evaluate the anti-interference ability of the model in complex legal reasoning scenarios.
[0004] 2. Lack of Chinese support: Most current general attack techniques are based on English corpus experiments, and the applicable language environment is mainly English. Due to the large differences in grammar structure, vocabulary selection, and expression methods between Chinese and English, existing methods cannot be directly applied to Chinese scenarios, and there is a lack of effective evaluation methods for the robustness of large language models in Chinese language environments.
[0005] 3. Lack of legal professionalism: Existing attack methods are mostly for general field evaluation, and lack of targeted attacks on professional field key words and concepts. In the legal field, reasoning and judgment require knowledge of specific legal terminology, logical structure, and four elements. Therefore, existing techniques cannot comprehensively evaluate the performance of large language models in the legal field.
[0006] 4. Unable to meet the evaluation needs of the legal field: Large language models in the legal field are knowledge-intensive, logical, and require high context understanding, while existing methods are usually limited to the evaluation of general field models, and cannot accurately evaluate the ability of legal field models to handle complex legal reasoning tasks, making it difficult to meet the actual application needs of the legal field. SUMMARY
[0007] The present application aims to at least partially solve one of the technical problems in the related art.
[0008] Therefore, the first objective of this application is to propose a robustness assessment method for large language models in the legal field based on knowledge injection attacks.
[0009] The second objective of this application is to propose a robustness assessment device for large language models in the legal field based on knowledge injection attacks.
[0010] The third objective of this application is to propose an electronic device.
[0011] The fourth objective of this application is to provide a computer-readable storage medium.
[0012] The fifth objective of this application is to provide a computer program product.
[0013] To achieve the above objectives, the first aspect of this application proposes a robustness assessment method for large language models in the legal domain based on knowledge injection attacks, comprising:
[0014] By using search-enhanced generation attacks and similarity-based crime attacks, the major premise judgments made by large language models in the legal field based on input legal provisions are interfered with;
[0015] The accuracy of the large language model's narrative based on the input case facts is interfered with through lexical attacks, element attacks, and narrative attacks.
[0016] By interfering with the final conclusion judgment of the large language model through prior behavioral attacks and expert opinion attacks, its resistance to interference in legal reasoning is assessed.
[0017] Optional, also includes:
[0018] Fine-grained annotations are applied to a legal domain knowledge dataset to support robustness evaluation of the large language model at the three levels of major premise, minor premise, and conclusion. The dataset contains input examples for legal reasoning, including case facts, relevant legal provisions, and question prompts. The annotations include annotations for similar crime names, annotations for conviction logic reasoning, and annotations for domain synonyms.
[0019] The annotation of the dataset was led by a team of legal experts and carried out by legal professionals. Trial annotation was conducted before the formal annotation, and the annotation was carried out in accordance with the designed Chinese annotation guidelines to ensure the consistency and accuracy of the annotation. The annotation guidelines include crime names, logical reasoning of the four elements, use of domain synonyms, and actual annotation examples.
[0020] Optionally, the interference with the major premise judgments made by the large language model in the legal field based on the input legal provisions through retrieval enhancement generation attacks and similarity charge attacks includes:
[0021] Introducing incorrect legal provisions unrelated to the facts of the case during the question prompting stage misleads the large language model when retrieving legal provisions.
[0022] By adding vocabulary interference related to similar crimes to the problem description, the wording used to determine whether a similar crime or other offense has been committed interferes with the major premise judgment of the large language model.
[0023] Optionally, the interference with the accuracy of the large language model's narrative based on the input case facts through lexical attacks, element attacks, and narrative attacks includes:
[0024] By replacing words in the facts of a case with synonyms, the logic of the large language model in judging the facts of the case is disrupted. This includes replacing random words in the facts of the case with common synonyms, replacing words related to the four elements in the facts of the case with common words, and replacing words related to the four elements in the facts of the case with synonymous legal element words.
[0025] By adding summary legal elements or provisions from similar crimes after the facts of the case, the logical judgment of the facts of the case by the large language model is disrupted.
[0026] Contextual or irrelevant narrative statements are added after the facts of the case to disrupt the logical judgment of the large language model on the facts of the case.
[0027] Optionally, the interference with the final conclusion judgment of the large language model through prior behavioral attacks and expert opinion attacks to assess its resistance to interference in legal reasoning includes:
[0028] Insert a description of the perpetrator's previous criminal behavior after the facts of the case to assess whether the large language model can ignore previous behavior that is irrelevant to the current case;
[0029] By incorporating the perspectives of specific individuals regarding the judgment of crimes into the question prompts, we can verify whether the large language model can be influenced by external opinions.
[0030] Optional, also includes:
[0031] Insert legal provisions most relevant to the facts of the case into the question prompts to enhance the robustness of the large language model in dealing with major premise attacks;
[0032] The problem prompt requires the large language model to reason step by step according to the logic of the four elements of criminal law, so as to improve the reasoning ability and anti-interference ability of the large language model;
[0033] Adding two similar case analysis examples to the question prompts enables the large language model to make accurate judgments based on the reasoning logic of similar cases.
[0034] To achieve the above objectives, a second aspect of this application proposes a robustness assessment device for a large language model in the legal domain based on knowledge injection attacks, comprising:
[0035] The major premise knowledge injection attack module is used to interfere with the major premise judgments made by the large language model in the legal field based on the input legal provisions through retrieval enhancement generation attacks and similar crime attacks.
[0036] The minor premise knowledge injection attack module is used to interfere with the accuracy of the large language model's narrative based on the input case facts through lexical attacks, element attacks, and narrative attacks.
[0037] The conclusion knowledge injection attack module is used to interfere with the final conclusion judgment of the large language model through prior behavior attacks and expert opinion attacks, so as to evaluate its anti-interference ability in legal reasoning.
[0038] To achieve the above objectives, a third aspect of this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0039] The memory stores computer-executed instructions;
[0040] The processor executes computer execution instructions stored in the memory to implement the method as described in any one of the first aspects.
[0041] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method as described in any one of the first aspects.
[0042] To achieve the above objectives, a fifth aspect of this application provides a computer program product that, when executed by a processor, implements the method described in any one of the first aspects.
[0043] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects:
[0044] 1. Compared with existing methods for attacking large models, this application proposes a logical attack framework based on Aristotle's syllogism, which attacks the semantic logic and professional logical reasoning of the problem. It more broadly covers the error types that often occur in real life, can truly evaluate the robustness of large models in the domain, and helps large model developers and related practitioners understand the defects in the model itself, thereby better promoting the development of large models.
[0045] 2. Currently, there is no robustness evaluation system specifically for large-scale models in the legal field worldwide. Attack methods and evaluation systems for general domains cannot be directly applied to the legal field. This application can be used to evaluate the robustness of general-purpose large-scale models in solving legal problems from various dimensions and to evaluate the robustness of legal large-scale models, thereby promoting the development of large language models and artificial intelligence in the legal field.
[0046] 3. This application integrates all existing common crimes and, based on real case circumstances, provides fine-grained annotations of common crimes and their four elements from the perspectives of crime names, logical reasoning for conviction, and legal synonyms.
[0047] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0048] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0049] Figure 1 A flowchart illustrating a robustness assessment method for a large language model in the legal domain based on knowledge injection attacks, provided as an embodiment of this application;
[0050] Figure 2 A schematic diagram of the overall framework of a robustness assessment method for a large language model in the legal domain based on knowledge injection attacks, provided for embodiments of this application;
[0051] Figure 3 A schematic diagram of the three-tiered legal knowledge attack framework provided for embodiments of this application;
[0052] Figure 4 A schematic diagram illustrating an attack premise provided in the embodiments of this application;
[0053] Figure 5 This is a schematic diagram of the structure of a robustness assessment device for a large language model in the legal field based on knowledge injection attacks, provided in an embodiment of this application. Detailed Implementation
[0054] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0055] To address the shortcomings of existing technologies, such as simplistic evaluation methods, inapplicability to Chinese, lack of legal expertise, and inability to cover the legal domain, this application provides a robustness evaluation method for a large language model based on knowledge injection attacks in the legal domain. (Refer to...) Figure 1 , Figure 2 and Figure 3 To address the robustness assessment of large-scale legal models, this paper introduces legal domain knowledge, abstract judgments of facts, and logical reasoning chains for attack. It proposes a knowledge attack framework specifically for the legal domain that is applicable to all language environments and sets an evaluation benchmark in the Chinese language environment. This solves the problems of the lack of an evaluation system for robustness assessment of large-scale legal models and the lack of evaluation methods for robustness assessment of Chinese large-scale legal models.
[0056] In current legal and regulatory-based systems, judges first retrieve relevant legal provisions and then infer the crime based on the legal facts recorded during the trial. This is a deductive reasoning method, which can be summarized as Aristotelian syllogism. An example of the traditional syllogistic reasoning method for conviction provided in this application's embodiments is as follows:
[0057] 1. Major premise and legal basis
[0058] Preliminary legal provision 1: Article 134, Paragraph 1 of the Criminal Law [Crime of Major Liability Accident]: Whoever violates the relevant safety management regulations in production or operation, thereby causing a major casualty accident or other serious consequences, shall be sentenced to fixed-term imprisonment of not more than three years or criminal detention; if the circumstances are particularly serious, he / she shall be sentenced to fixed-term imprisonment of not less than three years but not more than seven years.
[0059] The corresponding four elements of criminal law are as follows:
[0060] Subject: People engaged in production
[0061] Subjective aspect: negligence
[0062] Object: Production safety
[0063] Objective factors: Causing work safety accidents
[0064] Premise Article 2: Article 115 of the Criminal Law [Crime of Arson]: Whoever commits arson, flooding, explosion, or releases toxic, radioactive, or infectious disease pathogens or other dangerous substances, or uses other dangerous methods to cause serious injury or death to others or to cause significant damage to public or private property, shall be sentenced to fixed-term imprisonment of not less than ten years, life imprisonment, or death.
[0065] The corresponding four elements of criminal law are as follows:
[0066] Subject: Natural person
[0067] Subjective aspect: negligence
[0068] Object: Public safety
[0069] Objective aspects: the actions that caused the fire.
[0070] 2. Minor premise: facts of the case
[0071] Case Facts: According to safety operating procedures, workshop drivers transporting high-temperature raw materials should ensure that the materials are not extinguished. Workshop workers, unaware that the materials were not completely extinguished, still poured over 40 tons of high-temperature raw materials into the cooling equipment. Some materials mixed into the conveyor belt, causing a dust explosion, ultimately resulting in direct economic losses of 740,422 yuan.
[0072] The corresponding four elements of criminal law are as follows:
[0073] Main body: Staff
[0074] Subjective aspect: negligence
[0075] Object: Production safety
[0076] Objective aspects: Violation of safety production regulations
[0077] 3. Conclusion: Guilty or Not Guilty
[0078] The case meets all four elements of the crime of "major liability accident" and is therefore subject to the crime of "major liability accident".
[0079] The subject, object, and objective aspects of this case do not meet the criteria for the crime of arson, and are therefore unsuitable for the crime of arson.
[0080] Therefore, staff member Yang is guilty of causing a major accident due to negligence.
[0081] This application is based on Aristotle's syllogism and attacks the knowledge of the "major premise," "minor premise," and "conclusion" separately.
[0082] Figure 1 This is a flowchart illustrating a method for evaluating the robustness of a large language model in the legal domain based on knowledge injection attacks, as provided in an embodiment of this application. The flowchart describes a method that attacks the large language model layer by layer based on input case facts, relevant legal provisions, and question prompts, from the three levels of major premise, minor premise, and conclusion. Figure 1 As shown, the method includes the following steps:
[0083] Step 101: By searching for enhanced generation attacks and similar crime attacks, the major premise judgments made by the large language model in the legal field based on the input legal provisions are interfered with.
[0084] This step describes the specific interference methods in major premise attacks, which interfere with the ability of large language models to make major premise judgments in legal text interpretation.
[0085] In this embodiment of the application, RAG (Retrieval Enhancement Generation) attack refers to introducing incorrectly related legal provisions that are unrelated to the facts of the case during the question prompting stage, thereby misleading the large language model when retrieving legal provisions.
[0086] This attack is used to insert legal clauses related to similar offenses, to test whether the model will be misled by a false major premise, and whether it can independently retrieve and apply the correct major premise based on the facts of the case.
[0087] Similarity attack refers to the insertion of similar crime terms into the problem description to interfere with the major premise judgment of a large language model by judging whether similar crimes or other crimes have been committed.
[0088] This attack is used to mention similar crimes in prompts, interfering with the accuracy of Large Language Models (LLM) in inferring major premises.
[0089] In one possible embodiment, Figure 4 This is a schematic diagram illustrating an attack premise provided in the embodiments of this application.
[0090] Reference Figure 4 The red part represents the location and method of the attack, and the green part represents the correct answer.
[0091] To address this problem, we introduce perturbations during the questioning phase (to determine whether a similar crime has been committed) and perturbations in the relevant legal provisions (to introduce incorrect relevant legal provisions). We attack the major premise of the logical judgment of this problem, evaluate whether the large model can eliminate interfering factors, and draw a conclusion based on rigorous deduction using syllogism.
[0092] Additionally, it should be noted that, following the example above, fine-grained annotations of legal knowledge are required to support robustness assessments of the large language model at the major premise, minor premise, and conclusion levels.
[0093] The dataset contains input examples for legal reasoning, including case facts, relevant legal provisions, and question prompts. Annotations include annotations for similar crime names, annotations for conviction logic reasoning, and annotations for domain synonyms.
[0094] Among them, similar crime name annotations are used to help the model identify and distinguish similar crimes in major premise attacks; conviction logic reasoning annotations are used to mark the logical reasoning paths in the facts of the case and relevant legal provisions to support the substitution of words and interference of elements in the facts of the case in minor premise attacks; domain synonym annotations are used to mark synonym substitutions and their application in the legal context in minor premise attacks.
[0095] Furthermore, the annotation of the dataset needs to be led by a team of legal experts and executed by legal professionals. All members must undergo an interview before joining the team to ensure they understand Chinese legal concepts and knowledge, and trial annotations must be conducted before formal annotation. Additionally, annotations must be carried out according to a designed Chinese annotation guideline to ensure consistency and accuracy. The guideline includes crime names, logical reasoning based on the four elements, the use of domain synonyms, and practical annotation examples.
[0096] Step 102 involves interfering with the accuracy of the large language model's narrative based on the input case facts through lexical attacks, element attacks, and narrative attacks.
[0097] This step describes the specific interference methods in minor premise attacks and the interference model's ability to judge minor premises in understanding the facts of a case.
[0098] In this embodiment of the application, a lexical attack refers to: attacking the vocabulary in the facts of a case by replacing words with synonyms in order to disrupt the logical judgment of the facts of a case by a large language model, including replacing random words in the facts of a case with common synonyms, replacing the vocabulary of the four elements in the facts of a case with common words, and replacing the vocabulary of the four elements in the facts of a case with synonymous legal element words.
[0099] This attack replaces words in the facts of the attack case with synonyms. Depending on whether the attack word and the candidate synonym are general words or legal element words, the attack methods are divided into general word to general word attack, legal element word to general word attack, and legal element word to legal element word attack.
[0100] Element attack refers to the act of adding summary legal elements or provisions from similar crimes after the facts of a case in order to disrupt the logical judgment of the facts of the case by the large language model.
[0101] This attack is used to insert adversarial legal elements from similar offenses at the end of the facts of a case. These similar elements are divided into factual elements summarized in the facts of the case and provisional elements summarized in legal clauses.
[0102] Narrative attacks refer to the insertion of contextual or irrelevant narrative statements after the facts of a case in order to disrupt the logical judgment of the large language model on the facts of the case.
[0103] This attack investigates the impact of subtle semantic changes on the final judgment. In criminal trials, judges typically make judgments based on reasoning about the four elements of a crime. Each crime has four elements: 1) the perpetrator, referring to the person who commits the crime; 2) the subjective aspect of the crime, referring to the perpetrator's mental state regarding their criminal act and its consequences; 3) the objective aspect of the crime, referring to the specific manifestation of the criminal act; and 4) the object of the crime, referring to the social relations protected by criminal law and violated by the criminal act.
[0104] Step 103 involves interfering with the final conclusion judgment of the large language model through prior behavioral attacks and expert opinion attacks in order to assess its resistance to interference in legal reasoning.
[0105] This step describes the specific interference methods used in conclusion attacks, which interfere with the reasoning ability of the model in judging case conclusions.
[0106] In this embodiment of the application, prior behavior attack refers to inserting a description of the perpetrator's previous criminal behavior after the facts of the case in order to assess whether the large language model can ignore prior behavior that is irrelevant to the current case.
[0107] This attack is used to insert a prior offense committed by the perpetrator into the prompt. According to Chinese criminal law, a perpetrator's prior offenses have no impact on the current criminal judgment, and the overall model should not allow the logical deduction of the perpetrator's current facts to be misled by other logical chains.
[0108] Expert opinion attack refers to the insertion of a specific identity's opinion on the crime in the question prompt to verify whether the large language model is influenced by external opinions.
[0109] This attack is used to insert opinions from different perspectives (from elementary school students to judges) regarding the appropriate crime for an action into a prompt. The larger model should ignore the influence of these different perspectives on the case's conclusion and rely solely on logical reasoning based on the facts themselves.
[0110] By summarizing the various attack methods in steps 101-103, this application also presents a summary table of three-level legal knowledge injection attacks, as shown in Table 1.
[0111] Table 1
[0112]
[0113] Furthermore, this application proposes three additional methods to enhance the robustness of large language models, as follows:
[0114] 1. Retrieval Enhancement Generation (RAG): Insert the legal provisions most relevant to the facts of the case into the question prompts to enhance the robustness of the large language model in dealing with major premise attacks.
[0115] This method inserts the criminal law clause closest to the fact into the prompt and then attacks again using all the methods in the attack framework.
[0116] 2. Thinking Chain (COT): The problem prompts require the large language model to reason step by step according to the logic of the four elements of criminal law, so as to improve the reasoning ability and anti-interference ability of the large language model.
[0117] The method explicitly states in the prompt, "Please reason step by step according to the reasoning logic of the four elements of criminal law."
[0118] 3. Few-shot: Add two similar case analysis examples to the question prompts, enabling the large language model to make accurate judgments based on the reasoning logic of similar cases.
[0119] This method is used to insert two typical cases and similar crimes into the prompts, allowing the model to make judgments based on the analysis logic of these two cases.
[0120] To achieve the above embodiments, this application also proposes a robustness assessment device for large language models in the legal domain based on knowledge injection attacks. Figure 5 This is a schematic diagram of the structure of a robustness assessment device 10 for a large language model in the legal domain based on knowledge injection attacks, provided as an embodiment of this application. Figure 5 As shown, the device includes:
[0121] The major premise knowledge injection attack module 100 is used to interfere with the major premise judgments made by the large language model in the legal field based on the input legal provisions through retrieval enhancement generation attacks and similar crime attacks.
[0122] The minor premise knowledge injection attack module 200 is used to interfere with the accuracy of the large language model's narrative based on the input case facts through lexical attacks, element attacks, and narrative attacks.
[0123] The conclusion knowledge injection attack module 300 is used to interfere with the final conclusion judgment of the large language model through prior behavior attacks and expert opinion attacks, in order to evaluate its anti-interference ability in legal reasoning.
[0124] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0125] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.
[0126] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.
[0127] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0128] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0129] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.
[0130] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0131] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0132] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0134] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0135] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0136] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0137] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
[0138] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0139] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A robustness assessment method for large language models in the legal domain based on knowledge injection attacks, characterized in that, Includes the following steps: By using search-enhanced generation attacks and similarity-based crime attacks, the major premise judgments made by large language models in the legal field based on input legal provisions are interfered with; The accuracy of the large language model's narrative based on the input case facts is interfered with through lexical attacks, element attacks, and narrative attacks. By interfering with the final conclusion judgment of the large language model through prior behavioral attacks and expert opinion attacks, its anti-interference ability in legal reasoning is evaluated. The aforementioned attacks, including retrieval enhancement generation attacks and similarity attack attacks, interfere with the major premise judgments made by the large language model in the legal field based on the input legal provisions. These include: introducing incorrectly related legal provisions that are unrelated to the facts of the case during the question prompting stage, thereby misleading the large language model when retrieving legal provisions; and adding vocabulary interference of similar crimes to the question description, thereby interfering with the major premise judgments of the large language model by judging whether similar crimes or other offenses have been committed. The method of interfering with the accuracy of the large language model's narrative based on the input case facts through lexical attacks, element attacks, and narrative attacks includes: attacking the vocabulary in the case facts through synonym substitution to disrupt the large language model's logical judgment of the case facts, including replacing random words in the case facts with common synonyms, replacing the four elements of the case facts with common words, and replacing the four elements of the case facts with synonymous legal element words; adding summary legal elements or legal provisions from similar crimes after the case facts to disrupt the large language model's logical judgment of the case facts; and adding situational or irrelevant narrative statements after the case facts to disrupt the large language model's logical judgment of the case facts. The method of interfering with the final conclusion judgment of the large language model through prior behavior attacks and expert opinion attacks to assess its resistance to interference in legal reasoning includes: inserting descriptions of the perpetrator's previous criminal behavior after the facts of the case to assess whether the large language model can ignore prior behavior that is irrelevant to the current case; and adding the judgment of the crime based on the perspective of a specific identity in the question prompt to verify whether the large language model will be interfered with by external opinions.
2. The method according to claim 1, characterized in that, Also includes: Fine-grained annotations are applied to a legal domain knowledge dataset to support robustness evaluation of the large language model at the three levels of major premise, minor premise, and conclusion. The dataset contains input examples for legal reasoning, including case facts, relevant legal provisions, and question prompts. The annotations include annotations for similar crime names, annotations for conviction logic reasoning, and annotations for domain synonyms. The annotation of the dataset was led by a team of legal experts and carried out by legal professionals. Trial annotation was conducted before the formal annotation, and the annotation was carried out in accordance with the designed Chinese annotation guidelines to ensure the consistency and accuracy of the annotation. The annotation guidelines include crime names, logical reasoning of the four elements, use of domain synonyms, and actual annotation examples.
3. The method according to claim 1, characterized in that, Also includes: Insert legal provisions most relevant to the facts of the case into the question prompts to enhance the robustness of the large language model in dealing with major premise attacks; The problem prompt requires the large language model to reason step by step according to the logic of the four elements of criminal law, so as to improve the reasoning ability and anti-interference ability of the large language model; Adding two similar case analysis examples to the question prompts enables the large language model to make accurate judgments based on the reasoning logic of similar cases.
4. A robustness assessment device for large language models in the legal domain based on knowledge injection attacks, characterized in that, include: The major premise knowledge injection attack module is used to interfere with the major premise judgments made by the large language model in the legal field based on the input legal provisions through retrieval enhancement generation attacks and similar crime attacks. The minor premise knowledge injection attack module is used to interfere with the accuracy of the large language model's narrative based on the input case facts through lexical attacks, element attacks, and narrative attacks. The conclusion knowledge injection attack module is used to interfere with the final conclusion judgment of the large language model through prior behavior attacks and expert opinion attacks, so as to evaluate its anti-interference ability in legal reasoning. The aforementioned attacks, including retrieval enhancement generation attacks and similarity attack attacks, interfere with the major premise judgments made by the large language model in the legal field based on the input legal provisions. These include: introducing incorrectly related legal provisions that are unrelated to the facts of the case during the question prompting stage, thereby misleading the large language model when retrieving legal provisions; and adding vocabulary interference of similar crimes to the question description, thereby interfering with the major premise judgments of the large language model by judging whether similar crimes or other offenses have been committed. The method of interfering with the accuracy of the large language model's narrative based on the input case facts through lexical attacks, element attacks, and narrative attacks includes: attacking the vocabulary in the case facts through synonym substitution to disrupt the large language model's logical judgment of the case facts, including replacing random words in the case facts with common synonyms, replacing the four elements of the case facts with common words, and replacing the four elements of the case facts with synonymous legal element words; adding summary legal elements or legal provisions from similar crimes after the case facts to disrupt the large language model's logical judgment of the case facts; and adding situational or irrelevant narrative statements after the case facts to disrupt the large language model's logical judgment of the case facts. The method of interfering with the final conclusion judgment of the large language model through prior behavior attacks and expert opinion attacks to assess its resistance to interference in legal reasoning includes: inserting descriptions of the perpetrator's previous criminal behavior after the facts of the case to assess whether the large language model can ignore prior behavior that is irrelevant to the current case; and adding the judgment of the crime based on the perspective of a specific identity in the question prompt to verify whether the large language model will be interfered with by external opinions.
5. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-3.
7. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-3.
Citation Information
Patent Citations
Multi-composite detection method, device and equipment for smart contract and storage medium
CN117725594A
Method and device for automatically mapping vulnerabilities to attack techniques and tactics based on large language model
CN118368103A