A legal artificial intelligence system risk assessment method
Patent Information
- Application Number
- CN202610775547.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]但是,以上针对在中国司法领域使用人工智能风险的研究均没有提出具体的风险评估方法
由上述实施例可知,本申请根据风险评估指标,将法律人工智能系统的风险分为高、中、低三类。每个合法的人工智能系统都可以根据这些指标进行风险评估,从而确定其风险等级。通过对法律人工智能系统进行风险分类,从而在促进技术创新和减轻潜在风险之间取得平衡,也进一步为法律人工智能的发展和使用提供了参考。
Smart Images

Figure CN122654722A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of legal artificial intelligence technology, and in particular relates to a risk assessment method for legal artificial intelligence systems. Background Technology
[0002] Research in the field of legal artificial intelligence has primarily focused on technological improvements and experimental verification, with insufficient research into the risks involved. Consequently, many legal AI technologies that perform well in controlled environments have ultimately been deemed too risky, preventing widespread acceptance or implementation. COMPAS software is a prime example. The paper "T. Brennan, W. Dieterich, and B. Ehret. 2008. Evaluating the Predictive Validity of the COMPAS Risk and Needs Assessment System. Criminal Justice & Behavior 36, 1(2008), 21–40)" demonstrates a discrepancy between laboratory and real-world feasibility. Therefore, merely verifying the effectiveness and reliability of new technologies is insufficient; risk assessments of AI systems must be conducted within the context of actual legal scenarios to ensure their proper and regulated use.
[0003] Currently, there are some studies on the risks of using artificial intelligence in the Chinese judicial field. The paper "Xiao Junyong, Liu Hao. Cultural Security Risks and Countermeasures in the Application of Generative Artificial Intelligence [J / OL]. Journal of China University of Political Science and Law, 1-25 [2026-04-27]. https: / / link.cnki.net / urlid / 11.5607.D.20260425.1514.002." argues that generative artificial intelligence, by aligning with human thinking patterns to achieve "natural human-machine dialogue," deeply embeds itself in the process of cultural production, dissemination, and development, while also triggering multiple cultural security risks. Based on its target and generative mechanism, these risks can be summarized as risks to cultural subjects, cultural content, and cultural ecology. Specifically, these risks manifest as the risk of stripping away and dissolving cultural identity, the risk of fostering historical nihilism, the risk of alienating cultural symbols, and the risk of disrupting the order of cultural production. Gu Yahui. Risks and Regulation of Generative Artificial Intelligence Empowering Criminal Investigation [J / OL]. Journal of China People's Police University, 1-8 [2026-04-27]. https: / / link.cnki.net / urlid / 13.1434.G4.20260421.0943.004. argues that generative artificial intelligence empowering criminal investigation faces a series of risks, including the risk of unduly diminishing the subjectivity of rational investigative practice, the credibility risk of biased generated content, and the risk of the technology's operation impacting due process.
[0004] However, none of the above studies on the risks of using artificial intelligence in China's judicial field have proposed specific risk assessment methods. Summary of the Invention
[0005] The purpose of this application is to address the problems existing in the prior art by providing a risk assessment method for legal artificial intelligence systems.
[0006] According to a first aspect of the embodiments of this application, a method for risk assessment of a legal artificial intelligence system is provided, comprising: (1) Determine risk assessment indicators, which shall include at least interpretability, the potential impact of adverse consequences, and the performance of the technology used; (2) For each risk assessment indicator, determine the grading criteria and divide them into three levels according to their intensity; (3) For the legal artificial intelligence system to be evaluated, determine its level under each risk assessment indicator, generate the corresponding risk assessment triplet, and thus complete the risk assessment of the system.
[0007] Furthermore, the intensity classification of the interpretability index is based on the following criteria: If the results generated by the system conform to the reasoning process of legal syllogism, then they are highly interpretable; If the system can display the path of "key fact nodes → legal elements → conclusion", but the path still has black box components, then the interpretability is moderate. If the results generated by the system are determined solely by the model's internal parameters or feature weights, then the interpretability is weak.
[0008] Furthermore, systems generated using symbolic methods such as decision trees, rule learning, and linear models are considered to have strong interpretability; systems generated using nonlinear SVMs and random forests are considered to have moderate interpretability; and systems based on deep learning are considered to have weak interpretability.
[0009] Furthermore, the intensity classification of the influencing indicators is based on the following criteria: If the results generated by the system directly harm human rights and threaten people's health, safety, and fundamental rights, then the impact is serious. If the system poses a risk of deception or manipulation, but can be regulated through transparency requirements, and only affects the parties' access to information and their interactive experience, then the impact is moderate. If the system is used to improve work efficiency and has no impact on human decision-making, then the impact is minor.
[0010] Furthermore, the intensity of performance indicators is determined based on the following criteria: for a specific legal task, the system's performance indicator value falls within the corresponding intensity range. These performance indicators include, but are not limited to, accuracy, precision, recall, and F1 score.
[0011] Furthermore, in step (3): If the interpretability of the system is moderate or weak, and the impact of the system is severe, then the system is judged to be high-risk. If a system has weak interpretability and moderate impact, and its performance is weak, then the system is considered high-risk; otherwise, it is considered medium-risk. If a system is highly interpretable, has a serious impact, and has moderate or weak performance, then the system is considered high-risk. If the system has weak interpretability, minor impact, and moderate or weak performance, the system is classified as medium risk. If the interpretability and impact of the system are moderate, then the system is classified as medium risk. If the system's interpretability is moderate or strong, its impact is minor, and its performance is weak, then the system is classified as medium risk. If the system has strong interpretability, moderate impact, and weak or moderate performance, then the system is classified as medium risk. If the system is highly interpretable, has a significant impact, and performs well, then the system is classified as medium risk. If the system has weak interpretability, minor impact, and strong performance, then the system is judged to be low-risk. If the system has moderate or strong interpretability, minor impact, and moderate or strong performance, then the system is considered low risk. If a system has strong interpretability, moderate impact, and strong performance, it is considered low-risk.
[0012] According to a second aspect of the embodiments of this application, a risk assessment device for a legal artificial intelligence system is provided, comprising: The indicator determination module is used to determine risk assessment indicators, which include at least interpretability, the potential impact of adverse consequences, and the performance of the technology used. The indicator grading module is used to determine the grading criteria for each risk assessment indicator and divide them into three levels according to their intensity. The risk assessment module is used to determine the level of the legal AI system under evaluation under each risk assessment indicator, generate the corresponding risk assessment triplet, and thus complete the risk assessment of the system.
[0013] According to a third aspect of the embodiments of this application, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the method described in the first aspect.
[0014] According to a fourth aspect of the embodiments of this application, an electronic device is provided, comprising: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in the first aspect.
[0015] According to a fifth aspect of the embodiments of this application, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method as described in the first aspect.
[0016] The technical solutions provided by the embodiments of this application may include the following beneficial effects: As can be seen from the above embodiments, this application classifies the risks of legal artificial intelligence systems into three categories: high, medium, and low, based on risk assessment indicators. Each legitimate artificial intelligence system can undergo risk assessment according to these indicators to determine its risk level. By classifying legal artificial intelligence systems by risk, a balance is achieved between promoting technological innovation and mitigating potential risks, and this also provides a reference for the development and use of legal artificial intelligence.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] Figure 1 This is a flowchart illustrating a risk assessment method for a legal artificial intelligence system according to an exemplary embodiment.
[0020] Figure 2 This is a heatmap of the risk assessment method for legal artificial intelligence systems on the interpretability index. (a) shows the risk assessment results corresponding to different combinations of "potential impact" and "performance" when the system's "interpretability" is weak; (b) shows the risk assessment results corresponding to different combinations of "potential impact" and "performance" when the system's interpretability is moderate; and (c) shows the risk assessment results corresponding to different combinations of "potential impact" and "performance" when the system's interpretability is strong.
[0021] Figure 3 This is a schematic diagram showing the risk level classification results of a portion of legal artificial intelligence systems after applying a risk assessment method for legal artificial intelligence systems according to this application.
[0022] Figure 4 This is a block diagram illustrating a risk assessment device for a legal artificial intelligence system according to an exemplary embodiment.
[0023] Figure 5 This is a schematic diagram of an electronic device according to an exemplary embodiment. Detailed Implementation
[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0025] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0026] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0027] Figure 1 This is a flowchart illustrating a risk assessment method for a legal artificial intelligence system according to an exemplary embodiment, such as... Figure 1 As shown, the method may include the following steps: (1) Determine the risk assessment indicators, which shall include at least the explainability, the potential impact of adverse consequences, and the performance of the technology used. (2) For each risk assessment indicator, determine the grading criteria and divide them into three levels according to their intensity; In one embodiment, arrow symbols “↑”, “→” and “↓” are used to represent the different intensities of each indicator, as defined in Table 1.
[0028] Table 1: Meaning of different arrow symbols in various indicators In this application, the determination of the system's interpretability and performance indicators are based on relevant literature, covering aspects such as: judgment outcome prediction (including civil and criminal cases), legal text generation (including large-scale models), legal document proofreading, extraction of points of contention, and similar case retrieval. These documents already provide detailed descriptions of the system's interpretability and performance, while the system's impact on humans can be evaluated manually.
[0029] In this application, the intensity of the three-dimensional indicators is divided into: interpretability intensity: weak, medium, strong; potential impact intensity: slight, medium, severe; performance intensity: weak, medium, strong. The criteria for classifying the indicator intensity levels in one embodiment are summarized in Table 2: Table 2 Reference Basis for Risk Assessment Indicator Strength Table 3 System performance evaluation standards used for different legal tasks (3) For the legal artificial intelligence system to be evaluated, determine its level under each risk assessment indicator, generate the corresponding risk assessment triplet, and thus complete the risk assessment of the system; In this application, the three indicators of the legal artificial intelligence system—explainability, impact, and performance—can be represented as a triple (E, I, P) for risk assessment. The specific assessment criteria are as follows: Figure 2 As shown: (I) High risk (i) When a system is poorly interpretable (“↓” or “→”) and is applied to legal scenarios that may have a serious impact on human rights (“↑”), it is classified as high risk regardless of its performance. (ii) When the interpretability of the system is weak ("↓") and it is applied to legal scenarios that do not have a serious impact on human rights ("→"), if its performance is also weak ("↓"), it is classified as high risk. (iii) When a system is highly interpretable ("↑") and is applied to legal scenarios that may have a serious impact on human rights ("↑"), it is also classified as high-risk if its performance is poor ("↓" or "→").
[0030] (II) Medium risk: (i) When the interpretability of a system is weak (“↓”) and it is used in legal scenarios that do not have a serious impact on human rights (“→”), it is classified as medium risk if its performance is not bad (“→” or “↑”). (ii) When the interpretability of a system is weak (“↓”) and it is used in a legal scenario that has no impact on human rights (“↓”), it is classified as medium risk if its performance is average or weak (“↓” or “→”). (iii) When a system has moderate interpretability (“→”) and is used in a legal scenario that has no serious impact on human rights (“→”), it is classified as a medium-risk application regardless of its performance. (iv) When a system is highly interpretable or generally interpretable (“↑” or “→”) and is used in a legal scenario that has no impact on human rights (“↓”), it is classified as medium risk if it exhibits weak performance (“↓”). (v) When a system is highly interpretable (“↑”) and used in legal scenarios that do not have a serious impact on human rights (“→”), it is classified as medium risk if its performance is weak (“↓” or “→”). (vi) When a system is highly interpretable (“↑”) and used in legal scenarios that have a potentially serious impact on human rights (“↑”), it is classified as medium risk if it is high in performance (“↑”).
[0031] (III) Low risk: (i) When a system has poor interpretability and is used in a legal scenario that has no impact on human rights (“↓”), it is classified as low risk if it has strong performance (“↑”). (ii) When a system has good interpretability (“→” or “↑”) and is used in a legal scenario that has no impact on human rights (“↓”), it is classified as low risk if its performance is good (“↑” or “→”). (iii) A system is classified as low risk if it is highly interpretable (“↑”) and used without serious impact on human rights (“→”) and if it is highly performant (“↑”).
[0032] The three-dimensional space mapped by arranging and combining risk assessment indicators according to different intensities was divided into three two-dimensional heat maps, such as... Figure 2 As shown, the risk level classification criteria above are also illustrated.
[0033] In addition, based on the risk assessment method for legal artificial intelligence systems described above, we have categorized the risks of some systems as follows: (1) High-risk systems: a. Legal Judgment Prediction System for Criminal Cases. Existing legal AI systems for predicting criminal case judgments are mostly based on deep learning. While they exhibit good performance, their interpretability is weak, and errors in criminal case prediction can lead to serious consequences. Therefore, the risk triplet (E, I, P) for such systems is considered to be (↓, ↑, ↑), as follows: Figure 2 As shown in (a).
[0034] b. Legal generative systems based on large models. While large models currently achieve satisfactory performance in various legal tasks, the generated results still lack interpretability. More importantly, large models are prone to illusions, which can lead to extremely serious consequences. For example, large models may generate non-existent legal provisions, outdated judicial interpretations, or fabricated case information. Therefore, the risk triplet (E, I, P) for this task is considered to be (↓, ↑, ↑), as follows... Figure 2 As shown in (a).
[0035] (2) Application of medium risk: a. Dispute Focus Mining System. This system automatically identifies and extracts interactive arguments with the same theme from the plaintiff's and defendant's statements, treating these arguments as a "sentence pair" classification problem. Current systems are primarily based on deep learning methods, offering good performance but lacking interpretability. Furthermore, since errors in these systems generally do not lead to serious consequences, the risk triplet (E, I, P) for this task is considered to be (↓, →, ↑), as follows... Figure 2 As shown in (a) in the figure.
[0036] b. Prediction of Civil Case Judgments. For civil cases, the predicted legal judgment is a response to the plaintiff's claims (support or non-support), essentially a legal text classification problem. It is argued that civil case legal judgment prediction should be considered a medium-risk level: compared to criminal cases, the consequences of errors in civil cases are relatively minor. Furthermore, most methods in the system are based on deep learning, which has strong performance but lacks interpretability; therefore, it is considered a medium-risk method, with the corresponding risk triplet (E, I, P) being (↓, →, ↑). Figure 2 (a) in the middle.
[0037] c. Legal document generation without using large models (excluding judicial judgments). Besides large models, there are currently two main methods for generating legal documents: generation based on pre-defined templates or using Seq2Seq models. Both methods perform well in this task, and template-based generation offers a degree of interpretability. Furthermore, errors in these applications generally do not lead to serious consequences. It is considered that using legal AI to generate legal documents should be viewed as a medium-risk level, with corresponding risk triples (E, I, P) of (↓, →, ↑) and (→, →, ↑). Figure 2 (a) and (b) in the example.
[0038] (3) Low-risk applications: a. Legal document summarization generation systems. Legal document summarization generation systems are considered low-risk. Currently, most methods are based on deep learning and statistical learning, and the error rate of the generated summaries is low; even if errors occur, they will not cause serious consequences. Therefore, the risk triplet (E, I, P) for this type of system is considered to be (↓, ↓, ↑), (→, ↓, ↑), and (↑, ↓, ↑), corresponding to... Figure 2 (a), (b), (c).
[0039] b. Similar Case Retrieval System. The use of legal AI systems for similar case retrieval is considered low-risk. Currently, most retrieval methods are based on deep learning models, which offer good performance but poor interpretability. Furthermore, retrieval errors generally do not lead to serious consequences. Therefore, its risk triplet (E, I, P) is considered to be (↓, ↓↑), corresponding to... Figure 2 (a)
[0040] c. Legal Document Proofreading Systems. Using legal AI systems for legal document proofreading is considered low-risk. These systems typically combine rules with probabilistic statistics or use N-gram models to detect grammatical structures and legal terminology in documents for automatic proofreading. Additionally, some neural network methods have performed well. Furthermore, errors in these systems generally do not have serious consequences. Therefore, the risk triplet (E, I, P) for these systems is considered to be (↑, ↓, ↑), (→, ↓, ↑), or (↓, ↓, ↑), corresponding to... Figure 2 (c), (b), (a).
[0041] The final risk assessment and classification results are as follows: Figure 3 As shown.
[0042] Corresponding to the aforementioned embodiments of the legal artificial intelligence system risk assessment method, this application also provides embodiments of the legal artificial intelligence system risk assessment device.
[0043] Figure 4 This is a block diagram illustrating a risk assessment device for a legal artificial intelligence system according to an exemplary embodiment. (Refer to...) Figure 4 The device may include: The indicator determination module M1 is used to determine risk assessment indicators, which include at least interpretability, the potential impact of adverse consequences, and the performance of the technology used. The indicator classification module M2 is used to determine the classification criteria for each risk assessment indicator and divide them into three levels according to their intensity. The risk assessment module M3 is used to determine the level of the legal artificial intelligence system under each risk assessment indicator and generate the corresponding risk assessment triplet to complete the risk assessment of the system.
[0044] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0045] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0046] Accordingly, this application also provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the legal artificial intelligence system risk assessment method described above.
[0047] Accordingly, this application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the legal artificial intelligence system risk assessment method described above. Figure 5 The diagram shown is a hardware structure diagram of any device with data processing capabilities, which is the risk assessment device of a legal artificial intelligence system provided in an embodiment of the present invention. Except for... Figure 5 In addition to the processor, memory, and network interface shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0048] Accordingly, this application also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the aforementioned risk assessment method for a legal artificial intelligence system. The computer-readable storage medium can be an internal storage unit of any data-processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data-processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data-processing device, and can also be used to temporarily store data that has been output or will be output.
[0049] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
Claims
1. A risk assessment method for a legal artificial intelligence system, characterized in that, include: (1) Determine risk assessment indicators, which shall include at least interpretability, the potential impact of adverse consequences, and the performance of the technology used; (2) For each risk assessment indicator, determine the grading criteria and divide them into three levels according to their intensity; (3) For the legal artificial intelligence system to be evaluated, determine its level under each risk assessment indicator, generate the corresponding risk assessment triplet, and thus complete the risk assessment of the system.
2. The method according to claim 1, characterized in that, The intensity classification for interpretability indicators is based on the following criteria: If the results generated by the system conform to the reasoning process of legal syllogism, then they are highly interpretable; If the system can display the path of "key fact nodes → legal elements → conclusion", but the path still has black box components, then the interpretability is moderate. If the results generated by the system are determined solely by the model's internal parameters or feature weights, then the interpretability is weak.
3. The method according to claim 1, characterized in that, Systems generated using symbolic methods such as decision trees, rule learning, and linear models are considered to have high interpretability; systems generated using nonlinear SVMs and random forests are considered to have moderate interpretability; and systems based on deep learning are considered to have low interpretability.
4. The method according to claim 1, characterized in that, The intensity classification of the influencing indicators is based on the following criteria: If the results generated by the system directly harm human rights and threaten people's health, safety, and fundamental rights, then the impact is serious. If the system poses a risk of deception or manipulation, but can be regulated through transparency requirements, and only affects the parties' access to information and their interactive experience, then the impact is moderate. If the system is used to improve work efficiency and has no impact on human decision-making, then the impact is minor.
5. The method according to claim 1, characterized in that, The intensity of performance metrics is determined as follows: for a specific legal task, the system's performance metric values fall within the corresponding intensity range. These performance metrics include, but are not limited to, accuracy, precision, recall, and F1 score.
6. The method according to claim 1, characterized in that, In step (3): If the interpretability of the system is moderate or weak, and the impact of the system is severe, then the system is judged to be high-risk. If a system has weak interpretability and moderate impact, and its performance is weak, then the system is considered high-risk; otherwise, it is considered medium-risk. If a system is highly interpretable, has a serious impact, and has moderate or weak performance, then the system is considered high-risk. If the system has weak interpretability, minor impact, and moderate or weak performance, the system is classified as medium risk. If the interpretability and impact of the system are moderate, then the system is classified as medium risk. If the system's interpretability is moderate or strong, its impact is minor, and its performance is weak, then the system is classified as medium risk. If the system has strong interpretability, moderate impact, and weak or moderate performance, then the system is classified as medium risk. If the system is highly interpretable, has a significant impact, and performs well, then the system is classified as medium risk. If the system has weak interpretability, minor impact, and strong performance, then the system is judged to be low-risk. If the system has moderate or strong interpretability, minor impact, and moderate or strong performance, then the system is considered low risk. If a system has strong interpretability, moderate impact, and strong performance, it is considered low-risk.
7. A risk assessment device for a legal artificial intelligence system, characterized in that, include: The indicator determination module is used to determine risk assessment indicators, which include at least interpretability, the potential impact of adverse consequences, and the performance of the technology used. The indicator grading module is used to determine the grading criteria for each risk assessment indicator and divide them into three levels according to their intensity. The risk assessment module is used to determine the level of the legal AI system under evaluation under each risk assessment indicator, generate the corresponding risk assessment triplet, and thus complete the risk assessment of the system.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method as described in any one of claims 1-6.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.
10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-6.