Legal big data model construction method, legal information analysis method, computer device

CN120973922BActive Publication Date: 2026-09-01BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511256898.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-09-01
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

[0005]有鉴于此,本公开实施例提供了一种法务大模型构建方法、法务信息分析方法、计算机装置,能够解决现有技术中存在的法律领域高质量标注数据获取难且成本高致训练遇数据瓶颈、未充分利用法律知识结构特性、少量数据微调易过拟合或遗忘原有能力以及低资源条件下难以适应多样化法律任务等问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973922B_ABST
    Figure CN120973922B_ABST
Patent Text Reader

Abstract

This application discloses a method for constructing a large-scale legal model, a method for analyzing legal information, and a computer device. The method for constructing the large-scale legal model includes: constructing a legal domain knowledge graph with a multi-layered organizational knowledge architecture; obtaining a preset number of manually labeled seed samples; extracting their structural features and generating target templates; filling the target templates with the legal domain knowledge graph to generate diverse training samples; configuring a multi-layered sample generation strategy based on the training samples and requirements; training a large language model layer by layer according to this strategy; and using the trained large language model as the large-scale legal model. This method automatically generates diverse, high-quality legal training samples using a legal domain knowledge graph and a small number of seed samples, effectively reducing reliance on large amounts of labeled data and computing resources. The multi-layered sample generation strategy ensures that the model can gradually master complex legal knowledge and reasoning, effectively improving data utilization and the accuracy of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method for constructing a legal big data model, a method for analyzing legal information, and a computer device. Background Technology

[0002] In recent years, large language model technology has shown rapid development. Open-source large models such as DeepSeek-R1, QWEN, and LLAMA have demonstrated powerful capabilities in natural language understanding and generation. However, the application of these general-purpose large models in specialized vertical fields such as law has been less than satisfactory. Specifically, this manifests in several ways: insufficient understanding of legal knowledge; while possessing some legal knowledge, the depth of understanding of legal terminology, concepts, and their complex relationships is insufficient, often leading to misinterpretations or inappropriate applications; a lack of legal reasoning ability; insufficient training in professional skills such as the application of legal provisions, case analysis, and legal reasoning; frequent professional errors; a tendency to generate "illusions" when dealing with legal issues, resulting in content that violates legal provisions; and poor adaptability, making it difficult to adapt to changes in legal systems across different countries and periods.

[0003] To address the challenges of applying general-purpose large-scale models in the legal field, various solutions have been proposed. Full-scale fine-tuning involves collecting a large legal corpus to fine-tune all parameters of the model, as seen in models like ChatLaw. However, this method requires a large amount of labeled data and computational resources, resulting in extremely high costs. Retrieval-enhanced generation leverages external knowledge bases to improve model output, but it relies heavily on retrieval quality and struggles to handle complex legal reasoning. Expert system integration combines traditional legal expert systems with large-scale models; however, rule maintenance is costly, and it struggles to handle boundary cases. Efficient parameter fine-tuning uses techniques like LoRA to fine-tune only a subset of parameters to reduce resource consumption, but simply performing efficient parameter fine-tuning is insufficient to meet the specialized needs of the legal field.

[0004] Current technologies still have several major shortcomings. First, obtaining high-quality labeled data in the legal field is difficult and costly, making it difficult for model training to overcome the data bottleneck. Second, legal knowledge has structured characteristics and strict logical relationships, while most existing methods use unstructured text for direct training, failing to fully utilize the structural characteristics of legal knowledge. Third, in scenarios where the model is fine-tuned with a small amount of data, it is prone to overfitting the training examples or forgetting its original capabilities. Fourth, legal tasks are diverse, covering consultation, drafting, review, dispute resolution, etc., but existing methods are difficult to adapt to multiple legal tasks simultaneously under low-resource conditions. Summary of the Invention

[0005] In view of this, the present disclosure provides a method for constructing a large legal model, a method for analyzing legal information, and a computer device, which can solve the problems in the prior art such as the difficulty and high cost of obtaining high-quality labeled data in the legal field, which leads to data bottlenecks in training, the failure to fully utilize the characteristics of the legal knowledge structure, the tendency for overfitting or forgetting of original capabilities when fine-tuning with a small amount of data, and the difficulty in adapting to diverse legal tasks under low resource conditions.

[0006] Firstly, this disclosure provides a method for constructing a large-scale legal model, including: Construct a knowledge graph in the legal field; the knowledge graph in the legal field includes a multi-layered organizational knowledge architecture; Obtain a preset amount of manually labeled seed samples; Extract the structural features of the seed sample, and generate several target templates based on the structural features; Based on the legal domain knowledge graph, several target templates are populated to generate diverse training samples; Based on the diverse training samples and model training requirements, a multi-level sample generation strategy is configured. The large language model is trained based on the multi-level sample generation strategy, and the trained large language model is used as the legal large model.

[0007] Secondly, this disclosure also provides a legal information analysis method, including: Identify the legal information to be analyzed; The legal information is input into the target large model to generate feedback information; The target large model is a legal vertical large model that has been fine-tuned using the legal large model construction method described above.

[0008] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution: The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform either the legal big data model construction method or the legal information analysis method described above.

[0009] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to execute any of the legal big data model construction methods or legal information analysis methods described above.

[0010] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0011] The legal big data model construction method disclosed in this application first constructs a legal domain knowledge graph including a multi-level organizational knowledge architecture; second, it obtains a preset amount of manually labeled seed samples, extracts the structural features of the seed samples, and generates several target templates based on the structural features; then, it fills the target templates according to the legal domain knowledge graph to generate diverse training samples; finally, it configures a multi-level sample generation strategy according to the diverse training samples and model training requirements, trains the large language model based on the multi-level sample generation strategy, and uses the trained large language model as the legal big data model. This method automatically generates diverse high-quality legal training samples through the legal domain knowledge graph and a small number of seed samples, effectively reducing the dependence on a large amount of labeled data and computing resources, and alleviating the data and resource bottleneck problem to a certain extent. The multi-level sample generation strategy allows the model to use samples of different levels at different stages for training, ensuring that the model can gradually master complex legal knowledge and reasoning. This method makes full use of limited data, effectively improves the efficiency of data utilization and the accuracy of model training, and obtains a high-precision model that can adapt to different legal tasks under low-resource conditions.

[0012] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart illustrating the legal big data model construction method provided in this embodiment of the disclosure.

[0015] Figure 2 This is a flowchart illustrating the method for constructing a legal domain knowledge graph as provided in this embodiment of the disclosure.

[0016] Figure 3 This is a flowchart illustrating the method for obtaining seed samples provided in an embodiment of the present disclosure.

[0017] Figure 4 This is a flowchart illustrating a method for generating diverse training samples according to embodiments of the present disclosure.

[0018] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation

[0019] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0020] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0021] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0022] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0023] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0024] Reference Figure 1The first aspect of this application discloses a method for constructing a legal big data model, including: S100, constructing a knowledge graph in the legal field; the knowledge graph in the legal field includes a multi-layered organizational knowledge architecture.

[0025] Constructing a knowledge graph in the legal field, which includes a multi-layered knowledge architecture, can systematically and structurally organize legal knowledge. At the same time, the multi-layered knowledge architecture makes legal knowledge more systematic and orderly, facilitating subsequent retrieval and use.

[0026] Furthermore, it can flexibly adapt to the legal systems of different countries and periods as needed, keeping pace with the times and meeting the latest needs of customers.

[0027] S200: Obtain a preset amount of manually labeled seed samples.

[0028] The seed sample is a high-quality seed sample, and the preset quantity refers to a small quantity.

[0029] Manually labeled seed samples ensure the accuracy and reliability of the data, providing a solid foundation for subsequent model training. Seed samples can guide the model to learn the correct legal knowledge and solutions, thereby improving the model's performance.

[0030] S300 extracts the structural features of the seed sample and generates several target templates based on the structural features.

[0031] The target templates include legal Q&A templates, case analysis templates, clause interpretation templates, contract drafting templates, or case reasoning templates.

[0032] Target templates can serve as the basis for generating diverse training samples, greatly improving the efficiency of sample generation. Templates can ensure that the generated training samples have a certain structure and format, enabling the model to better learn the expression rules of legal knowledge.

[0033] S400 uses a legal knowledge graph to populate several target templates, generating diverse training samples.

[0034] The samples generated through this step can reflect the structural characteristics of legal knowledge, enabling the model to learn the inherent logical relationships of legal knowledge during training and fully utilize the structural information of legal knowledge. At the same time, the supplemented training data can increase the diversity and richness of model training, allowing the model to learn more knowledge and patterns in different detailed scenarios, improve the model's generalization ability, and optimize the model's performance in the legal field. When dealing with tasks in the legal field, the model can more accurately understand the questions, generate high-quality answers, and improve the quality of task completion.

[0035] Diverse training samples allow the model to learn more legal knowledge and solutions, improving its generalization ability. With fine-tuning on a small amount of data, the risk of overfitting is reduced while retaining the model's original capabilities. At the same time, it makes full use of the rich knowledge in the legal domain knowledge graph to provide more information for model training.

[0036] S500 configures a multi-level sample generation strategy based on diverse training samples, trains a large language model based on the multi-level sample generation strategy, and uses the trained large language model as a legal big model.

[0037] Multi-level sample generation strategies allow existing mature large language models to gradually learn legal knowledge of different difficulties and types, improving the training effect of the models. After training and optimization, the large models can have a certain depth of understanding of legal terms, concepts and their complex relationships, possess strong legal reasoning ability, and have professional abilities such as application of legal provisions, case analysis, and legal reasoning. They can better handle various legal issues and provide users with accurate and reliable legal answers.

[0038] The legal big data model construction method disclosed in this application provides the model with comprehensive and accurate legal knowledge and high-quality training data by constructing a legal domain knowledge graph and using manually labeled seed samples, thereby improving the model's accuracy and reliability. Diverse training samples and multi-level sample generation strategies enable the model to learn more legal knowledge and solution methods, enhancing its generalization ability and enabling it to better handle various types of legal problems. The use of the legal domain knowledge graph makes the model's decision-making process more transparent, providing users with legal basis and reasoning processes, thus improving the model's interpretability. By generating diverse training samples, the reliance on a large amount of manually labeled data is reduced, lowering the model training cost. Compared to traditional full-scale fine-tuning methods, this solution achieves comparable or even better performance in the legal domain using only 5-10% of the labeled data.

[0039] Reference Figure 2 The methods for constructing legal domain knowledge graphs, as described in S100, include: S110, Determine the constituent elements of the legal field's map.

[0040] The graph consists of elements related to legal principles, legal provisions, and cases. Elements related to legal principles include basic concepts in the legal field, elements related to legal provisions include legal provisions in the legal field, and elements related to cases include related cases in the legal field and the relationships between legal subjects in the legal field.

[0041] Specifically, basic concepts in the legal field include legal terminology, legal subjects (natural persons, legal persons, etc.), and legal objects (property rights, creditor's rights, etc.); legal provisions in the legal field include the Constitution, laws, administrative regulations, local regulations, judicial interpretations, and attribute information for each provision; attribute information includes one or more of the following: the effective date, level of validity, and revision history of the provision; related cases in the legal field include typical cases, judicial precedents, and other practical application scenarios; and legal subject relationships in the legal field include one or more of the following: application relationship, interpretation relationship, reference relationship, and priority relationship.

[0042] This step clearly defines the scope of the legal knowledge graph, providing a clear direction for subsequent knowledge architecture construction. It also covers elements at three levels: legal theory, legal provisions, and cases, and can comprehensively and systematically reflect the knowledge system of the legal field.

[0043] S120 constructs a jurisprudential knowledge architecture based on basic concepts in the legal field. The jurisprudential knowledge architecture includes several basic concept knowledge nodes and the relationship edges between different basic concept knowledge nodes.

[0044] For example, basic legal concepts collected can be used as knowledge nodes, such as "property rights," "ownership," and "usufruct rights." Each knowledge node contains information such as the name, definition, and related attributes of the concept. The relationships between different basic concepts can be analyzed. For example, since "ownership" is a specific type of "property rights," a "belonging" relationship edge can be established between the two knowledge nodes of "ownership" and "property rights." In this way, a network structure that reflects the logical relationships of legal principles can be constructed.

[0045] This step graphically illustrates the inherent logical relationships between basic legal concepts, which helps to deepen the understanding of the legal theory system and provides a foundation for legal reasoning. For example, relevant legal principles and rules can be inferred from the relationships between concepts.

[0046] S130, construct a legal provision-level knowledge architecture based on legal provisions in the legal field. The legal provision-level knowledge architecture includes several legal provision-type knowledge nodes and the relationship edges between different legal provision-type knowledge nodes.

[0047] Specifically, each legal provision is treated as a knowledge node, containing information such as its number, name, and specific content. The relationships between the provisions are analyzed; for example, some provisions may supplement, interpret, or restrict others. For instance, the provisions on special contracts in the Civil Code may supplement the provisions on general contracts, thus establishing a corresponding relationship between the knowledge nodes of these two provisions.

[0048] This step clearly demonstrates the interrelationships between legal provisions, making it easier to find and understand relevant provisions. When dealing with legal issues, it allows for the quick location and application of relevant provisions based on the relationships between them.

[0049] S140. Construct a case-level knowledge architecture based on related cases in the legal field and the relationships between legal subjects in the legal field. The case-level knowledge architecture includes several case-type knowledge nodes and the relationship edges between different case-type knowledge nodes.

[0050] Specifically, each related legal case is treated as a knowledge node, containing basic case information such as case number, parties, cause of action, and judgment. The similarities and connections between different cases are analyzed. For example, two cases involving the same legal issues and relationships between legal entities can establish a "similar" relationship edge. Simultaneously, based on the relationships between legal entities in the cases, connections are established between the case knowledge nodes and related entity knowledge nodes.

[0051] By examining the relationships between cases, we can summarize some common legal problem-solving patterns and adjudication rules; when dealing with new legal issues, we can quickly find relevant cases for reference and learning.

[0052] S150 is a legal domain knowledge graph based on a multi-layered organizational knowledge architecture encompassing jurisprudential, statute, and case-level knowledge structures.

[0053] Specifically, the jurisprudential knowledge architecture, the legal provision-level knowledge architecture, and the case-level knowledge architecture are integrated to establish connections between knowledge nodes at different levels. For example, jurisprudential concepts are associated with relevant legal provisions, and legal provisions are associated with cases involving those provisions, forming a knowledge graph with a multi-level structure. The top layer can be jurisprudential knowledge, the middle layer is legal provision-level knowledge, and the bottom layer is case-level knowledge. Different levels are interconnected through relational edges to form an organic whole.

[0054] Furthermore, the legal-related elements in the legal theory-level knowledge architecture, the legal-related elements in the legal provision-level knowledge architecture, and the case-related elements in the case-level knowledge architecture are each treated as independent knowledge nodes. Knowledge nodes at different levels are interconnected through relational edges, and each relational edge corresponds to a correlation weight, which is used to describe the relationship between different nodes.

[0055] By integrating knowledge from three levels—legal theory, legal provisions, and case law—a complete and systematic legal knowledge framework is formed. Users can query the knowledge graph from different dimensions, such as searching for relevant legal provisions and cases based on legal concepts, or searching for related legal theories and provisions based on cases. Simultaneously, cross-level knowledge reasoning is also possible, improving the efficiency of resolving legal issues.

[0056] The method disclosed in this embodiment systematically organizes and integrates complex legal knowledge by constructing a multi-layered organizational knowledge architecture legal domain knowledge graph, making the organization and management of legal knowledge more efficient. The clear relationship structure and rich knowledge information in the knowledge graph provide strong support for the analysis and resolution of legal issues, enabling rapid location of relevant legal principles, provisions, and cases, and improving the accuracy of legal reasoning and decision-making. The knowledge graph presents legal knowledge in a visual way, facilitating sharing and communication among different users, and contributing to the popularization and dissemination of legal knowledge. It also provides a foundation for intelligent applications in the legal field, such as intelligent legal consultation and automatic generation of legal documents, promoting the digital transformation of the legal industry.

[0057] Furthermore, this application also includes: establishing a knowledge acquisition and updating mechanism. Specifically, this includes: performing automated extraction, using natural language processing technology to automatically extract structured knowledge from publicly available data such as legal texts and judicial judgments.

[0058] Furthermore, expert review can be implemented, with legal experts reviewing and supplementing the automatically extracted knowledge to ensure its accuracy. In the expert review step, legal experts review and supplement the automatically extracted knowledge based on laws and regulations, legal principles, and judicial practice experience.

[0059] Furthermore, dynamic updates can be designed, and an incremental update mechanism can be built to automatically update relevant knowledge nodes and relationships when laws and regulations change. In the dynamic update process, the official release channels of laws and regulations are monitored, and the incremental update mechanism is triggered when new laws and regulations are released or existing laws and regulations are revised.

[0060] Furthermore, version control can be implemented to maintain historical versions of the knowledge graph to support legal consultations at specific points in time. In the version control process, version numbers are used to record different versions of the knowledge graph, and each version contains knowledge nodes and relationship information at the corresponding point in time.

[0061] Furthermore, multimodal representation can be performed to transform legal knowledge into various forms such as vectors, symbolic rules, and text descriptions, facilitating interaction with different levels of the model; deep learning models can be used to transform legal knowledge into vector form, and logical reasoning rules can be used to transform legal knowledge into symbolic rule form.

[0062] Furthermore, in the hierarchical organization process, legal principles can be used as the highest level, legal provisions as the intermediate level, and case studies as the lowest level for knowledge organization.

[0063] Furthermore, association strength quantification can be performed, assigning weights to the relationship edges in the knowledge graph to represent the association strength between different legal concepts. In the association strength quantification step, the weights of the relationship edges are determined based on factors such as the frequency of citations and the degree of dependence between legal concepts.

[0064] Reference Figure 3 The method for S200 "obtaining a preset amount of manually labeled seed samples", i.e., the method for obtaining seed samples, includes: S210, randomly select a first preset amount of original samples from the legal database.

[0065] The first preset value is N1, where N1 > 500.

[0066] Random sampling ensures that the sample is broadly representative of the entire legal database, avoiding bias caused by human selection. The sample drawn in this way can reflect the overall characteristics of the legal database, making subsequent analysis and processing more universal. Randomly selecting samples from a large amount of legal data can cover various types of legal scenarios and cases, providing rich material for subsequent manual annotation and screening.

[0067] S220 involves manually labeling the extracted raw samples.

[0068] Specifically, the labeled content consists of the structural elements in the original sample.

[0069] Specifically, a professional annotation team can be assembled, comprising legal professionals and data annotators. Detailed annotation rules and guidelines should be developed, such as clearly defining the annotation types of legal clauses, case results, and legal entities. Annotators should then annotate each of the extracted original samples according to these rules. During the annotation process, specialized annotation tools, such as Label Studio, can be used to mark and categorize key information in the samples. For example, in a contract dispute case, annotators would annotate the type of contract, the cause of the dispute, and the judgment result.

[0070] Manual annotation can fully utilize the legal knowledge and experience of professionals to ensure the accuracy and reliability of the annotated information. Compared with automated annotation, manual annotation can better understand the complex semantics and context of legal texts, avoiding erroneous annotations caused by algorithmic misunderstandings. Different legal issues may require different annotation methods, and manual annotation can be flexibly adjusted according to specific needs and business scenarios to meet diverse annotation requirements.

[0071] S230, selects seed samples from the manually labeled original samples according to preset quality standards.

[0072] Among them, the diversity index corresponding to the uniformity of sample distribution in the seed sample is greater than 0.7, the sample covers at least 10 mainstream legal categories and the sample coverage is greater than 80%.

[0073] By screening through preset quality standards, we can ensure that the final seed samples are of high quality and representativeness. These samples meet certain requirements in terms of distribution uniformity, legal category coverage, and sample coverage, and can provide more valuable information for subsequent model training and data analysis. The number of seed samples selected is relatively small and of high quality, which can reduce the workload of subsequent processing and improve the efficiency of model training and data analysis.

[0074] The number of seed samples is N², where 100 ≤ N² ≤ 500. A preset range of 100-500 seed samples is set, with a typical value of 200. This is used to design a leveraged knowledge graph (KG), avoiding massive manual annotation and reducing workload by 95% compared to entirely manual annotation. This approach is suitable for bootstrapping processes.

[0075] Specifically, the preset quality standards include: 1) Accuracy: The labeling accuracy of seed samples must be greater than 95%, and the labeling must be ensured to be error-free through manual secondary verification; 2) Coverage: The samples must cover mainstream legal categories, such as civil law, criminal law, and administrative law, covering at least 10 categories, and the sample coverage must be greater than 80%; 3) Diversity index: The Shannon diversity index is used to calculate the uniformity of sample distribution, and the diversity index must be greater than 0.7; 4) Consistency: The Kappa coefficient of multiple experts is used to evaluate the consistency of sample labeling, and the consistency must be greater than 90% to ensure the reliability of seed samples and support the generation of high-quality templates.

[0076] This method efficiently guides data augmentation by identifying a small number of high-quality seed samples, reducing the cost of manual annotation (which is expensive), ensuring the accuracy of generated templates, and avoiding noise propagation. Experiments show that high-quality seeds can improve the accuracy of the final model by 10%. It draws on the principles of semi-supervised learning, leveraging a small amount of labeled data to augment unlabeled data, and has higher resource efficiency compared to fully supervised methods.

[0077] The method for "extracting structural features from seed samples" in S300 specifically includes: A100 performs word segmentation on each seed sample to obtain several independent sub-words.

[0078] Suppose we have a seed sample: "Zhang San sued Li Si in the Chaoyang District People's Court of Beijing, the cause of action being a contract dispute." After segmenting the text using a word segmentation tool, we get the independent sub-words: "Zhang San," "in," "Beijing," "Chaoyang District," "People's Court," "sued," "Li Si," ",," "cause of action," "is," and "contract dispute." Word segmentation is a fundamental step in natural language processing. Breaking the text down into independent sub-words facilitates more detailed analysis of each sub-word, such as extracting entity information and performing semantic understanding. Different sub-words carry different semantic information, and word segmentation can more clearly separate this information, providing convenience for subsequent processing.

[0079] A200 utilizes the named entity recognition function of NLP tools to extract entity information from several independent sub-words. The entity information includes one or more of the following: personal names, place names, organization names, legal provisions, and case types.

[0080] Based on the independent sub-words obtained in the previous step, we can identify the following entities: personal names as "Zhang San" and "Li Si"; place names as "Beijing Municipality" and "Chaoyang District"; organization names as "Chaoyang District People's Court"; and case type as "contract dispute". Specifically, we can use Named Entity Recognition (NER) technology, such as a pre-trained model (e.g., BERT combined with a CRF model), to accomplish this task. Entity information is a crucial and significant part of the text; extracting entity information helps us quickly locate the core content of the text and understand the key elements such as people, places, and organizations involved. In legal text analysis, entity information such as personal names, place names, and organization names is very important for case identification, association, and analysis, and can improve the efficiency of information retrieval and analysis.

[0081] A300 parses the text structure of seed samples and determines the logical feature type corresponding to each seed sample.

[0082] For example, the text structure of the seed sample mentioned above can be viewed as a logical structure of "person's behavior (lawsuit) + location (court) + cause of action". We can analyze the text structure by defining some rules or using machine learning models. For example, based on the keyword "lawsuit", we can determine that this is a text related to legal proceedings and determine that its logical feature type is "description of litigation behavior".

[0083] Understanding the logical feature types of text helps us grasp its content and intent as a whole. Different logical feature types correspond to different semantics and uses. By identifying the logical feature types, we can classify and organize text, facilitating subsequent information retrieval, reasoning, and decision-making.

[0084] A400 determines the complexity feature type corresponding to each seed sample based on the text structure of the seed sample and several independent sub-words.

[0085] The complexity feature type can be determined based on factors such as text length, vocabulary specialization, and sentence structure complexity. For the seed sample mentioned above, its sentence structure is relatively simple, and the vocabulary is mostly common legal terms, so its complexity feature type can be determined as "low complexity". If the text contains a large number of professional legal citations, complex legal reasoning, and nested sentence structures, its complexity feature type can be determined as "high complexity".

[0086] Determining the complexity feature type can help us better understand the difficulty and processing cost of text. Different processing strategies can be adopted for texts with different complexities.

[0087] The A500 uses a topic modeling algorithm to analyze seed samples and identify potential topics in each seed sample.

[0088] Multiple seed samples are analyzed using topic modeling algorithms (such as Latent Dirichlet Allocation, LDA). Suppose we have a set of seed samples containing descriptions of multiple legal cases. The LDA algorithm can categorize these samples into different topics, such as "contract disputes," "tort disputes," and "labor disputes." For the aforementioned seed samples, by analyzing their independent sub-words and text content, the LDA algorithm can determine their topic as "contract disputes."

[0089] Topic modeling can help us discover hidden topic structures in text, summarize and generalize large amounts of textual information; by identifying potential topics, we can classify and cluster texts by topic, making it easier for users to quickly find texts of interest.

[0090] The method for "generating target templates based on structural features" in S300 specifically includes: B100 uses the K-means clustering algorithm to group seed samples based on their entity information, logical feature type, complexity feature type, and potential topic, grouping seed samples with similar features into the same cluster. B200 extracts a general framework from each cluster and replaces specific entity information with placeholders; B300 generates target templates based on preset rule engines.

[0091] In one specific embodiment, we collected a series of legal-related seed samples. Each sample contains entity information, logical feature types, complexity feature types, and potential topics. The entity information involves legal subjects (such as plaintiff Zhang San and defendant Li Si), legal provisions (such as Article 107 of the Contract Law), and legal events (such as breach of contract and tort). The logical feature types can be divided into liability determination, compensation calculation, and evidence acceptance. The complexity feature types include simple single legal questions (such as "What liability should be borne for breach of contract?") and complex multi-legal relationship issues (such as "How should liability be divided and compensation calculated in the case of breach of contract and tort?"). The potential topics include contract disputes, tort disputes, and labor disputes. First, the K-means clustering algorithm calculates the similarity between samples based on these features, potentially forming the following clusters: the first cluster contains all samples related to liability determination in contract breach; the second cluster contains samples for compensation calculation in tort disputes; the third cluster contains samples for evidence acceptance in labor disputes; and the fourth cluster contains samples of complex issues involving multiple legal relationships. Given the vast amount of information in the legal field, clustering similar seed samples into one category effectively integrates large amounts of scattered data, facilitating subsequent centralized processing and helping to quickly focus on different types of legal issues. For example, contract breach issues are concentrated in one cluster, making it easier to conduct in-depth analysis and processing of such issues later.

[0092] Next, taking the first cluster (samples of liability determination related to contract breach) as an example, samples within the cluster might include: "In the sales contract signed between Zhang San and Li Si, Zhang San breached the contract; what liability should Zhang San bear?" or "In the lease contract between Wang Wu and Zhao Liu, Wang Wu breached the contract; what liability should Wang Wu bear?" We replace specific entity information (such as "Zhang San," "Li Si," "sales contract," and "lease contract") with placeholders to obtain a general framework: "In the [contract type] signed between [Party A] and [Party B], [Party A] breached the contract; [Party A] should bear [responsibility inquiry method]." Abstracting specific legal entities and contract types into placeholders allows the general framework to cover more similar legal scenarios, enhancing the template's adaptability and facilitating the generalization of legal knowledge. It is no longer limited to specific cases but can be applied to a wider range of contract breach liability determination issues.

[0093] Finally, several target templates are generated based on the preset rule engine. For example, Rule 1 includes: if the general framework involves determining liability for breach of contract and includes a method for inquiring about liability, then a target template for "determining the liability of [the breaching party] in breach of [contract type]" is generated. Rule 2 includes: if the general framework is for calculating compensation in tort disputes and includes statements related to compensation, then a target template for "calculating the amount of compensation for [the infringing party] in [tort type]" is generated. For the general framework obtained in step B200, "[Party A] breached the contract in [contract type] signed between [Party A] and [Party B], and [Party A] should bear [method of inquiring about liability]", according to Rule 1, the target template "determining the liability of [Party A] in breach of [contract type]" is generated. The rule engine ensures the standardization and normalization of target template generation, avoids inconsistencies that may occur when manually generating templates, and can quickly and accurately generate target templates based on the general framework, improving the efficiency of legal issue processing and analysis.

[0094] In situations involving numerous and complex legal cases, this solution significantly reduces the time and effort legal professionals spend dealing with problems and improves work efficiency through automated clustering and template generation processes. The generated target templates can be reused in different legal scenarios, promoting the effective reuse of legal knowledge and avoiding repetitive work. Standardized target templates help legal professionals analyze and handle legal issues more clearly, providing strong support for legal decision-making.

[0095] Specifically, when the preset rule engine is a legal question-and-answer rule engine, the method for generating the target template based on the preset rule engine includes: B311, when the preset rule engine is legal question-and-answer rules, performs detailed analysis on question-and-answer pairs in the same cluster, identifying the question content, question type (such as factual inquiry, legal application inquiry, liability determination inquiry, etc.), answer structure, and answer content (such as simple statement, step-by-step solution, legal provision citation, etc.). B312. Based on the type of question, abstract the content of the question to form a general question pattern; (for example, abstract "Does Zhang San's behavior constitute fraud?" into "Does [the subject's] [behavior] constitute [crime]?" B313. Based on the answer structure and content, a general solution framework can be formed; (For example, the answer "According to Article 266 of the Criminal Law of the People's Republic of China, anyone who defrauds public or private property of a relatively large amount shall be sentenced to fixed-term imprisonment of not more than three years, criminal detention, or public surveillance, and may also be fined. Zhang San's behavior meets the provisions of this article, therefore constituting the crime of fraud." can be abstracted as "According to [legal article number], [behavior description], [legal consequences]. The [subject's] [behavior] meets the provisions of this article, therefore constituting [crime].") B314 combines general question patterns and general answer frameworks into a legal Q&A template.

[0096] When the preset rule engine is a case analysis rule, the specific methods for generating the target template based on the preset rule engine include: B321, when the preset rule engine is case analysis rule, performs structural analysis on sample cases in the same cluster to obtain the logical relationship between each part; B322 outlines the logical relationships between the various parts, forming a general case analysis framework; for example, "In the context of the case, [Subject 1] and [Subject 2] have reached a [point of contention]; [Subject 1] provided [Evidence 1], and [Subject 2] provided [Evidence 2]. Based on the provisions of [Legal Article Number] and considering the actual circumstances of this case, the court rendered a [Judgment]." When the preset rule engine is a text interpretation rule, the specific methods for generating the target template based on the preset rule engine include: B331, when the preset rule engine is the clause interpretation rule, analyzes the clause structure in the same cluster to obtain clause association information; the clause association information includes the division of clauses, paragraphs and items, as well as the main idea, scope of application and specific provisions of the clauses; B332 extracts key elements for interpreting the provisions from related information, such as the legislative purpose of the provisions, the definition of core concepts, applicable conditions, and exceptions. B333 defines the logical relationships between interpretive elements and constructs a logical framework for interpretation. For example, it first states the legislative purpose, then explains the core concepts, and finally elaborates on the applicable conditions and exceptions. B334 replaces the specific article content and explanatory information with placeholders to generate an article explanation template. For example, "[Article Number] has the legislative purpose of [Purpose]. Among them, [Core Concept] refers to [Definition]. This article applies to the situation of [Applicable Conditions], except for [Exceptions]." When the preset rule engine is for contract drafting rules, the specific methods for generating the target template based on the preset rule engine include: B341, when the preset rule engine is the contract drafting rule, analyzes the contract structure in the same cluster, which generally includes the contract header (contract name, party information), contract body (clause content), and contract footer (signature, date, etc.).

[0097] B342, Clause Classification and Abstraction: Classify the clauses in the main body of the contract, such as subject matter clauses, price clauses, performance clauses, and liability for breach of contract clauses; replace the specific content in each clause with placeholders to form a general clause template, for example, "[Party A] shall deliver the [subject matter] to [Party B] in accordance with the [delivery method] at the [delivery time]. B343, Overall Contract Framework Construction: Combine common clause templates into a contract drafting template according to a certain logical order, while retaining the basic structure of the contract's preamble and conclusion.

[0098] When the preset rule engine is a case-based reasoning rule, the specific methods for generating the target template based on the preset rule engine include: B351, when the preset rule engine is case-based reasoning rule, analyzes sample cases in the same cluster and extracts key elements, including case facts, legal disputes, reasoning, and judgment results.

[0099] B352, Logical Reasoning Analysis: This section analyzes the logical reasoning in case precedents, from identifying the facts of the case to the legal issues, then to the explanation of the reasoning behind the judgment and the conclusion of the judgment. For example, by analyzing the facts of the case, the legal issues are identified, and then reasoning is conducted based on relevant legal provisions and jurisprudence to arrive at the judgment.

[0100] B353, Construction of a General Reasoning Pattern: Replace the specific elements and reasoning processes in case law with placeholders to construct a general case law reasoning template, such as, "Given the [facts of the case], the legal issue in this case is [issue]; based on [legal provision number] and [legal basis], and in light of the actual circumstances of this case, the court has rendered [judgment]." Reference Figure 4 The S400 method of "populating several target templates based on a legal domain knowledge graph to generate diverse training samples," i.e., the method for generating diverse training samples, includes: S410, based on the type of each target template, obtain at least two related subgraphs from the legal domain knowledge graph; S420, Filter out several similar entities that match the placeholder type from each associated subgraph; S430, randomly select one from several similar entities as the target filling entity, and fill the corresponding placeholder in the target template based on the target filling entity; S440: After all the associated subgraphs have been filled for each target template, obtain all the relationship chain path information associated with the target filling entity based on all the associated subgraphs. S450, add all relationship chain path information to the corresponding target template after filling to obtain the optimized sample; All optimized samples form a diverse training sample.

[0101] Suppose we have a target template: "[Plaintiff] sues [Defendant] on the grounds of [Dispute Type]". This is a template about legal proceedings. A legal domain knowledge graph is a graph structure containing various legal entities (such as plaintiff, defendant, dispute type, etc.) and their relationships. Based on the type of this template (legal proceedings type), we can obtain related subgraphs from the legal domain knowledge graph. For example, one related subgraph could be centered on "contract disputes," containing entities such as plaintiffs and defendants involved in contract disputes and the relationships between them; another related subgraph could be centered on "tort disputes." Different related subgraphs represent different legal scenarios and knowledge sets. By obtaining multiple related subgraphs, we can provide a rich and diverse source of entities for subsequent template filling, thereby generating diverse training samples and avoiding sample homogeneity.

[0102] Continuing with the template above, in the subgraph centered on "contract disputes," for the placeholder "[Plaintiff]," we filter out similar entities matching the plaintiff type, such as "Company A" and "Enterprise B." For the placeholder "[Defendant]," we filter out "Company C" and "Individual D." For the placeholder "[Dispute Type]," since the subgraph is centered on contract disputes, we filter out "Sales Contract Dispute" and "Lease Contract Dispute." In the subgraph centered on "Tort Disputes," similar filtering is performed, such as "Individual E" as the plaintiff, "Company F" as the defendant, and "Portrait Right Infringement Dispute" and "Reputation Right Infringement Dispute" as the dispute type. This step ensures that the entities filled into the template match the type of the placeholders, guaranteeing the rationality and logic of the generated training samples. This results in samples that better reflect real-world legal scenarios, helping to improve the accuracy of the model in practical applications.

[0103] Then, Company A is randomly selected from "Company A" and "Enterprise B" as the target entity for "[Plaintiff]"; Company C is randomly selected from "Company C" and "Individual D" as the target entity for "[Defendant]"; and "Sales Contract Dispute" is randomly selected from "Sales Contract Dispute" and "Lease Contract Dispute" as the target entity for "[Dispute Type]". These target entities are then filled into the template to obtain "Company A sues Company C on the grounds of a sales contract dispute". Randomly selecting entities increases the randomness and diversity of the samples. Different random combinations can generate a large number of different training samples, enabling the model to learn more different legal scenarios and expressions, thereby improving the model's generalization ability.

[0104] After completing the template filling, based on the association subgraph centered on "contract dispute," we discovered a relationship chain path between "Company A" and "Company C"—"contract signing - contract breach - lawsuit." Simultaneously, other related relationship information may exist in the association subgraph, such as a client relationship between "Company A" and a law firm, or an association between "Company C" and a witness. All of this relationship chain path information was extracted. This relationship chain path information provides the model with richer contextual information; it not only tells the model the surface relationships between entities but also reveals the underlying logic and causal relationships, helping the model to understand the legal scenario more deeply and improving its reasoning and judgment capabilities.

[0105] Finally, the relationship chain information, such as "Company A signs a contract with Company C - Company C breaches the contract - Company A hires a law firm to sue Company C," is added to the filled template "Company A sues Company C on the grounds of a sales contract dispute," resulting in the optimized sample: "Company A sues Company C on the grounds of a sales contract dispute; Company A signs a contract with Company C - Company C breaches the contract - Company A hires a law firm to sue Company C." All optimized samples form a diverse training sample. The optimized sample contains richer information, enabling the model to learn more comprehensive legal knowledge and logical relationships. Such samples can improve the model's performance, allowing it to make more accurate judgments and decisions when dealing with real-world legal issues.

[0106] The method disclosed in this embodiment generates highly diverse training samples by randomly selecting entities from multiple related subgraphs for filling and adding relationship chain path information. This enables the model to learn more different legal scenarios and expressions, avoids overfitting, and improves the model's generalization ability. The addition of relationship chain path information provides the model with richer context and logical relationships, which helps the model to understand legal scenarios more deeply and improves the model's reasoning and judgment abilities. The diversified and optimized training samples enable the model to perform better when dealing with actual legal issues, improve the model's accuracy and reliability, and provide more effective support for applications in the legal field.

[0107] Furthermore, this application also includes: 1) Logical consistency check: checking whether the filled sample is logically consistent. For example, in a question-and-answer sample, checking whether the logic between the answer and the question is coherent, and whether the cited legal provisions are applicable to the situation described in the question. 2) Semantic reasonableness check: verifying whether the filled sample is semantically reasonable. For example, checking whether the description of the subject's behavior is reasonable, and whether the determination of legal consequences complies with legal provisions.

[0108] Furthermore, in key legal areas, legal experts can be introduced to conduct sampling reviews of the generated samples, continuously optimize the generation strategy, and then train the large language model based on the optimized strategy, using the trained large language model as the legal big model.

[0109] For example, within the legal system, there are certain areas that have a significant impact on society, involve complex legal relationships, or are emerging and can be defined as key legal areas. For instance, financial law involves numerous complex financial transaction rules and regulatory requirements; intellectual property law faces new challenges as technology advances; and medical law concerns public health and medical industry standards. These key areas can be identified based on the importance, complexity, and development trends of the legal matters involved.

[0110] Specifically, methods such as random sampling and stratified sampling can be employed. For example, in stratified sampling, the generated samples are stratified according to case type, legal clause category, etc., and then a certain number of samples are drawn from each stratum. Legal experts develop detailed review criteria, including the legality, accuracy, logic, and completeness of the sample content. For example, they review whether the legal provisions cited in the samples are accurate and whether the legal reasoning is reasonable. A review team is formed, composed of legal experts with extensive experience and expertise in key legal areas; these experts may include lawyers, judges, law professors, etc. Legal experts carefully review the generated samples according to the sampling review plan, recording any problems found, such as factual errors, improper application of law, and reasoning loopholes, and providing corresponding modification suggestions. The problems discovered by the legal experts are categorized and statistically analyzed, for example, by frequency of occurrence, severity, and the legal field to which the problems belong. The causes of the problems are analyzed: are they defects in the generation strategy itself, insufficient training data, or an inaccurate understanding of legal knowledge? Based on the review results and problem analysis, the sample generation strategy is adjusted. For example, if misinterpretations of certain legal provisions are frequently found in the generated samples, the generation rules can be adjusted to enhance the learning and application of those provisions. Supplementing or revising the training data can also improve the quality of the generated samples. For instance, adding typical cases or the latest legal regulations as training data can be helpful.

[0111] This approach allows for a focus on critical and complex legal issues, improving the relevance and efficiency of the review process and ensuring that the generated samples better meet the needs of actual legal practice. Scientific sampling methods and clear review standards guarantee the objectivity and accuracy of the review results, providing a reliable basis for subsequent analysis and optimization. Legal experts, leveraging their professional knowledge and practical experience, can identify potential problems in the generated samples and provide high-quality modification suggestions, enhancing the quality and credibility of the samples. In-depth analysis of the review results can accurately pinpoint the root causes of problems, providing a clear direction for optimizing the generation strategy. Continuously adjusting the generation rules and updating training data enables the generation strategy to adapt to changes and needs in the legal field, improving the quality and diversity of the generated samples.

[0112] Through review by legal experts and optimization of generation strategies, the generated samples will be significantly improved in terms of legality, accuracy, and logic, providing high-quality data support for the training of the legal big data model. The generated samples are closer to actual legal business, enabling the trained legal big data model to better adapt to different legal tasks and scenarios, improving the model's practicality and generalization ability. This fully leverages the professional knowledge of legal experts and the generation capabilities of artificial intelligence, achieving complementary advantages between humans and machines and promoting the intelligent development of the legal field.

[0113] The S500 method, which "configures a multi-level sample generation strategy based on diverse training samples and model training requirements, trains a large language model based on the multi-level sample generation strategy, and uses the trained large language model as a legal big model," specifically includes: C100 categorizes diverse training samples into easy, medium, and high difficulty levels based on model training requirements.

[0114] For samples of simple difficulty, basic interpretations of common legal provisions can be selected, such as the interpretation of the basic provisions on marriage in the Civil Code. These samples are clear in content and explicit in concepts, making them relatively easy to understand and handle. Samples of medium difficulty can be more complex case analyses, such as traffic accident cases involving the determination of multiple parties' liability. These cases involve the cross-application of various legal provisions and require a certain level of logical reasoning and comprehensive application of legal knowledge. Samples of high difficulty can be cutting-edge cases with legal controversies, such as legal issues arising from emerging technologies (e.g., the copyright ownership of works created by artificial intelligence). These samples require a deep understanding of legal principles and a forward-looking judgment of legal development trends.

[0115] By grading the difficulty of training samples, the model can gradually learn different levels of knowledge and skills. Easy-difficulty samples help the model establish basic legal concepts and knowledge systems, laying a solid foundation for learning more complex content later. Medium-difficulty samples train the model's logical reasoning and knowledge integration abilities, enabling it to handle common complex legal issues. High-difficulty samples challenge the model's limits, prompting it to continuously evolve to cope with various complex and ever-changing legal scenarios in reality.

[0116] C200 introduces noise and mutation operations into samples of medium difficulty to generate intermediate layer samples.

[0117] The noise can be randomly inserted words or symbols that do not affect the overall semantics, and the mutation operations include synonym replacement, sentence recombination, etc. For example, replacing "contract" with "entity" in the sample, or adjusting the word order of the sentence, can increase the diversity of the sample and the robustness of the model.

[0118] Introducing noise and mutation operations can enhance the robustness and generalization ability of the model. In real life, the description of legal issues often contains a lot of irrelevant information, and the model needs to be able to accurately extract key content from this information. At the same time, the specific circumstances of legal issues are also ever-changing. By performing mutation operations on the samples, the model can learn the handling methods in different situations and improve its ability to cope with complex and ever-changing legal scenarios.

[0119] C300 uses easy-difficulty samples and high-difficulty samples as basic layer samples and advanced layer samples, respectively. Determine the training proportions of base layer samples, intermediate layer samples, and high-level layer samples.

[0120] Determining the training proportion of samples at different levels allows for the rational allocation of training resources, enabling the model to receive targeted training at different stages. A certain proportion of basic layer samples ensures that the model consistently consolidates and reinforces its fundamental knowledge, preventing forgetting. A larger proportion of intermediate layer samples allows the model sufficient opportunities to comprehensively apply knowledge and improve its capabilities. An appropriate proportion of advanced layer samples can stimulate the model's potential, enabling it to continuously tackle more challenging tasks, thereby improving the overall performance of the model.

[0121] C400 determines the sample set for each training session based on the training ratio, performs hierarchical training on the large language model according to the sample hierarchy, and uses the trained large language model as the legal large model.

[0122] Assuming we use 100 samples per training session, with a training ratio of 3:5:2, each training sample set will contain 30 basic layer samples, 50 intermediate layer samples, and 20 advanced layer samples. In this tiered training process, the basic layer samples are used first to train the model, allowing it to learn basic legal knowledge and concepts. Then, intermediate layer samples are used for training. At this point, the model has a certain foundation and can begin to handle more complex problems. Training with intermediate layer samples further enhances its logical reasoning and knowledge integration capabilities. In this stage, the model needs to learn to extract key information from samples with interfering information and correctly understand and process modified legal expressions. Finally, advanced layer samples are used for training, pushing the model to its limits to handle highly difficult legal problems. Advanced layer samples typically involve the integrated application of multiple legal concepts and complex logical reasoning. Through training with these samples, the model can better handle complex problems in real-world legal scenarios. After multiple rounds of such tiered training, a well-trained legal model is finally obtained.

[0123] Layered training allows the model to learn step by step from easy to difficult, which aligns with human cognitive patterns. Through training at different levels, the model can continuously consolidate and strengthen its learned knowledge while gradually improving its ability to handle complex problems. This targeted training method can improve training efficiency and avoid overfitting or learning difficulties caused by excessive difficulty during training. This enables the model to better adapt to legal scenarios of varying difficulty, thereby improving the practicality and accuracy of the legal big data model.

[0124] The method disclosed in this embodiment, through a multi-level sample generation strategy and hierarchical training, enables the model to comprehensively learn legal knowledge and skills at different difficulty levels, from basic concepts to complex problem handling, gradually improving its comprehensive capabilities and enabling it to perform well in various legal scenarios. Introducing noise and mutation operations to generate intermediate-level samples, and rationally allocating the training proportion of samples at different levels, all contribute to enhancing the model's robustness and generalization ability, allowing it to better adapt to complex and ever-changing legal issues in real life. The hierarchical training method aligns with human cognitive patterns, allowing the model to learn in a targeted manner, avoiding the waste of resources and inefficiency caused by blind training, thereby improving training efficiency and shortening the training cycle. A well-trained legal model can more accurately handle legal issues of varying difficulty, providing more reliable legal support and services for legal professionals and ordinary users, thus possessing higher practicality and application value.

[0125] During model training, continuously monitor the model's performance metrics, such as accuracy, recall, and F1 score. Dynamically adjust the multi-level sample generation strategy based on the model's performance. If the model's accuracy is low during training, it indicates that the model may be insufficient in handling complex situations; in this case, the proportion of high-level layer samples can be appropriately increased. Conversely, if the model exhibits overfitting on simple tasks, the proportion of basic layer samples can be appropriately reduced.

[0126] Secondly, this application discloses a legal information analysis method, including: Identify the legal information to be analyzed; Input legal information into the target large model to generate feedback information; The target large model is a legal vertical large model that has been fine-tuned using the legal large model construction method disclosed in the first aspect of this application.

[0127] Thirdly, this application discloses a legal big data model construction system for executing the legal big data model construction method disclosed in the first aspect of this application. The system includes: The building module is used to construct a knowledge graph in the legal field; the knowledge graph in the legal field includes a multi-layered knowledge architecture. The seed sample acquisition module is used to acquire a preset amount of manually labeled seed samples; The target template generation module is used to extract the structural features of the seed sample and generate several target templates based on the structural features. The fill module is used to fill in several target templates based on a legal domain knowledge graph to generate diverse training samples; The training module is used to configure multi-level sample generation strategies based on diverse training samples and model training requirements. The large language model is trained based on the multi-level sample generation strategies, and the trained large language model is used as the legal big model.

[0128] A computer device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0129] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the legal big data model construction method or legal information analysis method of the foregoing embodiments of this disclosure.

[0130] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.

[0131] like Figure 5 This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 5 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0132] like Figure 5 As shown, a computer device may include a processor (such as a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0133] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 5 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.

[0134] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the legal big data model construction method or legal information analysis method of embodiments of this disclosure are performed.

[0135] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0136] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When these non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the legal big data model construction method or legal information analysis method described in the foregoing embodiments of the present disclosure are performed.

[0137] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0138] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0139] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0140] In this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, devices, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used herein refer to the terms "and / or," and are used interchangeably with them unless the context clearly indicates otherwise. The term "such as" as used herein refers to the phrase "such as but not limited to," and is used interchangeably with it.

[0141] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0142] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.

[0143] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0144] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0145] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for constructing a large-scale legal model, characterized in that, include: Constructing a knowledge graph in the legal field; The legal knowledge graph includes a multi-layered organizational knowledge architecture; Obtain a preset amount of manually labeled seed samples; Extract the structural features of the seed sample, and generate several target templates based on the structural features; The step of generating several target templates based on the structural features includes: grouping seed samples according to the entity information, logical feature type, complexity feature type, and potential topic corresponding to each seed sample using the K-means clustering algorithm, and grouping seed samples with similar features into the same cluster; extracting a general framework from each cluster and replacing the specific entity information with placeholders; and generating several target templates according to a preset rule engine. The method involves filling several target templates with the legal domain knowledge graph to generate diverse training samples. Specifically, this includes: obtaining at least two related subgraphs from the legal domain knowledge graph based on the type of each target template; selecting several similar entities matching the placeholder type from each related subgraph; randomly selecting one of the similar entities as the target filling entity and filling the corresponding placeholder in the target template based on the target filling entity; after filling all the related subgraphs for each target template, obtaining all relationship chain path information associated with the target filling entity based on all the related subgraphs; adding all the relationship chain path information to the corresponding filled target template to obtain optimized samples; and forming diverse training samples from all optimized samples. Based on the diverse training samples and model training requirements, a multi-level sample generation strategy is configured. The large language model is trained based on the multi-level sample generation strategy, and the trained large language model is used as the legal large model.

2. The legal large-scale model construction method according to claim 1, characterized in that, The construction of the legal domain knowledge graph includes: The constituent elements of the legal field's map are determined. These constituent elements include jurisprudential-level related elements, legal provision-level related elements, and case-level related elements. The jurisprudential-level related elements include basic concepts in the legal field, the legal provision-level related elements include legal provisions in the legal field, and the case-level related elements include related cases in the legal field and the relationships between legal subjects in the legal field. A jurisprudential knowledge architecture is constructed based on the basic concepts in the legal field. The jurisprudential knowledge architecture includes several basic concept knowledge nodes and the relationship edges between different basic concept knowledge nodes. A legal provision-level knowledge architecture is constructed based on the legal provisions of the legal field. The legal provision-level knowledge architecture includes several legal provision-type knowledge nodes and the relationship edges between different legal provision-type knowledge nodes. A case-level knowledge architecture is constructed based on the related cases in the legal field and the legal subject relationships in the legal field. The case-level knowledge architecture includes several case-type knowledge nodes and the relationship edges between different case-type knowledge nodes. A legal domain knowledge graph is constructed based on the aforementioned jurisprudential knowledge architecture, the aforementioned legal provision-level knowledge architecture, and the aforementioned case-level knowledge architecture, forming a multi-layered organizational knowledge architecture.

3. The method for constructing a large legal model according to claim 1, characterized in that, The acquisition of a preset amount of manually labeled seed samples includes: Randomly select a first preset number of original samples from the legal database; the first preset number is N1, N1 > 500; The extracted original samples were manually labeled; Seed samples are selected from the manually labeled original samples according to preset quality standards; The seed sample must have a diversity index greater than 0.7 corresponding to the uniformity of sample distribution, cover at least 10 mainstream legal categories, and have a sample coverage greater than 80%. The number of seed samples is N2, where 100 ≤ N2 ≤ 500.

4. The method for constructing a large legal model according to claim 3, characterized in that, The extraction of structural features from the seed sample includes: Each seed sample is segmented to obtain several independent sub-words; Entity information is extracted from several independent sub-words, and the entity information includes one or more of the following: personal name entities, place name entities, organization name entities, legal provision number entities, and case type entities; The text structure of the seed samples is parsed to determine the logical feature type corresponding to each seed sample; Based on the text structure of the seed samples and several independent sub-words, determine the complexity feature type corresponding to each seed sample; A topic modeling algorithm is used to analyze the seed samples and identify potential topics in each seed sample.

5. The legal large-scale model construction method according to claim 4, characterized in that, The target templates are legal Q&A templates, case analysis templates, clause interpretation templates, contract drafting templates, or case reasoning templates.

6. The method for constructing a large legal model according to claim 5, characterized in that, When the preset rule engine is a legal question-and-answer rule engine, the step of generating the target template according to the preset rule engine includes: When the preset rule engine is a legal question-and-answer rule, the question-and-answer pairs in the same cluster are analyzed in detail to identify the question content, question type, answer structure, and answer content. The problem content is abstracted based on the problem type to form a general problem pattern; Based on the answer structure and the answer content, a general solution framework is formed; The general question pattern and the general answer framework are combined into a legal Q&A template.

7. The method for constructing a large legal model according to claim 5, characterized in that, The process involves configuring a multi-level sample generation strategy based on diverse training samples and model training requirements, training a large language model based on this strategy, and using the trained large language model as a legal big model. This includes: Based on the model training requirements, the diverse training samples are divided into easy difficulty samples, medium difficulty samples, and high difficulty samples. Noise and mutation operations are introduced into the medium-difficulty samples to generate intermediate layer samples; The easy difficulty class samples and the high difficulty class samples are respectively used as the base layer samples and the advanced layer samples; Determine the training proportions of the base layer samples, the intermediate layer samples, and the high-level layer samples; The sample set for each training session is determined based on the training ratio. The large language model is then trained hierarchically according to the sample hierarchy, and the trained large language model is used as the legal large model.

8. A method for analyzing legal information, characterized in that, include: Identify the legal information to be analyzed; The legal information is input into the target large model to generate feedback information; The target large model is a legal vertical large model that has been fine-tuned using the legal large model construction method described in any one of claims 1-7.

9. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the legal big data model construction method according to any one of claims 1-7 or the legal information analysis method according to claim 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the legal big data model construction method of any one of claims 1-7 or the legal information analysis method of claim 8.

11. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the legal big data model construction method according to any one of claims 1-7 or the legal information analysis method according to claim 8.

Citation Information

Patent Citations

  • Large language model reliable legal question and answer generation method based on knowledge fine tuning

    CN118210891A

  • Data enhancement method and system, electronic equipment and computer storage medium

    CN119322847A