Target vertical large model fine-tuning method and system, computer device, and medium
Patent Information
- Application Number
- CN202511146938.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-08-15
AI Technical Summary
[0005]有鉴于此,本公开实施例提供了一种目标垂类大模型微调方法及系统、计算机装置、介质,能够解决现有技术中存在的模型微调复杂、与目标领域的适配性差等问题
[0011]The method for fine-tuning a large-scale target domain model disclosed in this application includes: constructing a multi-level target domain knowledge graph, which can systematically and structurally organize the knowledge of the target domain; then, inputting several target domain concept words obtained from the target domain knowledge graph into the large language model and tracking the activation state of the model to establish an explicit correspondence between target domain concepts and model neurons, allowing the model to accurately focus on target domain knowledge, avoiding wasting computational resources on irrelevant information, and greatly improving knowledge utilization efficiency; secondly, obtaining the neuronal regions in the large language model that are sensitive to target domain knowledge based on the explicit correspondence, focusing on sensitive neuronal regions allows the model to process target domain knowledge more accurately; and then, determining a sparse activation strategy based on the complexity of the task to be analyzed, a suitable sparse activation strategy allows the model to better capture target domain knowledge. The model identifies the characteristics of domain knowledge. Finally, based on a sparse activation strategy, the actual activation ratio of neurons in the neuron regions is dynamically adjusted. The adjusted large language model is then used as the target vertical category model. During graph-guided model fine-tuning, the activation ratio of parameters in the neuron regions is dynamically adjusted based on a preset sparse activation strategy. This allows the model to better capture the characteristics of target domain knowledge and enhance its ability to represent and understand knowledge. By using graph-guided model fine-tuning, neuron regions sensitive to target domain knowledge are identified in advance, and the parameter activation ratio is dynamically adjusted. This allows the model to update parameters more specifically during training, effectively avoiding blind adjustments to the entire model's parameters. This optimizes the model's performance in the target domain, enabling it to more accurately understand questions, generate high-quality answers, and improve the quality of task completion when handling tasks in the target domain.
Smart Images

Figure CN121094050B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method and system for fine-tuning a large target vertical model, a computer device, and a medium. Background Technology
[0002] In recent years, large-scale legal modeling technology has experienced rapid development. Open-source models such as DeepSeek-R1, QWEN, and LLAMA have demonstrated powerful capabilities in natural language understanding and generation. However, the application of these general-purpose models in specialized vertical fields such as law has been less than satisfactory. Specifically, this manifests in several ways: insufficient understanding of legal knowledge; while possessing some legal knowledge, the depth of understanding of legal terminology, concepts, and their complex relationships is inadequate, often leading to misinterpretations or inappropriate applications; a lack of legal reasoning ability; insufficient training in professional skills such as the application of legal provisions, case analysis, and legal reasoning; frequent professional errors; a tendency to generate "illusions" when dealing with legal issues, resulting in content that violates legal provisions; and poor adaptability, making it difficult to adapt to changes in legal systems across different countries and periods.
[0003] To address the challenges of applying general-purpose large-scale models in the legal field, various solutions have been proposed. Full-scale fine-tuning involves collecting a large legal corpus to fine-tune all parameters of the model, as seen in models like ChatLaw. However, this method requires a large amount of labeled data and computational resources, resulting in extremely high costs. Retrieval-enhanced generation leverages external knowledge bases to improve model output, but it relies heavily on retrieval quality and struggles to handle complex legal reasoning. Expert system integration combines traditional legal expert systems with large-scale models; however, rule maintenance is costly, and it struggles to handle boundary cases. Efficient parameter fine-tuning uses techniques like LoRA to fine-tune only a subset of parameters to reduce resource consumption, but simply performing efficient parameter fine-tuning is insufficient to meet the specialized needs of the legal field.
[0004] Most existing methods use unstructured text for direct training, failing to fully utilize the structural characteristics of legal knowledge; third, in scenarios with limited data for fine-tuning, the model is prone to overfitting training examples or forgetting its original capabilities; fourth, legal tasks are diverse, covering consultation, drafting, review, dispute resolution, etc., but existing methods are difficult to adapt to multiple legal tasks simultaneously under low resource conditions. Summary of the Invention
[0005] In view of this, the present disclosure provides a method and system for fine-tuning a large model of a target vertical category, a computer device, and a medium, which can solve the problems of complex model fine-tuning and poor adaptability to the target domain in the prior art.
[0006] In a first aspect, embodiments of this disclosure provide a method for fine-tuning a large target vertical model, comprising: constructing a multi-level target domain knowledge graph; inputting several target domain concept words obtained from the target domain knowledge graph into a large language model and tracking the activation state within the model to establish an explicit correspondence between target domain concepts and model neurons; obtaining neuronal regions in the large language model that are sensitive to target domain knowledge based on the explicit correspondence; determining a sparse activation strategy based on the complexity of the task to be analyzed; dynamically adjusting the actual activation ratio of neurons in the neuronal regions based on the sparse activation strategy, and using the adjusted large language model as the large target vertical model.
[0007] Secondly, this disclosure also provides a legal information analysis method, including: determining the legal information to be analyzed; inputting the legal information into a target large model to generate feedback information; the target large model is a large model fine-tuned using the target vertical large model fine-tuning method.
[0008] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution: The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform either the target vertical category large model fine-tuning method or the legal information analysis method described above.
[0009] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to execute any of the target vertical category large model fine-tuning methods or the legal information analysis methods described above.
[0010] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.
[0011] The method for fine-tuning a large-scale target domain model disclosed in this application includes: constructing a multi-level target domain knowledge graph, which can systematically and structurally organize the knowledge of the target domain; then, inputting several target domain concept words obtained from the target domain knowledge graph into the large language model and tracking the activation state of the model to establish an explicit correspondence between target domain concepts and model neurons, allowing the model to accurately focus on target domain knowledge, avoiding wasting computational resources on irrelevant information, and greatly improving knowledge utilization efficiency; secondly, obtaining the neuronal regions in the large language model that are sensitive to target domain knowledge based on the explicit correspondence, focusing on sensitive neuronal regions allows the model to process target domain knowledge more accurately; and then, determining a sparse activation strategy based on the complexity of the task to be analyzed, a suitable sparse activation strategy allows the model to better capture target domain knowledge. The model identifies the characteristics of domain knowledge. Finally, based on a sparse activation strategy, the actual activation ratio of neurons in the neuron regions is dynamically adjusted. The adjusted large language model is then used as the target vertical category model. During graph-guided model fine-tuning, the activation ratio of parameters in the neuron regions is dynamically adjusted based on a preset sparse activation strategy. This allows the model to better capture the characteristics of target domain knowledge and enhance its ability to represent and understand knowledge. By using graph-guided model fine-tuning, neuron regions sensitive to target domain knowledge are identified in advance, and the parameter activation ratio is dynamically adjusted. This allows the model to update parameters more specifically during training, effectively avoiding blind adjustments to the entire model's parameters. This optimizes the model's performance in the target domain, enabling it to more accurately understand questions, generate high-quality answers, and improve the quality of task completion when handling tasks in the target domain. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating the target vertical category large model fine-tuning method provided in the embodiments of this disclosure.
[0014] Figure 2 This is a flowchart illustrating the method for establishing an explicit correspondence between target domain concepts and model neurons, as provided in embodiments of this disclosure.
[0015] Figure 3 This is a flowchart illustrating the method for obtaining the neuron activation pattern feature vector corresponding to each sub-concept provided in this embodiment of the disclosure.
[0016] Figure 4This is a flowchart illustrating a method for obtaining neuronal regions sensitive to target domain knowledge in a large language model provided in an embodiment of this disclosure.
[0017] Figure 5 This is a flowchart illustrating a method for determining neuronal regions sensitive to target domain knowledge in a large language model based on real-time activation values, as provided in an embodiment of this disclosure.
[0018] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation
[0019] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0020] Reference Figure 1 This application discloses a method for fine-tuning a large target vertical model, including: S100 constructs a multi-layered target domain knowledge graph.
[0021] The target domains include vertical fields such as law, finance, and government affairs. Specifically, the target domain knowledge graph comprises a multi-layered organizational knowledge architecture. This knowledge graph can integrate knowledge in the target domain in a structured manner, clearly demonstrating the relationships between entities. This helps subsequent steps to accurately obtain the conceptual terms of the target domain, providing a solid knowledge foundation for building a large-scale model of the target vertical; at the same time, the multi-layered structure can comprehensively cover domain knowledge from macro to micro levels, facilitating analysis and application at different levels.
[0022] S200 inputs several target domain concept words obtained from the target domain knowledge graph into the large language model and tracks the activation state inside the model to establish an explicit correspondence between target domain concepts and model neurons.
[0023] In large-scale language models (LLMs, such as Transformer-based models), parameters (weight matrices) and knowledge representation are intrinsically linked. Not all parameters in an LLM are involved in a specific task (e.g., legal reasoning). Within a large language model, specific neurons (i.e., specific parameters) or sets of neurons (i.e., parameter sets) are particularly sensitive to domain-specific concepts. When processing text containing target domain concepts, these specific neurons exhibit distinct activation patterns. These activation patterns can be identified, quantified, and mapped to concept nodes in a legal knowledge graph. Furthermore, large language models can provide powerful natural language understanding and generation capabilities for open-source large language models such as DeepSeek-R1 or LLAMA. Taking the DeepSeek-R1-Chat-7B open-source large language model as an example, it is based on a decoder-only architecture and has 7 billion parameters (i.e., neurons).
[0024] By establishing this explicit correspondence, it is possible to identify which neurons in the large language model respond to specific concepts in the target domain. This provides a basis for subsequently locating neuronal regions sensitive to target domain knowledge, enabling the model to process target domain information more accurately.
[0025] S300: Based on explicit correspondences, obtain the neuronal regions in the large language model that are sensitive to target domain knowledge.
[0026] Focusing on neural regions sensitive to target domain knowledge avoids unnecessary computations and adjustments across the entire large language model, thus improving the model's efficiency and accuracy in handling target domain tasks and reducing interference from irrelevant information. In this step, an explicit correspondence is established between legal knowledge and model parameters, providing guidance for subsequent selective parameter fine-tuning.
[0027] S400 determines the sparse activation strategy based on the complexity of the task to be analyzed.
[0028] Dynamically adjusting the sparse activation strategy based on task complexity enables the model to allocate computational resources reasonably when handling tasks of varying difficulty. For simple tasks, increasing sparsity can reduce computation and speed up processing; for complex tasks, decreasing sparsity can increase the model's computational resource investment and improve processing accuracy.
[0029] S500 dynamically adjusts the actual activation ratio of neurons in neuronal regions according to the sparse activation strategy, and uses the adjusted large language model as the target vertical category large model.
[0030] This step uses legal knowledge graph information as explicit guidance, selectively activating neural network paths related to legal knowledge for fine-tuning, significantly improving the efficiency of parameter fine-tuning.
[0031] Furthermore, during fine-tuning, only about 0.1%-1% of the model parameters need to be adjusted, which reduces computational resource consumption by about 95% compared to full fine-tuning, and effectively shortens training time by about 90%. Different sparse activation strategies are adopted for parameters at different levels of the model, with the lower level retaining general language capabilities and the higher level focusing on optimizing the reasoning capabilities for target domain tasks.
[0032] By dynamically adjusting the actual activation ratio of neurons, the model can better adapt to the task requirements of the target domain, improving its performance and professionalism in the target domain, enabling it to handle problems in the target domain more accurately and output results that better meet professional requirements.
[0033] Furthermore, when laws and regulations are updated, only the knowledge graph and a few related parameters need to be updated, without retraining the entire model, thus significantly reducing maintenance costs.
[0034] The target vertical category large model fine-tuning method disclosed in this application constructs a multi-level target domain knowledge graph and establishes an explicit correspondence between target domain concepts and model neurons. This enables the large language model to more accurately understand the professional knowledge of target domains such as law. This helps the large model to have a deeper understanding of legal terms, concepts and their complex relationships in the application of the legal professional vertical field, reducing misinterpretations or inappropriate applications, and improving its adaptability and accuracy in the professional field. By acquiring neuronal regions sensitive to target domain knowledge and determining sparse activation strategies according to the complexity of the task to be analyzed, the actual activation ratio of neurons is dynamically adjusted. This approach allows the model to more specifically call upon relevant knowledge and capabilities when handling tasks such as legal reasoning, thereby enhancing legal reasoning ability and better performing tasks such as legal application and case analysis. Because the model can more accurately understand and apply legal knowledge, and can dynamically adjust neuron activation according to the task, it can effectively reduce the occurrence of "illusions" and the generation of content that violates legal provisions when dealing with legal issues, and reduce the frequency of professional errors. The multi-level target domain knowledge graph can cover the legal system knowledge of different countries and different periods. In this way, the model can better adapt to the changes in the legal system of different countries and different periods, and improve its adaptability in different legal environments.
[0035] In scenarios involving fine-tuning with limited data, dynamically adjusting the activation ratio of neurons using a sparse activation strategy can prevent the model from overfitting to training examples or forgetting its original capabilities. Furthermore, compared to methods like full-scale fine-tuning that require large amounts of labeled data and computational resources, this knowledge graph-guided approach utilizes resources more efficiently and reduces costs. This method can dynamically adjust the model based on the complexity of the task being analyzed, enabling the model to better adapt to the diverse nature of legal tasks, such as consultation, drafting, review, and dispute resolution, even under low-resource conditions.
[0036] Taking the legal field as an example, the target vertical category large-scale model fine-tuning method disclosed in this application mainly addresses problems such as the difficulty and high cost of acquiring large-scale models in the legal field, insufficient understanding of legal knowledge, and high resource consumption in handling various legal tasks (consultation, drafting, review, dispute resolution, etc.). It proposes a set of efficient parameter fine-tuning techniques based on knowledge graph enhancement. This solution organically combines explicit legal knowledge structures with deep learning models, effectively improving the understanding and application capabilities of open-source large language models such as DeepSeek-R1 in legal scenarios with only a small amount of labeled data, achieving precise processing of legal text understanding, question answering, analysis, and reasoning. The innovation of this solution lies in proposing a knowledge graph-guided sparse parameter efficient fine-tuning architecture, effectively solving the adaptability and accuracy problems of existing large-scale legal models in scenarios with few samples.
[0037] This application discloses a target vertical category large-scale model fine-tuning method that proposes an explicit mapping mechanism between legal knowledge and model parameters, providing a theoretical basis and implementation method for sparse parameter fine-tuning. Knowledge graph-guided sparse activation innovatively combines knowledge graphs and parameter activation to achieve efficient and targeted model fine-tuning. Compared to full-scale fine-tuning which requires significant computational resources, this invention only requires fine-tuning 0.1%-1% of the parameters, reducing resource consumption by over 95%. Compared to simple retrieval augmentation generation (RAG), this invention achieves deep integration of legal knowledge and model parameters, resulting in stronger reasoning capabilities. Compared to dedicated models optimized for single tasks, this invention can uniformly handle multiple legal tasks, broadening its application scenarios. Compared to black-box fine-tuning methods, this invention provides traceable reasoning paths through knowledge graphs, enhancing interpretability. Compared to traditional methods requiring frequent retraining, this invention only requires local adjustments during legal updates, significantly reducing maintenance costs.
[0038] The methods for "constructing a knowledge graph for the target domain" in S100 specifically include: S110, Determine the constituent elements of the target domain map.
[0039] The components of the map include basic concepts of the target domain, legal provisions of the target domain, related cases of the target domain, and legal relationships between legal entities in the target domain.
[0040] Taking the legal field as an example, the basic concepts of the target domain include legal concepts, i.e., conceptual entities in the target domain. These conceptual entities include basic concepts such as legal terms, legal subjects (natural persons, legal persons, etc.), and legal objects (property rights, creditor's rights, etc.); legal provisions in the target domain include various legal norms such as the Constitution, laws, administrative regulations, local regulations, and judicial interpretations, as well as the attribute information of each provision, including the effective date, level of validity, and revision history; related cases in the target domain include typical cases, judicial precedents, and other practical application scenarios; and legal subject relationships in the target domain include one or more of the following: application relationship, interpretation relationship, reference relationship, and priority relationship. Legal knowledge is highly specialized and complex. Knowledge graphs can integrate scattered legal knowledge, clearly present legal concepts and relationships, provide a comprehensive and accurate knowledge foundation for subsequent model construction, and facilitate rapid location and understanding of legal information.
[0041] S120, construct a knowledge graph for the target domain based on the constituent elements of the graph.
[0042] The target domain knowledge graph includes several knowledge nodes and the relationship edges between different knowledge nodes. The knowledge nodes are basic concept nodes, legal provisions nodes, related case nodes, or legal subject relationship nodes in the target domain. The relationship edges are used to describe the relationships between different nodes.
[0043] The target domain knowledge graph constructed in this embodiment can integrate scattered legal concepts, provisions, cases, and other information in the target domain into a unified knowledge system. This helps break down information silos and improve the efficiency of sharing and utilizing legal information. It also facilitates the updating and maintenance of legal information. When new legal provisions are promulgated or new cases occur, relevant information can be added to the knowledge graph in a timely manner, and the edges between knowledge nodes can be updated to ensure the timeliness and accuracy of the knowledge graph.
[0044] Furthermore, when the target domain is the legal domain, it also includes: assigning a correlation weight to each relation edge in the target domain (legal) knowledge graph, the correlation weight being used to represent the strength of the association between different legal concepts; and organizing the knowledge architecture in a multi-level manner, including jurisprudential level, legal provision level, and case level.
[0045] Furthermore, the construction of the target domain (law) knowledge graph also includes: establishing a knowledge acquisition and updating mechanism, specifically including: 1) Automated extraction, using natural language processing technology to automatically extract structured knowledge from publicly available data such as legal texts and judicial judgments; 2) Expert review, where legal experts review and supplement the automatically extracted knowledge to ensure its accuracy. In the expert review step, legal experts review and supplement the automatically extracted knowledge based on laws and regulations, legal principles, and judicial practice experience; 3) Dynamic updating, constructing an incremental update mechanism. When laws and regulations change, relevant knowledge nodes and relationships are automatically updated. That is, in the dynamic update step, by monitoring the official release channels of laws and regulations, the incremental update mechanism is triggered when new laws and regulations are released or existing laws and regulations are revised; 4) Version control, maintaining historical versions of the knowledge graph to support legal consultation at specific points in time. That is, in the version control step, version numbers are used to record different versions of the knowledge graph, and each version contains knowledge nodes and relationship information at the corresponding point in time.
[0046] Furthermore, it also includes multimodal representation, specifically transforming legal knowledge into various forms such as vectors, symbolic rules, and text descriptions, to facilitate interaction with different levels of the model. In this embodiment, a deep learning model can be used to transform legal knowledge into vector form, and logical reasoning rules can be used to transform legal knowledge into symbolic rule form.
[0047] Furthermore, it also includes hierarchical organization, organizing knowledge according to the hierarchy of "jurisprudence-legal provisions-cases" to reflect the inherent logical structure of legal knowledge. That is, in this hierarchical organization step, jurisprudence is the highest level, legal provisions are the middle level, and cases are the lowest level for knowledge organization.
[0048] Furthermore, it also includes quantifying the association strength, assigning weights to the relationship edges in the knowledge graph to represent the association strength between different legal concepts. That is, in the association strength quantification step, the weights of the relationship edges can be determined based on factors such as the frequency of reference and the degree of dependence between legal concepts.
[0049] Reference Figure 2 The S200 method, which "inputs several target domain concept words obtained from the target domain knowledge graph into a large language model and tracks the activation state within the model to establish an explicit correspondence between target domain concepts and model neurons," specifically includes the following: S210, based on the target domain knowledge graph, divide the target domain knowledge into several major categories of knowledge scenarios; obtain several sub-concepts contained in each major category of knowledge scenario.
[0050] Each sub-concept corresponds to a category of conceptual terms. For example, in the legal field, knowledge can be divided into major categories such as criminal law, civil law, and administrative law based on a legal knowledge graph. Taking criminal law as an example, its sub-concepts may include "theft," "robbery," and "fraud." Classifying legal knowledge makes subsequent processing more organized and systematic; breaking down the vast legal knowledge system into different major categories of knowledge scenarios facilitates targeted analysis and processing of sub-concepts within each scenario, improving processing efficiency and accuracy.
[0051] S220 involves inputting the sub-concepts into the large language model to obtain the feature vector of neuron activation patterns corresponding to each sub-concept. The large language model is an open-source model.
[0052] The feature vector of the neuron activation pattern corresponding to each sub-concept can reflect how the sub-concept is represented within the large language model. This helps to gain a deeper understanding of how the large language model processes and represents specific concepts in the legal field, providing a foundation for subsequent analysis and the establishment of correspondences.
[0053] S230, construct the neuron activation pattern feature vector of the large language model based on the obtained neuron activation values when processing historical text.
[0054] S240, obtain the similarity between the neuron activation pattern feature vector corresponding to each sub-concept and the neuron activation pattern feature vector of the large language model.
[0055] Specifically, the cosine similarity algorithm can be used to obtain the similarity between the neuron activation pattern feature vector corresponding to each sub-concept and the neuron activation pattern feature vector of the large language model. S250: Obtain all neurons in each sub-concept whose similarity is greater than a preset similarity threshold, and denote them as key neurons.
[0056] By setting a similarity threshold, key neurons that are highly relevant to sub-concepts can be selected, while neurons that are not closely related to sub-concepts can be excluded, reducing the complexity of subsequent processing.
[0057] S260 stores key neurons and their corresponding activation thresholds into a neuron set.
[0058] For example, the key neurons for the crime of "theft" and their corresponding activation thresholds can be stored as data records in a set, such as (neuron 1, activation threshold 1), (neuron 2, activation threshold 2), etc. Storing key neurons and activation thresholds in a neuron set facilitates unified management and use later, providing a data foundation for establishing explicit correspondences.
[0059] S270, group all sub-concepts into a set of target domain sub-concepts.
[0060] Integrating all sub-concepts into a single set facilitates unified management and processing of sub-concepts across the entire legal field, providing a clear list of concepts for establishing the mapping relationship between sub-concepts and neurons.
[0061] S280, establish an explicit correspondence between the set of sub-concepts in the target domain and the set of neurons.
[0062] For example, by clearly identifying the neuron sets corresponding to "theft" (neuron 1, activation threshold 1), (neuron 2, activation threshold 2), etc., each sub-concept is mapped one-to-one with its corresponding key neuron and activation threshold, forming a mapping relationship. Establishing explicit correspondences allows for a clear understanding of which neurons in the large language model respond to specific legal concepts and at what activation threshold. This helps explain the internal mechanisms of the large language model in processing legal knowledge, improving the model's interpretability.
[0063] In this embodiment, the explicit correspondence between the target domain sub-concept set and the neuron set is the mapping relationship, which can clearly identify the key neurons and activation thresholds corresponding to each sub-concept.
[0064] In this embodiment, by establishing an explicit correspondence between target domain concepts and model neurons, it is possible to intuitively understand which neurons are activated when the large language model processes legal knowledge, and the reasons for their activation, thereby explaining the model's decision-making process. This enables the large language model to better understand and process specific knowledge in the legal domain, improving the model's performance in legal tasks such as legal question answering and case analysis. After clarifying the correspondence between concepts and neurons, the model can be optimized and debugged specifically for particular legal concepts, improving the model's accuracy and stability.
[0065] The method S230, "Constructing a neuron activation pattern feature vector of a large language model based on the obtained neuron activation values when processing historical text," specifically includes: 1) obtaining the activation value of each neuron corresponding to all types of historical text when the large language model processes them; 2) taking the mean of all activation values corresponding to each neuron and adding a preset multiple of the standard deviation to determine the activation threshold; 3) obtaining the neurons corresponding to activation values greater than the activation threshold, and constructing a neuron activation pattern feature vector of the large language model based on the obtained neurons.
[0066] Suppose we have a large language model with 1000 neurons, and we have collected 100 historical texts from three categories: Category 1, Category 2, and Category 3. When we input these 300 historical texts into the large language model for processing, the model performs a series of calculations on each text. During this process, each neuron generates an activation value. For example, for the first Category 1 text, neuron 1 might have an activation value of 0.2, neuron 2 might have an activation value of 0.3, and so on. After processing all 300 texts, we obtain a 1000×300 matrix, where each row represents a neuron, each column represents a text, and the elements in the matrix are the activation values of the neurons when processing the corresponding text. By obtaining the activation values when processing all types of historical text, we can gain a comprehensive understanding of the response of each neuron in the large language model under different types of text input. Different types of text may stimulate different semantic understanding and processing methods of the model. Collecting this information helps us discover the commonalities and differences of the model when processing various types of text. A large amount of historical text data provides rich information for subsequent analysis, making the conclusions we draw based on this data more representative and reliable.
[0067] Continuing with the example above, for neuron 1, it has 300 activation values. We first calculate the mean of these 300 activation values, assuming a mean of 0.5; then we calculate the standard deviation of these 300 activation values, assuming a standard deviation of 0.1. We preset the multiplier to 2, so the activation threshold for neuron 1 is 0.5 + 2 × 0.1 = 0.7. Using the same method, we can calculate the activation thresholds for the remaining 999 neurons. Using the mean plus a preset multiplier of the standard deviation to determine the activation threshold allows for dynamic adjustment of the threshold based on the distribution of activation values for each neuron. Different neurons may perform different functions in the model, and their activation value distributions will also differ. This method avoids the inaccuracies caused by using a uniform, fixed threshold. The standard deviation reflects the dispersion of the data; adding a preset multiplier to the standard deviation sets the threshold at a position that highlights relatively abnormally high activation values. This allows us to filter out neurons that are truly active when processing historical text, excluding some accidental low-level activations.
[0068] Furthermore, for neuron 1, we have calculated its activation threshold to be 0.7. Among the previously obtained 300 activation values, we check if each activation value is greater than 0.7. Assuming 20 activation values are greater than 0.7, we consider neuron 1 to be an active neuron processing a portion of the historical text. Performing this check on all 1000 neurons will eventually yield a set of active neurons. Assuming we have 100 active neurons, we can construct a feature vector of length 1000, where each element is either 0 or 1. If a neuron at a given position is active, the value at that position is 1; otherwise, it is 0. For example, if neuron 1 is active, the first element of the feature vector is 1; if neuron 2 is not active, the second element of the feature vector is 0. Simplifying the complex neuron activation patterns of a large language model into a single feature vector makes the model's activation patterns easier to understand and analyze. This feature vector can serve as a concise representation of the model for subsequent tasks such as classification and clustering. By filtering out neurons with activation values greater than a threshold, we can focus on those neurons that truly play an important role in processing historical text, while ignoring inactive neurons. This effectively captures the key information of the model and reduces the dimensionality and noise of the data.
[0069] By constructing neuron activation pattern feature vectors, we can better understand the internal workings of large language models when processing historical text. These feature vectors reveal the model's activation patterns under different text types, helping us explain the model's decision-making process and semantic understanding. The feature vectors can serve as an effective feature representation applied to various downstream tasks, such as text classification and sentiment analysis. In these tasks, the model's neuron activation patterns may be correlated with the text's category or sentiment tendency; utilizing these feature vectors can improve the performance of downstream tasks. Different large language models may have different neuron activation pattern feature vectors. Comparing these feature vectors allows us to evaluate the similarities and differences between different models when processing historical text, providing a reference for model selection and optimization.
[0070] Reference Figure 3 The method S220, "inputting sub-concepts into a large language model to obtain the neuron activation pattern feature vector corresponding to each sub-concept," specifically includes the following: S221, Configure several types of historical texts associated with sub-concepts in the target domain.
[0071] Each type of historical text is further categorized into basic conceptual texts, composite conceptual texts, or context-ambiguous conceptual texts.
[0072] Different types of historical texts can cover relevant information of sub-concepts from multiple dimensions; basic texts allow the model to grasp the basic characteristics and elements of contract breach; composite texts help the model understand the relationship and mutual influence between contract breach and other legal concepts; contextually ambiguous texts can improve the model's ability to handle complex legal semantics and ambiguous situations, making the subsequently acquired neuron activation pattern feature vectors more comprehensive and accurate.
[0073] Assuming the target field is law, the sub-concept is "breach of contract." Basic concept texts refer to texts about the basic definition and constituent elements of breach of contract, such as "breach of contract refers to the act of one party failing to perform its contractual obligations or performing them in a manner that does not conform to the agreement; its constituent elements include the existence of a valid contract and the existence of a breach." Composite concept texts refer to texts that combine breach of contract with other legal concepts, such as "breach of contract may lead to tort liability, such as situations where the breach causes personal injury to the other party." Contextually ambiguous concept texts refer to texts with semantic ambiguity, such as "In this contract, one party delays delivery of goods; whether this constitutes a breach of contract needs to be determined based on the specific clauses in the contract and the actual circumstances."
[0074] S222, a neuron detection system is built in the large language model. The neuron detection system is used to capture the activation values of all neurons in the model in milliseconds.
[0075] Constructing a neuron detection system can comprehensively and timely obtain the activation values of all neurons in the model. This helps to deeply analyze the internal operating mechanism of the model when processing legal texts and provides rich data support for accurately analyzing the neuron activation patterns corresponding to sub-concepts such as "breach of contract".
[0076] Taking a large language model with a multi-layer neural network architecture as an example, this model has 20 layers, each with 16 attention heads, and a hidden vector dimension of 1024. Specifically, a dedicated monitoring component can be built as a neuron detection system. During the model's forward propagation, this system will monitor the neurons corresponding to each layer, each attention head, and each hidden vector in real time. For example, when the model processes input text, if the neuron corresponding to the 200th hidden vector of the 8th attention head in the 5th layer generates an activation value, the system will quickly record this value. To achieve millisecond-level data retrieval, efficient data acquisition algorithms can be used, along with high-performance servers and optimized memory management, to ensure that data can be collected quickly and accurately.
[0077] Furthermore, the neuron detection system is a system developed to monitor the activation values of neurons in each layer of a large language model in real time. Its core function is to accurately capture the activation information of key parts during model operation and efficiently store and process this information for subsequent analysis. This tool mainly consists of hardware components, software algorithms, and data storage modules. It monitors the internal state of the model by inserting "probe layers" at specific locations within the large language model.
[0078] The specific deployment methods of dedicated probe tools include: 1) Inserting a "probe layer": Specifically, a "probe layer" is inserted between the front-end and back-end of a trained and parameter-frozen large language model. The front-end of the large language model is typically responsible for receiving input data and performing preliminary processing, while the back-end is responsible for generating the final output. The "probe layer" acts as an "intermediary," positioned between these two parts to intercept and acquire intermediate information during model operation. From a technical implementation perspective, the "probe layer" is a specially designed piece of code or module that seamlessly integrates with other layers of the large language model. During model runtime, input data enters from the front-end, passes through the "probe layer," and is then transmitted to the back-end. The "probe layer" intercepts and analyzes the data during this process to obtain the required activation value information. 2) Capturing activation values: The "probe layer" has the ability to capture activation values of key parts of the model within a millisecond timescale. Specifically, it captures the activation values of each layer, each head attention, and each latent vector of the model. For each layer: Large language models typically consist of multiple layers, each with its specific function and role. The "probe layer" traverses each layer of the model, recording the activation values of neurons in that layer. These activation values reflect the level of activity of that layer when processing input data, which is crucial for understanding the model's working mechanism and performance. For each attention head: In large language models employing attention mechanisms, the attention head is a key component. Different attention heads can focus on different parts of the input data, thereby helping the model better understand and process information. The "probe layer" records the activation values of each attention head separately to analyze the model's attention allocation in different aspects. For each latent vector: Latent vectors are intermediate representations generated by the model during input data processing. They contain the model's understanding and abstraction of the input data. The "probe layer" captures the activation values of each latent vector, which can reflect the model's encoding and transformation of the input data at different stages.
[0079] 3) Mapping to a Dedicated Cache: After capturing activation values, the "probe layer" maps these data to a dedicated cache. This dedicated cache is a storage space specifically designed for storing activation value data, featuring high-speed read / write capabilities to ensure timely and accurate saving of the activation value data. During the mapping process, the "probe layer" assigns a unique identifier to each activation value and associates it with the corresponding model location (such as layer number, attention head number, latent vector index, etc.). This allows for quick location and extraction of the required activation value data during subsequent analysis. Simultaneously, the dedicated cache can also perform some preprocessing and compression on the data to reduce storage space usage and improve data processing efficiency.
[0080] S223, input each type of historical text into the large language model, and collect the activation value of each neuron through the neuron detection system.
[0081] By inputting different types of historical text and collecting activation values, we can observe the response of neurons in the model when processing different legal semantic information. This helps to discover the impact of different types of text on the activation of model neurons, thereby more accurately capturing the representation of sub-concepts such as "contract breach" in the model.
[0082] Specifically, basic, compound, and contextually ambiguous historical texts are sequentially input into the large language model. When the basic concept text "Contract breach refers to the act of one party to a contract failing to perform its contractual obligations or performing them in a manner that does not conform to the agreement, the constituent elements of which include the existence of a valid contract and the existence of a breach of contract" is input, the neuron detection system records the activation values of all neurons in the model when processing the text. For compound and contextually ambiguous texts, the same input and activation value collection operations are performed, and the activation values of each neuron are stored in a dedicated data storage structure each time text is input.
[0083] S224, obtain the average activation value of all activation values corresponding to each neuron.
[0084] The average activation value can eliminate random fluctuations in a single input text and more stably reflect the response characteristics of neurons to sub-concepts. It can summarize the overall activation of neurons when processing multiple related legal texts, making the subsequently constructed feature vectors more representative.
[0085] S225, construct the neuron activation pattern feature vector corresponding to each sub-concept based on the average activation value of all neurons.
[0086] In this embodiment, each dimension of the neuron activation pattern feature vector corresponding to each sub-concept is the average activation value of each neuron. That is, in this embodiment, each average activation value is used as a dimension of the feature vector to construct the neuron activation pattern feature vector corresponding to the sub-concept.
[0087] Suppose that the large language model has 15,000 neurons, and each neuron has a corresponding average activation value. Arrange these average activation values in the order of the neuron numbers to form a vector of length 15,000. This vector is the feature vector of the neuron activation pattern corresponding to the sub-concept such as "contract breach". Each dimension in the vector corresponds to the average activation value of a neuron.
[0088] In this embodiment, by integrating different types of legal historical texts and comprehensive collection of neuron activation values, a deeper and more accurate understanding of the representation of legal sub-concepts in the large language model can be achieved. This helps improve the model's semantic understanding of legal concepts and better handle legal-related natural language texts. Transforming legal sub-concepts into neuron activation pattern feature vectors realizes the quantitative representation of legal concepts, facilitating various legal calculations and analyses. This provides more effective feature inputs for subsequent legal natural language processing tasks, such as legal reasoning and legal document generation. The neuron activation pattern feature vectors can provide some explanatory basis for the decision-making process of the legal model, helping legal professionals understand how the model processes and represents legal concepts, thereby improving the interpretability of the legal model and enhancing its credibility in the legal field. These feature vectors can be applied to multiple legal natural language processing tasks, such as legal case analysis, legal provision recommendation, and legal risk assessment, providing strong support for performance improvement of related legal tasks and promoting the intelligent development of the legal field.
[0089] Reference Figure 4 The method for S300, "obtaining neuronal regions sensitive to target domain knowledge in a large language model based on explicit correspondences," specifically includes: S310, Configure test text for the target domain.
[0090] The test texts include basic concept texts, compound concept texts, and context-ambiguous concept texts. The number of basic concept texts is no less than 20 times the number of context-ambiguous concept texts, the number of compound concept texts is no less than 3 times the number of context-ambiguous concept texts, and the number of context-ambiguous concept texts is no less than 10. Each type of concept text includes several positive texts, several negative texts, and several synonymous distractor texts. Positive texts are texts related to the target concept, while distractor texts are texts unrelated to the target concept or that have a distracting effect.
[0091] Furthermore, the basic concept texts cover the most fundamental and core legal concepts in the legal field. For example, "a contract is an agreement between equal parties to establish, modify, or terminate civil rights and obligations," and "a crime is an act that violates the provisions of criminal law and is subject to criminal punishment." If the number of context-ambiguous concept texts is set at 10, the required number of basic concept texts is at least 200. Each category of concept texts includes positive, negative, and near-synonymous interfering texts. Positive texts include "a contract is an agreement between equal parties to establish, modify, or terminate civil rights and obligations"; negative texts include "a contract is not an agreement between equal parties to establish, modify, or terminate civil rights and obligations"; and near-synonymous interfering texts include "a contract is an agreement between equal parties to establish, modify, or terminate civil rights and obligations" (contract and agreement are similar in meaning to some extent, but may have subtle differences in legal context).
[0092] Composite legal texts are legal descriptions composed of multiple basic legal concepts. Examples include "A contract concluded due to a significant misunderstanding may be modified or rescinded by one party upon request to the People's Court or an arbitration institution," and "In a joint crime, the principal offender shall be punished according to all the crimes he participated in, organized, or directed." Since the number of composite legal texts should be no less than three times the number of context-ambiguous legal texts, there should be at least 30 such texts. There are also positive, negative, and near-synonymous interfering texts. A positive text is like "A contract concluded due to a significant misunderstanding may be modified or rescinded by one party upon request to the People's Court or an arbitration institution"; a negative text is like "A contract concluded due to a significant misunderstanding may not be modified or rescinded by one party upon request to the People's Court or an arbitration institution"; and a near-synonymous interfering text is like "A contract concluded due to fraud may be modified or rescinded by one party upon request to the People's Court or an arbitration institution" (both fraud and significant misunderstanding can lead to contract modification or rescission, but the specific circumstances differ).
[0093] Ambiguous legal concepts can be interpreted differently in different legal contexts. For example, phrases like "If legitimate self-defense exceeds the necessary limits and causes significant harm, criminal liability shall be borne, but the punishment shall be mitigated or exempted" raise questions about the definition of "necessary limits" and "Where are the boundaries of 'freedom of speech' in the online environment," totaling at least 10 instances. Positive interpretations include "If legitimate self-defense exceeds the necessary limits and causes significant harm, criminal liability shall be borne, but the punishment shall be mitigated or exempted"; negative interpretations include "If legitimate self-defense exceeds the necessary limits and causes significant harm, no criminal liability shall be borne"; and near-synonymous texts include "If emergency avoidance exceeds the necessary limits and causes undue harm, criminal liability shall be borne, but the punishment shall be mitigated or exempted" (emergency avoidance and legitimate self-defense share some similarities, but the legal provisions differ).
[0094] By configuring basic, composite, and contextually ambiguous concept texts, the model can comprehensively cover various types of knowledge in the legal field, from basic legal concepts to complex combinations of legal provisions, as well as legal issues that are prone to ambiguity, ensuring a comprehensive evaluation of the model in the legal field. Each type of concept text contains positive, negative, and near-synonymous interference texts, which help the model better distinguish between correct and incorrect legal information, as well as similar but different legal concepts, thereby improving the model's accuracy in the legal field.
[0095] S320: Input the test text into the large language model, use the neuron detection system to explicitly correspond to the real-time activation value of each neuron in the relation, and determine the neuron regions in the large language model that are sensitive to the target domain knowledge based on the real-time activation value.
[0096] Further reference Figure 5 Methods for determining neuronal regions sensitive to target domain knowledge in large language models based on real-time activation values include: S321, input the concept text of each category into the large language model in batches, obtain the activation value of each neuron through the neuron detection system, and compare the first type of activation difference between the forward text and the reverse text, and the second type of activation difference between the forward text and the synonymous interference text, token by token.
[0097] Specifically, positive sentences, antonyms, and near-synonyms from the same batch are simultaneously input into the model. The activation differences between positive sentences and antonyms, and between positive sentences and near-synonyms, are compared token by token. For positive and antonyms, which describe the positive and negative aspects of a concept, comparing activation differences highlights the differences in neuron activation when the model encounters correct and incorrect expressions of the concept. For example, for the concept of "bona fide acquisition," comparing the activation values of the neuron corresponding to each token when inputting the positive sentence "The buyer meets the requirements for bona fide acquisition…" and the antonym "The buyer maliciously colluded… and does not constitute bona fide acquisition…", reveals the neurons sensitive to correct and incorrect expressions of the concept. For positive and near-synonyms, although near-synonyms have similar expressions to the target concept, their actual meanings differ. Comparing activation differences helps identify the key neurons that distinguish the target concept from the distracting concepts. For example, comparing "Bona fide adoption does not belong to bona fide acquisition in the sense of property law" with the positive sentence analyzes the differences in neuron activation.
[0098] Taking basic class concept texts as an example, compare token-by-token the first type of activation difference between forward texts and reverse texts, and the second type of activation difference between forward texts and near-synonym interference texts. For example, for the forward text "A contract is an agreement establishing, modifying and terminating the relationship of civil rights and obligations between equal parties" and the reverse text "A contract is not an agreement establishing, modifying and terminating the relationship of civil rights obligations between equal parties", after being input into the model, the neuron activation values corresponding to each token (such as "contract", "is", "equal parties", etc.) are compared, and the first type of activation difference is calculated. Similarly, for the forward text "A contract is an agreement establishing, modifying and terminating the relationship of civil rights and obligations between equal parties" and the near-synonym interference text "A covenant is an agreement establishing, modifying and terminating the relationship of civil rights and obligations between equal parties", the second type of activation difference is calculated.
[0099] By comparing the activation differences between forward texts and reverse texts, and between forward texts and near-synonym interference texts token by token, neurons responsive to legal domain knowledge can be identified more accurately, since this comparison can highlight the differences between different legal texts, thereby reflecting the sensitivity of neurons to different legal information; analyzing each token enables an in-depth understanding of the activation status of neurons when processing different legal vocabularies and semantics, which facilitates a more detailed study of the internal processing mechanism of the model in the legal field.
[0100] Herein, a token is a basic processing unit in natural language processing, which can be a word, a subword or a character, specifically depending on the tokenization method. For example, for the English sentence "I love apples", tokenized by words, there are 3 tokens: "I", "love", "apples"; while for the Chinese sentence "Wo ai pingguo (I love apples)", it may be tokenized into 3 tokens: "Wo (I)", "ai (love)", "pingguo (apples)".
[0101] In S322, a significance threshold of 95% confidence interval is set, and background neurons are eliminated according to the first type of activation difference and the second type of activation difference.
[0102] Background neurons refer to neurons whose activation values do not change significantly when forward, reverse and interference texts are input, and which contribute little to concept discrimination. Through threshold setting and screening, only neurons with significant changes in activation values under different types of text input are retained.
[0103] Specifically, a significance threshold of 95% confidence interval is set, which can be statistically derived from a large amount of experimental data in the legal field. For example, after multiple experiments, it was found that the activation value of a certain neuron fluctuates within a specific range when processing legal text. When the activation difference exceeds this range (with a 95% confidence level), the activation change of that neuron is considered significant. Based on the first and second types of activation differences, neurons whose activation differences do not reach the significance threshold are considered background neurons and are eliminated. For example, if a neuron has a small activation difference when comparing forward and reverse text, or forward text and similar interfering text, and does not exceed the significance threshold, it is eliminated.
[0104] Changes in activation values reflect the degree to which neurons respond to the input test text. Setting a significant threshold and removing background neurons can eliminate neurons that are insensitive to legal knowledge, fluctuate randomly, or are affected by other factors, thereby reducing noise interference. This makes the subsequently selected high-response neurons more representative of the model's processing of legal knowledge, while reducing the number of neurons requiring further analysis, improving the efficiency of selecting high-response neurons, and saving time and computational resources.
[0105] S323, randomly select a preset number of neurons from the remaining neurons as high-response neurons for the corresponding class of concept text.
[0106] All high-response neurons are neuronal regions in the large language model that are sensitive to target domain knowledge. The preset number is 0.002%-0.005% of the total parameters of the neural network in the large language model. Furthermore, for each concept category, 120-300 "high-response neurons" are retained. In this embodiment, high-response neurons are neurons whose activation values change significantly and play a key role in concept differentiation during comparisons of positive, antonymous, and distracting text inputs.
[0107] By randomly selecting a predetermined number of high-response neurons, we can quickly identify neuronal regions in a large language model that are sensitive to legal knowledge. This provides a specific target for further research on the model's performance in the legal field and for optimizing the model. Furthermore, we can record information such as the neuron's number, layer number, and average activation threshold, which facilitates in-depth research on these high-response neurons. For example, we can analyze how neurons in different layers process legal knowledge, and this also provides a foundation for developing applications based on the legal field.
[0108] This solution allows for in-depth understanding of which neuronal regions play a key role in the legal knowledge processing process of large language models. It helps reveal the model's internal mechanisms, supporting the interpretability of legal models and enabling legal practitioners to better understand the model's decision-making process. Once the neuronal regions sensitive to legal knowledge are identified, the model can be optimized in a targeted manner, such as adjusting the parameters of these neurons or adding relevant legal training data. This improves the model's performance and accuracy in the legal field, enabling more accurate interpretation of legal provisions and case analysis. For legal intelligence applications, such as intelligent legal consultation and legal document generation, this solution can help identify neuronal regions sensitive to legal knowledge, providing a foundation for developing more professional and accurate legal intelligence applications and promoting the intelligent development of the legal industry.
[0109] The S400 method for "determining the sparse activation strategy based on the complexity of the task to be analyzed" specifically includes: S410, Determine the complexity of the task to be analyzed.
[0110] Specifically, this includes: S411, analyzing the input task to be analyzed from multiple dimensions to determine the multi-dimensional score of the task to be analyzed; S412, obtaining the complexity score corresponding to the task to be analyzed based on the multi-dimensional weights and multi-dimensional scores. In this embodiment, the multi-dimensional aspects include five dimensions: knowledge scope, reasoning depth, concept density, ambiguity, and inter-document dependency.
[0111] For example, the method for obtaining a multi-dimensional score for the knowledge scope dimension includes: counting the number of different knowledge domains involved in the target input information. If only 1-2 knowledge domains are involved, the score is 1 point; if 3-4 knowledge domains are involved, the score is 2 points; if 5-6 knowledge domains are involved, the score is 3 points; and if 7 or more knowledge domains are involved, the score is 4 points.
[0112] For example, the method for obtaining a multi-dimensional score for the reasoning depth dimension includes: determining the number of reasoning steps required to draw a conclusion from the input information. If only 1-2 simple reasoning steps are required, the score is 1 point; if 3-4 reasoning steps are required, the score is 2 points; if 5-6 reasoning steps are required, the score is 3 points; and if 7 or more complex reasoning steps are required, the score is 4 points.
[0113] For example, the method for obtaining a multi-dimensional score for the concept density dimension includes: calculating the number of professional concepts appearing in a unit length of text; 1-2 professional concepts per 100 words, 1 point; 3-4 professional concepts, 2 points; 5-6 professional concepts, 3 points; and 7 or more professional concepts, 4 points. Similarly, the method for obtaining a multi-dimensional score for the fuzziness dimension includes: obtaining a score based on the proportion of fuzzy words (such as "probably," "possibly," "maybe," etc.) in the input information; less than 10% fuzzy words, 1 point; 10%-20%, 2 points; 20%-30%, 3 points; and more than 30%, 4 points.
[0114] For example, the method for obtaining a multi-dimensional score for the document dependency dimension includes: 1 point if the target input information can be understood without referring to other documents; 2 points if it requires referring to a few (1-2) other documents; 3 points if it requires referring to a moderate number (3-4) other documents; and 4 points if it requires referring to a large number (5 or more) other documents. This step analyzes the target input information from multiple dimensions, enabling a comprehensive and detailed assessment of the information's complexity and avoiding the bias of a single-dimensional evaluation. The complexity score is obtained through weighted calculation, comprehensively considering the importance of each dimension, making the complexity assessment more accurate and scientific.
[0115] S420 determines the complexity level based on complexity.
[0116] Specifically, the complexity level is determined as follows: when the complexity meets the first threshold condition (preferably less than 0.3), the complexity level is determined as low complexity; when the complexity meets the second threshold condition (preferably 0.3-0.7), the complexity level is determined as medium complexity; and when the complexity meets the third threshold condition (preferably 0.7 or higher), the complexity level is determined as high complexity. Dividing complexity into different levels makes the description of complexity more intuitive and clear, facilitating the adoption of corresponding strategies based on different levels.
[0117] S430 determines the sparse activation strategy based on the complexity level.
[0118] Specifically, the strategy includes: when the complexity level is low, the sparse activation strategy is determined to be the first preset strategy, which includes updating the current sparsity to 150% of the initial sparsity (i.e., increasing the sparsity by 50%); when the complexity level is medium, the sparse activation strategy is determined to be the second preset strategy, which includes maintaining the initial sparsity; and when the complexity level is high, the sparse activation strategy is determined to be the third preset strategy, which includes updating the current sparsity to 70% of the initial sparsity (i.e., reducing the sparsity by 30%).
[0119] Different sparse activation strategies are determined based on varying levels of complexity, allowing for dynamic adjustment of sparsity according to the actual complexity of the task. This improves the efficiency and performance of the model in handling tasks of varying complexity. For low-complexity tasks, increasing sparsity reduces computation and speeds up processing; for high-complexity tasks, decreasing sparsity increases computational resource investment and improves accuracy. By dynamically adjusting sparsity, the model can better adapt to tasks of varying complexity, improving overall performance and generalization ability. This avoids wasting excessive computational resources on low-complexity tasks while ensuring sufficient resources for high-complexity tasks, achieving a rational allocation of resources.
[0120] Furthermore, the S500 method, which "dynamically adjusts the actual activation ratio of neurons in a neuron region according to a sparse activation strategy, and uses the adjusted large language model as the target vertical category large model," specifically includes: The initial sparsity is denoted as S0, and the initial activation ratio is denoted as A0 = 1 − S0. When the sparse activation strategy is the first preset strategy, for each neuron in the neuron region, its importance score in processing target domain knowledge is calculated. The importance score can be determined based on indicators such as the average activation value of the neuron and the connection strength with other neurons; neurons are ranked according to their importance scores; the top Q neurons are selected as activated neurons in descending order, while the remaining neurons are inactive. Here, Q = N × (1 − 1.5 × S0), where N is the total number of neurons in the neuron region.
[0121] When the sparse activation strategy is the third preset strategy, the updated sparsity S2 = 0.7 × S0 is calculated. The updated activation ratio A2 = 1 − S2 is then obtained. Similarly, the importance score of the neurons is calculated. Neurons are ranked according to their importance scores. The top A2 × N neurons are selected as activated neurons in descending order, while the remaining neurons remain inactive.
[0122] The target vertical category large model fine-tuning method disclosed in this application constructs a target domain knowledge graph and establishes a correspondence between it and the model's neurons, and dynamically adjusts the activation ratio of neurons. This enables the large language model to understand and process target domain knowledge more deeply, becoming a specialized target vertical category large model that outperforms general-purpose large language models in the target domain. The method determines a sparse activation strategy based on task complexity, avoiding overcomputation on simple tasks and insufficient resources on complex tasks, thus improving the model's task processing efficiency and reducing computational resource waste. This scheme can be flexibly adjusted according to different target domains and task complexities, exhibiting strong adaptability and applicability to multiple different professional fields, enabling the construction of diverse target vertical category large models.
[0123] Taking the legal field as an example, high-quality labeled data in the legal field is difficult to obtain and costly in the existing technology. This application constructs a multi-level target domain knowledge graph. The knowledge graph can integrate legal knowledge from multiple sources and does not necessarily rely on a large amount of high-quality labeled data. By guiding the model learning through the knowledge graph, the data bottleneck problem is alleviated to a certain extent, the dependence on a large amount of labeled data is reduced, and the data acquisition cost is lowered.
[0124] Legal knowledge possesses structured characteristics and strict logical relationships, while existing methods mostly employ direct training with unstructured text. This application constructs a multi-layered target domain knowledge graph. Knowledge graphs inherently possess structured characteristics, fully reflecting the logical relationships between legal knowledge. By inputting target domain concepts from the knowledge graph into a large language model and establishing explicit correspondences, the model can leverage the structural characteristics of legal knowledge to learn, thereby improving its understanding and application of legal knowledge.
[0125] In scenarios involving fine-tuning with limited data, existing methods are prone to causing the model to overfit training examples or forget its original capabilities. This application determines a sparse activation strategy based on the complexity of the task being analyzed, dynamically adjusting the actual activation ratio of neurons in the neuron region. This dynamic adjustment mechanism allows the model to learn and apply knowledge more flexibly when fine-tuning with limited data, avoiding over-reliance on training examples and thus reducing the risk of overfitting, while also better preserving its original capabilities.
[0126] Legal tasks are diverse, and existing methods struggle to adapt to multiple tasks simultaneously under low-resource conditions. This application's sparse activation strategy dynamically adjusts the activation ratio of neurons in the model based on the complexity of different tasks. For different types of legal tasks, such as consultation, drafting, examination, and dispute resolution, the model can selectively leverage its capabilities by adjusting neuron activation, thus better adapting to multiple legal tasks under low-resource conditions.
[0127] The target vertical category large-scale model fine-tuning method disclosed in this application has broad prospects for industrial application, mainly including: 1) Intelligent legal services: providing intelligent auxiliary tools for law firms and legal consulting institutions to improve service efficiency and quality; 2) Automated corporate legal affairs: helping corporate legal departments to automate contract review and legal risk management, reducing compliance costs; 3) Judicial decision support: providing courts, procuratorates and other judicial institutions with auxiliary functions such as case retrieval and legal provision recommendation to improve trial efficiency; 4) Legal education and training: serving as an intelligent assistant for legal education and training, providing learning support for law school students and legal practitioners; 5) Intelligent legal Q&A: building a public-oriented intelligent legal Q&A system to improve the legal literacy of the entire population.
[0128] In the practical application of the vertical-specific large-scale model in this application, the relationship between "neurons" and "models" is not an abstract concept, but a mapping relationship from micro-computational units to macro-system capabilities, specifically reflected in the following three aspects: 1) Neurons are the "knowledge perception units" of the vertical-specific large-scale model; vertical-specific large-scale models (such as specialized models for medicine, finance, and manufacturing) achieve the encoding and understanding of industry knowledge by simulating the connection strength and activation mechanism between neurons. 2) Neuron density determines the "professional depth" of the vertical-specific model; the performance of the vertical-specific model is directly related to the number of neurons and hierarchical structure: general-purpose large-scale models like GPT-3 contain 2048 neurons per layer, with a total of 96 layers and 175 billion parameters, possessing broad generalization capabilities. Vertical-specific models, by reducing general-purpose neurons and increasing domain-specific neurons, achieve higher accuracy and lower resource consumption in specific tasks. 3) The interpretability of neurons affects the "compliance and trustworthiness" of the vertical-specific model; in highly compliant scenarios such as medicine and finance, vertical-specific models need to explain their decision-making logic; visualization of neuron activation paths (such as which neurons respond to "sepsis indicators") becomes a key requirement. Compared to general large models, vertical models reduce the risk of "illusions" and improve the credibility of results by limiting the scope of neuronal connections (e.g., only associating nodes in the medical knowledge graph).
[0129] Traditional full-parameter fine-tuning is inefficient (due to the massive number of parameters) and prone to overfitting or resource waste. In large language models (LLMs, such as Transformer-based models), parameters (weight matrices) and knowledge representation are intrinsically linked. Not all parameters in an LLM are involved in a specific task (such as legal reasoning); instead, only a subset of "neurons" (or parameter subsets) are associated with a specific knowledge domain. Therefore, by guiding "sparse activation" through a knowledge graph (KG), efficiency can be improved by fine-tuning only relevant parameters for legal subgraphs (such as entity-relationship in contract law).
[0130] In this embodiment, inspired by the human brain's "selective activation of neurons" (e.g., specific neurons in the visual cortex responding to specific stimuli), we treat LLM parameters as "artificial neural networks." The legal knowledge graph is used as a "guiding signal," mapped to the model's attention layer or FFN (feedforward network) submodule to achieve precise activation.
[0131] For example, in the legal field, the construction of a large-scale legal model (i.e., a target vertical-category model) requires the efficient integration of professional knowledge (such as provisions of the Civil Code). Existing research shows that model parameters can "locate" specific facts / knowledge (such as the Rome project), which prompts us to design a "KG-guided sparse path" to avoid the inefficiency of full-network fine-tuning. In the scheme disclosed in this application, through activation analysis, the number of fine-tuning parameters is reduced by 70%, and performance is improved by 15%.
[0132] Based on neural network theory, Transformer architecture, and knowledge localization research, the following demonstrates that the association between "neurons" (generally referring to model parameters or sub-modules) and specific knowledge / tasks is quantifiable and operable. The solution uses a knowledge graph (KG) to map these associations, enabling efficient fine-tuning.
[0133] 1. Basic Concepts: The Role of Neurons in LLM. 1.1) Extended Definition of Neuron: In deep learning, a "neuron" traditionally refers to the basic unit of a fully connected layer (e.g., input weighted summation + activation function). In the Transformer model, this extends to attention heads, feedforward network (FFN) layers, or embedding vectors. These "neurons" store and process distributed knowledge: for example, one attention head might be dedicated to handling "causal relationships," while another handles "entity recognition." 1.2) Association Mechanism: LLM learns from massive amounts of data through pre-training, forming a parameter-knowledge mapping. Research shows that specific subsets of parameters "encode" specific knowledge (e.g., factual associations, grammatical rules). During fine-tuning, without guidance, full parameter updates can easily disrupt irrelevant knowledge, leading to catastrophic forgetting.
[0134] 2. Expertise and Research Evidence. 2.1) Knowledge Localization and Neuron Activation: Project Rome (Locating and Editing Factual Associations in GPT): Research from MIT and Harvard University (2022, published in NeurIPS) demonstrates that in the GPT model, specific intermediate-layer FFN neurons store factual knowledge (e.g., "The Eiffel Tower is located in Paris"). Through causal tracing, these "knowledge neurons" can be located and edited to update the model's facts without affecting the overall model. This research uses gradient analysis to locate associations, supporting the mapping of our legal knowledge graph (e.g., "contract-breach-indemnity" triples) to similar neurons, guiding sparse fine-tuning. 2.2) Attention Head Specialization in Transformers: OpenAI research (2020) shows that Transformer attention heads tend to be "specialized" in specific tasks (e.g., dependency resolution or coreference resolution). For example, in legal texts, a head might be activated in a "legal citation" mode. Our approach extends this: KG subgraphs are used as input embeddings, activation scores are computed (e.g., via a Saliency Map), and only highly correlated neurons are fine-tuned. 2.3) Sparse activation and parameter efficiency: Mixture of Experts (MoE) model: Google's SwitchTransformer (2021, ICLR) introduces the MoE architecture, dynamically activating "expert" subnetworks based on the input (each expert is equivalent to a subset of parameters or a "group of neurons"). This reduces computation by 90%, proving the efficiency of sparse paths. 2.4) Parameter-efficient fine-tuning (PEFT) techniques: such as LoRA (Low-Rank Adaptation, Microsoft 2021), which only fine-tunes low-rank matrices, rather than all parameters. Combining KG guidance, we further "sparse" the model: using graph neural networks (GNNs) to compute the similarity between the embeddings and model parameters (e.g., only neurons with cosine similarity > 0.7 are activated), avoiding irrelevant updates.
[0135] 3. Quantitative Experiment Support. 3.1) Neuron Activation Analysis: On the BERT model, research (Stanford NLP Group, 2019) used the Probe task to demonstrate that neurons in specific layers are highly correlated with semantic roles (accuracy >85%). Our approach is similar: running an activation heatmap on the legal dataset visualizes the distribution of "contract law neurons," supporting precise fine-tuning guided by KG. 3.2) Forgetting Resistance and Efficiency: EleutherAI's research (2023) shows that knowledge-guided sparse fine-tuning can retain 90% of the original knowledge while improving the performance of downstream tasks.
[0136] 4. Scientific justification of the solution: The aforementioned knowledge originates from the intersection of neuroscience and AI, such as "distributed representation theory" (Hinton, 1986): knowledge is distributed in neural networks and can be located / activated by external signals (such as KG). This is not a random assumption, but rather a consensus based on thousands of papers (e.g., more than 10,000 results related to "Neuron Activation in LLM" on Google Scholar).
[0137] Secondly, this application discloses a legal information analysis method, including: S10, determining the legal information to be analyzed; S20, inputting the legal information into a target large model and generating feedback information; wherein, the target large model is a legal vertical large model fine-tuned using the aforementioned legal vertical large model fine-tuning method.
[0138] The benefits of the legal information analysis method provided in this application are detailed below from several aspects, including comparison of legal accuracy, comparison of multi-task adaptability, knowledge update adaptability experiment, and ablation experiment.
[0139] Table 1: Comparison of legal accuracy of different methods (%) The method of this invention outperforms the comparative methods in all dimensions of legal accuracy. In particular, in terms of legal reasoning accuracy, this invention improves the accuracy by 6.2 percentage points compared to the strongest baseline, indicating that knowledge graph-guided sparse parameter fine-tuning can effectively enhance the model's legal reasoning ability.
[0140] Table 2: Comparison of F1 scores (%) for different methods on various legal tasks By comparing the F1 scores of different methods across various legal tasks, it can be seen that the method of application disclosure is also quite excellent in terms of multi-task adaptability.
[0141] Table 3: Comparison of Model Adaptability under Legal and Regulatory Update Scenarios In scenarios involving updates to laws and regulations, this method demonstrates excellent adaptability, enabling rapid model updates while maintaining high levels of old knowledge retention and new knowledge application accuracy. Compared to the RAG method, this invention better integrates new knowledge into model parameters, providing more consistent and accurate answers.
[0142] To verify the effectiveness of each component of this invention, an ablation experiment was conducted, and the results are as follows: Table 4: Ablation Experiment Results (Overall Accuracy %) Ablation experiments show that all components of this invention contribute positively to the final performance. Among them, knowledge graph guidance has the greatest impact on performance; removing it reduces the average accuracy by 6.4 percentage points, which verifies the core role of knowledge-parameter mapping. Multi-granularity knowledge injection is also a key component; removing it reduces performance by 5.2 percentage points, indicating that the integration of legal knowledge at different levels is crucial for the model's understanding of legal concepts and reasoning.
[0143] Furthermore, a detailed description is provided using a legal consultation case as an example. The legal information to be analyzed includes: I have a dispute with my neighbor over a water leak in their house. The leak has damaged my walls, and I am demanding compensation, but they disagree. What should I do? Inputting legal information into the target model generates feedback information including: You are facing a property damage compensation dispute arising from neighborly relations. According to Article 286 of the Civil Code, "Adjacent property owners shall properly handle neighborly relations in accordance with the principles of benefiting production, facilitating life, mutual assistance, fairness, and reasonableness." And Article 288 states, "When property owners use adjacent real estate for water supply, drainage, passage, laying pipelines, etc., they shall try their best to avoid causing damage to adjacent property owners."
[0144] If your neighbor's house leaks and causes damage to your walls, this constitutes a tort. According to Article 1165 of the Civil Code, a person who infringes upon the civil rights and interests of others due to fault and causes damage shall bear tort liability.
[0145] We suggest you take the following steps: 1) Collect evidence: Take photos to document the damage to the wall and invite the property management or a third party to witness the source of the leak; 2) Negotiate a solution: Communicate with your neighbors, explaining the legal provisions and compensation requirements; 3) Apply for mediation: You can apply for mediation with the neighborhood committee, village committee, or property management company; 4) Resolve through litigation: If negotiation and mediation fail, you can file a lawsuit in the People's Court, demanding compensation for repair costs and other reasonable losses. The statute of limitations is three years, calculated from the date you knew or should have known that your rights had been infringed.
[0146] The method of this invention accurately cites relevant provisions of the Civil Code, including both general provisions on neighborly relations and specific tort liability provisions. The analysis is reasonable, and the suggestions are specific and feasible, demonstrating a precise grasp and application of legal knowledge.
[0147] Furthermore, a detailed description is provided using a legal reasoning case as an example. The legal information to be analyzed includes: Company A and Company B signed a sales contract stipulating that Company A would supply Company B with 500 tons of steel at a price of 3,000 yuan per ton, with delivery within 30 days of contract signing. After the contract was signed, the market price of steel rose to 4,500 yuan per ton. Company A, citing the significant increase in market prices, demanded that Company B increase the price or terminate the contract, which Company B refused. Please analyze whether Company A's request has a legal basis. Inputting legal information into the target large model generates feedback information including: analyzing whether Company A's request for increased price or termination of contract has a legal basis, and examining the provisions of the Civil Code regarding the principle of changed circumstances.
[0148] According to Article 563 of the Civil Code, "A party may terminate a contract under any of the following circumstances: (1) the purpose of the contract cannot be achieved due to force majeure; (2) before the expiration of the performance period, one party expressly states or indicates by its conduct that it will not perform its principal obligations; (3) one party delays performance of its principal obligations and fails to perform within a reasonable period after being urged to do so; (4) one party delays performance of its obligations or commits other breaches of contract that render the purpose of the contract impossible to achieve; or (5) other circumstances stipulated by law." Regarding the doctrine of changed circumstances, Article 533 of the Civil Code stipulates: "After a contract is established, if a significant change occurs in the basic conditions of the contract that was unforeseeable to the parties at the time of entering into the contract and does not constitute a commercial risk, and continued performance of the contract would be manifestly unfair to one party, the adversely affected party may renegotiate with the other party; if the renegotiation fails within a reasonable period, the parties may request the People's Court or an arbitration institution to modify or terminate the contract. The People's Court or arbitration institution shall, in light of the actual circumstances of the case, decide whether to modify or terminate the contract based on the principle of fairness." Analyzing this case: Does the increase in steel prices constitute a change of circumstances? Price increases are a normal phenomenon in a market economy and fall under the category of commercial risk, unless the increase reaches a level that renders the purpose of the contract impossible to achieve. In this case, the price increase of 50%, while significant, has not reached the point of fundamentally altering the basic conditions of the contract.
[0149] Who should bear the price risk? In commercial transactions, price risk is usually borne by the seller, which is also one of the sources of the seller's profit. Company A, as the operator, should have anticipated possible changes in market prices.
[0150] Is fulfilling the contract manifestly unfair to Company A? Although Company A may not be able to obtain the expected profit, or may even incur some losses, it does not constitute manifest unfairness.
[0151] Conclusion: Company A's request lacks sufficient legal basis. According to the principle of changed circumstances in the Civil Code, price fluctuations are a normal commercial risk, and Company A should fulfill its contractual obligations. It is recommended that Company A negotiate with Company B to seek a mutually acceptable solution, such as partial price adjustments or phased delivery. If negotiations fail, Company A should still fulfill its obligations under the contract; otherwise, it may constitute a breach of contract. This invention demonstrates sophisticated legal reasoning, accurately citing Civil Code provisions, conducting an in-depth analysis of changed circumstances and contract termination, evaluating the case from multiple perspectives, and providing well-reasoned conclusions and recommendations, reflecting the model's profound understanding of legal concepts, rules, and reasoning.
[0152] Thirdly, this application discloses a target vertical category large-scale model fine-tuning method system, used to execute the knowledge graph-guided vertical category large-scale model fine-tuning method disclosed in the first aspect of this application. The system specifically includes: The knowledge graph layer is used to construct a multi-layered target domain knowledge graph. The knowledge-parameter mapping layer is used to input several target domain concept words obtained from the target domain knowledge graph into the large language model and track the activation state inside the model to establish an explicit correspondence between target domain concepts and model neurons. The sparse parameter fine-tuning layer is used to obtain the neuron regions in the large language model that are sensitive to the target domain knowledge based on the explicit correspondence. The sparse activation strategy is determined according to the complexity of the task to be analyzed. Based on the sparse activation strategy, the actual activation ratio of neurons in the neuron regions is dynamically adjusted, and the adjusted large language model is used as the target vertical category large model.
[0153] A computer device according to an embodiment of this disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Figure 6 This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 6 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein. Figure 6 As shown, a computer device may include a processor, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from a storage device into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0154] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 6 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.
Claims
1. A method for fine-tuning a large target vertical category model, characterized in that, include: Construct a multi-layered target domain knowledge graph; the target domain is the legal domain. Several concept texts associated with sub-concepts of the target domain, obtained from the target domain knowledge graph, are input into the large language model, and the activation state inside the model is tracked to establish an explicit correspondence between target domain concepts and model neurons. Based on the explicit correspondence, obtain the neuronal regions in the large language model that are sensitive to target domain knowledge; Determine the sparse activation strategy based on the complexity of the task to be analyzed; According to the sparse activation strategy, the actual activation ratio of neurons in the neuronal region is dynamically adjusted, and the adjusted large language model is used as the target vertical category large model. The step of inputting several concept texts associated with sub-concepts of the target domain, obtained from the target domain knowledge graph, into the large language model and tracking the activation state within the model to establish an explicit correspondence between target domain concepts and model neurons includes: Based on the target domain knowledge graph, the target domain knowledge is divided into several major categories of knowledge scenarios; Obtain several sub-concepts contained in each of the major categories of knowledge scenarios; Configure several types of concept texts associated with the sub-concept in the target domain, where each type of concept text is a basic concept text, a composite concept text, or a contextually ambiguous concept text; The concept text is input into a large language model to obtain the neuron activation pattern feature vector corresponding to each sub-concept; Constructing neuron activation pattern feature vectors for a large language model; Obtain the similarity between the neuron activation pattern feature vector corresponding to each sub-concept and the neuron activation pattern feature vector of the large language model; Obtain all neurons in each sub-concept whose similarity is greater than a preset similarity threshold, and denote them as key neurons; The key neurons and their corresponding activation thresholds are stored in a neuron set; Organize all the aforementioned sub-concepts into a target domain sub-concept set; Establish an explicit correspondence between the target domain sub-concept set and the neuron set.
2. The method for fine-tuning the target vertical category large model according to claim 1, characterized in that, The construction of a multi-layered target domain knowledge graph includes: Determine the constituent elements of the target domain's map, which include the target domain's concepts, legal provisions, related cases, and legal relationships among the target domain's entities. Construct a target domain knowledge graph based on the graph's constituent elements; The target domain knowledge graph includes several knowledge nodes and the relationship edges between different knowledge nodes.
3. The method for fine-tuning the target vertical category large model according to claim 2, characterized in that, The large language model mentioned is an open-source large language model.
4. The method for fine-tuning the target vertical category large model according to claim 3, characterized in that, The step of inputting the conceptual text into a large language model to obtain the neuron activation pattern feature vector corresponding to each sub-concept includes: A neuron detection system is constructed in a large language model, which is used to capture the activation values of all neurons in the model in milliseconds; Each type of concept text is input into a large language model, and the activation value of each neuron is collected through the neuron detection system. Obtain the average activation value of all activation values corresponding to each neuron; A neuron activation pattern feature vector is constructed for each sub-concept based on the average activation value of all neurons; wherein each dimension of the neuron activation pattern feature vector for each sub-concept is the average activation value for each neuron.
5. The method for fine-tuning the target vertical category large model according to claim 4, characterized in that, The step of obtaining the neuronal regions in the large language model that are sensitive to target domain knowledge based on the explicit correspondence includes: Configure test text for the target domain; The test text is input into the large language model. The real-time activation value of each neuron in the explicit correspondence of the neuron detection system is used to determine the neuron regions in the large language model that are sensitive to target domain knowledge based on the real-time activation value.
6. The method for fine-tuning the target vertical category large model according to claim 5, characterized in that, The test texts include basic concept texts, compound concept texts, and contextually ambiguous concept texts. Each type of concept text includes several positive texts, several negative texts, and several synonymous interference texts. The process of inputting the test text into a large language model, using the real-time activation value of each neuron in the explicit correspondence of the neuron detection system, and determining the neuron regions in the large language model sensitive to target domain knowledge based on the real-time activation value includes: The concept text of each category is input into the large language model in batches. The activation value of each neuron is obtained through the neuron detection system. The first type of activation difference between the positive text and the negative text, and the second type of activation difference between the positive text and the synonymous interference text are compared token by token. Set a significance threshold for the 95% confidence interval, and remove background neurons based on the first type of activation difference and the second type of activation difference; A predetermined number of neurons are randomly selected from the remaining neurons to serve as high-response neurons for the corresponding class of conceptual text; All of the high-response neurons are neuronal regions in the large language model that are sensitive to target domain knowledge; The preset quantity is 0.002%-0.005% of the total number of parameters in the neural network of the large language model.
7. The method for fine-tuning the target vertical category large model according to claim 1, characterized in that, The step of determining the sparse activation strategy based on the complexity of the task to be analyzed includes: Determine the complexity of the task to be analyzed; Based on the stated complexity, determine the complexity level; Based on the complexity level, a sparse activation strategy is determined.
8. The method for fine-tuning the target vertical category large model according to claim 7, characterized in that, Determining the complexity of the task to be analyzed includes: The input task to be analyzed is analyzed from multiple dimensions to determine the multi-dimensional score of the task to be analyzed. Based on the multi-dimensional weights and the multi-dimensional scores, the complexity score corresponding to the task to be analyzed is obtained.
9. The method for fine-tuning the target vertical category large model according to claim 7, characterized in that, The step of determining the complexity level based on the complexity includes: when the complexity meets a first threshold condition, determining the complexity level as a low complexity level; When the complexity meets the second threshold condition, the complexity level is determined to be medium complexity. When the complexity meets the third threshold condition, the complexity level is determined to be a high complexity level.
10. The method for fine-tuning the target vertical category large model according to claim 7, characterized in that, The step of determining the sparse activation strategy based on the complexity level includes: When the complexity level is low, the sparse activation strategy is determined to be the first preset strategy, which includes updating the current sparsity to 150% of the initial sparsity; When the complexity level is medium complexity, the sparse activation strategy is determined to be the second preset strategy, which includes maintaining the initial sparsity. When the complexity level is high complexity level, the sparse activation strategy is determined to be the third preset strategy, which includes updating the current sparsity to 70% of the initial sparsity.
11. A method for analyzing legal information, characterized in that, include: Identify the legal information to be analyzed; The legal information is input into the target large model to generate feedback information; The target large model is the large model after fine-tuning using the target vertical category large model fine-tuning method as described in any one of claims 1-10.
12. A system for fine-tuning a large target vertical model, characterized in that, include: The knowledge graph layer is used to construct a multi-layered knowledge graph for the target domain; the target domain is the legal domain. The knowledge-parameter mapping layer is used to input several concept texts associated with sub-concepts of the target domain, obtained from the target domain knowledge graph, into the large language model and track the internal activation state of the model, establishing an explicit correspondence between target domain concepts and model neurons. Specifically, it includes: dividing target domain knowledge into several major categories of knowledge scenarios based on the target domain knowledge graph; obtaining several sub-concepts contained in each major category of knowledge scenario; configuring several types of concept texts associated with the sub-concepts in the target domain, where each type of concept text is a basic concept text, a composite concept text, or a context-ambiguous concept text; and configuring the concept texts... The input is fed into a large language model to obtain the neuron activation pattern feature vector corresponding to each sub-concept; the neuron activation pattern feature vector of the large language model is constructed; the similarity between the neuron activation pattern feature vector corresponding to each sub-concept and the neuron activation pattern feature vector of the large language model is obtained; all neurons in each sub-concept whose similarity is greater than a preset similarity threshold are obtained and recorded as key neurons; the key neurons and their corresponding activation thresholds are stored in a neuron set; all the sub-concepts are combined into a target domain sub-concept set; an explicit correspondence is established between the target domain sub-concept set and the neuron set. A sparse parameter fine-tuning layer is used to obtain neuronal regions in the large language model that are sensitive to target domain knowledge based on the explicit correspondence. A sparse activation strategy is determined based on the complexity of the task to be analyzed. The actual activation ratio of neurons in the neuronal regions is dynamically adjusted according to the sparse activation strategy. The adjusted large language model is then used as the target vertical category large model.
13. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the target vertical category large model fine-tuning method according to any one of claims 1-10 or the legal information analysis method according to claim 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the target vertical category large model fine-tuning method as described in any one of claims 1-10 or the legal information analysis method as described in claim 11.
15. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the target vertical category large model fine-tuning method as described in any one of claims 1-10 or the legal information analysis method as described in claim 11.
Citation Information
Patent Citations
Machine learning model training method and device
CN116992972A
Domain knowledge-driven large language model fine tuning method, system and equipment and storage medium
CN117056531A