A Fine-tuning Method and Device for Large Language Models Based on Causal Relationship Perception

The large language model is fine-tuned through the method of causal perception, which solves the problem of over-specialization of the model, improves the model's performance in traditional Chinese translation and ancient Chinese reasoning tasks, and achieves efficient model updates and extensive task capabilities.

CN119443182BActive Publication Date: 2025-07-08SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510025727.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-07-08
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

When fine-tuning large language models, the prior art can easily lead to over-specialization of the model, which damages its ability to generalize on unseen tasks, and lacks effective methods to prevent this phenomenon, while ensuring that the model maintains strong general capabilities over a wide range of tasks.

Method used

The large language model is fine-tuned through the method of causality perception, including building a traditional ancient text data set, adjusting the word segmentation strategy of the pre-trained model, incremental training, identifying causal relationships, prioritizing the parameters of the causal feature enhancement layer, setting dynamic learning rates, and performing multiple rounds of iterative optimization to avoid excessive fine-tuning of irrelevant parameters.

Benefits of technology

It significantly improves the performance of the model in traditional Chinese translation and ancient Chinese reasoning tasks, improves training efficiency, ensures that the model maintains strong general capabilities within a wide range of tasks, and reduces computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119443182B_ABST
    Figure CN119443182B_ABST
Patent Text Reader

Abstract

The present invention discloses a fine-tuning method and device for large language models based on causal relationship perception, which relates to the technical field of natural language processing; for the application scenario of ancient Chinese understanding, fine-tuning the large language model, including: Step 1: Prepare data: construct a traditional Chinese ancient text dataset; Step 2: Prepare a pre-trained basic model; Step 3: According to the traditional Chinese ancient text dataset, form a fine-tuning dataset, and expand and process the data in the fine-tuning dataset; Step 4: Load the pre-trained model, and fine-tune the pre-trained Chinese Llama model based on causal relationship according to the fine-tuning dataset; Step 5: Evaluate and optimize the iteratively fine-tuned Chinese Llama model to obtain the performance indicators of the model; Step 6: Select the best model according to the performance indicators of the model. The present invention identifies the causal relationships in the training data through causal analysis and uses this as a basis to guide the efficient update of the model parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a fine-tuning method and device for large language models based on causal relationship perception, which relates to the technical field of natural language processing. Background Art

[0002] Large language models (LLMs) have demonstrated excellent capabilities in various tasks such as question answering and text summarization, and have been widely applied in practical scenarios such as customer service and code assistance. However, when these models are deployed to specific domains, such as the field of classical Chinese understanding, they usually need to be adjusted to meet specific requirements. Currently, fine-tuning large language models on specific datasets is a common practice to improve the performance of target tasks. However, fine-tuning LLMs usually leads to the over-specialization of the model for the fine-tuning task and damages the generalization ability of the model on unseen tasks through in-context learning. There is currently no perfect method to prevent the over-specialization of the fine-tuned model on a few tasks while ensuring that the model still has strong general capabilities within a wide range of tasks. Summary of the Invention

[0003] Aiming at the problems of the existing technology, the present invention provides a fine-tuning method and device for large language models based on causal relationship perception. For the application scenario of classical Chinese understanding, by performing causal analysis on the training data, identifying the causal relationships in the data, and guiding the update of model parameters based on these causal relationships, more efficient fine-tuning can be achieved, and it can prevent the over-specialization of the fine-tuned model on a few tasks while ensuring that the model still has strong general capabilities within a wide range of tasks.

[0004] The specific solution proposed by the present invention is as follows:

[0005] The present invention provides a fine-tuning method for large language models based on causal relationship perception. For the application scenario of classical Chinese understanding, fine-tuning the large language model includes:

[0006] Step 1: Prepare data: Construct a traditional Chinese classical text dataset.

[0007] Step 2: Prepare a pre-trained base model: Use the Chinese Llama model as the pre-trained base model, adjust the tokenization strategy of the Chinese Llama model according to the characteristics of traditional Chinese classical texts, and incrementally train the Chinese Llama model.

[0008] Step 3: Based on the traditional Chinese classical text dataset, form a fine-tuning dataset, and expand and process the data in the fine-tuning dataset.

[0009] Step 4: Load the pre-trained model, and fine-tune the pre-trained Chinese Llama model based on causal relationships according to the fine-tuning dataset.

[0010] Step 41: Establish a causal relationship model, and perform causal analysis of traditional Chinese ancient text sentences and time series analysis of the preceding and following sentences through the causal relationship model.

[0011] Step 42: Construct a structural equation model SEM, and use the structural equation model SEM to determine the causal relationships between multiple features.

[0012] Step 43: Construct a causal relationship network, and realize the causal path between the input variables of traditional Chinese ancient text components and the output variables of modern Chinese translations through the causal relationship network.

[0013] Step 44: Use the causal relationship model and the structural equation model SEM to identify the causal relationships between the data in the fine-tuning dataset, extract the causal features of the fine-tuning dataset, adjust the weights of the causal features according to the strength of the causal relationships, and then adjust the weights of the parameters related to the causal features in the Chinese Llama model.

[0014] Step 45: Select layers in the Chinese Llama model for fine-tuning: preferentially update the parameters of the causal feature enhancement layer, and set the parameter update strategy.

[0015] Do not adjust the layers in the Chinese Llama model with weak or irrelevant causality.

[0016] Step 46: Establish an adaptive fine-tuning strategy: dynamically adjust the parameter update strategy, and perform multiple rounds of iterative optimization for fine-tuning.

[0017] Step 5: Evaluate and optimize the Chinese Llama model after iterative fine-tuning to obtain the performance indicators of the model.

[0018] Step 6: Select the best model according to the performance indicators of the model.

[0019] Furthermore, in step 1 of the method for fine-tuning a large language model based on causal relationship perception, constructing the traditional Chinese ancient text dataset includes:

[0020] S11: Integrate open-source datasets: Collect existing open-source traditional Chinese ancient text datasets, and perform data cleaning and formatting processing.

[0021] S12: Screen general datasets: Screen out the traditional Chinese ancient text parts from existing general Chinese datasets.

[0022] S13: Perform simplified-to-traditional conversion: Convert some general Chinese data into traditional Chinese ancient texts.

[0023] Furthermore, in step 2 of the method for fine-tuning a large language model based on causal relationship perception, incrementally training the Chinese Llama model includes:

[0024] General corpus pre-training: Use a general corpus containing simplified and traditional Chinese ancient texts for pre-training.

[0025] Domain-specific training for traditional Chinese classical texts: Use a traditional Chinese classical text dataset for in-domain training to obtain a pre-trained Chinese Llama model.

[0026] Furthermore, in step 3 of the method for fine-tuning a large language model based on causal relationship perception, according to the traditional Chinese classical text dataset, a fine-tuning dataset is formed, including:

[0027] Construct Q&A data based on the traditional Chinese classical text dataset:

[0028] Design questions: Design questions at different levels, and the question types include: questions about word interpretation, questions about sentence meaning understanding, questions about causal relationship analysis, and questions about time sequence inference.

[0029] Construct Q&A pairs: Based on the modern Chinese translations corresponding to the traditional Chinese classical texts, construct Q&A pairs.

[0030] Furthermore, in step 3 of the method for fine-tuning a large language model based on causal relationship perception, the data in the fine-tuning dataset is expanded and processed, including:

[0031] Screening and auditing: Preliminarily screen the traditional Chinese classical texts in the fine-tuning dataset that have no modern Chinese translations, conduct preliminary translations on the selected traditional Chinese classical texts, select translations with high confidence, and audit the selected translations with high confidence.

[0032] Data augmentation: Perform operations such as synonym replacement, sentence pattern transformation, and context reconstruction on the data in the fine-tuning dataset to achieve data augmentation.

[0033] Furthermore, in step 41 of the method for fine-tuning a large language model based on causal relationship perception, a causal relationship model is established, including:

[0034] Step 411: Conduct causal analysis on traditional Chinese classical text sentences through the causal relationship model to capture the causal relationships in the sentences or the dependencies in text generation.

[0035] Step 412: Conduct time series analysis through the causal relationship model. By analyzing the logical relationships between the previous and subsequent sentences, infer the order of events, and then speculate on the causal relationships.

[0036] Furthermore, in step 45 of the method for fine-tuning a large language model based on causal relationship perception, the parameters of the causal feature enhancement layer are preferentially updated, including: Dynamically adjusting the parameter update rate according to the causal feature importance of the causal feature enhancement layer;

[0037] When setting the parameter update strategy, set the learning rate based on the strength of the causal relationship. If the strength of the causal relationship is high, set a higher learning rate; if the strength of the causal relationship is low, set a lower learning rate.

[0038] Further, in step 46 of the method for fine-tuning a large language model based on causal relationship perception, dynamically adjusting the parameter update strategy includes: verifying the fine-tuned Chinese Llama model using a validation set, obtaining feedback, and adjusting the learning rate according to the feedback.

[0039] Further, in step 5 of the method for fine-tuning a large language model based on causal relationship perception, evaluating and optimizing the fine-tuned Chinese Llama model includes:

[0040] Step 51: Task evaluation:

[0041] Conducting a standardized test on the fine-tuned Chinese Llama model: comprehensively testing the fine-tuned Chinese Llama model using a standardized test set.

[0042] Conducting a cross-task test on the fine-tuned Chinese Llama model: in addition to conducting a classical Chinese understanding task test, also testing the fine-tuned Chinese Llama model on modern Chinese translation tasks, text summarization tasks, and sentiment analysis tasks.

[0043] Step 52: Conducting result comparison and analysis:

[0044] Conducting a comparison before and after fine-tuning: comparing and analyzing the Chinese Llama model before and after fine-tuning.

[0045] Conducting a comparison using an automated tool: using an automated tool to compare the differences between the answers generated by the Chinese Llama model and the standard answers.

[0046] The present invention also provides a device for fine-tuning a large language model based on causal relationship perception. The device fine-tunes the large language model for the application scenario of classical Chinese understanding, including a data preparation module, a model management module, a model fine-tuning module, and an evaluation and optimization module.

[0047] The data preparation module prepares data: constructing a traditional Chinese classical text dataset.

[0048] The model management module prepares a pre-trained base model: using the Chinese Llama model as the pre-trained base model, adjusting the tokenization strategy of the Chinese Llama model according to the characteristics of traditional Chinese classical texts, and incrementally training the Chinese Llama model.

[0049] The data preparation module forms a fine-tuning dataset based on the traditional Chinese classical text dataset, and expands and processes the data in the fine-tuning dataset.

[0050] The model fine-tuning module loads the pre-trained model and fine-tunes the pre-trained Chinese Llama model based on causal relationships according to the fine-tuning dataset.

[0051] Step 41: Establish a causal relationship model, and perform causal analysis on traditional Chinese ancient text sentences and time series analysis on the front and back sentences through the causal relationship model.

[0052] Step 42: Construct a structural equation model SEM, and use the structural equation model SEM to determine the causal relationship between multiple features.

[0053] Step 43: Construct a causal relationship network, and realize the causal path between the input variables of traditional Chinese ancient text components and the output variables of modern Chinese translations through the causal relationship network.

[0054] Step 44: Use the causal relationship model and the structural equation model SEM to identify the causal relationships between the data in the fine-tuning dataset, extract the causal features of the fine-tuning dataset, adjust the weights of the causal features according to the strength of the causal relationship, and then adjust the weights of the parameters related to the causal features in the Chinese Llama model.

[0055] Step 45: Select layers in the Chinese Llama model for fine-tuning: preferentially update the parameters of the causal feature enhancement layer and set the parameter update strategy.

[0056] Do not adjust the layers in the Chinese Llama model with weak or irrelevant causality.

[0057] Step 46: Establish an adaptive fine-tuning strategy: dynamically adjust the parameter update strategy and perform multiple rounds of iterative optimization fine-tuning.

[0058] The evaluation and optimization module evaluates and optimizes the Chinese Llama model after iterative fine-tuning to obtain the performance indicators of the model.

[0059] The model management module selects the best model according to the performance indicators of the model.

[0060] The advantages of the present invention are:

[0061] Identify the causal relationships in the training data through causal analysis, and use this as a basis to guide the efficient update of model parameters. The method of the present invention focuses on updating the features and parameters related to causal relationships, avoiding excessive fine-tuning of irrelevant parameters, significantly improving the training efficiency, being able to perform excellently in traditional Chinese ancient text translation and ancient text reasoning tasks, achieving efficient fine-tuning with less computational cost, and improving work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0063] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.

[0064] Example 1: The present invention provides a method for fine-tuning a large language model based on causal relationship perception. For the application scenario of classical Chinese understanding, the large language model is fine-tuned, including:

[0065] Step 1: Prepare data: Construct a traditional Chinese classical text dataset.

[0066] Among them, constructing a traditional Chinese classical text dataset includes:

[0067] S11: Integrate open-source datasets: Collect existing open-source traditional Chinese classical text datasets, and perform data cleaning and formatting processing.

[0068] S12: Screen general datasets: Screen out the traditional Chinese classical text part from existing general Chinese datasets.

[0069] S13: Perform simplified-traditional conversion: Convert some general Chinese data into traditional Chinese classical texts, such as converting modern general Chinese corpora, ancient literary corpora such as poems and classical Chinese texts into traditional Chinese classical texts.

[0070] Step 2: Prepare a pre-trained base model: Use the Chinese Llama model as the pre-trained base model, adjust the tokenization strategy of the Chinese Llama model according to the characteristics of traditional Chinese classical texts, and incrementally train the Chinese Llama model.

[0071] Among them, incrementally training the Chinese Llama model includes:

[0072] General corpus pre-training: Use a general corpus containing both simplified and traditional Chinese classical texts for pre-training.

[0073] Traditional Chinese classical text domain refinement training: Use the traditional Chinese classical text dataset for in-domain training to obtain a pre-trained Chinese Llama model.

[0074] Step 3: Form a fine-tuning dataset according to the traditional Chinese classical text dataset, and expand and process the data in the fine-tuning dataset.

[0075] Among them, forming a fine-tuning dataset according to the traditional Chinese classical text dataset may further include:

[0076] Construct question-and-answer data according to the traditional Chinese classical text dataset:

[0077] Design problems: Design problems at different levels. The types of problems include: word meaning interpretation problems, sentence meaning comprehension problems, causal relationship analysis problems, and chronological order inference problems. For example, word meaning interpretation problems: What is the meaning of a key word in an ancient classical Chinese text in modern Chinese? For example, sentence meaning comprehension problems: What meaning does this passage express after being translated into modern Chinese? For example, causal relationship analysis problems: What is the cause or consequence of a certain event? For example, chronological order inference problems: In the context, which events occurred first and which occurred later?

[0078] Construct question-answer pairs: Based on the modern Chinese translations corresponding to the ancient classical Chinese texts, construct question-answer pairs. For example, design questions for the paragraphs in the ancient texts, and the explanations in the modern translations are used as answers. This can simulate the process of understanding and translating ancient classical Chinese texts.

[0079] Among them, the expansion and processing of the data in the fine-tuning dataset can include:

[0080] Screening and auditing: Conduct a preliminary screening of the ancient classical Chinese texts in the fine-tuning dataset that do not have modern Chinese translations, conduct a preliminary translation of the selected ancient classical Chinese texts, select translations with high confidence, and audit the selected translations with high confidence;

[0081] Data augmentation: Perform operations such as synonym replacement, sentence pattern transformation, and context reconstruction on the data in the fine-tuning dataset to achieve data augmentation. When performing synonym replacement in the question-answer data, generate more variants by replacing synonyms in the questions or answers. When transforming sentence patterns, adjust the sentence patterns in the questions and answers, such as changing the word order of the sentences or adopting different expressions, while retaining the original meaning. When reconstructing the context, in the translation of ancient texts, the adjustment of the order of some context sentences may affect the sentence meaning. Different context scenarios can be designed to enhance the model's ability to handle complex syntactic structures.

[0082] Step 4: Fine-tune the pre-trained Chinese Llama model based on the causal relationship according to the fine-tuning dataset:

[0083] Step 41: Establish a causal relationship model, and conduct causal analysis of ancient classical Chinese sentences and time series analysis of the front and back sentences through the causal relationship model.

[0084] Among them, establishing a causal relationship model can include:

[0085] Step 411: Conduct causal analysis of ancient classical Chinese sentences through the causal relationship model to capture the causal relationship in the sentences or the dependency relationship in text generation.

[0086] For example: For the ancient classical Chinese translation task, it is necessary to understand the causal relationship in the ancient text translation task. In ancient texts, the sentence structure is different from that in modern Chinese, and it is often more compact and context-dependent.

[0087] Dependency parsing tools can be used to parse dependency relationships in traditional Chinese sentences, for example, to identify the subject, predicate, object, and other sentence components.

[0088] It can also construct causal chains: in the task of translating ancient Chinese texts, it can identify the causal relationships within and between sentences. For example, in ancient Chinese texts, expressions such as "because...so..." and "so..." indicate causal relationships.

[0089] Step 412: Perform time series analysis through the causal relationship model, and infer the order of events by analyzing the logical relationship between the previous and subsequent sentences, and then infer the causal relationship.

[0090] The time sequence of events can be inferred from the context. For example, certain words or phrases, such as "first", "after", "since", and "then", can imply the order of events. Analyze the logical relationship within the sentence, infer the possible time sequence, and thus identify the cause and effect relationship.

[0091] Step 42: Construct a structural equation model (SEM) and use the structural equation model (SEM) to determine the causal relationship between multiple characteristics.

[0092] In the text data, each sentence can be regarded as an observation unit, which contains multiple variables, such as vocabulary, phrases, sentence components, etc. For the traditional Chinese classical Chinese translation task, a structural equation model (SEM) is constructed, in which the input variables are the various components of the classical Chinese sentence, such as the subject-verb-object structure, and the output variable is the modern Chinese translation. Through the structural equation model (SEM), the causal relationship strength between the various components can be quantified and the model parameters can be adjusted accordingly.

[0093] Step 43: Construct a causal relationship network, and use the causal relationship network to realize the causal path between the input variables of the traditional Chinese components and the output variables of the modern Chinese translation.

[0094] Step 44: Use the causal relationship model and the structural equation model (SEM) to identify the causal relationship between the data in the fine-tuning dataset and extract the causal features of the fine-tuning dataset. For example, based on the results of the causal analysis, common causal phrases or structures in traditional Chinese texts can be identified and used as features.

[0095] The weight of the causal feature is adjusted according to the strength of the causal relationship, and then the weight of the parameters involving the causal feature in the Chinese LLama model is adjusted. For example, if a causal structure, such as "因……故……", appears frequently in the dataset and has a significant impact on the translation quality, then the parameter corresponding to the structure will be assigned a higher weight.

[0096] Step 45: Select layers in the Chinese Llama model for fine-tuning: Prioritize updating the parameters of the causal feature enhancement layer and set the parameter update strategy, where the parameter update rate is dynamically adjusted according to the importance of the causal features in the causal feature enhancement layer;

[0097] When setting the parameter update strategy, set the learning rate based on the strength of the causal relationship. If the strength of the causal relationship is high, set a higher learning rate; if the strength of the causal relationship is low, set a lower learning rate.

[0098] Do not adjust the layers with weak or irrelevant causality in the Chinese Llama model, such as freezing the basic vocabulary mapping layer: If the model has already well mastered the basic vocabulary mapping, such as the mapping from a single Chinese character to a modern Chinese word, these layers can be frozen to focus on learning more complex causal relationships.

[0099] For example, freeze the static feature layer: For those features that are stable in different tasks, such as part-of-speech tagging, named entity recognition, etc., the corresponding layers can be frozen to avoid the parameters of these layers being disturbed by unnecessary updates.

[0100] Step 46: Establish an adaptive fine-tuning strategy: Dynamically adjust the parameter update strategy and perform multiple rounds of iterative optimization fine-tuning, where dynamically adjusting the parameter update strategy includes: Using the validation set to verify the fine-tuned Chinese Llama model, obtaining feedback, and adjusting the learning rate according to the feedback.

[0101] The recursive causal chain can also be optimized. In traditional Chinese classical texts, nested causal structures often exist. During the recursive process, the model will first identify simple causal relationships, such as "because...therefore...", and then gradually delve into complex, nested causal structures. For example, in a classical Chinese sentence, there may be multiple events dependent on a prior event. This nested causal relationship needs to be gradually analyzed and optimized through recursion.

[0102] For example: Translate the classical Chinese sentence in traditional characters "Those who wish to govern their country must first regulate their families; those who wish to regulate their families must first cultivate their personal lives; those who wish to cultivate their personal lives must first rectify their minds", which contains a recursive causal chain. First, those who wish to govern their country well must first manage their families, which is a direct causal relationship; then, those who wish to manage their families must first cultivate their personal qualities; those who wish to cultivate their personal qualities must first rectify their minds, which is a further recursive causal relationship. Through recursive optimization, the model first captures the causal relationship that those who wish to govern their country well must first manage their families, and then identifies the causal chain that those who wish to manage their families must first cultivate their personal qualities; those who wish to cultivate their personal qualities must first rectify their minds, and finally recursively updates the relevant parameters to ensure that the causal logic at these levels is accurately conveyed during translation.

[0103] Step 5: Evaluate and optimize the fine-tuned Chinese Llama model to obtain the performance metrics of the model. Among them, evaluating and optimizing the fine-tuned Chinese Llama model includes:

[0104] Step 51: Task evaluation:

[0105] Conduct a standardized test on the fine-tuned Chinese Llama model: Use the standardized test set to comprehensively test the fine-tuned Chinese Llama model.

[0106] Conduct cross-task tests on the fine-tuned Chinese Llama model: In addition to conducting tests on classical Chinese understanding tasks, also conduct tests on modern Chinese translation tasks, text summarization tasks, and sentiment analysis tasks on the fine-tuned Chinese Llama model.

[0107] Step 52: Conduct result comparison and analysis:

[0108] Conduct a comparison before and after fine-tuning: Compare and analyze the Chinese Llama models before and after fine-tuning.

[0109] Conduct a comparison using an automated tool: Use an automated tool to compare the differences between the answers generated by the Chinese Llama model and the standard answers.

[0110] Step 6: Select the best model according to the performance metrics of the model.

[0111] Embodiment 2: The present invention also provides a large language model fine-tuning device based on causal relationship perception. The device fine-tunes the large language model for the application scenario of classical Chinese understanding, including a data preparation module, a model management module, a model fine-tuning module, and an evaluation and optimization module.

[0112] The data preparation module prepares data: Construct a traditional Chinese classical text dataset.

[0113] The model management module prepares a pre-trained basic model: Use the Chinese Llama model as the pre-trained basic model, adjust the tokenization strategy of the Chinese Llama model according to the characteristics of traditional Chinese classical texts, and incrementally train the Chinese Llama model.

[0114] The data preparation module forms a fine-tuning dataset according to the traditional Chinese classical text dataset, and expands and processes the data in the fine-tuning dataset.

[0115] The model fine-tuning module fine-tunes the pre-trained Chinese Llama model based on causal relationships according to the fine-tuning dataset:

[0116] Step 41: Establish a causal relationship model, and conduct causal analysis of traditional Chinese classical text sentences and time series analysis of front and back sentences through the causal relationship model.

[0117] Step 42: Construct a structural equation model (SEM), and use the SEM to determine the causal relationships among multiple features.

[0118] Step 43: Construct a causal relationship network, and implement the causal path between the input variables of traditional Chinese ancient text components and the output variables of modern Chinese translations through the causal relationship network.

[0119] Step 44: Use the causal relationship model and the SEM to identify the causal relationships among the data in the fine-tuning dataset, extract the causal features of the fine-tuning dataset, adjust the weights of the causal features according to the strength of the causal relationships, and then adjust the weights of the parameters related to the causal features in the Chinese Llama model.

[0120] Step 45: Select layers in the Chinese Llama model for fine-tuning: preferentially update the parameters of the causal feature enhancement layer and set the parameter update strategy.

[0121] Do not adjust the layers with weak or irrelevant causality in the Chinese Llama model.

[0122] Step 46: Establish an adaptive fine-tuning strategy: dynamically adjust the parameter update strategy and perform multiple rounds of iterative optimization for fine-tuning.

[0123] The evaluation and optimization module evaluates and optimizes the Chinese Llama model after iterative fine-tuning to obtain the performance metrics of the model.

[0124] The model management module selects the best model according to the performance metrics of the model.

[0125] Regarding the information interaction, execution process, etc. among the above-mentioned modules in the device, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.

[0126] Similarly, the device of the present invention identifies the causal relationships in the training data through causal analysis and uses this as a basis to guide the efficient update of model parameters. The method of the present invention focuses on updating the features and parameters related to causal relationships, avoids over-fine-tuning of irrelevant parameters, significantly improves the training efficiency, can perform excellently in traditional Chinese ancient text translation and ancient text reasoning tasks, achieves efficient fine-tuning with less computational cost, and improves work efficiency.

[0127] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structures described in the above embodiments can be physical structures or logical structures. That is, some modules may be implemented by the same physical entity, or some modules may be implemented separately by multiple physical entities, or some components in multiple independent devices may be jointly implemented.

[0128] The above-described embodiments are merely preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A fine-tuning method for large language models based on causal relationship perception, characterized by For the application scenario of ancient Chinese understanding, fine-tune the large language model, including: Step 1: Prepare data: Construct a traditional Chinese ancient text dataset, Step 2: Prepare a pre-trained base model: Use the Chinese Llama model as the pre-trained base model, adjust the tokenization strategy of the Chinese Llama model according to the characteristics of traditional Chinese ancient texts, and incrementally train the Chinese Llama model, Step 3: Based on the traditional Chinese ancient text dataset, form a fine-tuning dataset, and expand and process the data in the fine-tuning dataset, Step 4: Load the pre-trained model, and fine-tune the pre-trained Chinese Llama model based on causal relationships according to the fine-tuning dataset: Step 41: Establish a causal relationship model, and perform causal analysis on traditional Chinese ancient text sentences and time series analysis of the preceding and following sentences through the causal relationship model, including: Step 411: Perform causal analysis on traditional Chinese ancient text sentences through the causal relationship model to capture the causality in the sentences or the dependency relationships in text generation, Step 412: Perform time series analysis through the causal relationship model. By analyzing the logical relationships between the preceding and following sentences, infer the order of events, and then speculate on the causal relationship, Step 42: Construct a structural equation model SEM, and use the structural equation model SEM to determine the causal relationships between multiple features, Step 43: Construct a causal relationship network, and implement the causal path between the input variables of traditional Chinese ancient text components and the output variables of modern Chinese translations through the causal relationship network, Step 44: Use the causal relationship model and the structural equation model SEM to identify the causal relationships between the data in the fine-tuning dataset, extract the causal features of the fine-tuning dataset, adjust the weights of the causal features according to the strength of the causal relationships, and then adjust the weights of the parameters involving causal features in the Chinese Llama model, Step 45: Select layers in the Chinese Llama model for fine-tuning: Prioritize updating the parameters of the causal feature enhancement layer and set the parameter update strategy, Do not adjust the layers in the Chinese Llama model with weak or irrelevant causality, Step 46: Establish an adaptive fine-tuning strategy: Dynamically adjust the parameter update strategy and perform multiple rounds of iterative optimization for fine-tuning, Step 5: Evaluate and optimize the Chinese Llama model after iterative fine-tuning to obtain the performance metrics of the model, Step 6: Select the best model according to the performance metrics of the model.

2. The fine-tuning method of a large language model based on causal relationship perception according to claim 1, characterized in that In Step 1, constructing the traditional Chinese ancient text dataset includes: S11: Integrate open-source datasets: Collect existing open-source traditional Chinese ancient text datasets, and perform data cleaning and formatting processing, S12: Screen general datasets: Screen out the traditional Chinese ancient text parts from existing general Chinese datasets, S13: Perform simplified-to-traditional conversion: Convert some general Chinese data into traditional Chinese ancient texts.

3. A fine-tuning method for large language models based on causal relationship perception according to claim 1, characterized in that In Step 2, incrementally training the Chinese Llama model includes: Pre-training with general corpus: Use a general corpus containing simplified and traditional Chinese ancient texts for pre-training, Domain-specific training for traditional Chinese ancient texts: Use the traditional Chinese ancient text dataset for in-domain training to obtain the pre-trained Chinese Llama model.

4. A method for fine-tuning a large language model based on causal relationship perception according to claim 1, characterized in that In Step 3, forming the fine-tuning dataset based on the traditional Chinese ancient text dataset includes: Construct question-and-answer data according to the traditional Chinese ancient text dataset: Design problems: Design problems at different levels. The types of problems include: word definition problems, sentence meaning understanding problems, causal relationship analysis problems, and time sequence inference problems. Construct question-answer pairs: Based on the modern Chinese translations corresponding to traditional Chinese ancient texts, construct question-answer pairs.

5. A method for fine-tuning a large language model based on causal relationship perception according to claim 1, characterized in that In step 3, expand and process the data in the fine-tuning dataset, including: Screening and auditing: Preliminarily screen the traditional Chinese ancient texts without modern Chinese translations in the fine-tuning dataset, conduct preliminary translations of the selected traditional Chinese ancient texts, select translations with high confidence, and audit the selected translations with high confidence. Data augmentation: Perform operations such as synonym replacement, sentence pattern transformation, and context reconstruction on the data in the fine-tuning dataset to achieve data augmentation.

6. A fine-tuning method for large language models based on causal relationship perception according to claim 1, characterized in that In step 45, preferentially update the parameters of the causal feature enhancement layer, including: dynamically adjusting the parameter update rate according to the importance of causal features in the causal feature enhancement layer. When setting the parameter update strategy, set the learning rate based on the strength of the causal relationship. If the strength of the causal relationship is high, set a higher learning rate; if the strength of the causal relationship is low, set a lower learning rate.

7. A fine-tuning method for large language models based on causal relationship perception according to claim 1, characterized in that In step 46, dynamically adjust the parameter update strategy, including: using the validation set to verify the fine-tuned Chinese Llama model, obtaining feedback, and adjusting the learning rate according to the feedback.

8. A fine-tuning method for large language models based on causal relationship perception according to claim 1, characterized in that In step 5, evaluate and optimize the fine-tuned Chinese Llama model, including: Step 51: Task evaluation: Conduct a standardized test on the fine-tuned Chinese Llama model: Use the standardized test set to comprehensively test the fine-tuned Chinese Llama model. Conduct a cross-task test on the fine-tuned Chinese Llama model: In addition to conducting tests on the ancient Chinese understanding task, also conduct tests on the modern Chinese translation task, text summarization task, and sentiment analysis task on the fine-tuned Chinese Llama model. Step 52: Conduct result comparison and analysis: Conduct a comparison before and after fine-tuning: Compare and analyze the Chinese Llama models before and after fine-tuning. Conduct a comparison using an automated tool: Use an automated tool to compare the differences between the answers generated by the Chinese Llama model and the standard answers.

9. A fine-tuning device for large language models based on causal relationship perception, characterized by The described device fine-tunes the large language model for the application scenario of ancient Chinese understanding, including a data preparation module, a model management module, a model fine-tuning module, and an evaluation and optimization module. The data preparation module prepares data: Construct a traditional Chinese ancient text dataset. The model management module prepares a pre-trained basic model: Use the Chinese Llama model as the pre-trained basic model, adjust the tokenization strategy of the Chinese Llama model according to the characteristics of traditional Chinese ancient texts, and incrementally train the Chinese Llama model. The data preparation module forms a fine-tuning dataset based on the traditional Chinese ancient text dataset, and expands and processes the data in the fine-tuning dataset. The model fine-tuning module loads the pre-trained model and fine-tunes the pre-trained Chinese Llama model based on causal relationships according to the fine-tuning dataset. Step 41: Establish a causal relationship model, and conduct causal analysis of traditional Chinese ancient text sentences and time series analysis of the previous and subsequent sentences through the causal relationship model, including: Step 411: Conduct causal analysis of traditional Chinese ancient text sentences through the causal relationship model to capture the causality in the sentences or the dependencies in text generation. Step 412: Conduct time series analysis through a causal relationship model. By analyzing the logical relationships between sentences before and after, infer the order of event occurrence, and then speculate on the causal relationships. Step 42: Construct a structural equation model SEM. Use the structural equation model SEM to determine the causal relationships between multiple features. Step 43: Construct a causal relationship network. Through the causal relationship network, achieve the causal path between the input variables of traditional Chinese ancient text components and the output variables of modern Chinese translations. Step 44: Use the causal relationship model and the structural equation model SEM to identify the causal relationships between data in the fine-tuning dataset, extract the causal features of the fine-tuning dataset, adjust the weights of the causal features according to the strength of the causal relationships, and then adjust the weights of the parameters related to the causal features in the Chinese Llama model. Step 45: Select layers in the Chinese Llama model for fine-tuning: Prioritize updating the parameters of the causal feature enhancement layer and set the parameter update strategy. Do not adjust the layers in the Chinese Llama model with weak or irrelevant causality. Step 46: Establish an adaptive fine-tuning strategy: Dynamically adjust the parameter update strategy and conduct multiple rounds of iterative optimization for fine-tuning. The evaluation and optimization module evaluates and optimizes the Chinese Llama model after iterative fine-tuning to obtain the performance metrics of the model. The model management module selects the best model according to the performance metrics of the model.

Citation Information

Patent Citations

  • Ancient Chinese language translation method, system and equipment based on large model and storage medium

    CN118862904A

  • Large language model time dimension optimization method, medium and system

    CN118916447A