Text entity extraction method based on LLaMA Factory large model fine tuning

Through the LLaMA Factory big model fine-tuning tool, the open source large language model Qwen2-7B-Instruct and LoRA method are used to generate training data sets in combination with manual evaluation and correction, and optimize model parameters, the problem of low efficiency of free text extraction in the existing technology is solved, and efficient and accurate text entity extraction is achieved.

CN120235145APending Publication Date: 2025-07-01UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510372730.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing open source information target extraction scheme is difficult to efficiently process massive free text data with non-fixed structures. The rules and dictionary-based methods are poor in universality. The statistical learning-based methods rely on high-cost annotation data. The deep learning-based methods lack flexibility and efficiency, making it difficult to adapt to complex and diverse application scenarios.

Method used

Through the LLaMA Factory big model fine-tuning tool, the open source large language model Qwen2-7B-Instruct is used to combine manual evaluation and correction to generate training data sets, and fine-tune the model through the LoRA method, set appropriate hyperparameters and real-time monitoring to optimize model performance.

Benefits of technology

It significantly improves the accuracy of text entity extraction in specific fields, reduces computing resources and development cycles, and improves the adaptability and flexibility of the model in specific tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235145A_ABST
    Figure CN120235145A_ABST
Patent Text Reader

Abstract

The invention discloses a text entity extraction method based on LLaMA Factory large model fine tuning, which comprises the following steps: step 1, performing target extraction on open source data to generate a training data set of a large language model; step 2, finishing fine adjustment of the open source large model through an LLaMA Factory tool; step 2-1, preparing training data: putting a training data set and an open source data set into a data directory; 2-2, an LLaMA Factory fine adjustment page is entered; step 2-3, setting specific parameters of fine tuning training; step 2-4, fine tuning and real-time monitoring and adjustment are carried out; and 2-5, completing training and exporting the model: realizing target extraction by utilizing the trained model. According to the method, the LLaMA Factory fine adjustment tool is used, the fine adjustment parameters can be rapidly and simply set, meanwhile, the parameter changes in the fine adjustment process can be observed, and corresponding adjustment can be made in time according to the changes in the fine adjustment process. The fine-tuned model can greatly improve the accuracy of extraction and recognition in a specific field, and the performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for text entity extraction based on fine-tuning of the LLaMA Factory large model, which is mainly applied to entity extraction of news texts. Background Art

[0002] With the advent of the information age, the quantity of open-source information has increased sharply, covering various types of data such as news, social media, academic papers, government reports, technical documents, etc. Especially unstructured data, such as open information in the form of text, pictures, videos, etc., contains a large amount of potential knowledge. How to accurately and efficiently extract information elements related to specific targets from massive unstructured data has become a key challenge in the field of modern information processing and an important direction in target extraction research. Target extraction refers to automatically identifying relevant entities, relationships, events, emotions, etc. from unstructured text and organizing them into a structured form for further analysis and application. In target extraction, semantic analysis plays a crucial role. It provides rich context information by analyzing context, grammatical structure, and semantic levels, supporting the accurate extraction of target information from complex natural language. This requires the model not only to identify single entities but also to accurately understand the relationships between entities and the implicit meanings in the context.

[0003] Modern target extraction methods mainly rely on machine learning and deep learning technologies. Especially the rapid development of deep learning has greatly promoted the progress of target extraction. Deep learning algorithms such as deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTMs), etc. have become the core technologies in the field of target extraction because they can process large-scale data, automatically learn complex patterns, and efficiently capture context relationships. In particular, pre-trained language models (such as BERT, GPT, T5, etc.) have powerful context modeling and language understanding capabilities through large-scale pre-training and fine-tuning processes, and have achieved remarkable breakthroughs in tasks such as named entity recognition and relation extraction. However, high-quality labeled data is the basis for achieving accurate target extraction. In deep learning methods, the training of models usually requires a large amount of labeled data, and the cost of obtaining high-quality labeled data is extremely high. Moreover, the scale, quality, and diversity of the data directly determine the performance of the model. Currently, several commonly used extraction methods are: information extraction techniques based on rules and dictionaries, information extraction techniques based on statistical learning, and information extraction techniques based on deep learning.

[0004] (1) Rule- and dictionary-based information extraction technology: Rule-based methods mainly rely on linguists to formulate rule templates for features such as punctuation marks, demonstrative words, locative words, and positional words. Based on the established knowledge base and dictionary, entity recognition tasks are achieved through pattern and string matching. The earliest methods used for named entity recognition are rule- and dictionary-based methods. They use named entity libraries, rely on manual rules, and assign weights to each rule. This method has high accuracy, but it is too dependent on the construction scale of the knowledge base and dictionary, as well as specific data. It requires business experts to spend a lot of effort in formulating rules, and the rules formulated are difficult to apply to other fields, with poor universality.

[0005] (2) Statistical learning-based information extraction technology: Statistical learning-based methods train by using labeled data to optimize the parameters of the entity recognition model, thereby achieving automatic recognition of named entities. This type of method requires preprocessing of the data, including preliminary processing steps such as text tokenization, part-of-speech tagging, and named entity recognition. Secondly, meaningful features need to be extracted from the original text, such as lexical features, context features, and syntactic features. Then, a machine learning model is trained based on the labeled data to minimize errors and optimize parameters. Common statistical learning models include the maximum entropy model, maximum entropy Markov model, hidden Markov model, support vector machine (SVM), and conditional random field (CRF), etc. Statistical learning methods can learn patterns from a large amount of labeled data and automatically adapt to different application scenarios. However, statistical learning methods rely on a large amount of high-quality labeled data, and the acquisition cost of labeled data is relatively high, especially in professional fields.

[0006] (3) Deep learning-based information extraction technology: Deep learning automatically mines complex semantic features between texts that are difficult for humans to directly describe based on large-scale neural networks, forming an abstract representation, overcoming the disadvantages of time-consuming and laborious feature engineering in traditional statistical machine learning methods. Models trained based on deep learning continuously refresh the results of previous models. For example, frameworks such as LSTM-CRF and BiLSTM-CRF, and models fine-tuned based on basic language representation models such as the Bidirectional Encoder Representation from Transformers (BERT) series and ERNIE series based on bidirectional Transformers have achieved good results in named entity recognition tasks and have been widely applied in actual production tasks. Deep learning-based methods require a large amount of labeled data for training, especially in professional fields such as medical and legal fields, where the acquisition cost of labeled data is high and the data is scarce. At the same time, deep learning models usually require a large amount of computing resources and time for training, and have certain requirements for computing resources.

[0007] Text data in open source information can generally be divided into three categories: structured text, semi-structured text, and free text. Among them, the extraction of structured text and semi-structured text is relatively simple because their formats are relatively fixed and the accuracy of information extraction is high. However, when facing free text, rule-based and dictionary-based information extraction techniques often perform poorly. The data organization structure in free text is irregular and lacks a standardized format, making it difficult for traditional methods to cope with this flexible and changeable data type. Methods based on statistical learning rely on labeled data for training. This method relies too much on data annotation and has the disadvantages of being time-consuming, laborious, and difficult to cope with complex and diverse application scenarios. Information extraction technology based on deep learning can automatically learn features from labeled data without human intervention and can automatically adapt to different fields and tasks. However, this extraction technology usually requires designing and training models from scratch, including selecting network architectures, tuning hyperparameters, etc. From data preparation to model design, training, and testing, the entire development cycle is long, and it lacks flexibility and is difficult to dynamically update knowledge.

[0008] Existing open source information target extraction solutions are difficult to efficiently process massive data with non-fixed structures. Target extraction of data content based on open source large language models can save the complex process of model pre-training, but it is difficult to cope with complex and diverse application scenarios by only setting prompt words. Simple prompt word strategies cannot fully capture the deep information in the data, especially when dealing with complex texts or domain-specific tasks, and often cannot accurately extract the required information. Summary of the invention

[0009] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a text entity extraction method based on LLaMA Factory large model fine-tuning. Using the LLaMA Factory fine-tuning tool, the fine-tuning parameters can be set quickly and concisely, and the parameter changes in the fine-tuning process can be observed, and corresponding adjustments can be made in time for the changes in the fine-tuning process. The fine-tuned model can greatly improve the accuracy of extraction and recognition in specific fields and improve performance.

[0010] The object of the present invention is achieved through the following technical scheme: a text entity extraction method based on LLaMA Factory large model fine-tuning, comprising the following steps:

[0011] Step 1: Generate a training data set for a large language model by performing target extraction on open source data. The specific steps are as follows:

[0012] Step 1-1: Select a high-quality text dataset related to the target field;

[0013] Step 1-2: Preprocess the text data in the dataset: Complete the cleaning operation of the text data through Python code to remove invalid characters; delete duplicate lines and text fragments; delete abnormally short and long texts;

[0014] Step 1-3: Conduct preliminary information extraction on the preprocessed text dataset by calling the open-source large language model Qwen2-7B-Instruct; Guide the model to extract target information by setting prompt words to initially achieve target information extraction in a specific domain;

[0015] Step 1-4: Evaluate and manually correct the extraction results: First, manually review the extraction results, compare them with the original text content, and judge whether the entities extracted in the extraction results are correct, whether the behavioral relationships are reasonable, and whether the extracted data is complete to evaluate the accuracy of the model extraction; Second, for the errors or omissions found in the evaluation, correct them through manual correction; The correct data generated after correction will be used as a new training dataset to provide labeled data for subsequent model fine-tuning;

[0016] Step 2: Complete the fine-tuning of the open-source large model through the LLaMA Factory tool; The specific steps are as follows:

[0017] Step 2-1: Prepare the training data: Put the training dataset obtained in Step 1-4 into the data directory in the LLaMA Factory file directory; At the same time, select an open-source dataset with a number of data entries similar to that of the training dataset as auxiliary training data and put it into the data directory together;

[0018] Step 2-2: Enter the LLaMA Factory fine-tuning page: Open the web page for LLaMAFactory model fine-tuning training through llamafactory-cli; On the web page, select the training stage as Supervised Fine-Tuning, that is, the fine-tuning stage; Configure the model fine-tuning training, including adjusting model parameters, selecting the fine-tuning training task, and selecting the fine-tuning training dataset;

[0019] On the web page of LLaMA Factory, select the model to be trained as Qwen2-7B-Instruct, select the training dataset and the open-source dataset as inputs in the data directory on the page; Select the training method and parameters on the web page; Select the LoRA method for the fine-tuning training method;

[0020] Step 2-3: Set the specific parameters for fine-tuning training. For the fine-tuning training of Qwen2-7B-Instruct, the learning rate is selected as 1e-4, the batch size is set to 32, the gradient accumulation is set to 8, the number of training epochs is selected as 5, the learning rate scheduler is set to cosine, the Warmup steps are set to 500, and the mixed precision training is selected as BF16;

[0021] For the selection of LoRA parameters, the rank of LoRA is selected as 8, the LoRA scaling factor is selected as 16, the LoRA random dropout is set to 0.1, and the LoRA + learning rate ratio is set to 1;

[0022] Step 2-4: Real-time monitoring and adjustment during fine-tuning. During the training process, the Web page will display the real-time training progress and performance metrics. If it is found that the model performance deteriorates or overfitting occurs during the training process, adjust the training parameters and retrain;

[0023] Step 2-5: Completion of training and model export. After the model training is completed, export the trained model file to the local, call it by writing Python code, use the trained model for inference and data processing, and achieve target extraction for text data.

[0024] The beneficial effects of the present invention are as follows: The present invention proposes an information extraction method for fine-tuning an open-source large language model based on LLaMA Factory. Fine-tuning can further train a pre-trained model for a specific task. Through fine-tuning, it is only necessary to train on a small-scale dataset, adjust some parameters of the model to adapt to the new task, skip the large workload of the initial pre-training stage of the model, and save a large amount of human and computing resources. Fine-tuning on a specific domain training dataset can effectively improve the focus of the large model on a specific task, thereby improving its accuracy and effect on the target task. The following beneficial effects can be finally brought:

[0025] 1. The present invention uses the open-source large language model Qwen2-7B-Instruct, and through the method of setting prompt words, conducts preliminary extraction processing on text data, evaluates and corrects the extraction results in combination with manual judgment, and quickly and conveniently generates a training dataset to prepare for subsequent fine-tuning training.

[0026] 2. By using the LLaMA Factory fine-tuning tool, the setting of fine-tuning parameters can be quickly and simply achieved, and at the same time, the parameter changes during the fine-tuning process can be observed, and corresponding adjustments can be made in a timely manner for the changes during the fine-tuning process. The fine-tuned model can greatly improve the accuracy of extraction recognition in a specific domain and improve the performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a flowchart of the text entity extraction method based on fine-tuning of the LLaMA Factory large model of the present invention. Specific implementation manners

[0028] Large model fine-tuning technology: Fine-tuning refers to retraining an existing pre-trained model using data from a specific domain, so that the model performs better on specific tasks. General large models are trained using large-scale general datasets and have powerful language understanding capabilities, but may not achieve optimal performance in specific domains or tasks. Therefore, fine-tuning technology endows large models with more customized functions by training on specific task data, enabling general large models to better adapt to the needs and applications of specific domains. The general process of large model fine-tuning is as follows: Based on the selected relevant dataset and pre-trained model, by setting appropriate hyperparameters and making necessary adjustments to the model, the model is trained using specific task data to optimize its performance. The advantage of large model fine-tuning is the ability to quickly adapt to different task requirements and reduce the time and computing resources required for training from scratch. Through fine-tuning, the model can not only maintain generality but also achieve higher accuracy and effectiveness in specific tasks or domains. Fine-tuning the open-source large language model can achieve accurate extraction of target information at low cost and high efficiency.

[0029] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0030] As Figure 1 shown, a text entity extraction method based on fine-tuning of the LLaMA Factory large model of the present invention includes the following steps:

[0031] Step 1: Perform target extraction on part of the open-source data, and achieve extraction of specific targets by setting prompt words. Evaluate and correct the extraction results of the open-source large language model to generate a training dataset for the large language model. The specific steps are as follows:

[0032] Step 1-1: Select a high-quality text dataset related to the target domain; the text dataset is closely related to the target task to ensure the representativeness and accuracy of the data. When the target task is news event extraction, the text dataset should include rich news reports, social media articles, etc. These texts should cover various types of information required for the target task, such as time, place, person, event, relationship, etc.

[0033] Step 1-2: Preprocess the text data in the dataset to ensure that the input data is clean, in a unified format, and suitable for subsequent processing. Complete the cleaning operation of the text data through Python code, such as formatting the data and removing invalid characters; delete duplicate lines and text fragments to ensure the diversity of the training dataset; delete abnormally short texts (such as overly brief descriptions) and long texts (such as long passages beyond the model's understanding and processing capabilities);

[0034] Step 1-3: Conduct preliminary information extraction on the preprocessed text dataset by calling the open-source large language model Qwen2-7B-Instruct; guide the model to extract target information by setting prompts (Prompts), and initially achieve target information extraction in a specific domain; for example, set the prompt for the large language model as:

[0035] prompt_template = """

[0036] You are an entity recognition expert. Please extract all entities and the relationships between entities from the following text and return the result in JSON format. The JSON contains three fields: Entity 1, Entity 2, and Behavioral Relationship.

[0037] Example input text:

[0038] input_text = """Elon Musk is the CEO of Tesla, Inc."""

[0039] Example output: {

[0040] "Entity 1": "Elon Musk",

[0041] "Entity 2": "the CEO of Tesla, Inc.",

[0042] "Behavioral Relationship": "is"

[0043] }

[0044] """

[0045] Through this prompt template, the model will perform information extraction based on the given input text. The result returned by the model is JSON-formatted data containing entities and their relationships.

[0046] Steps 1 - 4: Evaluate and manually correct the extraction results to ensure their accuracy and consistency. The extraction results are evaluated and corrected by combining the text content of the original dataset with the extraction results of the large model through manual judgment. First, the extraction results are manually reviewed and compared with the original text content to determine whether the entities extracted in the extraction results are correct, whether the behavioral relationships are reasonable, and whether the extracted data is complete, so as to evaluate the accuracy of the model extraction; second, for the errors or omissions found in the evaluation, manual correction is used to make corrections; for example, if the model fails to correctly extract entities or relationships, the manual can supplement and correct the results manually. The correct data generated after correction will be used as a new training dataset to provide labeled data for subsequent model fine-tuning.

[0047] Step 2: Complete the fine-tuning of the open-source large model through the LLaMA Factory tool; Through the above steps, a part of the data is extracted and processed to generate a certain amount of training dataset, and then the fine-tuning of the open-source large model is completed through the LLaMA Factory tool. LLaMA Factory provides a simple and efficient framework for model fine-tuning. Combining the needs of specific domains can significantly improve the task adaptation ability of large language models. The specific steps of fine-tuning are as follows:

[0048] Step 2 - 1: Prepare the training data: Put the training dataset obtained in Step 1 - 4 into the data directory under the LLaMA Factory file directory; At the same time, select an open-source dataset with a number of data entries similar to that of the training dataset as auxiliary training data and put it into the data directory together; The role of this open-source dataset is to help prevent overfitting and ensure that the model can handle a wider range of scenarios during the fine-tuning process, preventing overfitting in subsequent model fine-tuning training.

[0049] Step 2 - 2: Enter the LLaMAFactory fine-tuning page: Open the web page for LLaMAFactory model fine-tuning training through llamafactory-cli; In the web page, select the training stage as Supervised Fine-Tuning, that is, the fine-tuning stage; At the same time, configure the model fine-tuning training, including adjusting model parameters, selecting the fine-tuning training task, and selecting the fine-tuning training dataset; Real-time monitoring can also be achieved to view the progress and performance metrics of the model fine-tuning training.

[0050] In the web page of LLaMA Factory, select the model to be trained as Qwen2-7B-Instruct, and select the training dataset and the open-source dataset in the data directory on the page as the input; select the training method and parameters in the web page; for the fine-tuning training method, select the LoRA method; the LoRA method can adaptively fine-tune the model parameters by introducing a low-rank matrix, which can reduce the computational amount and resource consumption required during training while maintaining the original performance of the model, and is suitable for the fine-tuning of large-scale pre-trained models.

[0051] Step 2-3: Set the specific parameters for fine-tuning training: For the fine-tuning training of Qwen2-7B-Instruct, the learning rate is selected as 1e-4. A moderate learning rate can, to a certain extent, avoid destroying the original knowledge representation of the model, prevent drastic fluctuations in model parameters, and also prevent the model loss from decreasing too slowly in the initial stage of training. Set the batch size to 32. A smaller batch size (such as 16 or smaller) may result in relatively few data samples in each batch, which may cause the model to be easily affected by the bias of individual samples during training, resulting in the model learning incomplete data features. Setting the batch size to 32, each batch contains more comprehensive data samples, which can better represent the characteristics of the entire dataset, thereby reducing data bias, helping the model to learn the laws and trends in the data more comprehensively, and not causing excessive video memory occupation, achieving a balance between performance and resource usage. Set the gradient accumulation to 8 to make up for the shortage of video memory with a larger gradient accumulation. Select 5 for the number of training epochs to avoid overfitting of the model. Set the learning rate scheduler (LR Scheduler) to cosine. This method adjusts the learning rate by simulating the periodic changes of the cosine function, making the learning rate decrease first and then increase as the training progresses. This method can help the model jump out of local optima and find better solutions. Set the warmup steps to 500 to prevent unstable updates of the model weights caused by the learning rate in the initial stage of training. Select BF16 for mixed precision training to reduce video memory occupation and accelerate training.

[0052] For the selection of LoRA parameters, the rank of LoRA is selected as 8, which can find a balance between performance and efficiency and reduce the consumption of the existing system. The LoRA Scaling Factor is selected as 16. A larger scaling factor can increase the impact of the low-rank matrix on the model parameters and improve the task performance of the model. The LoRA Random Dropout is set to 0.1. Discarding a certain proportion of the low-rank matrix weights can improve the generalization ability of the model. The LoRA+Learning Rate Ratio is set to 1 to prevent overfitting.

[0053] Step 2-4, Fine-tuning Real-time Monitoring and Adjustment: During the training process, the Web page will display real-time training progress and performance metrics (such as loss, accuracy, model evaluation results, etc.); if it is found that the model performance deteriorates or overfitting occurs during the training process, adjust the training parameters (such as learning rate, batch size, etc.) and retrain.

[0054] Step 2-5, Training Completion and Model Export: After completing the training of the model, export the trained model file to the local machine, call it by writing Python code, and use the trained model for inference and data processing to achieve target extraction for a large amount of news text data.

[0055] Local Deployment of the Fine-tuned Large Model and Target Entity Extraction: Based on the fine-tuning operation of LLaMA Factory in the previous step, export the fine-tuned model to the local machine and make a new call in the way of Python code.

[0056] The specific steps are as follows:

[0057] (1) Select the model export directory on the LLaMA Factory fine-tuning page and save the trained model files (fine-tuned model weights, configuration files, vocabulary files) to the local machine.

[0058] (2) Write Python code for calling, set the model path to the storage path of the fine-tuned model, and at the same time write the prompt instruction of the model. The following is the Python code for the calling part:

[0059]

[0060]

[0061] (3) In the above code, four parameters are correspondingly passed in, namely the path of the fine-tuned model, the API address for model deployment, the prompt words for model extraction, and the text input data. Finally, by using the trained model and the Python call code, a large amount of text data can be input to obtain the target extraction result.

[0062] In the information extraction method based on the fine-tuning of the open-source large language model (Qwen2-7B) of LLaMA Factory proposed by the present invention, the open-source large language model directly provides a pre-trained model, supports few-shot or even zero-shot learning, reduces the time and cost of model design and training, and at the same time also supports fine-tuning training on large-scale general data, and can well capture context semantics and domain knowledge. The open-source large model provides ready-to-use interfaces and models, and developers can quickly deploy them, significantly shortening the development cycle. At the same time, the LLaMA Factory fine-tuning tool provides an easy-to-use and highly modular framework, and users can customize the model fine-tuning process according to their own needs. The LLaMA Factory fine-tuning tool provides an automated hyperparameter optimization function, which can automatically adjust key hyperparameters such as the learning rate, batch size, and number of training steps according to the goals set by the user (such as minimizing the loss function or improving model accuracy), and helps users find the best parameter combination among multiple candidate configurations. The LLaMA Factory fine-tuning tool supports running fine-tuning tasks in a distributed environment, including multi-GPU and cross-machine training. It uses a distributed training framework and can effectively utilize multiple computing resources to accelerate model training. The LLaMA Factory fine-tuning tool provides real-time training monitoring and performance analysis tools to help users understand the performance changes, resource usage, and training progress of the model during the training process. Key metrics such as the loss curve and accuracy curve can be viewed through a visualization panel. The fine-tuning method based on the LLaMA Factory fine-tuning tool can significantly improve the effect of the open-source large language model in the target information extraction task. Especially when dealing with free-structured text, it can provide higher extraction accuracy and flexibility. At the same time, the efficiency and automation of the fine-tuning process help developers reduce manual intervention, shorten the development cycle, and reduce development costs.

[0063] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on these technical revelations disclosed by the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A text entity extraction method based on fine-tuning of the LLaMA Factory large model, characterized in that: The following steps are involved: Step 1: Generate a training data set for a large language model by performing target extraction on open source data. The specific steps are as follows: Step 1-1: Select a high-quality text dataset related to the target field; Step 1-2: Preprocess the text data in the dataset: Use Python code to clean the text data, remove invalid characters, delete duplicate lines and text fragments, and delete abnormal short and long texts. Step 1-3: Perform preliminary information extraction on the preprocessed text dataset by calling the open source large language model Qwen2-7B-Instruct; By setting prompt words, the model is guided to extract target information, and the extraction of target information in specific fields is initially achieved; Steps 1-4: Evaluate and manually correct the extraction results: First, manually review the extraction results and compare them with the original text content to determine whether the extracted entities in the extraction results are correct, whether the behavioral relationships are reasonable, and whether the extracted data is complete, so as to evaluate the accuracy of the model extraction; secondly, any errors or omissions found in the evaluation are corrected through manual correction; the correct data generated after correction will be used as a new training data set to provide labeled data for subsequent model fine-tuning; Step 2: Use the LLaMA Factory tool to fine-tune the open source large model; the specific steps are as follows: Step 2-1, prepare training data: put the training data set obtained in step 1-4 into the data directory under the LLaMA Factory file directory; at the same time, select an open source data set with a similar number of data entries as the training data set as auxiliary training data, and put it into the data directory; Step 2-2, enter the LLaMA Factory fine-tuning page: open the LLaMA Factory model fine-tuning training web page through llamafactory-cli; on the web page, select the training stage as Supervised Fine-Tuning, that is, the fine-tuning stage; configure the model fine-tuning training, including adjusting model parameters, selecting fine-tuning training tasks, and selecting fine-tuning training data sets; In the LLaMA Factory web page, select the model to be trained as Qwen2-7B-Instruct, select the training dataset and open source dataset as input in the data directory on the page; select the training method and parameters on the web page; select the LoRA method for fine-tuning the training method; Step 2-3, set the specific parameters of fine-tuning training: For the fine-tuning training of Qwen2-7B-Instruct, select the learning rate as 1e-4, set the batch size to 32, set the gradient accumulation to 8, select the training cycle to 5, set the learning rate scheduler to cosine, set the number of warmup steps to 500, and select BF16 for mixed precision training; For the selection of LoRA parameters, the rank of LoRA is selected as 8, the LoRA scaling factor is selected as 16, the LoRA random dropout is set to 0.1, and the LoRA+ learning rate ratio is set to 1; Step 2-4: Fine-tune real-time monitoring and adjustment: During the training process, the web page will display the real-time training progress and performance indicators; if the model performance deteriorates or overfitting occurs during training, adjust the training parameters and retrain; Step 2-5, training completion and model export: After completing the model training, export the model file to the local computer, call it by writing Python code, use the trained model for reasoning and data processing, and achieve target extraction of text data.