Lightweight information extraction method and system based on knowledge distillation and thinking chain
By introducing knowledge distillation and thinking chain mechanisms into information extraction technology, the problems of poor generalization capabilities and high resource demand in the existing technology are solved, and efficient deployment of lightweight models and the accuracy of information extraction are improved.
Patent Information
- Application Number
- CN202411932410.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing information extraction technologies have shortcomings in terms of poor generalization capabilities, strong data dependence, and difficulty in dealing with complex contexts, especially in specific fields or resource-constrained environments.
The lightweight information extraction method based on knowledge distillation and thinking chain is adopted to generate pseudo-data by introducing a small sample learning technology to enhance the generalization ability of the model; the knowledge distillation technology is used to compress the large language model into a lightweight small parameter model, reducing the deployment cost and improving the operability of the model.
It effectively solves the problems of huge model, slow inference speed and high resource demand, and achieves efficient operation on resource-constrained terminal devices while ensuring inference accuracy, improving the accuracy, timeliness and robustness of information extraction.
Smart Images

Figure CN120011533A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information extraction in natural language processing, and in particular to a lightweight information extraction method and system based on knowledge distillation and thought chaining. Background Art
[0002] With the development of the Internet and social media, a large amount of unstructured text data continues to emerge, such as news, social media posts, and internal corporate reports. Extracting useful structured information (such as entities, relationships, events, etc.) from these text data has become an important task in information processing. Especially in specific fields (such as law, medicine, etc.), how to quickly and accurately extract key information from a large amount of text is of great significance for decision support and intelligence analysis.
[0003] Existing information extraction technologies are mostly based on specific rules or models in specific fields for extraction, but the limitations of these methods are: 1. Scarcity of data in specific fields: Traditional information extraction technologies usually rely on a large amount of labeled data for training, but in some specific fields, especially emerging or niche fields, it is often very difficult to obtain high-quality labeled data, which limits the training effect and practical application of the model. 2. High computing resource requirements and stringent deployment requirements: The complexity of existing methods usually requires high computing resources and storage capabilities, especially when dealing with large-scale data sets, which makes their deployment and application in resource-constrained environments challenging. 3. Lack of domain knowledge and timeliness knowledge: Current methods often fail to make full use of domain-specific knowledge and real-time updated information. This lack of knowledge may lead to a decrease in the accuracy and reliability of the model in information extraction, especially in rapidly changing fields.
[0004] In recent years, with the widespread application of large language models (LLM, such as GPT, BERT, etc.), information extraction technology based on LLM has gradually shown its advantages. LLM not only performs well in understanding context, but also can quickly adapt to data in new fields through a small number of examples (few-shot learning) and improve generalization ability. At the same time, combined with technologies such as knowledge distillation, RAG (Retrieval-Augmented Generation) and CoT (Chain-of-Thought), efficient small models can also be generated to reduce deployment costs while ensuring model effects. However, there is still a lack of an information extraction method that integrates large language models with efficient deployment capabilities, which can achieve flexible structured information extraction in different fields. Summary of the invention
[0005] In order to solve the above problems, the purpose of the present invention is to provide a lightweight information extraction technology based on knowledge distillation and thought chain. In view of the problems of existing information extraction methods in poor generalization ability, strong data dependence, and difficulty in handling complex contexts, the present invention introduces few-shot learning technology and generates pseudo data using a small number of labeled examples to enhance the generalization ability and adaptability of the model in different fields, especially when the amount of data is limited or the scene is complex. In addition, the present invention uses knowledge distillation technology to compress large language models into lightweight small parameter models to reduce deployment costs and improve the operability of the model in actual terminal environments.
[0006] In order to achieve the above technical objectives, this application provides a lightweight information extraction method based on knowledge distillation and thought chain, including the following processes:
[0007] Data preprocessing process: preprocess the given text data, including fine-tuning the model using general domain datasets, using few-shot learning for data enhancement, and converting the text data into a digital form that can be understood by computers;
[0008] Knowledge distillation process: select a pre-trained language model for knowledge distillation, generate high-confidence pseudo labels through reasoning, define the distillation loss function, and build the teacher model and student model;
[0009] Deployment and reasoning process: The model after knowledge distillation is deployed to the terminal device, and external knowledge enhancement technology is used to improve the accuracy of information extraction. At the same time, the thinking chain is used to gradually process complex information extraction tasks.
[0010] Preferably, during the data preprocessing process, diverse general domain text data is collected, and after determining the size of the data set, noise and irrelevant information are removed to complete the general domain data set processing.
[0011] Preferably, during data preprocessing, few-sample learning is used for data enhancement, and several representative samples are selected for training. At the same time, the teacher model is used to generate pseudo data related to the target task, and the pseudo data is used as additional training samples to complete the processing of specific domain data sets.
[0012] Preferably, during the data preprocessing process, the text data is converted into a digital form understandable by a computer through character segmentation, case conversion, text digitization, sentence length unification and text length unification.
[0013] Preferably, in the knowledge distillation process, a large pre-trained language model is selected as the teacher model, and a lightweight model is selected as the student model;
[0014] Based on the pre-fine-tuned teacher model, high-confidence pseudo labels and intermediate step information in the thought chain process are generated through reasoning to form a hierarchical knowledge transfer path for the student model to perform knowledge distillation, wherein the distillation loss function is defined by calculating the difference between the student model output and the teacher model output, and the cross entropy loss between the student model CoT reasoning ideas and the teacher reasoning CoT.
[0015] Preferably, during the deployment and reasoning process, the lightweight model after knowledge distillation is deployed to the terminal device to support real-time or offline information extraction tasks;
[0016] Improve the accuracy of the model when performing information extraction tasks by querying external knowledge bases;
[0017] Use multi-step reasoning to gradually process complex information extraction tasks, ensuring the coherence and logic of the reasoning chain;
[0018] By combining external knowledge and step-by-step deduction, ambiguity in information is eliminated, thereby generating accurate information extraction results.
[0019] Preferably, in the process of improving the accuracy of the model, specific information needs are identified by analyzing the content of the input text and generating corresponding query requests;
[0020] Send the generated query request to the relevant external knowledge base to obtain relevant data;
[0021] The information retrieved from the external knowledge base is fused with the model’s own predictions to form the final output.
[0022] Preferably, in the process of ensuring the coherence and logic of the reasoning chain, external knowledge is combined to guide the model to perform step-by-step reasoning by constructing clear prompts, wherein the direction of reasoning is made clear by prompts, and the task is divided into smaller steps;
[0023] Combine the input context with external knowledge and perform logical deduction step by step, where key entities and relationships in the input text are identified and the type of information to be extracted is determined. The RAG module is used to extract external knowledge related to the current task to ensure that the latest facts are processed;
[0024] At each reasoning step, the accuracy and consistency of the previous derivation is ensured by verifying the previous derivation results. Among them, the ambiguity in the information is eliminated by structuring the identified relationships so that the output results are consistent with the context.
[0025] In the last step of the reasoning chain, the final answer is formed by combining all the deduction results.
[0026] The present invention also discloses a lightweight information extraction system based on knowledge distillation and thought chain, which is used to implement the above-mentioned lightweight information extraction method based on knowledge distillation and thought chain, including:
[0027] The data preprocessing module is used to preprocess the given text data, including fine-tuning the model using general domain datasets, performing data enhancement using few-shot learning, and converting the text data into a digital form that can be understood by computers;
[0028] The knowledge distillation module is used to select a pre-trained language model for knowledge distillation, generate high-confidence pseudo labels through reasoning, define the distillation loss function, and build the teacher model and student model;
[0029] The deployment and reasoning module is used to deploy the model after knowledge distillation to the terminal device, use external knowledge enhancement technology to improve the accuracy of information extraction, and use the thinking chain to gradually process complex information extraction tasks.
[0030] The present invention discloses the following technical effects:
[0031] (1) The present invention effectively solves the common problems of large models, slow reasoning speed, and high resource requirements in complex reasoning tasks based on large models. Through knowledge distillation and lightweight model deployment technology, the present invention can run efficiently on resource-constrained terminal devices while ensuring reasoning accuracy. Compared with traditional model compression methods, the hash coding and model pruning technology used in the present invention significantly reduces the memory usage of the model and improves reasoning efficiency.
[0032] (2) The present invention combines external knowledge enhancement (RAG) with chain of thought (CoT) mechanism, which can not only deal with the accurate acquisition of domain-specific or emerging information, but also gradually handle complex information extraction tasks through a multi-step reasoning process. Compared with the traditional static knowledge base query method, the present invention uses dynamic external knowledge fusion and step-by-step reasoning chain construction to enable the model to effectively eliminate ambiguity in the text, thereby improving the timeliness, accuracy and robustness of information.
[0033] (3) This paper designs a complete process from data preprocessing, model training to deployment and reasoning. In particular, in areas where data is scarce, few-shot learning is used for data enhancement, which greatly improves the performance of the model in small sample conditions. Compared with traditional data augmentation methods, pseudo data generated based on the teacher model helps to improve the generalization ability of the model and further enhances the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0035] Figure 1 It is a schematic diagram of the overall framework of the method described in the present invention;
[0036] Figure 2 This is a schematic diagram of the framework of data preprocessing according to the present invention;
[0037] Figure 3 It is a schematic diagram of the framework of the knowledge distillation described in the present invention;
[0038] Figure 4 A schematic diagram of the deployment and reasoning framework of the present invention;
[0039] Figure 5 This is a schematic diagram of the Llama series model described in the present invention. DETAILED DESCRIPTION
[0040] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0041] like Figure 1-5 As shown, the present invention provides a lightweight information extraction technology based on knowledge distillation and thought chain, and the specific steps are as follows:
[0042] Data preprocessing: Preprocess the given text data and perform corresponding operations according to the needs of different fields. When there is less data in a specific field, data expansion is performed through few-shot learning technology to generate pseudo data to ensure the quality and consistency of input data. The specific process includes cleaning and annotation preparation of general field data, as well as few-shot learning and pseudo data generation for specific field data sets.
[0043] Fine-tune the teacher model: For the preprocessed data, use a pre-trained large language model (such as LLaMA3.1-70B) to fine-tune it to adapt to specific tasks and fields, and enhance the model's ability to extract specific information.
[0044] Knowledge distillation technology: In order to reduce the cost and computational overhead of localized deployment, the present invention adopts adaptive knowledge distillation technology, uses a large language model as a teacher model, and obtains a small-parameter student model (such as LLaMA3.2-3B) through distillation to achieve model lightweighting. This process includes reasoning to generate high-confidence pseudo-labels, and supplementing the intermediate steps in the reasoning process through the Chain of Thought (CoT) reasoning chain.
[0045] Structured information extraction and reasoning: In the reasoning stage, the retrieval-augmented generation (RAG) and chain reasoning (CoT) technologies are combined to further improve the effect of structured information extraction. The model obtains the latest information by querying the external knowledge base, and integrates the retrieved external knowledge with its own prediction results to ensure the timeliness and accuracy of the information. At the same time, through the way of thinking chain, the text and external information are gradually analyzed to ensure the accuracy and consistency of the output information. It is suitable for a variety of information extraction tasks such as relationship extraction, event detection and entity linking.
[0046] As a preferred technical solution, the specific steps of data preprocessing include:
[0047] Select general domain datasets: Use rich datasets in general domains to directly fine-tune the model to improve its adaptability to common tasks.
[0048] Data enhancement for specific fields: When the amount of data in a specific field is small, few-shot learning technology is used to enhance the data through a small number of examples, generate pseudo data for training, and improve the model's domain generalization ability.
[0049] Noise removal: such as removing irrelevant content such as HTML tags, special symbols, stop words, etc., and retaining key information related to the task.
[0050] Text segmentation: Based on the characteristics of the target language, use a dictionary or a segmentation algorithm (such as BERT segmentation, Jieba segmentation) to divide the text into words or sub-word units.
[0051] Part-of-speech tagging: Label each word or subword after segmentation with the corresponding part of speech (such as noun, verb, adjective, etc.) to better understand the text semantics in subsequent processing steps.
[0052] Text vectorization: Use word embedding models (such as Word2Vec and BERT embedding) to convert words or subwords in the text into vector representations, which are used as input features for the model. In particular, for data in specific fields, word embedding models in specific fields (such as pre-trained models for fields such as medicine and law) are used to convert vocabulary in specific fields into vector representations, thereby maintaining a strong correlation with knowledge in that field.
[0053] Sample expansion: For domain-specific small sample data, pseudo data is generated through data augmentation techniques (such as random insertion, synonym replacement, random deletion, etc.).
[0054] Context extension: Incorporating domain knowledge, the context of existing data is expanded through generative models (such as GPT and T5) to increase the breadth and depth of the data and ensure that the model can capture the underlying complex relationships in the domain.
[0055] Unify text length: For text samples of different lengths, unify the length to ensure the consistency of input data. For texts that are too long, truncate them; for texts that are too short, use special filling symbols to fill them.
[0056] Case standardization: Convert all text to a uniform lowercase format (this can be ignored for some languages) to reduce the noise interference of the model caused by case differences.
[0057] As a preferred technical solution, the steps of constructing and training the knowledge distillation model include:
[0058] Model selection: In the knowledge distillation process, we first select the appropriate teacher model and student model. The teacher model is a large pre-trained language model such as LLaMA3.1-70B, which has strong contextual understanding and reasoning capabilities and has been fine-tuned in specific fields to adapt to actual application scenarios. The student model is a lightweight model with fewer parameters, such as LLaMA3.2-3B, so that it can run efficiently on resource-constrained devices.
[0059] Reasoning to generate pseudo labels: The teacher model is used to reason about the input data and generate high-confidence pseudo labels containing reasoning chains. In order to further improve the knowledge transfer effect, the Chain of Thought (CoT) reasoning chain is used to generate the reasoning results of the intermediate steps. Such pseudo labels contain both the final output results and the logical level of each step of reasoning, which helps the student model learn more complex reasoning tasks.
[0060] Temperature scaling: Temperature scaling is used in the generated soft labels. By adjusting the temperature parameter T in the Soft max function, the predicted probability distribution becomes smoother. When T>1, the difference between categories decreases, and the student model can better understand the deep information of the teacher model. The common temperature parameter is set to 5 or 10.
[0061] Distillation loss calculation: Distillation loss consists of two parts:
[0062] a. KL divergence loss: Calculate the KL divergence between the soft labels output by the teacher model and the probability distribution output by the student model to measure the degree to which the student model absorbs the knowledge of the teacher model;
[0063] b. CoT loss: Calculate the cross entropy loss between the student model output and the true label to ensure that the student model not only learns the reasoning steps of the teacher model, but also effectively learns the actual task objectives.
[0064] Balanced loss weight adjustment: Use weight factors α and β to adjust the relative weights of KL divergence loss and CoT loss to ensure that the student model can learn the rich knowledge of the teacher model while maintaining efficient performance on specific tasks. Typically, the α value is set to 0.9 to emphasize knowledge absorption, while the β value is set to 0.1 to ensure task completion.
[0065] Model fine-tuning and optimization: Lora technology is used to fine-tune the student model, and the Adam optimizer is used to update the parameters. During the training process, the learning rate is gradually reduced through dynamic learning rate adjustment strategies (such as Cosine Annealing or Learning Rate Decay) to prevent the model from falling into the local optimal solution and improve the efficiency and effect of training. The performance of the model is regularly evaluated on the validation set to ensure that the final distillation effect meets expectations.
[0066] As a preferred technical solution, the steps of model deployment include:
[0067] Lightweight model deployment: In the deployment phase, the goal is to successfully apply the lightweight model trained by knowledge distillation to the terminal device to perform real-time or offline tasks. The specific steps are as follows:
[0068] Model optimization: In view of the computing resource limitations of terminal devices, model pruning and quantization techniques are used to reduce model parameters and computing overhead to ensure model operation efficiency. In order to further reduce memory usage, parameter compression technology based on hash coding is used to map the original model parameters into a more compact form.
[0069] Real-time application: The deployed lightweight model can quickly respond to input data and is suitable for various scenarios such as public opinion analysis, ensuring the efficient application of the model in edge devices or terminal environments.
[0070] As a preferred technical solution, the model reasoning step includes:
[0071] External knowledge augmentation (RAG): During the information extraction process, the reasoning ability of the model is enhanced by querying the external knowledge base. The specific steps are as follows:
[0072] a. Query external knowledge base: The model generates query requests based on the input text and dynamically connects to relevant external knowledge bases, such as official websites or real-time data sources. By acquiring the latest knowledge, the model integrates this external information with its own prediction results to improve the accuracy of the output.
[0073] b. Information alignment: By querying external knowledge bases, the model can use the latest information to correct and supplement the input data to ensure the accuracy and relevance of the output. This process can effectively eliminate ambiguity in the text and enhance the timeliness of reasoning.
[0074] Chain of Thought (CoT) mechanism: It processes complex tasks through multi-step reasoning, which is especially suitable for cross-sentence relation extraction and event detection. The specific steps are as follows:
[0075] a. Prompt construction: Introduce the thought chain mechanism and guide the model to reason step by step by constructing detailed prompts. For example, use the latest information retrieved from the external knowledge base to provide background for the prompt, and generate reasoning results step by step through logical deduction.
[0076] b. Multi-step reasoning: The model combines the input context with external knowledge and derives the output results step by step. In each step, the model analyzes the key entities and relationships in the text, and verifies the correctness of the reasoning by combining external knowledge, gradually reducing the uncertainty in the information and ensuring that the output information is clear and accurate.
[0077] The present invention adopts different processing methods by designing data preprocessing steps and combining the data characteristics of general and specific fields. For general field data, it is directly used for fine-tuning of large language models to improve the performance of models for common tasks; for specific field data, few-shot learning technology is used to generate pseudo data through a small number of examples for data enhancement to ensure the generalization ability of the model in small data scenarios. In addition, the present invention also combines knowledge distillation technology, takes advantage of the large model (teacher model), generates a small parameter model through knowledge distillation, and facilitates efficient deployment on client terminals. In the reasoning stage, the present invention further integrates RAG (Retrieval-Augmented Generation) technology to enhance the accuracy of information extraction through an external knowledge base. When performing information extraction tasks, the model automatically generates query requests for specific information needs by analyzing the input text, dynamically connects with the external knowledge base, obtains the latest data and eliminates ambiguity in the text. At the same time, combined with the CoT (Chain-of-Thought) reasoning mechanism, multi-step reasoning is performed step by step to ensure accurate information extraction in complex contexts. This method not only improves the model's reasoning ability in complex contexts, but also enhances stability and efficiency in real-time applications. Compared with existing similar technologies, this invention not only greatly improves the accuracy and robustness of information extraction, but also reduces the deployment cost of the model through knowledge distillation and model optimization technology, making it more practical.
[0078] Embodiment: The present invention provides a lightweight information extraction method based on knowledge distillation and thought chain, the method comprising the following steps:
[0079] S1: Data preprocessing:
[0080] like Figure 2 As shown, the present invention preprocesses the given text data, mainly including fine-tuning the model using a general domain dataset, and performing data enhancement using few-shot learning for a specific domain dataset to generate more pseudo data.
[0081] The input data preprocessing of step S1 specifically includes the following sub-steps:
[0082] S11: General domain dataset processing
[0083] In the general domain dataset, use the rich text data to perform preliminary fine-tuning of the model. This step includes the following:
[0084] a. Collect a large amount of general domain text data to ensure data diversity and coverage to improve the generalization ability of the model, such as IEPile, ACE05, and NYT. Set the data set size to N, that is
[0085]
[0086] b. Data cleaning: Clean the original data to remove noise and irrelevant information. The data set obtained after cleaning is D clean .
[0087] c. Labeling preparation: If supervised learning is required, prepare the corresponding labeled data L.
[0088] S12: Domain-specific dataset processing
[0089] For datasets in specific fields, due to data scarcity, few-shot learning is used for data enhancement. The specific steps are as follows:
[0090] a. Few-sample learning: select k representative samples for training, and achieve better learning results on a small amount of data through the transfer learning ability of the model.
[0091] b. Pseudo data generation: Use the teacher model to generate pseudo data D related to the target task pseudo , these pseudo data will serve as additional training samples.
[0092] S13: Data Representation Conversion
[0093] In order for the model to process the data, the text data needs to be converted into a form suitable for computers to understand. This step includes:
[0094] a. Character set design: Design a character set C that includes basic characters, Arabic numerals, punctuation marks, and special symbols in the target language to form a dictionary V.
[0095] b. Sentence division: According to the sentence terminators of the target language, the text is divided into multiple sentences to form a set S of sentences.
[0096] c. Text digitization: Using the designed dictionary, the characters in each sentence are converted into corresponding subscript sequences. The final generated digitized text is represented as D numeric .
[0097] Step S2: Establish knowledge distillation model:
[0098] like Figure 3 As shown, the knowledge distillation model of the present invention includes the construction and training of a teacher model and a student model, and the specific process is as follows:
[0099] Step S2 establishes a knowledge distillation model, which specifically includes the following sub-steps:
[0100] S21: Model selection:
[0101] In the process of knowledge distillation, we first select a suitable model:
[0102] a. Teacher model: Select the large pre-trained language model LLaMA3.1-70B. This type of model has strong context understanding and information extraction capabilities, and is fine-tuned on domain data for specific tasks to adapt to specific application scenarios.
[0103] b. Student model: Choose a lightweight model, such as LLaMA3.2-3B, which has fewer parameters and is suitable for deployment on edge devices or terminals. This choice ensures that the student model can run effectively in resource-constrained environments while maintaining high performance.
[0104] S22: Data Preparation
[0105] Data preparation is a key step in knowledge distillation, including:
[0106] Reasoning generation: The teacher model generates high-confidence pseudo-labels by reasoning about specific tasks, and supplements the intermediate steps in the reasoning process through a recursive Chain of Thought (CoT) reasoning chain to form a hierarchical knowledge transfer path. The generated CoT pseudo-label can be expressed as:
[0107] CoT teacher ={s1, s2, ..., sT};
[0108] By explicitly supervising each step of the reasoning process, the stability of the student model in complex reasoning tasks is ensured.
[0109] S23: Training process
[0110] During the training process, a distillation loss function is defined to promote the student model to learn the teacher model knowledge. The specific steps are:
[0111] a. Calculation of distillation loss:
[0112] KL divergence loss: Calculates the KL divergence between the probability distribution of the student model output and the soft label output by the teacher model to measure the degree to which the student model absorbs the knowledge of the teacher model.
[0113] The specific calculation is:
[0114]
[0115] CoT loss: Calculate the cross entropy loss between the student model output and the true label to ensure that the student model closely follows the task goal during the learning process. Its formula is:
[0116]
[0117] Among them, yi is the true label, is the predicted output of the student model.
[0118] b. Comprehensive loss function:
[0119] Loss = α·KL(P teacher , P student )+β·Loss CoT
[0120] Among them, α and β are hyperparameters used to adjust the weights of different losses to balance the goals of knowledge absorption and task learning.
[0121] S24: Model training:
[0122] In the model training phase, the LLaMAFactory platform is used to fine-tune Lora, and the optimization algorithm (such as Adam) is used to train the student model. This process achieves efficient knowledge distillation by gradually adjusting the model parameters to minimize the comprehensive loss function. Specifically, the model is gradually optimized through multiple iterative training so that the performance of the student model is close to the level of the teacher model. To ensure the effectiveness of the training, a dynamic learning rate adjustment strategy (such as CosineAnnealing or Learning Rate Decay) is adopted to dynamically adjust the learning rate according to the progress of training to avoid falling into the local optimal solution and accelerate convergence. The performance of the student model is evaluated on the validation set, and the model training effect is monitored by calculating indicators (such as accuracy, F1-score) so that the training strategy can be adjusted according to the evaluation results to further improve the robustness and generalization ability of the model.
[0123] S3: Deployment and Inference
[0124] like Figure 4 As shown, the deployment and reasoning process of the present invention includes the application of lightweight models, external knowledge enhancement and thinking chain mechanism, and the specific process is as follows:
[0125] Step S3 deployment and reasoning, specifically including the following sub-steps:
[0126] S31: Lightweight model deployment
[0127] In the deployment phase, the goal is to successfully deploy the lightweight model after knowledge distillation to the terminal device in order to perform real-time or offline information extraction tasks. The specific steps are:
[0128] a. Model optimization: For terminal devices in edge computing environments, model pruning and quantization technology is used to ensure the computational efficiency and resource usage of lightweight models. During terminal deployment, memory usage is further optimized through a parameter compression scheme based on hash coding:
[0129] θ compressed =H(θ original )
[0130] Where H(·) is a hash mapping function.
[0131] b. Real-time application: The deployed model can directly process the input data and quickly return the results. It is suitable for various application scenarios, such as public opinion monitoring and text analysis.
[0132] S32: External Knowledge Enhancement (RAG)
[0133] External knowledge enhancement is to improve the accuracy of information extraction by querying external knowledge bases, especially when dealing with domain-specific or newly emerging entities and relations. The specific steps are:
[0134] a. Query external knowledge base: When performing information extraction tasks, the model dynamically connects with external knowledge bases through the following process to enhance the accuracy and timeliness of information extraction. First, the model analyzes the content of the input text and automatically generates query requests for specific information needs. These query requests are sent to relevant external knowledge bases, such as official websites or real-time news sources, to obtain the latest and most relevant data. For example, when processing text related to the 2024 Nobel Prize in Literature, the model queries official information to confirm the latest developments of the winners. Next, the model integrates the retrieved external knowledge with its own prediction results to form the final information output. Specifically, this process can be expressed as:
[0135] External Knowledge=Fuse(RAG retrieved , Prediction model );
[0136] b. Information alignment: By querying the external knowledge base, the model can obtain the latest information related to the Nobel Prize winners and use it as part of the prompt to participate in the next task reasoning. This not only eliminates ambiguity in the original data, but also ensures that the information it outputs is accurate and up-to-date, ensuring the timeliness and relevance of the information, thereby enhancing the overall reasoning ability.
[0137] S33: Chain of Thought (CoT) Mechanism
[0138] Chain of Thought (CoT) gradually handles complex information extraction tasks through multi-step reasoning, which is especially important in cross-sentence relations and event detection. The specific steps are:
[0139] a. Prompt construction: In order to ensure the coherence of the reasoning chain, the present invention introduces a thinking chain mechanism, combined with the external knowledge in the RAG module, to gradually eliminate ambiguity in the information. Prompt helps the model clarify the direction of reasoning, divides the task into smaller tasks step by step through logical deduction, and generates accurate information extraction results step by step to ensure that the output information is clear and accurate. In practical applications, the following Prompt can be used to guide the model to perform relationship extraction tasks: "###External knowledge: <Query [target information], the information obtained is: "[external knowledge content]">###Extract relations from the following text: <[text to be analyzed]>###Please think step by step and show the reasoning process of each step."
[0140] The Prompt construction in the CoT framework can be generated recursively through multi-step reasoning. The CoT reasoning chain is as follows:
[0141]
[0142] b. Multi-step reasoning: The model combines the input context with external knowledge and performs logical deduction step by step to form the final answer. First, the model identifies the key entities and relationships in the input text and determines the type of information to be extracted, such as entities, relationships, and attributes. Next, the RAG module extracts external knowledge related to the current task to ensure that the latest facts can be processed. The model then performs a series of reasoning steps, gradually analyzing the text and external information to find out the associations and potential conflicts between the information. In each step, the model verifies the previous deduction results to ensure their accuracy and consistency, and structures the identified relationships. Through this iterative reasoning and verification process, the model can effectively eliminate ambiguity in the information and ensure that the output information is both accurate and contextual, suitable for various information extraction tasks, such as relationship extraction, event detection, and entity linking. Ultimately, the application of multi-step reasoning not only improves the accuracy of information extraction, but also enhances the model's ability to handle complex tasks.
[0143] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.
[0144] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0145] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A lightweight information extraction method based on knowledge distillation and thought chain, characterized in that: The process includes: Data preprocessing process: preprocess the given text data, including fine-tuning the model using general domain datasets, using few-shot learning for data enhancement, and converting the text data into a digital form that can be understood by computers; Knowledge distillation process: select a pre-trained language model for knowledge distillation, generate high-confidence pseudo labels through reasoning, define the distillation loss function, and build the teacher model and student model; Deployment and reasoning process: The model after knowledge distillation is deployed to the terminal device, and external knowledge enhancement technology is used to improve the accuracy of information extraction. At the same time, the thinking chain is used to gradually process complex information extraction tasks.
2. According to claim 1, a lightweight information extraction method based on knowledge distillation and thought chaining is characterized by: During the data preprocessing process, diverse general domain text data is collected. After determining the size of the data set, noise and irrelevant information are removed to complete the general domain data set processing.
3. According to claim 2, a lightweight information extraction method based on knowledge distillation and thought chaining is characterized by: During the data preprocessing process, few-sample learning is used for data enhancement. Several representative samples are selected for training. At the same time, the teacher model is used to generate pseudo data related to the target task, and the pseudo data is used as additional training samples to complete the processing of specific domain data sets.
4. According to claim 3, a lightweight information extraction method based on knowledge distillation and thought chaining is characterized by: During the data preprocessing process, the text data is converted into a digital form that can be understood by a computer through character segmentation, case conversion, text digitization, unified sentence length and unified text length.
5. According to claim 4, a lightweight information extraction method based on knowledge distillation and thought chaining is characterized by: In the knowledge distillation process, a large pre-trained language model is selected as the teacher model, and a lightweight model is selected as the student model; Based on the pre-fine-tuned teacher model, high-confidence pseudo labels and intermediate step information in the thought chain process are generated through reasoning to form a hierarchical knowledge transfer path for the student model to perform knowledge distillation, wherein the distillation loss function is defined by calculating the difference between the student model output and the teacher model output, and the cross entropy loss between the student model CoT reasoning ideas and the teacher reasoning CoT.
6. According to claim 5, a lightweight information extraction method based on knowledge distillation and thought chaining is characterized by: During the deployment and reasoning process, the lightweight model after knowledge distillation is deployed to the terminal device to support real-time or offline information extraction tasks; Improve the accuracy of the model when performing information extraction tasks by querying external knowledge bases; Use multi-step reasoning to gradually process complex information extraction tasks, ensuring the coherence and logic of the reasoning chain; By combining external knowledge and step-by-step deduction, ambiguity in information is eliminated, thereby generating accurate information extraction results.
7. According to claim 6, a lightweight information extraction method based on knowledge distillation and thought chaining is characterized by: In the process of improving the accuracy of the model, the content of the input text is analyzed to identify specific information needs and generate corresponding query requests; Send the generated query request to the relevant external knowledge base to obtain relevant data; The information retrieved from the external knowledge base is fused with the model’s own predictions to form the final output.
8. According to claim 7, a lightweight information extraction method based on knowledge distillation and thought chaining is characterized by: In the process of ensuring the coherence and logic of the reasoning chain, external knowledge is combined to guide the model to perform step-by-step reasoning by constructing clear prompts. The prompts clarify the direction of reasoning and divide the task into smaller steps. Combine the input context with external knowledge and perform logical deduction step by step, where key entities and relationships in the input text are identified and the type of information to be extracted is determined. The RAG module is used to extract external knowledge related to the current task to ensure that the latest facts are processed; At each reasoning step, the accuracy and consistency of the previous derivation is ensured by verifying the previous derivation results. Among them, the ambiguity in the information is eliminated by structuring the identified relationships so that the output results are consistent with the context. In the last step of the reasoning chain, the final answer is formed by combining all the deduction results.
9. A lightweight information extraction system based on knowledge distillation and thought chain, used to implement a lightweight information extraction method based on knowledge distillation and thought chain as described in any one of claims 1 to 8, characterized in that: include: The data preprocessing module is used to preprocess the given text data, including fine-tuning the model using general domain datasets, performing data enhancement using few-shot learning, and converting the text data into a digital form that can be understood by computers; The knowledge distillation module is used to select a pre-trained language model for knowledge distillation, generate high-confidence pseudo labels through reasoning, define the distillation loss function, and build the teacher model and student model; The deployment and reasoning module is used to deploy the model after knowledge distillation to the terminal device, use external knowledge enhancement technology to improve the accuracy of information extraction, and use the thinking chain to gradually process complex information extraction tasks.
Citation Information
Patent Citations
Text information extraction method based on large language model and efficient parameter fine tuning
CN118132674A
Medical knowledge relation extraction method and system based on large language model fine tuning and retrieval enhancement generation
CN118569263A
Multitask electric power data entity identification method based on knowledge distillation
CN118940757A
Knowledge distillation method, device, equipment, storage medium and computer program product
CN119005175A
Cited By
Baselevel contradictory dispute risk assessment method, system and device and storage medium
CN120355250A
Heat stroke auxiliary judgment system based on DeepSeek large model
CN120636777A
Vehicle defense diagnosis system and method based on knowledge distillation and long thinking chain
CN120744379A
Large language model compression method for power business and related device
CN121303372A
Large Language Model Compression Method and Related Devices for Power Business
CN121303372B