Domain Knowledge Fusion Data Augmentation Method for Text Classification

By building a basic corpus from the target domain literature library and optimizing data enhancement large language models with low-rank adaptive technology, the shortcomings of text data enhancement in the existing technology are solved, and high-quality text samples with strong domain adaptability are generated, and the accuracy and generalization ability of text classification models are improved.

CN119646223BActive Publication Date: 2025-07-22DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411803742.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-07-22
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

Existing text data augmentation technologies are unable to effectively combine domain knowledge, resulting in insufficient semantic depth and domain adaptability of the generated samples, and limited data quality and availability.

Method used

By connecting the target domain literature library to build a basic corpus, defining data augmentation prompt words, and using the data augmentation large language model to generate an enhanced corpus, using low-rank adaptive technology for fusion feedback learning, and optimizing the model to generate high-quality text samples with strong domain adaptability.

Benefits of technology

The accuracy and generalization capabilities of the text classification model are improved, the availability and coverage of generated data are enhanced, and the generated samples are significantly improved in terms of semantic depth and domain adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646223B_ABST
    Figure CN119646223B_ABST
Patent Text Reader

Abstract

The domain knowledge fusion data augmentation method for text classification provided by this application relates to the field of data augmentation technology. After connecting to the target domain literature library, it extracts and constructs a basic corpus according to the text classification task, defines data augmentation prompt words, inputs the prompt words and the basic corpus into a data augmentation large language model for data augmentation to generate an augmented corpus, constructs an instruction data set, uses the low-rank adaptation technology to perform fusion feedback learning on the data augmentation large language model, obtains an optimized augmented corpus according to the optimized model, and further obtains a classification result, solving the problem that it is impossible to effectively combine domain knowledge for deep text data augmentation, resulting in insufficient semantic depth and domain adaptability of the generated samples, achieving the effect of effectively generating high-quality text samples with strong domain adaptability, improving the accuracy and generalization ability of the text classification model, and enhancing the usability of the generated data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data augmentation, and specifically to a domain knowledge fusion data augmentation method for text classification. Background Art

[0002] Data augmentation is a common technique in the fields of machine learning and deep learning, used to improve the generalization ability of models, especially under the condition of scarce data. It increases the diversity and scale of the training set by modifying existing data samples or generating new data samples. In the field of computer vision, data augmentation includes operations such as rotation, flipping, and scaling; while in the field of natural language processing (NLP), it includes techniques such as synonym replacement, back translation, and sentence rearrangement to generate new texts that are grammatically correct and semantically consistent.

[0003] Although data augmentation has been widely applied in improving model performance, existing text data augmentation techniques mainly focus on simple text operations, such as synonym replacement or rule-based text transformation. These methods often fail to fully consider the language characteristics and deep semantic structures of specific domains. In addition, these methods have limitations in the diversity and quality of the generated text samples, which may lead to the inability of the augmented data to effectively cover the complexity of the domain, thus limiting the effectiveness and robustness of the model in practical applications.

[0004] In summary, the prior art has technical problems that it cannot effectively combine domain knowledge for in-depth text data augmentation, the richness and coverage of data are poor, resulting in insufficient semantic depth and domain adaptability of the generated samples, and the quality and usability of the generated data are limited. Summary of the Invention

[0005] The present application provides a domain knowledge fusion data augmentation method for text classification, aiming to solve the technical problems existing in the prior art that it cannot effectively combine domain knowledge for in-depth text data augmentation, the richness and coverage of data are poor, resulting in insufficient semantic depth and domain adaptability of the generated samples, and the quality and usability of the generated data are limited.

[0006] The domain knowledge fusion data augmentation method for text classification provided by the present application includes the following steps:

[0007] Connect to the target domain literature library, extract from the target domain literature library according to the text classification task, and construct a basic corpus; define data augmentation prompt words according to the target text classification task, where the data augmentation prompt words include task descriptions, seed sentences, and generation requirements; input the data augmentation prompt words and the basic corpus into a data augmentation large language model for data augmentation to generate an augmented corpus; construct an instruction dataset, and use the low-rank adaptation technology based on the instruction dataset to perform fusion feedback learning on the data augmentation large language model to obtain an optimized data augmentation large language model; according to the optimized data augmentation large language model, obtain an optimized augmented corpus, and obtain the classification result of the text classification task according to the optimized augmented corpus.

[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0009] The domain knowledge fusion data augmentation method for text classification provided in this application constructs a basic corpus by connecting to the target domain literature library and extracting from the target domain literature library according to the text classification task; defines data augmentation prompt words according to the target text classification task, where the data augmentation prompt words include task descriptions, seed sentences, and generation requirements; inputs the data augmentation prompt words and the basic corpus into a data augmentation large language model for data augmentation to generate an augmented corpus; constructs an instruction dataset, and uses the low-rank adaptation technology based on the instruction dataset to perform fusion feedback learning on the data augmentation large language model to obtain an optimized data augmentation large language model; according to the optimized data augmentation large language model, obtains an optimized augmented corpus, and obtains the classification result of the text classification task according to the optimized augmented corpus, solving the technical problem existing in the prior art that it is impossible to effectively combine domain knowledge for in-depth text data augmentation, the richness and coverage of the data are poor, resulting in insufficient semantic depth and domain adaptability of the generated samples, and the quality and usability of the generated data are limited, achieving the technical effect of effectively generating high-quality text samples with strong domain adaptability, improving the accuracy and generalization ability of the text classification model, and enhancing the usability of the generated data. Brief Description of the Drawings

[0010] Figure 1 It is a schematic flowchart of the domain knowledge fusion data augmentation method for text classification provided in this application.

[0011] Figure 2 It is a schematic flowchart of the construction of the instruction dataset in the domain knowledge fusion data augmentation method for text classification provided in this application. Detailed Embodiments

[0012] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0013] To make the solution of this application more comprehensively described, the implementation of data augmentation and its scenario application in text classification tasks are now described:

[0014] Currently, data augmentation in natural language processing mainly uses linguistic knowledge and statistical methods to expand and enrich training data. The main process is as follows: First, perform linguistic analysis and understanding on the original text data. Then, based on predefined language rules and corpus statistical information, identify the parts of the text that can be replaced, recombined, or transformed. Next, perform specific augmentation operations, usually including methods such as synonym replacement, back translation, and sentence recombination. Among them, during the synonym replacement process, according to the context and semantics of the words, select appropriate candidate words from the synonym library for replacement to ensure that the sentence after replacement retains its original meaning. During the back translation process, first translate the original text into another language, and then translate it back to the original language, thereby generating sentences with the same meaning but different expressions. The sentence recombination method mainly creates new expressions by adjusting the sentence structure or word order while maintaining the original semantics.

[0015] As a key technology in text classification tasks, data augmentation plays an important role in improving model performance and solving data-related problems. As an effective means of training data expansion, it can significantly improve the learning process of the model. Specifically, in existing text classification tasks, the performance of the model is usually directly related to the amount of available training data. By applying augmentation methods such as synonym replacement, back translation, or sentence recombination, the model can generate a large number of new training samples based on limited original data. These newly generated samples provide diverse expression forms while maintaining the original semantics, thus expanding the learning scope of the model. Secondly, data augmentation can improve the generalization ability of the model, which helps to improve the accuracy and stability of the classification results. In the specific text classification process, the model needs to learn to recognize the same semantics under different expression forms. The diverse samples generated by the augmentation technology enable the model to be exposed to a wider range of language variants, thereby enhancing its understanding ability of different expression forms. In addition, data augmentation can effectively alleviate the problem of unbalanced sample distribution. In many text classification tasks, the number of samples in different categories may vary significantly. By specifically augmenting the samples in the minority categories, the sample distribution among different categories can be balanced, thereby improving the model's recognition ability for the minority categories and increasing the overall classification accuracy.

[0016] Based on the above, an overview of the data augmentation process for text classification tasks and its implementation and application are introduced. Next, a detailed introduction to the domain knowledge fusion data augmentation method for text classification in this application will be presented.

[0017] Embodiment, such as Figure 1 As shown, this application provides a domain knowledge fusion data augmentation method for text classification, and the method includes:

[0018] Connect to the target domain literature library, extract from the target domain literature library according to the text classification task, and construct a basic corpus.

[0019] Specifically, first establish a data interface with the selected literature database, which usually involves using APIs (Application Programming Interfaces) or web crawler technologies. Access to existing literature resources in the target field is achieved by establishing a data interface connecting to the literature database in the target field. This literature database contains a wide range of professional literature, covering various research papers, case studies, review articles, etc. in the target field. The literature database in the target field is a collection of literature specifically established for a certain specific field (such as medicine, law, or engineering, etc.). Through this database, the aim is to extract content directly relevant to the text classification task. Analyze the document structure in the literature database in the target field according to a specific text classification task, such as topic recognition. This analysis process includes identifying key parts in the literature, such as abstracts, keywords, paper titles, research results, and discussions, etc., so as to accurately locate the text paragraphs relevant to the classification task. For example, in an analysis task for medical literature, special attention will be paid to the research results section to extract descriptions about the effects of drugs. Next, extract positive corpus samples and negative corpus samples from these structurally parsed contents. Positive corpus samples refer to text paragraphs closely related to the text classification task, and these samples directly correspond to the target categories of the classification task; while negative corpus samples are usually texts that are close in context position to the positive samples in the same document but do not belong to the target classification in terms of content. In this application, taking the text classification task of scientific value sentences in the biological field as an example, an artificial annotation and a generative data augmentation method based on a large model are used to construct the corpus. To balance the training dataset, reduce model bias, and improve the effectiveness of evaluation metrics, the ratio of positive and negative corpus samples constructed is set to 1:1. Specifically, the corpus construction process is as follows: First, use the method of artificial annotation to select sentences describing scientific value from the full texts of 1,213 scientific and technological literatures as positive samples of the training set, totaling 551 sentences, and select non-research question sentences from the contexts near these sentences as negative samples of the training set, totaling 551 sentences. Then, use the annotated artificial value sentences and prompt words to jointly construct the Prompt required for generating by the large language model, that is, the basic corpus. Finally, these extracted positive and negative samples are jointly constructed into a basic corpus, which will be used as the primary input for training the data augmentation model, providing rich and highly targeted training materials for subsequent data augmentation and model training processes. In this way, it is ensured that high-quality, field-related training data is obtained in the initial stage of data augmentation, providing a solid foundation for achieving efficient and accurate text classification.

[0020] Define data augmentation prompt words according to the text classification task of the target, and the data augmentation prompt words include task descriptions, seed sentences, and generation requirements.

[0021] Among them, the task description is used to define the goal of generative augmentation in the target field, the seed sentence is an example of the text classification task in the target field, and the generation requirements include the style, grammar, and semantic relevance of the generated text.

[0022] Optionally, defining data augmentation prompting words involves creating specified texts for guiding the generation tasks of data augmentation large language models. This process includes three main components: task description, seed sentences, and generation requirements. Each part is designed for specific goals to ensure that the generated data can effectively support the requirements of text classification tasks. The task description is set to clarify the goal of generative augmentation in the target field. It summarizes the overall task purpose of data augmentation, which is to generate sentences with different expression forms but semantically related to each seed sentence to enrich the dataset. This part tells the model the specific tasks to be completed. For example, in a task of classifying medical research literature, the task description may be "generate descriptive texts that can reflect the treatment effects of different drugs", which helps the model understand that the core purpose of its generation task is to highlight the effects of drugs. The seed sentences are example sentences for text classification tasks in the target field, which serve as the context and style benchmark for the data augmentation model. The seed sentences should be carefully selected from high-quality texts in the target field to ensure that they can represent the typical language usage and professional knowledge expression in the field. For example, if the target field is environmental science, the seed sentence may be "This study demonstrates the potential of new solar cells in reducing carbon emissions", such a seed sentence provides a professional context for scientific research and environmental protection technologies. The generation requirements detail the style, grammar, and semantic relevance standards that the generated text should follow. These requirements ensure that the generated text not only conforms to professional standards in form but also shows a high degree of relevance to the seed sentences in content and is innovative. The generation requirements guide the large language model to make innovative expressions while maintaining the original semantics. For example, the requirements may include "use a formal scientific paper style, maintain precise grammar, and at the same time, must introduce scientific concepts that are different but related to the original seed sentence". The data augmentation prompting words obtained by following the above steps can guide the large language model to effectively generate texts that meet the requirements of the professional field and have high originality and adaptability, thus supporting complex text classification tasks. The implementation of this method not only improves the performance of the classification model but also enhances the model's understanding and processing ability of texts in specific fields.

[0023] Input the data augmentation prompting words and the basic corpus into the data augmentation large language model for data augmentation to generate an augmented corpus.

[0024] Exemplarily, the data augmentation process is carried out by inputting the defined data augmentation prompting words and the basic corpus into the data augmentation large language model. This process aims to generate a more abundant and diverse augmented corpus to support more efficient text classification tasks. First, the combination of the data augmentation prompting words and the basic corpus provides a comprehensive input set, which includes specific task instructions and actual text instances. The task description, seed sentence, and generation requirements included in the data augmentation prompting words are all used to guide the model on how to generate and modify text. For example, the task description may specify generating text describing the therapeutic effect of a drug, the seed sentence provides a specific text sample, and the generation requirements detail the specific language style and semantic quality that the generated text should follow. Then, these prompting words and the text samples extracted from the basic corpus are input into the data augmentation large language model, which is an advanced machine learning model, usually based on deep learning techniques such as Transformer networks, and can understand and generate human language. In this step, the model uses the input prompting words as a guide and generates new text by learning the language patterns and context information in the basic corpus. In this process, the model not only repeats the original text but also creatively generates new text to meet the style and semantic relevance criteria specified in the generation requirements, which may include varying, expanding, or re-expressing the seed sentence to generate diverse text outputs, and these outputs are intended to expand the original corpus and increase its diversity. For example, if the seed sentence describes the therapeutic effect of a drug, the augmentation model may generate multiple sentences with different expressions that convey the same concept from different perspectives or using different languages, thus enriching the content and forms of expression of the corpus. Finally, the generated augmented corpus contains the text in the original corpus and the newly generated text, which are more diverse and creative, and helps to improve the performance and accuracy of subsequent text classification tasks. This augmented corpus can be directly used to train the text classification model or as a basic dataset for further research and analysis. Through this method, the data augmentation large language model is effectively used to combine theory and practice, not only improving the model's adaptability to texts in a specific field but also significantly enhancing the quality of the dataset and the training effect of the model.

[0025] Construct an instruction dataset, and based on the instruction dataset, adopt the low-rank adaptation technology to perform fusion feedback learning on the data augmentation large language model to obtain an optimized data augmentation large language model.

[0026] Furthermore, constructing an instruction dataset involves screening high-quality text samples from an augmented corpus and formulating specific learning tasks for each sample. The instruction dataset includes task descriptions, input data, and expected output data. The task description defines the goals and requirements of model training. For example, generating text descriptions of treatment effects related to specific medical conditions. The input data is the selected high-quality text samples, while the output data is the text that the model is expected to generate, usually an improved or modified version, demonstrating different expressions or deeper information. For example, to further improve the quality of generated scientific value sentences, an instruction dataset is constructed for model fine-tuning by combining manually annotated corpora and screened high-quality generated corpora. Specifically, continuing with the biological field as an example, 1134 scientific value sentences are generated using 551 manually annotated scientific value sentences as seed sentences, and then 623 high-quality generated scientific value sentences are obtained through manual screening. These seed sentences cover the theoretical importance of research questions, the innovation of methods, and the applicability of results, etc. An instruction fine-tuning dataset is constructed based on 1174 high-quality scientific value sentences. The instruction fine-tuning dataset consists of three parts: task description, input, and output. The task description clearly states what the goal of the generation task is, what tasks the model needs to complete, what specific requirements and constraints there are, etc., providing clear guidance for subsequent input and output. The input provides the context information or prompts for the model to generate. The input data matches the task description, providing the necessary background information for the model to complete the task. The output is the target text, image, etc. generated by the model according to the task description and input. The output should meet the requirements in the task description and achieve the expected effect. Next, the low-rank adaptation technique (LoRA) is used to train and optimize the data-augmented large language model. The LoRA technique adjusts and optimizes the weight distribution between layers by introducing low-rank matrices in each network layer of the model, which helps the model to more precisely adapt to specific task requirements without adding excessive computational burden. During implementation, the weights of each network layer are optimized with low-rank perturbations, which is achieved through low-rank projection matrices that capture and enhance signals particularly important for specific tasks. For example, if the task is to generate descriptions of the treatment effects of specific diseases, certain layers of the model may need to have a stronger response ability to terms or sentence patterns related to the mechanism of action of drugs. In this case, LoRA can adjust the weights of these layers so that they are more inclined to generate expressions related to the mechanism of action of drugs. During the optimization process, the low-rank projection matrices are continuously adjusted by comparing the model output before and after training with the expected output until the model output meets the preset quality standards. This method not only enhances the specific task adaptability of the model but also improves the accuracy and relevance of the generated text.Finally, the data-augmented large language model optimized by this low-rank adaptation technique can more effectively process and generate text closely related to the target task, greatly improving the quality of the corpus and the practicality of the model. This optimized model provides more accurate and diverse training data for subsequent text classification tasks, thus achieving higher performance and better results in practical applications.

[0027] According to the optimized data-augmented large language model, an optimized augmented corpus is obtained, and the classification result of the text classification task is obtained according to the optimized augmented corpus.

[0028] Specifically, the process of generating the augmented corpus by the optimized data-augmented large language model involves applying the model previously fine-tuned by the low-rank adaptation technique (LoRA) to the base corpus. This model has been adjusted for a specific task to make it more suitable for generating task-related text. For example, if the model is trained to generate medical-related text, at this stage, the model will be able to produce detailed text describing different diseases and treatment effects, which not only has the same language style as the original text but also is richer and more diverse in content. Then, the generated augmented corpus contains the original text data and the text newly generated by the model. This library now has the depth and width that the original corpus does not have, providing a more abundant training basis for the text classification task. The text in the augmented corpus is optimized by the model to better meet the requirements for data diversity and complexity in practical applications. For example, in legal text classification, it may include more variant descriptions of specific legal terms, improving the accuracy of automatic classification of legal documents. Finally, this optimized augmented corpus is used to obtain the classification result of the text classification task. At this stage, a classification algorithm suitable for the specific classification task, such as support vector machine (SVM), deep neural network, etc., is used to train the classification model according to the data in the augmented corpus. For example, in a classification task aimed at distinguishing the research types in scientific research articles, the model will learn how to judge whether it belongs to theoretical research, experimental research or review type according to the language use and content in the article. During the training process, the algorithm will learn various expressions and the use of terms from the augmented corpus, so as to be able to perform the classification task more accurately. Through the above process, not only the accuracy of text classification is significantly improved, but also the generalization ability and practicality of the model are enhanced by generating more abundant and relevant training data.

[0029] Furthermore, as Figure 2 shown, an instruction dataset is constructed, including:

[0030] Construct a quality screening module. Among them, the quality evaluation indicators of the quality screening module include the importance indicator, innovation indicator, and application indicator for characterizing the corpus. Evaluate the enhanced corpus according to the quality screening module, and screen out the screened corpus with a quality greater than or equal to the preset quality. According to the screened corpus, construct the instruction dataset, and the instruction dataset includes a task description, as well as input data and output data corresponding to the task description.

[0031] Furthermore, the purpose of constructing the instruction dataset is to ensure that the training data used by the data augmentation model meets the highest quality standards. This process involves three main steps: constructing a quality screening module, evaluating and screening the enhanced corpus, and constructing the instruction dataset based on the screening results. First, the quality screening module is constructed to systematically evaluate the quality of each text data in the enhanced corpus. This module sets several key quality evaluation indicators, including the importance indicator, innovation indicator, and application indicator for characterizing the corpus. Among them, the importance indicator for characterizing the corpus is used to evaluate the centrality or criticality of the text content in its field. For example, in medical literature, the importance of descriptions of specific diseases or treatment methods may be concerned; the innovation indicator is used to evaluate the novelty of the text content relative to known information. For example, an article proposing a new method for synthesizing drugs will score higher on this indicator; while the application indicator is used to evaluate the practicality and actual application potential of the text content, such as a detailed description of the application plan of a new technology in a technical document.

[0032] Next, use the quality screening module to evaluate the enhanced corpus. This step involves comprehensively scoring each piece of data in the library and screening out eligible corpus according to the preset quality standard to ensure that only the highest-quality texts are selected into the final training set, thereby improving the learning efficiency of the data augmentation model and the quality of the final output. For example, the screening process may exclude texts with low scores on the innovation indicator and retain texts with rich content and unique insights. Finally, construct the instruction dataset based on the screened corpus. This dataset not only includes the texts selected from the screened corpus as input data, but also designs corresponding output data, and these output data reflect the expected model responses. Each piece of input data is accompanied by a clear task description to guide the model on how to process this data. The task description may be specific to how to rewrite a sentence to increase its persuasiveness, or how to reorganize information to improve the logic of the text. For example, if the input data is a description of a treatment method for a certain disease, the corresponding output data may require the model to generate a more detailed and comprehensive version of the treatment plan. By this method, it is ensured that each piece of data used for training the data augmentation model is finely screened and optimized, thereby maximizing the training effect and the practical application value of the model.

[0033] Furthermore, based on the instruction dataset, a low-rank adaptation technique is used to perform fusion feedback learning on the data-augmented large language model. The method includes:

[0034] Parse multiple network layers of the data-augmented large language model; introduce a group of low-rank projection matrices on each of the multiple network layers to obtain multiple groups of low-rank projection matrices corresponding to the multiple network layers, where each group of low-rank projection matrices includes a first projection matrix and a second projection matrix; perform low-rank perturbation optimization on the weights of the multiple network layers according to the multiple groups of low-rank projection matrices until the optimization objective is met, and obtain the optimized data-augmented large language model.

[0035] Specifically, perform a structural analysis on the used data-augmented large language model to determine the functions and characteristics of each network layer, including identifying the specific tasks (such as feature extraction, semantic analysis, etc.) undertaken by each layer, and their roles in data processing. For example, the primary layer may mainly be responsible for processing basic language structures, while the more advanced layers process complex semantic associations. Subsequently, introduce a set of low-rank projection matrices, including the first and second projection matrices, on each parsed network layer. These matrices are designed to redefine and optimize the information transfer path between layers so that the model can more sensitively adjust its response to adapt to specific task requirements. Among them, the first projection matrix is responsible for capturing the key aspects of the input features, while the second projection matrix adjusts how these features affect the processing of subsequent layers. Next, use the group of low-rank projection matrices to optimize the weights of each network layer. This process iteratively adjusts the values in the matrix to minimize the objective function, usually the prediction error or classification error. The optimization process continues until the model output reaches the preset performance standard or optimization objective. For example, in a task of automatic classification of legal documents, the optimization may focus on improving the model's ability to analyze the relationships between different legal concepts. Through the above steps, the low-rank adaptation technique enables the data-augmented large language model to be more precisely optimized for specific tasks while maintaining its original complexity and depth. This method not only improves the operation efficiency and output quality of the model, but also makes it show better adaptability and accuracy in practical applications. In addition, due to the use of low-rank matrices, the computational resources required in the training and adjustment process of the model can be effectively controlled, further enhancing the practicality of the model.

[0036] Furthermore, perform low-rank perturbation optimization on the weights of the multiple network layers according to the multiple groups of low-rank projection matrices. Among them, the expression for performing low-rank perturbation optimization on the weight of any network layer includes:

[0037] h i ′=f i (x i )+A i B i fi (x i ) = f i (x i ) + Δ i ; where h i ′ is the forward calculation formula after correction for the i-th network layer of, h i = f i (x i ), f i (x i ) is the original forward calculation formula for the i-th network layer, A i is the first projection matrix for the i-th network layer, B i is the second projection matrix for the i-th network layer, Δ i = A i B i f i (x i ) represents the introduced low-rank perturbation correction term.

[0038] Optionally, by introducing a specific increment in each network layer of the model to improve the overall performance of the model, this method is particularly suitable for data-augmented large language models that require fine-grained adjustment. Specifically, h i ′ = f i (x i ) + A i B i f i (x i ) = f i (x i ) + Δ i describes how to adjust the output h i for each layer. Here, f i (x i ) represents the output of the i-th layer of the original model, A i and B i are two low-rank projection matrices specifically designed for this layer. These two matrices generate an increment Δ i (x i ), i.e., the original output of this layer. This increment is a modification based on the original output and is used to fine-tune the behavior of the model. Specifically, A i is designed to capture and enhance the information most critical to the specific task for this layer. Next, by applying these low-rank projection matrices, the increment Δ i is calculated and added to the original output f i (x i ) to form the adjusted output h i (x i ) + Δ i'. The purpose of this process is to enhance the model's response ability to specific types of inputs through subtle adjustments without significantly increasing computational complexity. In this way, each layer of the entire model is gradually optimized, making the final model output more suitable for specific application requirements, such as more accurately identifying and classifying complex text data. The implementation of this method not only improves the model's accuracy but also enhances the model's ability to handle specific tasks, thus achieving better performance in practical applications.

[0039] Furthermore, the optimization objective is constructed by minimizing the loss function, where the expression of the loss function includes:

[0040] where θ is the fixed parameter of the data-augmented large language model before optimization, represents multiple introduced groups of low-rank projection matrices, L is the total number of multiple groups of low-rank projection matrices, is the training dataset based on the instruction dataset, (x, y) ∈ D represents a single sample x and the corresponding label y in the training dataset, l is the cross-entropy loss based on the text classification task, y is the true label, used to measure the error between the predicted value and the true label, is the prediction function optimized based on the data-augmented large language model and the group of low-rank projection matrices.

[0041] Furthermore, by introducing a specific optimization strategy, the performance of the data-augmented large language model is further improved. This strategy mainly defines a loss function for the parameters θ in the model and each layer's low-rank projection matrices A i and B i for optimization. The expression of the loss function is The calculation of the loss function is based on the prediction error on the entire dataset Specifically, this function accumulates the prediction errors of each sample pair in the dataset, where represents the predicted output of the input x by the optimized model, and y is the corresponding true label. This method evaluates the model's performance by comparing the differences between the predicted output and the true output, with the goal of minimizing these differences. The optimization objective is to adjust the model parameters θ and the inter-layer projection matrices A i and B i to ensure that the entire model exhibits the best performance on specific tasks. Each low-rank projection matrix A i and B i is responsible for fine-tuning the internal representation of the model to more effectively process specific types of data features. By adjusting these matrices, the model can more precisely adapt to the complex patterns and changes in the data. The optimization process involves iteratively adjusting θ and on the entire training set Value. In each iteration step, calculate the value of the loss function under the current model parameters, and update the parameters through gradient descent or other optimization algorithms to gradually reduce the prediction error. This includes adjusting each layer's A i and B i to ensure that these matrices can effectively adjust the layer's output to meet the specific requirements of the training data. In this way, the model can not only more accurately identify the sentiment tendency in the text, but also better understand the subtle differences in sentiment expression. Through the above methods, the performance of the data-augmented large language model can be effectively improved, enabling it to provide more accurate and reliable outputs in specific applications.

[0042] Furthermore, extract from the target domain literature library according to the text classification task to construct a basic corpus. The method includes:

[0043] Perform structured parsing on each literature in the target domain literature library and output the structured parsing content; extract from the structured parsing content according to the text classification task to obtain positive corpus samples and negative corpus samples. Among them, the positive corpus samples are a corpus set that matches the text classification task, and the negative corpus samples are a corpus set of the context corresponding to the positive corpus samples; construct a basic corpus according to the positive corpus samples and the negative corpus samples.

[0044] Exemplarily, each document in the target domain literature library is subjected to structured parsing to automatically extract structured information from a large number of documents. The purpose of structured parsing is to convert the unstructured text in the document (such as the article body) into structured data that is easy to process (such as titles, abstracts, keywords, section titles and contents). For example, natural language processing tools and parsing algorithms are used to identify and label different parts of the article, such as introductions, methods, results, and discussions. After completing the structured parsing of the document, next, positive and negative corpus samples are extracted from the parsed content according to specific text classification requirements. Positive corpus samples refer to those parts that are directly related to the text classification task. For example, in a task aimed at classifying the efficacy of medical treatments, positive samples may include paragraphs describing specific treatment methods and treatment results. In contrast, negative corpus samples include texts that are adjacent to the positive samples in content but do not directly reflect the classification target, such as pathological discussions that have no direct relation to the treatment. Furthermore, positive and negative corpus samples are used to construct a basic corpus. This corpus will serve as the basis for data augmentation and machine learning model training, supporting more complex data processing and analysis tasks. The construction process includes integrating these samples, formatting the data to meet the requirements of the training model, and performing necessary data cleaning and preprocessing, such as removing noise and standardizing text formats. For example, if the task is to identify and classify literature on a specific drug treatment, then positive samples may contain key information such as drug names, usage methods, and clinical trial results, while negative samples may include introductions to disease backgrounds or discussions of non-target effects of the drug. In this way, the basic corpus will contain a wide range of relevant and irrelevant texts, providing a solid foundation for in-depth learning and accurate classification. By this method, not only the efficiency and accuracy of data processing are improved, but also it is ensured that the text classification model can perform excellently in challenging practical applications. This systematic method provides reliable technical support for the efficient processing and analysis of large-scale literature data.

[0045] Furthermore, to obtain the classification result of the text classification task according to the optimized augmented corpus, the method further includes:

[0046] Performing corpus fusion based on the optimized augmented corpus and the basic corpus to obtain a fused corpus; and obtaining the classification result of the text classification task according to the fused corpus.

[0047] Specifically, the optimized enhanced corpus is fused with the original basic corpus. The optimized enhanced corpus contains texts generated by the data augmentation large language model, which have been subject to targeted adjustment and optimization to better reflect the requirements of specific text classification tasks. The basic corpus contains the original texts directly extracted from the target domain literature library. The purpose of fusing these two corpora is to combine the extensiveness of the original texts and the specificity of the enhanced texts, thereby creating a more comprehensive and detailed training and test data set. Based on the fused corpus, a suitable machine learning or deep learning model is used to perform the text classification task. This process includes selecting a suitable classification algorithm (such as support vector machine, random forest or neural network) to train the model and finally evaluating its performance in the classification task. During the training process, the model learns how to identify and distinguish different categories from the text features in the fused corpus. After training, the model can effectively classify newly input texts and provide decision support on which category they belong to. For example, in a task aimed at classifying legal documents, the fused corpus allows the model to not only learn the basic usage of legal terms but also understand their specific applications in different legal cases. Through the above process, the fused corpus composed of the optimized enhanced corpus and the basic corpus can be effectively utilized to improve the accuracy and efficiency of the text classification task. This method not only enhances the model's understanding ability of specific text types but also improves its generalization ability in practical applications.

[0048] Through the technical solutions of the above embodiments, the domain knowledge fusion data augmentation method for text classification provided by the present application ensures the domain relevance of the augmented data by extracting representative positive and negative samples from the professional literature in the target domain as the basic data set; by analyzing the metadata elements in the text classification task, constructs data augmentation prompt words integrating domain knowledge to guide the large model to generate high-quality augmented samples with correct grammar, coherent semantics and rich content. At the same time, the generated diverse and innovative sentences enhance the richness and coverage of the data, improving the generalization ability and adaptability of the model in the text classification task. In addition, fine-tuning is also performed on the corpus of a specific domain or style to generate enhanced sentences that conform to the target domain or style, enhancing the quality and usability of the generated data. It solves the technical problems existing in the prior art that it is impossible to effectively combine domain knowledge for in-depth text data augmentation, the richness and coverage of the data are poor, resulting in the lack of semantic depth and domain adaptability of the generated samples, and the quality and usability of the generated data are limited, and achieves the technical effects of effectively generating high-quality text samples with strong domain adaptability, improving the accuracy and generalization ability of the text classification model, and enhancing the usability of the generated data.

[0049] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0050] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.

Claims

1. A domain knowledge fusion data augmentation method for text classification, characterized in that, The method includes: Connect to the target domain literature library, extract from the target domain literature library according to the text classification task, and construct a basic corpus; Define data augmentation prompt words according to the text classification task, where the data augmentation prompt words include a task description, a seed sentence, and generation requirements; Input the data augmentation prompt words and the basic corpus into a data augmentation large language model for data augmentation to generate an augmented corpus; Construct an instruction data set, and use the low-rank adaptation technology based on the instruction data set to perform fusion feedback learning on the data augmentation large language model to obtain an optimized data augmentation large language model; According to the optimized data augmentation large language model, obtain an optimized augmented corpus, and obtain the classification result of the text classification task according to the optimized augmented corpus.

2. The domain knowledge fusion data augmentation method for text classification according to claim 1, wherein Construct an instruction data set, including: Construct a quality screening module, where the quality evaluation indicators of the quality screening module include an importance indicator, an innovation indicator, and an application indicator for characterizing the corpus; Evaluate the augmented corpus according to the quality screening module, and screen out a screened corpus with a quality greater than or equal to a preset quality; According to the screened corpus, construct the instruction data set, where the instruction data set includes a task description, and input data and output data corresponding to the task description.

3. The domain knowledge fusion data augmentation method for text classification according to claim 1, wherein Use the low-rank adaptation technology based on the instruction data set to perform fusion feedback learning on the data augmentation large language model. The method includes: Parse multiple network layers of the data augmentation large language model; Introduce a low-rank projection matrix group on each of the multiple network layers to obtain multiple low-rank projection matrix groups corresponding to the multiple network layers, where each low-rank projection matrix group includes a first projection matrix and a second projection matrix; Perform low-rank perturbation optimization on the weights of the multiple network layers according to the multiple low-rank projection matrix groups until the optimization target is met, and obtain an optimized data augmentation large language model.

4. The domain knowledge fusion data augmentation method for text classification according to claim 3, wherein, Perform low-rank perturbation optimization on the weights of the multiple network layers according to the multiple low-rank projection matrix groups, where the expression for performing low-rank perturbation optimization on the weight of any network layer includes: h i ′ = f i (x i ) + A i B i f i (x i ) = f i (x i ) + Δ i ; Among them, h i ′ is the forward calculation formula after correction of the i-th network layer, h i = f i (x i ), f i (x i ) is the original forward calculation formula of the i-th network layer, A i is the first projection matrix of the i-th network layer, B i is the second projection matrix of the i-th network layer, Δ i = A i B i f i (x i ) represents the introduced low-rank perturbation correction term.

5. The domain knowledge fusion data augmentation method for text classification according to claim 4, characterized in that, Construct the optimization target by minimizing the loss function, where the expression of the loss function includes: Among them, θ is a fixed parameter of the data augmentation large language model before optimization, represents multiple introduced low-rank projection matrix groups, and L is the total number of multiple low-rank projection matrix groups, is the training dataset based on the instruction dataset, and (x, y) ∈ D represents a single sample x and the corresponding label y in the training dataset, is the cross-entropy loss based on the text classification task, where y is the true label, used to measure the error between the predicted value and the true label, is the prediction function optimized based on the data augmentation large language model and the low-rank projection matrix groups.

6. The domain knowledge fusion data enhancement method for text classification according to claim 1, wherein The data augmentation prompt words include a task description, a seed sentence, and generation requirements; Among them, the task description is used to define the target of generative augmentation in the target domain, the seed sentence is an example of the text classification task in the target domain, and the generation requirements include the style, grammar, and semantic relevance of the generated text.

7. The domain knowledge fusion data augmentation method for text classification according to claim 1, wherein Extract from the target domain literature library according to the text classification task to construct a basic corpus. The method includes: Perform structured parsing on each literature in the target domain literature library and output the structured parsing content; Extract from the structured parsing content according to the text classification task to obtain a corpus positive sample and a corpus negative sample, where the corpus positive sample is a corpus set matching the text classification task, and the corpus negative sample is a corpus set of the context corresponding to the corpus positive sample; Construct a basic corpus according to the corpus positive sample and the corpus negative sample.

8. The domain knowledge fusion data augmentation method for text classification according to claim 1, wherein Obtaining the classification result of the text classification task according to the optimized enhanced corpus, the method further includes: Performing corpus fusion based on the optimized enhanced corpus and the basic corpus to obtain a fused corpus; Obtaining the classification result of the text classification task according to the fused corpus.

Citation Information

Patent Citations

  • Classification method for enhancing multi-label intentions of large language model based on remote supervision algorithm

    CN119025679A

  • Prediction explanation in machine learning classifiers

    WO2021223882A1