Text classification method based on text data synthesis and label semantic interaction
Patent Information
- Application Number
- CN202510294678.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing technology has problems such as scarcity of data and low inference efficiency in text classification tasks, especially in application scenarios in specific business fields. Problems such as lack of focus on text semantic features, few samples and unbalanced categories are difficult to effectively solve.
Using a method based on text data synthesis and tag semantic interaction, a synthetic data set is generated through a large language model, the training data set is expanded, and a text classification model based on the semantic attention mechanism is constructed, and the attention weight is determined using tag continuous semantic text, which improves the classification task effect of the model in specific business fields.
It improves the classification task effect of the text classification model in specific business fields, improves data scarcity and inference efficiency issues, and enhances the model's sensitivity to text classification labels and the accuracy of classification results.
Smart Images

Figure CN120216686A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and in particular, to a text classification method based on text data synthesis and label semantic interaction, a text classification system based on text data synthesis and label semantic interaction, and a computer-readable storage medium. Background Art
[0002] The field of Natural Language Processing (NLP) is an important branch in the fields of computer science and artificial intelligence, focusing on enabling computers to understand, process, and generate human language.
[0003] In the existing text classification tasks in the field of natural language processing, traditional machine learning-based methods include Naive Bayes (NB), Support Vector Machine (SVM), Decision Tree (DT), and Random Forest (RF). These methods have problems such as complex feature engineering, the curse of dimensionality, and data sparsity. While deep learning methods (e.g., Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), and Graph Convolutional Network (GCN)) no longer require feature engineering, they still have deficiencies such as a large data requirement and a complex model.
[0004] With the emergence of the Transformer architecture, Large Language Models (LLMs) and Pre-trained Language Models (PLMs) have emerged one after another. A large language model is a large-scale pre-trained language model built based on deep learning technology and has powerful data generation capabilities. A pre-trained language model is a model that is pre-trained on a large-scale corpus to provide a basis for subsequent fine-tuning for specific tasks.
[0005] However, there is still a gap between the application effect of large language models in professional tasks and the supervised fine-tuning effect of pre-trained models. Moreover, deploying large language models for application inference also faces situations such as high resource requirements and slow running speed. In application scenarios such as specific business domain classification, there are still problems such as unfocused text semantic features, few samples, and class imbalance when using large language models or pre-trained models to achieve text classification.
[0006] To overcome the above-mentioned defects existing in the prior art, there is an urgent need in this field for a text classification method based on text data synthesis and label semantic interaction, which can improve the effect of the constructed text classification model in the classification tasks of specific business fields and improve the data scarcity and inference efficiency problems in classification tasks. Summary of the Invention
[0007] A brief overview of one or more aspects is given below to provide a basic understanding of these aspects. This overview is not an exhaustive survey of all contemplated aspects, and is neither intended to identify key or decisive elements of all aspects nor to attempt to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that follows.
[0008] To overcome the above-mentioned defects existing in the prior art, the present invention provides a text classification method based on text data synthesis and label semantic interaction, a text classification system based on text data synthesis and label semantic interaction, and a computer-readable storage medium, which can improve the effect of the constructed text classification model in the classification tasks of specific business fields and improve the data scarcity and inference efficiency problems in classification tasks.
[0009] Specifically, the above-mentioned text classification method based on text data synthesis and label semantic interaction provided by the first aspect of the present invention includes the steps of: generating a synthetic data set through a large language model based on structured text data and a prompt learning template, where the synthetic data set is used to augment the training data set; training a text classification model through the training data set and testing it through a test data set to obtain the text classification model that reaches the expected optimization goal, where the text classification model is a text classification model based on a semantic attention mechanism, and the attention layer of the text classification model determines attention weights based on label continuous semantic text; and implementing text classification of input data through the text classification model.
[0010] Preferably, in an embodiment of the present invention, the step of determining the attention weight based on the tag continuous semantic text includes: respectively converting a plurality of the tag continuous semantic texts corresponding to a plurality of text classification tags into a plurality of tag semantic vectors; and determining a corresponding dynamic attention weight matrix through attention calculation according to the feature semantic vector and the plurality of tag semantic vectors, where the feature semantic vector is generated based on the data input to the text classification model.
[0011] Preferably, in an embodiment of the present invention, the prompt learning template is constructed based on the classification task objective, the set context text description, the task details, and the data constraint conditions, and the classification task objective, the set context text description, the task details, and the data constraint conditions are designed based on the data attributes of the structured text data.
[0012] Preferably, in an embodiment of the present invention, the generation of the synthetic data set further includes the steps of: evaluating the synthetic data set generated by the large language model; in response to the synthetic data set not meeting the requirements, adjusting the prompt learning template; generating a synthetic data set through the large language model based on the structured text data and the adjusted prompt learning template; and repeating the above steps until the synthetic data set meets the requirements, and supplementing the synthetic data set to the training data set.
[0013] Preferably, in an embodiment of the present invention, the method for evaluating the synthetic data set includes semantic similarity evaluation and semantic dispersion evaluation.
[0014] Preferably, in an embodiment of the present invention, the semantic similarity evaluation includes the steps of: matching each first data in the synthetic data set with all second data in the reference data set one by one to form a plurality of data pairs of the first data; converting the first data of the data pair into a first semantic vector, and converting the second data into a second semantic vector; calculating the similarity of each data pair based on the first semantic vector and the second semantic vector of each data pair according to the cosine similarity formula; calculating the average value based on the similarities of all data pairs to determine the semantic similarity of the synthetic data set; in response to the semantic similarity being greater than or equal to a preset threshold, determining that the synthetic data set meets the accuracy requirements; and in response to the semantic similarity being less than the preset threshold, determining that the synthetic data set does not meet the accuracy requirements.
[0015] Preferably, in an embodiment of the present invention, the semantic dispersion evaluation includes the steps of: matching each first data in the synthetic data set with all second data in the reference data set one by one to form multiple data pairs of each first data; converting the first data of the data pair into a first semantic vector and converting the second data into a second semantic vector; calculating the semantic entropy of each data pair according to the semantic entropy formula based on the first semantic vector and the second semantic vector of each data pair; calculating an average value based on the semantic entropy of all data pairs to determine the semantic dispersion of the synthetic data set; determining that the synthetic data set meets the diversity requirement in response to the semantic dispersion being greater than or equal to a preset threshold; and determining that the synthetic data set does not meet the diversity requirement in response to the semantic dispersion being less than the preset threshold.
[0016] Preferably, in an embodiment of the present invention, the text classification model includes a feature representation layer, a feature extraction layer, the attention layer, and a classification output layer. The feature extraction layer is a combined model structure of an adaptive BiLSTM and CNN, wherein the parameters of the BiLSTM part are adjusted by the semantic complexity of the data, and the parameters of the CNN part are adjusted by the distribution density of the local features of the data.
[0017] In addition, the above-mentioned text classification system based on text data synthesis and label semantic interaction provided by the second aspect of the present invention includes a memory and a processor. Computer instructions are stored on the memory. The processor is connected to the memory and is configured to execute the computer instructions stored on the memory to implement the text classification method based on text data synthesis and label semantic interaction provided by any one of the above embodiments.
[0018] In addition, computer instructions are stored on the above-mentioned computer-readable storage medium provided by the third aspect of the present invention. When the computer instructions are executed by a processor, the text classification method based on text data synthesis and label semantic interaction provided by any one of the above embodiments is implemented. Description of the Drawings
[0019] After reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings, the above features and advantages of the present invention can be better understood. In the drawings, the components are not necessarily drawn to scale, and components with similar related characteristics or features may have the same or similar reference numerals.
[0020] Figure 1 Shows a schematic diagram of a text classification system based on text data synthesis and label semantic interaction provided by some embodiments of the present invention;
[0021] Figure 2The flowchart of a text classification method based on text data synthesis and label semantic interaction provided according to some embodiments of the present invention is shown;
[0022] Figure 3 The algorithm framework diagram of the text classification system provided according to Embodiment 1 of the present invention is shown; and
[0023] Figure 4 The structural diagram of the text classification model provided according to Embodiment 2 of the present invention is shown.
[0024] Reference numerals:
[0025] 100: A text classification system based on text data synthesis and label semantic interaction;
[0026] 110: Memory;
[0027] 111: Computer-readable storage medium;
[0028] 120: Processor;
[0029] 200: Text classification method;
[0030] 310, 320, 330, 340: Phased task objectives;
[0031] 311: Classification task objective;
[0032] 312: Set context text description;
[0033] 313: Task details;
[0034] 314: Data limitation conditions;
[0035] 321: Name of the business entity;
[0036] 322: Industry where located;
[0037] 323: Industry annotation;
[0038] 324: Business scope;
[0039] 325: Industry to which it belongs;
[0040] 326: Other dimensional data;
[0041] 331: Evaluation index;
[0042] 3311: Semantic similarity;
[0043] 3312: Semantic dispersion;
[0044] 332: Evaluation data;
[0045] 3321: Synonym replacement;
[0046] 3322: Transposition of adjacent Chinese characters;
[0047] 3323: Random addition and deletion of characters;
[0048] 3324: Synthesis by large language models;
[0049] 341: Text classification model;
[0050] 3421: Artificial intelligence;
[0051] 3422: Biomedicine;
[0052] 3423: Integrated circuits;
[0053] 3424: Other categories;
[0054] 400: Text classification model;
[0055] 410: Text input layer;
[0056] 411: Name of the business entity;
[0057] 412: Industry where located;
[0058] 413: Industry notes;
[0059] 414: Scope of business;
[0060] 420: Feature representation layer;
[0061] 430: Feature extraction layer;
[0062] 431: Combined model structure;
[0063] 440: Attention layer;
[0064] 441, 442, 443: Label semantic vectors;
[0065] 450: Classification output layer;
[0066] 451: Artificial intelligence;
[0067] 452: Biomedicine;
[0068] 453: Integrated circuits;
[0069] 454: Other categories; and
[0070] S210~S230: Steps. Detailed implementation manners
[0071] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the aspects described below in conjunction with the accompanying drawings and specific embodiments are merely exemplary and should not be construed as imposing any limitation on the protection scope of the present invention.
[0072] In the description of the present invention, it should be noted that, unless otherwise clearly defined and limited, the terms "mounted", "connected", and "coupled" should be construed in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0073] In addition, the "upper", "lower", "left", "right", "top", "bottom", "horizontal", and "vertical" used in the following description should be understood as the orientations shown in this section and the related drawings. This relative term is only for convenience of description and does not mean that the device described needs to be manufactured or operated in a specific orientation, so it should not be construed as a limitation to the present invention.
[0074] It can be understood that although terms such as "first", "second", and "third" can be used herein to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first component, region, layer, and / or part discussed below can be referred to as the second component, region, layer, and / or part without departing from some embodiments of the present invention.
[0075] As described above, there is still a gap in the application effect of large language models in professional tasks compared with the supervised fine-tuning effect of pre-trained models. Moreover, deploying large language models for application inference also faces situations such as high resource requirements and slow running speed. In application scenarios such as specific business domain classification, using large language models or pre-trained models to achieve text classification also has problems such as unfocused text semantic features, few samples, and class imbalance.
[0076] In order to overcome the above-mentioned defects existing in the prior art, the present invention provides a text classification method based on text data synthesis and label semantic interaction, a text classification system based on text data synthesis and label semantic interaction, and a computer-readable storage medium, which can improve the effect of the constructed text classification model in the classification tasks of specific business domains and improve the data scarcity and inference efficiency problems in classification tasks.
[0077] In some non - restrictive embodiments, the above - mentioned text classification method based on text data synthesis and label semantic interaction provided by the first aspect of the present invention can be implemented via the above - mentioned text classification system based on text data synthesis and label semantic interaction provided by the second aspect of the present invention.
[0078] Please refer to Figure 1 , Figure 1 which shows a schematic diagram of a text classification system based on text data synthesis and label semantic interaction provided according to some embodiments of the present invention.
[0079] As Figure 1 shown, in the text classification system 100 based on text data synthesis and label semantic interaction, a memory 110 and a processor 120 can be configured. The memory 110 includes but is not limited to the above - mentioned computer - readable storage medium 111 provided by the third aspect of the present invention, on which computer instructions are stored. The processor 120 is connected to the memory 110 and is configured to execute the computer instructions stored on the memory 110 to implement the text classification method based on text data synthesis and label semantic interaction provided by the first aspect of the present invention.
[0080] The working principle of the above - mentioned text classification system based on text data synthesis and label semantic interaction will be described below in combination with some embodiments of the text classification method based on text data synthesis and label semantic interaction. Those skilled in the art can understand that these embodiments of the text classification method based on text data synthesis and label semantic interaction are only some non - restrictive implementation manners provided by the present invention, aiming to clearly show the main concept of the present invention and provide some specific solutions convenient for the public to implement, rather than limiting all functions or all working modes of the text classification system based on text data synthesis and label semantic interaction. Similarly, the text classification system based on text data synthesis and label semantic interaction is also a non - restrictive implementation manner provided by the present invention, and does not limit the execution subject and execution order of each step in these text classification methods based on text data synthesis and label semantic interaction.
[0081] Please refer to Figure 2 , Figure 2 which shows a flowchart of a text classification method based on text data synthesis and label semantic interaction provided according to some embodiments of the present invention.
[0082] As Figure 2 shown, the text classification method 200 based on text data synthesis and label semantic interaction may include step S210: generating a synthetic data set based on structured text data and a prompt learning template through a large - language model, and the synthetic data set is used to augment the training data set.
[0083] Large language models can learn templates based on the prompts provided by the text classification system, perform feature learning and expansion on data attributes in structured text data, and thus generate synthetic data that meets the requirements. Based on the generated synthetic data, the text classification system can construct a synthetic dataset.
[0084] Herein, the structured text data can be formed based on historical data or manually input data. In some embodiments, the text classification system can select, sort, and collect from historical data or manually input data according to the designed data attributes to form structured text data. The designed data attributes include clear business characteristics in specific business fields. For example, when the text classification system is applied to text classification tasks related to business entities, the text classification system can design the data attributes of the structured text data as business entity name, industry, industry annotation, business scope, and affiliated industry.
[0085] The prompt learning template is a complex and highly targeted template for generating data specifically designed for classification tasks, including classification task objectives, set context text descriptions, task details, and data constraints. The classification task objectives, set context text descriptions, task details, and data constraints of the prompt learning template can be designed by the text classification system according to the clear business characteristics included in the structured text data of the actual business field.
[0086] Those skilled in the art can understand that the large language model can be a large language model with powerful text generation capabilities such as GPT (Generative Pretrained Transformer). According to the actual business scenario, the text classification system can also replace the large language model with a large language model that shows better performance or efficiency in a specific field.
[0087] The following is the preferred Embodiment 1, and the text classification system proposed by the present invention based on text data synthesis and label semantic interaction will be described in detail according to Embodiment 1.
[0088] Please refer to Figure 3 , Figure 3 which shows the algorithm framework diagram of the text classification system provided by Embodiment 1 of the present invention.
[0089] As Figure 3 shown, in Embodiment 1, the specific business field to which the text classification system is applied can be text classification related to business entities.
[0090] The text classification system can first execute the phased task objective 310 and design a prompt learning template according to the specific business field.
[0091] When constructing a prompt learning template, the text classification system can consider multi-dimensional data attributes of structured text data including explicit business characteristics, so that the constructed prompt learning template can be adapted to a specific business domain. For example, when the text classification system is applied to a text classification task related to business entities, the prompt learning template can fully consider data attributes of structured text data such as business entity names, industries, industry annotations, business scopes, and affiliated industries.
[0092] Based on the above considerations, the text classification system can design the classification task objective 311, the set context text description 312, the task details 313, and the data restriction conditions 314 respectively, and then construct a prompt learning template based on the classification task objective 311, the set context text description 312, the task details 313, and the data restriction conditions 314.
[0093] In Embodiment 1, the text classification system can set the classification of the affiliated industry of the business entity accurately as the classification task objective 311, so as to determine the direction for subsequent work.
[0094] Based on the classification task objective 311, the text classification system can add the set context text description 312. The set context text description 312 can be the context related to the task objective, or the context of the data attributes of each dimension of the structured text data. For example, the text classification system can use the definition of the business entity name in the data attributes as the set context text description.
[0095] Through the task details 313 and the data restriction conditions 314, the text classification system can specify in detail the multi-dimensional data format and requirements of the synthetic data template.
[0096] The task details 313 can include the designed synthetic data template. The designed template can cover the data attributes of the structured text data, such as information like business entity names, industries, etc., to ensure that the generated synthetic data content is comprehensive and well-organized. The task details 313 can also include requirements for data attributes. For example, the text classification system can clearly stipulate the integrity requirement of the name in the task details 313, that is, in the generated synthetic data, the business entity name needs to be expressed completely and accurately to avoid classification deviation caused by inaccurate key information in the future. For another example, the text classification system can also clearly stipulate the standard classification code expression of the industry in the task details 313.
[0097] The applied data constraint conditions 314 may include guiding conditions and constraint conditions for generating synthetic data. The text classification system can guide the large language model to focus on key information by highlighting key information in structured text data when generating synthetic data, so as to improve the data validity of the generated synthetic data and the processing efficiency and accuracy of the large language model. For example, when applying the data constraint conditions 314, the description of the core business of the business scope can be highlighted. In addition, the constraint conditions may include constraints on the length, content, or scope of the generated synthetic data. For example, the text classification system can limit the length of the template designed in the task details 313 within a certain range, so as to reasonably limit the length of the synthetic data.
[0098] Based on the above-determined classification task objective 311, rich and set context text descriptions 312, task details 313, and data constraint conditions 314, the text classification system can construct a prompt learning template. The content of the prompt learning template can be the text structure and can be modified so that the text classification system can adjust in a timely manner according to the actual situation subsequently.
[0099] The text classification system uses the constructed prompt learning template that is adapted to the large language model data synthesis and designed specifically for the classification task as the input basis for the large language model to generate synthetic data. Thus, through the phased task objective 310, the text classification system completes the preparatory work before generating the synthetic data set.
[0100] In this way, designing the prompt learning template based on the multi-dimensional data attributes of structured text data including clear business characteristics can enable the large language model to comprehensively understand the characteristics and background information of the structured text data when generating synthetic data, guide the large language model to generate data more closely related to the actual business, thereby improving the pertinence of the classification task, enhancing the accuracy and reliability of the subsequent text classification model, and effectively coping with the challenges brought by various complex and diverse text data in the business scenario.
[0101] As Figure 3 shown, the text classification system can continue to execute the phased task objective 320 to generate a synthetic data set through the large language model.
[0102] Based on the structured text data and the prompt learning template designed in the phased task objective 310, the text classification system generates a synthetic data set through the large language model.
[0103] The structured text data can include multi-dimensional data attributes with clear business characteristics. In the first embodiment, the data attributes of the structured text data may include the business entity name 321, the industry where it is located 322, the industry annotation 323, the business scope 324, the industry to which it belongs 325, and other dimensional data 326.
[0104] Specifically, the text classification system can determine structured text data based on the acquired historical data. For example, select "XX Technology Co., Ltd." from the historical data as the target business entity name of the historical data; clarify "software and information technology service industry" as the industry where the historical data is located; at the same time, collect the corresponding annotation content of this industry; sort out the business scope of the business entity of the historical data, such as "software development, system integration" in the text; determine "artificial intelligence industry" as the industry to which the historical data belongs. Thus, the text classification system can determine the business entity name 321, the industry where it is located 322, the industry annotation 323, the business scope 324, the industry to which it belongs 325, and other dimensional data 326 of the historical data. Collect and organize these data to form structured text data corresponding to the historical data, and use it as the input basis for the subsequent large language model to generate synthetic data.
[0105] In some embodiments, the text classification system can omit some data attributes in combination with the actual situation, thereby reducing the data processing volume of the large language model. For example, in Embodiment 1, when the business entity name, the industry where it is located, the industry annotation, and the business scope can provide sufficient information for the large language model to learn the key features of the business entity and generate synthetic data with acceptable quality, the data attribute of the industry to which it belongs can be omitted when forming the structured text data. Or, when the information of the industry where the business entity is located can clearly imply the industry to which it belongs, the importance of the data attribute of the industry to which it belongs will be relatively reduced, and the data attribute of the industry to which it belongs can be omitted when forming the structured text data. Or, when the information of the industry where the business entity is located can fully cover the information expressed by the industry annotation, the data attribute of the industry annotation can be omitted when forming the structured text data. In addition, when the name of the business entity itself already has a clear and detailed industry definition and scope description, considering that the contribution of the industry annotation information to feature extraction and classification is small, the data attribute of the industry annotation can be omitted when forming the structured text data.
[0106] In this way, the text classification system inputs the prepared structured text data into the large language model. The large language model conducts feature learning on the structured text data according to the pre-set prompt learning template, and deeply analyzes the business features contained in each data attribute of the structured text data. Then, on the basis of the large language model completing feature learning, the text classification system conducts feature expansion through the large language model and generates synthetic data based on the data features of the existing structured text data.
[0107] The synthetic data generated by the large language model in the text classification system can not only be text data similar to the original data, but also expand data with slightly different businesses based on the associations between data attributes, thereby enriching the diversity of data. For example, if the large language model can identify the relationship between the industry and business scope of the original business entity data, when the industry of the generated synthetic data is computer and the business scope is software development, the generated business entity name will reflect the words "Software Technology Co., Ltd.".
[0108] Moreover, through the powerful data generation ability and attribute expansion function of the large language model, the text classification system can generate synthetic data according to the quantity of data in each category, effectively solving the problems of few samples and class imbalance in the business field. For example, for the situation where the business entity data under a certain emerging industry category is scarce, the text classification system can generate more data of this emerging industry category through the large language model, so as to balance the data volume under each industry category, provide sufficient and balanced training data for the subsequent text classification model, and thus significantly improve the generalization ability of the text classification model.
[0109] In the first embodiment, the categories of text classification may include artificial intelligence, biomedicine, integrated circuits, and other categories. The text classification system can generate synthetic data according to the quantity of data in each category, so that the ratio of the data volume of the training data under artificial intelligence, biomedicine, integrated circuits, and other categories is 1:1:1:1, thereby providing sufficient and balanced training data for the subsequent text classification model and improving the generalization ability of the classification task for different business entities.
[0110] Furthermore, the text classification method may further include a process of evaluating the synthetic data set to ensure that the quality of the generated synthetic data set meets the standards. Please continue to refer to Figure 3 The text classification system can execute the phased task objective 330 to evaluate the synthetic data set.
[0111] In the first embodiment, the evaluation metrics 331 of the synthetic data set may include semantic similarity 3311 and semantic dispersion 3312. The semantic similarity 3311 can represent the certainty of the semantics of the synthetic data. The semantic dispersion 3312 can quantify the dispersion degree of the semantics of the synthetic data to reflect the uncertainty of the semantics.
[0112] The text classification system can use the cosine distance (i.e., the normalized Euclidean distance) to measure the semantic similarity 3311, and this measurement method can effectively process the embedded vectors with language non-linear characteristics in the high-dimensional vector space.
[0113] The text classification system evaluates the synthetic dataset through semantic similarity 3311. First, each first data in the synthetic dataset can be matched one by one with all the second data in the reference dataset to form multiple data pairs for each first data.
[0114] The reference dataset is a pre-prepared standard dataset. In some embodiments, the reference dataset can be formed based on the internal data of the large model synthetic dataset or a large-scale real dataset. For example, in Embodiment 1, the reference dataset can be formed by a large-scale real business entity dataset.
[0115] Then, through the pre-trained model, the first data of the data pair is transformed into a high-dimensional first semantic vector, and the second data is transformed into a high-dimensional second semantic vector. Here, the pre-trained model can be BERT (Bidirectional Encoder Representations from Transformers).
[0116] After that, based on the first semantic vector and the second semantic vector of each data pair, the similarity of each data pair is calculated according to the cosine similarity formula. Then, the average value is calculated based on the similarities of all data pairs to determine the semantic similarity of the synthetic dataset.
[0117] When the semantic similarity of the synthetic dataset is greater than or equal to the preset threshold, the text classification system can determine that the synthetic dataset meets the accuracy requirements. When the semantic similarity of the synthetic dataset is less than the preset threshold, the text classification system can determine that the synthetic dataset does not meet the accuracy requirements.
[0118] The text classification system can use semantic entropy to describe semantic uncertainty, thereby quantifying semantic dispersion 3312.
[0119] Evaluating the synthetic dataset through semantic dispersion 3312 can first match each first data in the synthetic dataset with all the second data in the reference dataset one by one to form multiple data pairs for each first data. Then, through the pre-trained model, the first data of the data pair is transformed into a high-dimensional first semantic vector, and the second data is transformed into a high-dimensional second semantic vector.
[0120] After that, based on the first semantic vector and the second semantic vector of each data pair, the semantic entropy of each data pair is calculated according to the semantic entropy formula. Then, the average value is calculated based on the semantic entropies of all data pairs to determine the semantic dispersion of the synthetic dataset.
[0121] When the semantic dispersion degree of the synthetic dataset is greater than or equal to the preset threshold, the text classification system can determine that the synthetic dataset meets the diversity requirement. When the semantic dispersion degree of the synthetic dataset is less than the preset threshold, the text classification system can determine that the synthetic dataset does not meet the diversity requirement.
[0122] Then, the semantic similarity 3311 is used to evaluate the accuracy of the synthetic data, and at the same time, the semantic dispersion degree 3312 is used to evaluate the diversity of the synthetic data. The evaluation results of these two indicators are combined to form the quality evaluation data of the comprehensive synthetic dataset.
[0123] When the result of a certain evaluation indicator does not meet the requirements, the text classification system can adjust the prompt learning template constructed in the phased task objective 310, and then generate a synthetic dataset based on the structured text data and the prompt learning template. Then, repeat the steps of evaluating and generating the generated synthetic dataset until the evaluation results of the two indicators of the generated synthetic dataset meet the requirements.
[0124] When the results of the two evaluation indicators of the synthetic data meet the requirements, the text classification system can use the synthetic dataset that meets the requirements to expand the training dataset of the subsequent text classification model, thereby preparing high-quality data for the optimized training of the subsequent text classification model.
[0125] Among the above evaluation indicators 331, the semantic similarity 3311 focuses on the accuracy at the semantic level of the text data. In the first embodiment, through the evaluation of the semantic similarity 3311, it can be ensured that the synthetic dataset that meets the requirements conforms to the logical requirements of the business entity classification in terms of semantic expression, and avoid classification errors caused by data semantic deviation. The semantic dispersion degree 3312 evaluates the diversity from the macroscopic perspective of the text data distribution. Through the evaluation of the semantic dispersion degree 3312, it can be prevented that the synthetic dataset that meets the requirements is concentrated in a specific pattern or range.
[0126] Through the precise evaluation system formed by the combination of the two, the text classification system can screen out high-quality synthetic data, provide high-quality data resources for the training of the subsequent text classification model, effectively improve the classification accuracy and stability of the text classification model, and reduce the performance fluctuation of the text classification model caused by data quality problems.
[0127] Those skilled in the art can understand that the calculation methods of semantic similarity 3311 and semantic dispersion 3312 in the above evaluation index 331 are only a non-limiting implementation manner provided by the present invention, aiming to clearly show the main concept of the present invention and provide some specific solutions for the public to implement, rather than limiting the protection scope of the present invention. Optionally, in some other embodiments, those skilled in the art can also, based on the above concept and relevant knowledge in the art, replace the calculation methods of semantic similarity 3311 and semantic dispersion 3312 according to the actual data characteristics to achieve the same technical effect of reflecting the semantic accuracy of text data or measuring the distribution difference of text data sets.
[0128] In the first embodiment, the phased task objective 330 for evaluating the synthetic data set may further include a comparative evaluation process with a specific evaluation data set. The synthetic data set can be compared and evaluated with the evaluation data 332 by means of synonym replacement 3321, adjacent Chinese character transposition 3322, random character addition and deletion 3323, and large language model synthesis 3324. Here, the evaluation data 332 can be a traditional text enhancement data set. Optionally, in some embodiments, if the quality of the synthetic data set can be effectively evaluated only through the evaluation index 331, the text classification system can omit the comparison step with the evaluation data 332 to simplify the evaluation process and improve the evaluation efficiency.
[0129] Please continue to refer to Figure 2 , the text classification method 200 may include step S220: training a text classification model with a training data set and testing it with a test data set to obtain a text classification model that achieves the expected optimization goal. The text classification model is a text classification model based on a semantic attention mechanism, and the attention layer of the text classification model determines attention weights based on label continuous semantic text.
[0130] In Figure 3 the first embodiment shown, the text classification system may execute the phased task objective 340 to optimize and train the text classification model 341. The text classification model 341 may be a text classification model based on a semantic attention mechanism.
[0131] Those skilled in the art can understand that the text classification model based on the semantic attention mechanism should be broadly understood as a text classification model that can learn the feature and category relationship of text data and can perform text classification tasks by combining the attention mechanism. In some embodiments, the text classification model based on the semantic attention mechanism may be some variant models in the Transformer architecture that can achieve the same technical effect, or other text classification models with similar structures.
[0132] After that, the synthetic dataset formed by the phased task objectives 310-330 is added to the original training dataset, and the original test dataset remains unchanged.
[0133] Then, the text classification system can input the augmented training dataset into the text classification model 341 for training. During the training process, the relevant parameters of the text classification model 341 are adjusted. The text classification system can focus on reasonably adjusting the attention weight parameters of the attention layer of the text classification model 341, so that the text classification model 341 can better focus on the key features in the input data.
[0134] In the first embodiment, through the training of the text classification model 341, the text classification model 341 is guided to learn the relationship between the business entity features and the industrial categories to which they belong in the business entity data, and gradually enhance the classification ability of the text classification model 341 itself for different categories of business entity data.
[0135] After that, the text classification model 341 is tested through the test dataset to obtain the text classification model 341 that meets the expected optimization goal.
[0136] For example, in the first embodiment, the text classification system can use the test dataset to perform performance tests on the trained text classification model 341 in the text classification tasks of four industrial categories: artificial intelligence 3421, biomedicine 3422, integrated circuits 3423, and other categories 3424.
[0137] Specifically, the text classification system can test the actual classification effect of the text classification model 341 through the preset performance metrics. The preset performance metrics can include the accuracy rate, recall rate, and F1 value of the four industrial categories in the text classification task.
[0138] Among the performance metrics, the accuracy rate can be as follows:
[0139]
[0140] The recall rate can be as follows:
[0141]
[0142] The F1 value can be as follows:
[0143]
[0144] Among them, TP is True Positives, that is, the number of samples that are actually positive classes and are predicted as positive classes by the text classification model 341; FP is False Positives, that is, the number of samples that are actually negative classes but are predicted as positive classes by the text classification model 341; FN is False Negatives, that is, the number of samples that are actually positive classes but are predicted as negative classes by the text classification model 341.
[0145] In the first embodiment, the actual classification effect of the text classification model 341 can be evaluated by analyzing the values of the above performance indicators. Then, according to the test results of the performance indicators, it is confirmed whether the text classification model 341 meets the expected optimization goal. If not, the parameters can be further adjusted and continue training until an optimized text classification model 341 that can meet the expected optimization goal is finally obtained.
[0146] The text classification model 341 constructed by the text classification system may include an attention layer. This attention layer can determine the attention weights based on the label continuous semantic text. The label continuous semantic text is the continuous semantic text including continuous semantic information within a specific business domain for defining text classification labels. The text classification label is the relevant label with a specific definition predicted by the text classification task.
[0147] For example, in the first embodiment, the categories with specific definitions predicted by the text classification task may be artificial intelligence 3421, biomedicine 3422, and integrated circuits 3423. The text classification system can obtain the authoritative or official definition texts of artificial intelligence 3421, biomedicine 3422, and integrated circuits 3423 in the specific business domain, so as to clarify the specific industrial category ranges covered by these three specific industrial categories. The obtained definition texts are the label continuous semantic texts corresponding to each industrial category.
[0148] After that, the text classification system can convert the obtained label continuous semantic text into the corresponding label semantic vector, so as to convert the unstructured definition text into a computable semantic constraint. The text classification model 341 can generate a feature semantic vector through feature representation and / or feature extraction based on the input data. Input the feature semantic vector into the attention layer of the text classification model 341, and introduce the label semantic vector into the attention layer.
[0149] The feature semantic vector combines with the label semantic vector introduced by the attention layer. Through attention calculation, the correlation between the feature semantic vector and the label semantic vector in each feature dimension of the feature semantic vector is determined, so as to determine the corresponding dynamic attention weight matrix. Then, based on the determined dynamic attention weight matrix, calculations are performed according to the label semantic vector and the feature semantic vector, changing each feature dimension of the feature semantic vector, enabling dynamic interaction between the feature semantic vector and the label semantic vector, so as to generate a feature semantic vector that incorporates the label semantic vector.
[0150] In this way, by dynamically introducing the label continuous semantic text into the feature semantic vector, when the text classification model implements the text classification task, it can pay more attention to the features that have a higher correlation with the defined text of the text classification label.
[0151] For example, in the first embodiment, for the industrial category of biomedicine 3422, the label continuous semantic text may include: "Focus on fields such as biological products, innovative chemical drugs, high-end medical devices, modern traditional Chinese medicine, and intelligent healthcare,... Promote in-depth integration of production and medicine, improve clinical research capabilities and transformation levels, support the joint construction of high-level research hospitals by medical enterprises, build several demonstration bases for the integration of production and medicine innovation, and promote the application and promotion of innovative drugs and innovative medical devices...". Based on this label continuous semantic text, the text classification model 341 will pay more attention to the descriptions related to biological products, medical devices, etc. in the data attributes regarding the business scope in the input data. This method effectively enhances the sensitivity of the text classification model to the text classification label, improves the discrimination ability of the text classification model in the text classification task, makes the classification result more in line with the actual classification needs, and improves the accuracy and credibility of the classification result.
[0152] The following is the preferred second embodiment. According to the second embodiment, the text classification model of the text classification system based on text data synthesis and label semantic interaction proposed by the present invention is elaborated.
[0153] Please refer to Figure 4 , Figure 4 which shows the structural diagram of the text classification model provided in the second embodiment of the present invention.
[0154] As Figure 4 shown, the text classification model 400 may include a text input layer 410, a feature representation layer 420, a feature extraction layer 430, an attention layer 440, and a classification output layer 450.
[0155] The text input layer 410 can be used to process the input data. The input data can be raw data or synthetic data. The text input layer 410 can determine the attributes that need to be collected based on the specific business field, and carry out data collection work for the input data based on the determined attributes.
[0156] In the second embodiment, the specific business field to which the text classification model 400 is applied may be text classification related to business entities. The attributes determined by the text input layer 410 may include the business entity name 411, the industry where it is located 412, the industry annotation 413, and the business scope 414. Based on the determined attributes, the text input layer 410 may select "XX Electronic Technology Company" from the input business entity data as the business entity name 411 of the target business entity, and elaborate the information of its "Manufacture of Computers, Communication and Other Electronic Equipment" as the industry where it is located 412. At the same time, collect the corresponding industry annotation 413 and the information of the detailed business scope 414 of the business entity corresponding to the input business entity data, and form a long text data containing the above four attribute contents.
[0157] More preferably, the text input layer 410 may also preprocess the long text data. In some embodiments, the preprocessing may include format unification and noise cleaning. The text input layer 410 can process the collected long text data into a standard state that meets the requirements of subsequent processing through preprocessing. Specifically, the text input layer 410 may select appropriate text format conversion, standardization processing and other methods according to the actual situation to pre-design preprocessing tools or algorithms, so as to achieve format unification and noise cleaning. For example, preprocessing is achieved by identifying and removing parts that interfere with normal text content such as extra spaces, special symbols, repeated and incorrect words in the text.
[0158] The text input layer 410 inputs the processed data into the subsequent feature representation layer 420, providing a standard and high-quality input text for the feature representation layer 420.
[0159] The text classification system may select a suitable pre-trained model to constitute the feature representation layer 420. Here, the selected pre-trained model has the ability to effectively represent the features of text data. The feature representation layer 420 performs semantic encoding on the text data input via the text input layer 410 based on the pre-trained model, converts the input text data into a semantic vector including multiple feature dimensions through vectorization operation, and inputs the semantic vector into the subsequent feature extraction layer 430. The feature extraction layer 430 is superimposed on the feature representation layer 420, and the text classification model 400 constructs an architecture capable of further extracting features based on the feature representation layer 420 and the feature extraction layer 430.
[0160] The feature extraction layer 430 can further extract global feature information and local feature information from the input semantic vector.
[0161] In the second embodiment, the feature extraction layer 430 may be a combined model structure 431 of an adaptive BiLSTM (Bidirectional Long Short-Term Memory) and a CNN.
[0162] By using the BiLSTM part, the feature extraction layer 430 can perform bidirectional processing on the sequence of input semantic vectors. Through the forward and backward calculation mechanisms of the BiLSTM part, long-distance semantic dependencies are captured from different directions, thereby extracting global feature information that can reflect the overall semantic situation. For example, in Example 2, the BiLSTM part can focus on capturing the impact of the semantic associations before and after the business scope of the input business entity data on the overall characteristics of the business entity.
[0163] Using the CNN part, the feature extraction layer 430 can use convolution kernels of different scales to slide on the sequence of input semantic vectors, and use the local perception characteristics of the convolution kernel to extract local feature information. For example, in the second embodiment, the CNN part can focus on the details such as the specific keyword combination features in the industry name of the input business entity data during processing, so as to mine the key features contained in different parts of the business entity data.
[0164] Here, the parameters of the above-mentioned BiLSTM part (such as the number of hidden layer nodes of the BiLSTM part and the weight parameters of the forget gate, input gate, and output gate) can be adjusted by the semantic complexity of the data. During the training process of the text classification model 400, the semantic complexity is determined by preliminary analysis of the different text data input. For example, in Example 2, the text classification system can increase the number of hidden layer nodes of the BiLSTM part for the business scope statement lengthy and involving multi-field business entity data text to better capture long-distance semantic dependencies.
[0165] The parameters of the CNN part (such as the convolution kernel size, step size, and pooling layer parameters) can be adjusted by the distribution density of the local features of the data. For example, in the second embodiment, for the business entity data with dense keyword combinations in the industry name, the text classification system can use a smaller convolution kernel and step size in the CNN part to finely extract local features.
[0166] Compared with the traditional combined model structure of BiLSTM and CNN with fixed parameter settings, the combined model structure 431 provided by the feature extraction layer 430 dynamically optimizes parameters through adaptive optimization, so that the text classification model 400 can adapt to the diversity of semantic complexity and local feature distribution of text data in specific business fields, and flexibly adjust its own structure and parameters according to the characteristics of different input text data.
[0167] After that, the combined model structure 431 can also perform a fusion integration operation (Concat) on the global feature information extracted by the BiLSTM part and the local feature information extracted by the CNN part, and organically combine the global feature information and the local feature information. Here, the fusion algorithm or method can include splicing, weighted summation, etc., and the fusion algorithm or method can also be selected according to the actual situation.
[0168] In this way, for complex semantic texts, the BiLSTM part can give full play to its long-distance dependence capture ability; for texts with dense local features, the CNN part can accurately extract key information. The feature semantic vector after the two are fused can more accurately reflect the internal features of the text data, effectively improving the accuracy and robustness of the text classification model 400 in the text classification task of different types of text data, reducing the classification error caused by data feature differences, and enhancing the adaptability of the text classification model 400 to changing text data.
[0169] Those skilled in the art can understand that the combined model structure 431 of the adaptive BiLSTM and CNN in the feature extraction layer 430 is only a non-limiting preferred implementation provided by the present invention. Optionally, in some other embodiments, those skilled in the art can also, based on the above concept, use other model combinations capable of extracting global and local feature information to achieve the same technical effect. For example, the feature extraction layer 430 can be a combination of GRU (Gated Recurrent Unit) and DenseNet (Densely Connected Convolutional Networks). The GRU part can be used to extract global feature information, and the DenseNet part can extract local feature information.
[0170] Based on the feature extraction and feature fusion of the semantic vector by the feature extraction layer 430, a feature semantic vector can be obtained. The determined feature semantic vector can be input into the subsequent attention layer 440. The attention layer 440 can determine the attention weight based on the label continuous semantic text introduced by the text classification system.
[0171] In the second embodiment, the text classification labels for the text classification task can include artificial intelligence, integrated circuits, and biomedicine. The text classification system can obtain authoritative or official definition texts regarding artificial intelligence, integrated circuits, and biomedicine, so as to clarify the specific industrial category scope covered by the text classification labels. For example, the definition text of integrated circuits in the text classification labels can cover the integrated circuit industry, etc.
[0172] In the second embodiment, the text classification system can use the obtained definition texts related to artificial intelligence, integrated circuits, and biomedicine as the label continuous semantic texts related to artificial intelligence, integrated circuits, and biomedicine corresponding to the text classification labels, and use pre-trained models such as BERT to convert the label continuous semantic texts related to artificial intelligence, integrated circuits, and biomedicine into the label semantic vector 441 corresponding to artificial intelligence, the label semantic vector 442 corresponding to integrated circuits, and the label semantic vector 443 corresponding to biomedicine respectively. Then, the text classification system introduces the label semantic vector 441, the label semantic vector 442, and the label semantic vector 443 into the attention layer 440 of the text classification model 400.
[0173] The attention layer 440 combines the feature semantic vector input by the feature extraction layer 430 and the label semantic vector introduced by the text classification system to perform attention calculation to determine the attention weights.
[0174] Specifically, the attention layer 440 can use the attention calculation mechanism to perform quantitative analysis on the correlation between the feature semantic vector and the label semantic vector in each feature dimension of the feature semantic vector, and measure the tightness of the semantic association between the feature semantic vector and the label semantic vector through calculation. Here, the attention calculation mechanism can include the dot-product-based attention mechanism and / or the additive attention mechanism.
[0175] After that, according to the correlation results between the above-mentioned feature semantic vector and the label semantic vector in each feature dimension of the feature semantic vector, a corresponding dynamic attention weight matrix is determined based on a preset weight determination rule. In some embodiments, the attention layer 440 can process the correlation scores through normalization and other methods to determine the weight values in the corresponding dynamic attention weight matrix.
[0176] Next, the attention layer 440 can calculate based on the dynamic attention weight matrix according to the label semantic vector and the feature semantic vector, change the values of each feature dimension of the feature semantic vector, so that the feature semantic vector and the label semantic vector interact dynamically, and achieve attention superposition on the feature semantic vector. In some embodiments, the attention layer 440 can change the values of each feature dimension of the feature semantic vector by multiplying each dimension one by one. In this way, through the feature semantic vector fused with the label semantic vector, the text classification system can enable the features with high semantic correlation with the label continuous semantic text to obtain higher attention, thereby highlighting the important features and guiding the text classification model 400 to focus more on the parts with strong relevance to the label continuous semantic text.
[0177] The feature semantic vector fused with the label semantic vector is input into the subsequent classification output layer 450 as the final feature vector.
[0178] Here, the classification output layer 450 can be a multi-layer neural network model. Specifically, the classification output layer 450 can enable the feature vectors determined by the attention layer 440 to be sequentially transmitted and calculated through the layers of the neural network until the feature vectors reach the output layer after being processed by the layers of the neural network. Then, according to the activation values of the neurons in the output layer, the category corresponding to the neuron with the highest activation value is determined as the classification result, thus completing the entire classification process and realizing the recognition and classification of the input text data.
[0179] In the second embodiment, the output layer may include four neurons, and the four neurons respectively correspond to the categories of the classification results of the text classification model 400, including artificial intelligence 451, biomedicine 452, integrated circuits 453, and other categories 454. After the feature vectors reach the output layer after being processed by the layers of the neural network, the activation values of the four neurons are compared, and the category corresponding to the neuron with the highest activation value is used as the classification result, that is, the industrial category corresponding to the input business entity data is determined.
[0180] In some other embodiments, the classification output layer 450 can also be replaced by a more efficient and accurate shallow neural network structure to achieve the same technical effect, such as some new neural network architectures that have been specially optimized and are applicable to text classification tasks.
[0181] Please continue to refer to Figure 2 , the text classification method 200 may include step S230: implementing text classification of the input data through the text classification model.
[0182] In the data synthesis stage of the text classification system, the large language model is used to learn and expand multi-dimensional data based on the prompt learning template to generate a large amount of synthetic data, effectively alleviating the problem of data scarcity. In the text classification model construction stage, a text classification model combined with a semantic attention mechanism is adopted, and the dataset augmented by the synthetic data is used for training and optimization. Through the obtained text classification model that reaches the expected optimization goal, text classification of the input data is realized.
[0183] The text classification method based on text data synthesis and label semantic interaction provided by the present invention cooperates between the large language model and the text classification model. The large language model creates rich and diverse synthetic data based on limited original data, making up for the data scarcity situation. Through the good reasoning efficiency of the text classification model combined with the semantic attention mechanism, the category to which the text data belongs is quickly and accurately discriminated. The two work together, not only solving the problem of insufficient training of the text classification model due to insufficient data in the specific business field, but also improving the reasoning speed and accuracy of the text classification model in actual classification, enabling the entire natural language processing system to show better performance and adaptability in the specific business scenario, and effectively coping with complex and changing data and classification requirements.
[0184] Although the methods described above are illustrated and described as a series of acts for simplicity of explanation, it should be understood and appreciated that the methods are not limited by the order of the acts, as some acts may occur in different orders and / or concurrently with other acts that are illustrated and described herein or other acts that are not illustrated and described herein but would be understood by those of ordinary skill in the art, in accordance with one or more embodiments.
[0185] Those of ordinary skill in the art will understand that information, signals, and data can be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0186] Those of ordinary skill in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0187] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal.
[0188] The foregoing description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A text classification method based on text data synthesis and label semantic interaction, characterized in that: Includes steps: Based on the structured text data and the prompt learning template, a synthetic data set is generated through a large language model, wherein the synthetic data set is used to expand the training data set; The text classification model is trained by the training data set and tested by the test data set to obtain the text classification model that achieves the expected optimization goal, wherein the text classification model is a text classification model based on a semantic attention mechanism, and the attention layer of the text classification model determines the attention weight based on the label continuous semantic text; as well as The text classification model is used to implement text classification of input data.
2. The text classification method according to claim 1, characterized in that: The step of determining the attention weight based on the label continuous semantic text comprises: Converting the plurality of label continuous semantic texts corresponding to the plurality of text classification labels into a plurality of label semantic vectors respectively; and According to the feature semantic vector and the multiple label semantic vectors, a corresponding dynamic attention weight matrix is determined through attention calculation, wherein the feature semantic vector is generated based on the data input into the text classification model.
3. The text classification method according to claim 1, characterized in that: The prompt learning template is constructed based on the classification task goal, the set context text description, the task details and the data restrictions, and the classification task goal, the set context text description, the task details and the data restrictions are designed based on the data attributes of the structured text data.
4. The text classification method according to claim 1, characterized in that: The generation of the synthetic data set also includes the steps of: evaluating the synthetic dataset generated by the large language model; In response to the synthetic data set not meeting the requirement, adjusting the prompt learning template; generating a synthetic dataset through a large language model based on the structured text data and the adjusted prompt learning template; as well as Repeat the above steps until the synthetic data set meets the requirements, and then add the synthetic data set to the training data set.
5. The text classification method according to claim 4, characterized in that: The method of evaluating the synthetic data set includes semantic similarity evaluation and semantic dispersion evaluation.
6. The text classification method according to claim 5, characterized in that: The semantic similarity evaluation comprises the steps of: Matching each first data in the synthetic data set with all second data in the reference data set one by one to form a plurality of data pairs of each first data; Converting the first data of the data pair into a first semantic vector, and converting the second data into a second semantic vector; Based on the first semantic vector and the second semantic vector of each data pair, calculating the similarity of each data pair according to a cosine similarity formula; Calculating an average value based on the similarities of all data pairs to determine the semantic similarity of the synthetic data set; In response to the semantic similarity being greater than or equal to a preset threshold, determining that the synthetic data set meets the accuracy requirement; as well as In response to the semantic similarity being less than a preset threshold, it is determined that the synthetic data set does not meet accuracy requirements.
7. The text classification method according to claim 5, characterized in that: The semantic discreteness evaluation comprises the steps of: Matching each first data in the synthetic data set with all second data in the reference data set one by one to form a plurality of data pairs of each first data; Converting the first data of the data pair into a first semantic vector, and converting the second data into a second semantic vector; Calculating the semantic entropy of each data pair according to a semantic entropy formula based on the first semantic vector and the second semantic vector of each data pair; Calculate an average value based on the semantic entropy of all data pairs to determine the semantic dispersion of the synthetic data set; In response to the semantic discreteness being greater than or equal to a preset threshold, determining that the synthetic data set meets the diversity requirement; as well as In response to the semantic discreteness being less than a preset threshold, it is determined that the synthetic data set does not meet the diversity requirement.
8. The text classification method according to claim 1, characterized in that: The text classification model includes a feature representation layer, a feature extraction layer, the attention layer and a classification output layer. The feature extraction layer is an adaptive combined model structure of BiLSTM and CNN, wherein the parameters of the BiLSTM part are adjusted by the semantic complexity of the data, and the parameters of the CNN part are adjusted by the distribution density of the local features of the data.
9. A text classification system based on text data synthesis and label semantic interaction, characterized in that: include: a memory having computer instructions stored thereon; as well as A processor is connected to the memory and is configured to execute computer instructions stored in the memory to implement the text classification method based on text data synthesis and label semantic interaction as described in any one of claims 1 to 8.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the text classification method based on text data synthesis and label semantic interaction as described in any one of claims 1 to 8 is implemented.