Major language model-based minority class sample generation method and system and storage medium

By using large language models to generate a minority class sample dataset in emotion classification task, the problem of uneven distribution in deep learning is solved and the ability to recognize minority class text samples is improved.

CN119917856APending Publication Date: 2025-05-02WUHAN FIBERHOME PUTIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411964307.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

In the emotional classification task of deep learning, the emotional data distribution of the training samples is uneven, which leads to the deep neural network biasing towards the majority class samples and making it difficult to identify a few class samples.

Method used

By dynamically optimizing the synthesized minority sample data sets of emotion classification, using large language models to generate minority model sample data sets, and combining them with minority platform sample data sets to perform emotional text classification.

Benefits of technology

It effectively narrows the distribution gap, improves the characteristics of a few types of samples, avoids overfitting problems in classification, and improves the neural network's ability to recognize a few types of text samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917856A_ABST
    Figure CN119917856A_ABST
Patent Text Reader

Abstract

The invention discloses a minority class sample generation method and system based on a large language model and a storage medium, and the method comprises the steps: carrying out the preprocessing and analysis of emotion data, so as to determine a minority class platform sample data set and a majority class platform sample data set; determining the number of majority class samples corresponding to the majority class platform sample data set; inputting the minority class sample data set into a large language model based on the minority class sample amplification factor and the majority class sample number to generate a minority class model sample data set; and combining the minority class platform sample data set and the minority class model sample data set to perform sentiment text classification through a sentiment classification model. Through data sample processing and minority sample data generation, a minority sample synthesis part based on a large language model samples a step-by-step data synthesis framework, and minority sample data is synthesized through dynamic optimization to efficiently reduce a sample data distribution gap, so that the accuracy of an emotion classification model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sentiment classification technology, and in particular to a method, system and storage medium for generating minority class samples based on a large language model. Background Art

[0002] In the deep learning sentiment classification task, the imbalanced distribution of sentiment data of training samples is a serious problem. In this case, the deep neural network will be biased towards the majority class samples and cannot learn the data features of the minority class samples well, making it difficult to identify the minority class samples.

[0003] The above contents are only used to assist in understanding the technical solution of the present invention and do not constitute an admission that the above contents are prior art. Summary of the invention

[0004] The main purpose of the present invention is to provide a method, system and storage medium for generating minority class samples based on a large language model, aiming to solve the technical problem of how to efficiently narrow the distribution gap by dynamically optimizing the minority class sample data set for synthetic sentiment classification.

[0005] To achieve the above object, the present invention provides a method for generating minority class samples based on a large language model, and the method for generating minority class samples based on a large language model comprises:

[0006] Acquire sentiment data through the public opinion information collection platform according to the platform labels, and preprocess and analyze the sentiment data to determine a minority platform sample data set and a majority platform sample data set;

[0007] Determine the number of majority class samples corresponding to the majority class platform sample data set;

[0008] Inputting the minority class sample data set into the large language model based on the minority class sample magnification factor and the number of majority class samples to generate a minority class model sample data set;

[0009] The minority class platform sample data set and the minority class model sample data set are merged, so that the merged minority class sentiment classification sample data set is subjected to sentiment text classification through the sentiment classification model.

[0010] Optionally, the step of preprocessing and analyzing the sentiment data to determine a minority platform sample data set and a majority platform sample data set includes:

[0011] Classifying the sentiment data based on the platform labels to obtain minority class sample data and majority class sample data;

[0012] Annotating the minority sample data, and constructing a minority platform sample data set according to the minority sample data and corresponding annotation information;

[0013] Determine the platform label accuracy corresponding to the majority class sample data;

[0014] If the platform label accuracy is greater than a preset label accuracy threshold, a majority class sample data set is constructed according to the majority class sample data and the corresponding platform labels.

[0015] Optionally, the step of inputting the minority class sample data set into the large language model based on the minority class sample magnification and the number of majority class samples to generate a minority class model sample data set includes:

[0016] Based on the large model text synthesis prompt project, text data is synthesized through a large language model according to the minority class sample data set to obtain a minority class text synthesis sample data set;

[0017] Based on the error analysis prompt project, obtaining a minority class error synthesis sample data set according to the minority class data sample set and the minority class text synthesis sample data set through the large language model;

[0018] Based on the minority class sample magnification and the number of majority class samples, the minority class sample data set, the minority class text synthesis sample data set and the minority class error synthesis sample data set are merged, and based on the large model text synthesis prompt project, the merged sample data set is passed through the large language model to generate a minority class model sample data set.

[0019] Optionally, the step of obtaining a minority class text synthesis sample data set based on the minority class sample data set through a large language model in the large model text synthesis prompt engineering includes:

[0020] Performing data cleaning on the minority class sample data in the minority class sample data set, and analyzing the cleaned minority class sample data to obtain element labels corresponding to the minority class sample data;

[0021] Extracting minority class sample labeled data from the minority class sample data, and determining the accuracy of the element labels corresponding to the minority class sample labeled data;

[0022] If the accuracy of the element label is greater than the preset label accuracy threshold, a large model text synthesis prompt project is generated according to the minority class sample data and the corresponding element label;

[0023] Based on the large model text synthesis prompt project, the minority class sample data and the corresponding element labels are passed through a large language model to obtain a minority class text synthesis sample data set.

[0024] Optionally, the step of obtaining a minority class error synthetic sample data set based on the error analysis prompting project according to the minority class data sample set and the minority class text synthetic sample data set through the large language model includes:

[0025] The minority class sample data set and the minority class text synthesis sample data set are used as a minority class training sample set, wherein the minority class training sample set includes minority class training sample data;

[0026] Determine the number of training samples corresponding to the minority class training sample set;

[0027] Selecting majority class training sample data from the majority class sample data set based on the number of training samples;

[0028] The sentiment classification model is trained according to the minority class training sample data and the majority class training sample data to obtain the minority class classification sample data category output by the sentiment classification model, and element label analysis and extraction are performed on the misclassified minority class classification sample data;

[0029] Based on the error analysis prompt engineering, the minority class classification sample data and the corresponding classification element labels are passed through a large language model to obtain a minority class error synthetic sample data set.

[0030] Optionally, before the step of obtaining a minority class error synthetic sample data set based on the error analysis prompting project through a large language model according to the minority class classification sample data and the corresponding classification element labels, the step further includes:

[0031] When the labeling accuracy of the minority classification sample data and the corresponding classification element labels meets the preset sampling accuracy inspection requirement, an error analysis prompt project is generated according to the minority classification sample data and the corresponding classification element labels.

[0032] In addition, to achieve the above purpose, the present invention also proposes a minority class sample generation system based on a large language model, and the minority class sample generation system based on a large language model includes:

[0033] A collection module is used to obtain sentiment data through a public opinion information collection platform according to platform labels, and preprocess and analyze the sentiment data to determine a minority platform sample data set and a majority platform sample data set;

[0034] A determination module, used to determine the number of majority class samples corresponding to the majority class platform sample data set;

[0035] A generating module, configured to input the minority class sample data set into a large language model based on the minority class sample magnification and the number of majority class samples to generate a minority class model sample data set;

[0036] The merging module is used to merge the minority class platform sample data set and the minority class model sample data set, so that the merged minority class sentiment classification sample data set is subjected to sentiment text classification through the sentiment classification model.

[0037] In addition, to achieve the above-mentioned purpose, the present invention also proposes a device for generating minority class samples based on a large language model, the device comprising: a memory, a processor, and a program for generating minority class samples based on a large language model stored in the memory and executable on the processor, the program for generating minority class samples based on a large language model being configured to implement the steps of the method for generating minority class samples based on a large language model as described above.

[0038] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which is stored a minority class sample generation program based on a large language model, and when the minority class sample generation program based on a large language model is executed by a processor, the steps of the minority class sample generation method based on a large language model as described above are implemented.

[0039] The present invention first obtains sentiment data through the public opinion information collection platform according to the platform label, and pre-processes and analyzes the sentiment data to determine the minority platform sample data set and the majority platform sample data set, and then determines the number of majority samples corresponding to the majority platform sample data set, and then inputs the minority sample data set into the large language model based on the minority sample magnification and the number of majority samples to generate a minority model sample data set, and finally merges the minority platform sample data set and the minority model sample data set, so that the merged minority sentiment classification sample data set is used to classify sentiment text through the sentiment classification model. Compared with the random oversampling operation in the prior art, the present invention generates new minority samples, improves the characteristics of minority samples and balances the training data set, avoids the overfitting problem in classification, and at the same time, uses the rich knowledge of the large language model to synthesize pseudo training samples for minority samples in sentiment classification tasks, achieves data and calculation efficiency, and improves the recognition ability of neural networks for minority text samples in sentiment classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a structural diagram of a device for generating minority class samples based on a large language model in a hardware operating environment involved in an embodiment of the present invention;

[0041] Figure 2It is a flowchart of a first embodiment of a method for generating minority class samples based on a large language model according to the present invention;

[0042] Figure 3 It is a schematic diagram of a framework for gradually synthesizing minority class samples for sentiment classification according to the first embodiment of the method for generating minority class samples based on a large language model of the present invention;

[0043] Figure 4 It is a structural block diagram of the first embodiment of the system for generating minority class samples based on a large language model of the present invention.

[0044] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0045] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0046] Reference Figure 1 , Figure 1 The present invention is a schematic diagram of a device structure for generating minority class samples based on a large language model in a hardware operating environment according to an embodiment of the present invention.

[0047] like Figure 1 As shown, the minority class sample generation device based on the large language model may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (Wireless-Fidelity, Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk storage. The memory 1005 may also be a storage system independent of the aforementioned processor 1001.

[0048] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the device for generating minority class samples based on a large language model, and may include more or fewer components than those shown in the figure, or combine certain components, or arrange the components differently.

[0049] like Figure 1 As shown, the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, and a minority class sample generation program based on a large language model.

[0050] exist Figure 1 In the minority class sample generation device based on a large language model shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the minority class sample generation device based on a large language model of the present invention can be set in the minority class sample generation device based on a large language model, and the minority class sample generation device based on a large language model calls the minority class sample generation program based on the large language model stored in the memory 1005 through the processor 1001, and executes the minority class sample generation method based on the large language model provided in the embodiment of the present invention.

[0051] The embodiment of the present invention provides a method for generating minority class samples based on a large language model. Figure 2 , Figure 2 It is a flowchart of the first embodiment of the method for generating minority class samples based on a large language model of the present invention.

[0052] In this embodiment, the method for generating minority class samples based on a large language model includes the following steps:

[0053] Step S10: Acquire sentiment data through the public opinion information collection platform according to the platform label, and pre-process and analyze the sentiment data to determine a minority platform sample data set and a majority platform sample data set.

[0054] It is easy to understand that the execution subject of this embodiment can be a minority class sample generation system based on a large language model with functions such as data processing, network communication and program running, or it can be other computer devices with similar functions, etc., and this embodiment is not limited.

[0055] In this embodiment, the platform tags include emotion-sensitive tags and emotion-insensitive tags, etc.

[0056] Furthermore, sentiment data is obtained through the public opinion information collection platform according to the platform labels, and the sentiment data is classified based on the platform labels to obtain minority sample data and majority sample data; the minority sample data is labeled, and a minority platform sample data set is constructed according to the minority sample data and the corresponding labeling information; the platform label accuracy corresponding to the majority sample data is determined; if the platform label accuracy is greater than a preset label accuracy threshold, a majority sample data set is constructed according to the majority sample data and the corresponding platform labels.

[0057] It should be noted that the annotation information is the corresponding label to determine whether the minority sample data is sensitive data or non-sensitive data. The platform label accuracy corresponding to the majority sample data can be understood as the average accuracy corresponding to the preset number of sample data in the majority sample data.

[0058] In the specific implementation, the sentiment data in the public opinion information platform is collected according to the platform label, and the collected unit time data (i.e., sentiment data within unit time) is manually labeled.

[0059] It should be understood that the number of collected minority class sample data needs to be set to m, and the number of collected majority class sample data needs to be set to n, where n>>m. The amount of majority class sample data collected by the data platform is often much larger than the amount of minority class sample data.

[0060] It should also be noted that in this embodiment, all minority class sample data need to be manually labeled; three m data samples are selected from the majority class sample data n, and the accuracy of these three m sample data is (Acc_m1, Acc_m2, Acc_m3). If the average accuracy of these three samples Acc_ave = (Acc_m1+Acc_m2+Acc_m3) / 3>98%, then the manual labeling of the majority class samples is terminated. If it is less than the preset label accuracy threshold (which can be 98%), the majority class sample data is manually labeled.

[0061] Step S20: Determine the number of majority class samples corresponding to the majority class platform sample data set.

[0062] In this embodiment, the formula for the magnification of the minority class samples can be set according to the number of majority class samples, and the formula for the magnification of the minority class samples is: N=(nm) / m, where N is the magnification of the minority class samples.

[0063] It should be understood that by using the minority sample generation method based on the large language model, the number of minority samples is enlarged to be substantially equal to the number of majority class samples.

[0064] Step S30: inputting the minority class sample data set into the large language model based on the minority class sample magnification and the number of majority class samples to generate a minority class model sample data set.

[0065] Furthermore, based on the large model text synthesis prompt project, text data is synthesized according to the minority class sample data set through the large language model to obtain the minority class text synthesis sample data set; based on the error analysis prompt project, the minority class data sample set and the minority class text synthesis sample data set are passed through the large language model to obtain the minority class error synthesis sample data set; based on the minority class sample magnification and the number of majority class samples, the minority class sample data set, the minority class text synthesis sample data set and the minority class error synthesis sample data set are merged, and based on the large model text synthesis prompt project, the merged sample data set is passed through the large language model to generate a minority class model sample data set.

[0066] It should also be noted that the large model-based text synthesis prompt project passes the minority sample data set through the large language model to obtain the minority text synthesis sample data set. The processing method is to clean the minority sample data in the minority sample data set, and analyze the cleaned minority sample data to obtain the feature labels corresponding to the minority sample data; extract the minority sample labeled data from the minority sample data, and determine the accuracy of the feature labels corresponding to the minority sample labeled data; if the accuracy of the feature label is greater than the preset label accuracy threshold, generate a large model text synthesis prompt project based on the minority sample data and the corresponding feature labels; based on the large model text synthesis prompt project, the minority sample data and the corresponding feature labels are passed through the large language model to obtain the minority text synthesis sample data set.

[0067] In this embodiment, reference Figure 3 , Figure 3 This is a schematic diagram of the framework for gradually synthesizing minority samples for sentiment classification according to the first embodiment of the method for generating minority samples based on a large language model of the present invention. Based on the large model text synthesis prompt project, the minority sample data set is passed through the large language model to obtain a minority text synthesis sample data set. The detailed process is as follows:

[0068] 1) Clean the minority sample data set A (i.e., seed data set) to be generated, remove meaningless symbols, and format them uniformly;

[0069] 2) Using a large language model, analyze and label the cleaned minority sample data from the perspectives of domain, scenario, text form, and speech tendency, and statistically analyze the proportion of each factor label;

[0070] 3) Sampling and labeling the minority sample data and their corresponding feature labels. It is necessary to correct the minority sample data with mismatched feature labels, and then continue random sampling until the sampling accuracy reaches a preset label accuracy threshold (e.g., 98%) in the minority samples, and then stop sampling. A list of minority sample data and feature labels is generated for generating a large model text synthesis prompt project;

[0071] 4) Based on the large model text synthesis prompt project, the minority class sample data and the corresponding feature labels are used to synthesize the minority class text synthesis sample data set (ie, the seed data set synthesized sample data) through the large language model, which is recorded as set C.

[0072] Furthermore, based on the error analysis prompting project, according to the minority class data sample set and the minority class text synthesis sample data set, a processing method for obtaining the minority class error synthesis sample data set through the large language model is to use the minority class sample data set and the minority class text synthesis sample data set as the minority class training sample set, and the minority class training sample set includes the minority class training sample data; determine the number of training samples corresponding to the minority class training sample set; select the majority class training sample data from the majority class sample data set based on the number of training samples; train the sentiment classification model according to the minority class training sample data and the majority class training sample data to obtain the minority class classification sample data and the corresponding classification element labels ; When the labeling accuracy of the minority class classification sample data and the corresponding classification element labels meets the preset sampling accuracy check requirements (that is, the accuracy of the classification element labels is greater than the preset label accuracy threshold), an error analysis prompt project is generated according to the minority class classification sample data and the corresponding classification element labels, and based on the error analysis prompt project, the minority class classification sample data and the corresponding classification element labels are passed through the large language model to obtain a minority class error synthetic sample data set; otherwise, when the preset sampling accuracy check requirements are not met, it is necessary to further manually label and proofread the minority class classification sample data and the corresponding classification element labels to ensure the quality of the synthetic sample data based on the large model.

[0073] In the specific implementation, set A and set C are used as minority class training sample set A1, and majority class samples whose number is equal to the number of samples in set A1 are randomly extracted from the majority class platform sample data set B, denoted as B1, and A1 and B1 are respectively divided into test set and training set according to the ratio of 1:q, and the sentiment classification model is trained through the training set; the test set is input into the trained sentiment classification model to obtain minority class classification sample data and corresponding classification element labels, and the minority class samples that are misclassified in the test set are collected; the large language model is used to analyze the misclassified minority class samples from the fields of domain, scene, text form, speech tendency, etc. and mark them with element labels. The method comprises the following steps: a. performing statistical analysis on the proportion of minority classification sample data and corresponding classification element labels; b. performing sampling and marking on the minority classification sample data and their corresponding labels, and continuing random sampling after correcting the data with mismatched element labels. The sampling is stopped when the sampling accuracy rate reaches more than 98% in the minority samples, and a list of minority classification sample data and corresponding classification element labels is generated. An error analysis prompt project is generated based on the minority classification sample data and the corresponding classification element labels; and a minority error synthetic sample data set is synthesized through a large language model based on the minority classification sample data and the corresponding classification element labels using the error analysis prompt project, which is recorded as data set D.

[0074] It should also be noted that the minority class sample sets A, C and D are subsequently mixed as the minority class platform sample data set A in step S10, and the above process is repeated until the minority class model sample data set (i.e., m*N minority class sample data) achieves a magnification N.

[0075] Step S40: merging the minority class platform sample data set and the minority class model sample data set, so that the merged minority class sentiment classification sample data set is subjected to sentiment text classification through the sentiment classification model.

[0076] It should be understood that the newly generated m*N minority class sample data are unified in data format with the m minority class sample data collected by the collection platform and manually annotated, and merged with the n majority class sample data collected by the data annotation platform and manually annotated to generate a minority class sentiment classification sample data set for sentiment text classification.

[0077] In this embodiment, sentiment data is first obtained through the public opinion information collection platform according to the platform label, and the sentiment data is preprocessed and analyzed to determine the minority platform sample data set and the majority platform sample data set, and then the number of majority samples corresponding to the majority platform sample data set is determined, and then the minority sample data set is input into the large language model based on the minority sample magnification and the number of majority samples to generate a minority model sample data set, and finally the minority platform sample data set and the minority model sample data set are merged, so that the merged minority sentiment classification sample data set is used for sentiment text classification through the sentiment classification model. Compared with the random oversampling operation in the prior art, this embodiment generates new minority samples, improves the characteristics of minority samples and balances the training data set, avoids the overfitting problem in classification, and at the same time, uses the rich knowledge of the large language model to synthesize pseudo training samples for minority samples in sentiment classification tasks, achieves data and computational efficiency, and improves the recognition ability of neural networks for minority text samples in sentiment classification tasks.

[0078] Reference Figure 4 , Figure 4 It is a structural block diagram of the first embodiment of the system for generating minority class samples based on a large language model of the present invention.

[0079] like Figure 4 As shown, the minority class sample generation system based on the large language model proposed in the embodiment of the present invention includes:

[0080] The collection module 4001 is used to obtain sentiment data through the public opinion information collection platform according to the platform label, and pre-process and analyze the sentiment data to determine the minority platform sample data set and the majority platform sample data set;

[0081] Determination module 4002, used to determine the number of majority class samples corresponding to the majority class platform sample data set;

[0082] A generating module 4003, configured to input the minority class sample data set into a large language model based on the minority class sample magnification and the number of majority class samples to generate a minority class model sample data set;

[0083] The merging module 4004 is used to merge the minority class platform sample data set and the minority class model sample data set, so that the merged minority class sentiment classification sample data set is subjected to sentiment text classification through the sentiment classification model.

[0084] Other embodiments or specific implementations of the system for generating minority class samples based on a large language model of the present invention may refer to the above-mentioned method embodiments and will not be described in detail here.

[0085] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0086] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0087] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory / random access memory, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0088] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for generating minority class samples based on a large language model, characterized in that: The method for generating minority class samples based on a large language model comprises the following steps: Acquire sentiment data through the public opinion information collection platform according to the platform labels, and preprocess and analyze the sentiment data to determine a minority platform sample data set and a majority platform sample data set; Determine the number of majority class samples corresponding to the majority class platform sample data set; Inputting the minority class sample data set into the large language model based on the minority class sample magnification factor and the number of majority class samples to generate a minority class model sample data set; The minority class platform sample data set and the minority class model sample data set are merged, so that the merged minority class sentiment classification sample data set is subjected to sentiment text classification through the sentiment classification model.

2. The method according to claim 1, characterized in that The step of preprocessing and analyzing the sentiment data to determine a minority platform sample data set and a majority platform sample data set includes: Classifying the sentiment data based on the platform labels to obtain minority class sample data and majority class sample data; Annotating the minority sample data, and constructing a minority platform sample data set according to the minority sample data and corresponding annotation information; Determine the platform label accuracy corresponding to the majority class sample data; If the platform label accuracy is greater than a preset label accuracy threshold, a majority class sample data set is constructed according to the majority class sample data and the corresponding platform labels.

3. The method according to claim 2, characterized in that The step of inputting the minority class sample data set into the large language model based on the minority class sample magnification and the number of majority class samples to generate a minority class model sample data set includes: Based on the large model text synthesis prompt project, text data is synthesized through a large language model according to the minority class sample data set to obtain a minority class text synthesis sample data set; Based on the error analysis prompt project, obtaining a minority class error synthesis sample data set according to the minority class data sample set and the minority class text synthesis sample data set through the large language model; Based on the minority class sample magnification and the number of majority class samples, the minority class sample data set, the minority class text synthesis sample data set and the minority class error synthesis sample data set are merged, and based on the large model text synthesis prompt project, the merged sample data set is passed through the large language model to generate a minority class model sample data set.

4. The method according to claim 3, characterized in that The step of obtaining a minority class text synthesis sample data set based on the large model text synthesis prompt engineering through a large language model according to the minority class sample data set includes: Performing data cleaning on the minority class sample data in the minority class sample data set, and analyzing the cleaned minority class sample data to obtain element labels corresponding to the minority class sample data; Extracting minority class sample labeled data from the minority class sample data, and determining the accuracy of the element labels corresponding to the minority class sample labeled data; If the accuracy of the element label is greater than the preset label accuracy threshold, a large model text synthesis prompt project is generated according to the minority class sample data and the corresponding element label; Based on the large model text synthesis prompt project, the minority class sample data and the corresponding element labels are passed through a large language model to obtain a minority class text synthesis sample data set.

5. The method according to claim 3, characterized in that The step of obtaining a minority class error synthesis sample data set based on the error analysis prompting project according to the minority class data sample set and the minority class text synthesis sample data set through the large language model comprises: The minority class sample data set and the minority class text synthesis sample data set are used as a minority class training sample set, wherein the minority class training sample set includes minority class training sample data; Determine the number of training samples corresponding to the minority class training sample set; Selecting majority class training sample data from the majority class sample data set based on the number of training samples; The sentiment classification model is trained according to the minority class training sample data and the majority class training sample data to obtain the minority class classification sample data category output by the sentiment classification model, and element label analysis and extraction are performed on the misclassified minority class classification sample data; Based on the error analysis prompt engineering, the minority class classification sample data and the corresponding classification element labels are passed through a large language model to obtain a minority class error synthetic sample data set.

6. The method according to claim 5, characterized in that Before the step of obtaining a minority class error synthetic sample data set based on the error analysis prompting project according to the minority class classification sample data and the corresponding classification element labels through a large language model, the step further includes: When the labeling accuracy of the minority classification sample data and the corresponding classification element labels meets the preset sampling accuracy inspection requirement, an error analysis prompt project is generated according to the minority classification sample data and the corresponding classification element labels.

7. A minority class sample generation system based on a large language model, characterized in that: The minority class sample generation system based on the large language model includes: A collection module is used to obtain sentiment data through a public opinion information collection platform according to platform labels, and to preprocess and analyze the sentiment data to determine a minority platform sample data set and a majority platform sample data set; A determination module, used to determine the number of majority class samples corresponding to the majority class platform sample data set; A generating module, configured to input the minority class sample data set into a large language model based on the minority class sample magnification and the number of majority class samples to generate a minority class model sample data set; The merging module is used to merge the minority class platform sample data set and the minority class model sample data set, so that the merged minority class sentiment classification sample data set is subjected to sentiment text classification through the sentiment classification model.

8. A storage medium, characterized in that: The storage medium stores a minority class sample generation program based on a large language model, and when the minority class sample generation program based on a large language model is executed by a processor, the steps of the minority class sample generation method based on a large language model as described in any one of claims 1 to 6 are implemented.