Data enhancement method and device, computer equipment, storage medium and program product

By analyzing the original data set, data enhancement information is generated and the enhancement process of data samples is guided according to the information, the problem of data enhancement in the existing technology is solved, and the intelligent generation and quality evaluation of data enhancement samples are realized, and the quality and reliability of data generation are improved.

CN120030351APending Publication Date: 2025-05-23BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184102.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing data augmentation technology lacks understanding of data samples, resulting in poor quality of generated data, and it is difficult to ensure the semantic rationality of data and the maintenance of business constraints.

Method used

By performing data analysis on the data samples in the original data set, data augmentation information is generated, and the data augmentation process is guided according to the data augmentation information, each type of data augmentation sample is generated, and the data augmentation sample is evaluated in quality to ensure the rationality and reliability of the generated data.

Benefits of technology

It realizes intelligent generation of data-enhanced samples, improves the quality and reliability of data generation, and reduces the cost and time of obtaining high-quality training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030351A_ABST
    Figure CN120030351A_ABST
Patent Text Reader

Abstract

The invention discloses a data enhancement method and device, computer equipment, a storage medium and a program product, and relates to the technical field of data processing.The method comprises the steps that a to-be-enhanced original data set is obtained, and the original data set comprises multiple different types of data samples; performing data analysis on the data samples in the original data set to generate data enhancement information; guiding enhancement processes for different types of data samples according to the data enhancement information, and generating various types of data enhancement samples; and performing quality evaluation on the data enhancement sample, and generating target enhancement data based on a quality evaluation result. By implementing the technical scheme of the invention, the understanding of the data sample is realized, the intelligent generation of the data enhancement sample is realized, the generation reasonability of the data enhancement sample is ensured, and meanwhile, the data reliability and the data generation quality of the data enhancement sample are improved through automatic quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data enhancement method, device, computer equipment, storage medium and program product. Background Art

[0002] In the field of machine learning, high-quality training data is crucial to model performance. However, in practical applications, it is costly to obtain a large amount of high-quality labeled data, the manual labeling process is prone to inconsistencies or errors, and the data set has the problem of uneven data distribution. Currently, data enhancement is mainly used to obtain high-quality training data. However, both traditional data enhancement technology and rule-based data enhancement technology lack an understanding of data samples, resulting in data enhancement being too mechanical and difficult to ensure the quality of the generated data. Summary of the invention

[0003] In view of this, the present disclosure provides a data enhancement method, apparatus, computer device, storage medium and program product to solve the problem of poor data generation quality.

[0004] In a first aspect, the present disclosure provides a data enhancement method, including: obtaining an original data set to be enhanced, the original data set including a plurality of different types of data samples; performing data analysis on the data samples in the original data set to generate data enhancement information; guiding an enhancement process for different types of data samples according to the data enhancement information to generate data enhancement samples of various types; performing quality assessment on the data enhancement samples, and generating target enhancement data based on the quality assessment results.

[0005] In a second aspect, the present disclosure provides a data enhancement device, including: an acquisition module, used to acquire an original data set to be enhanced, the original data set including multiple different types of data samples; an analysis module, used to perform data analysis on the data samples in the original data set to generate data enhancement information; an enhancement module, used to guide the enhancement process for different types of data samples according to the data enhancement information to generate data enhancement samples of various types; an evaluation module, used to perform quality evaluation on the data enhancement samples and generate target enhancement data based on the quality evaluation results.

[0006] In a third aspect, the present disclosure provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the data enhancement method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0007] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the data enhancement method of the first aspect or any corresponding embodiment thereof.

[0008] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, which are used to enable a computer to execute the data enhancement method of the first aspect or any corresponding embodiment thereof.

[0009] The data enhancement method, device, computer equipment, storage medium and program product provided by the present disclosure perform data analysis on the data samples in the original data set to achieve understanding of the data samples and obtain corresponding data enhancement information, thereby guiding the enhancement process of the data samples according to the data enhancement information, ensuring the rationality of the generation of the data enhancement samples, and realizing the intelligent generation of the data enhancement samples. At the same time, the quality of the data enhancement samples is evaluated, and the data reliability and data generation quality of the data enhancement samples are improved through automated quality evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] Figure 1 is a flowchart of a data enhancement method according to an embodiment of the present disclosure;

[0012] Figure 2 is a flowchart of another data enhancement method according to an embodiment of the present disclosure;

[0013] Figure 3 is a flowchart of another data enhancement method according to an embodiment of the present disclosure;

[0014] Figure 4 is a structural block diagram of a data enhancement device according to an embodiment of the present disclosure;

[0015] Figure 5 It is a schematic diagram of the hardware structure of the computer device of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.

[0017] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0018] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0019] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the computer device.

[0020] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0021] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0022] In the field of machine learning, there are problems in obtaining high-quality training data: the cost of obtaining a large amount of high-quality labeled data is high (i.e. manual labeling requires a large number of professionals, the labeling process is time-consuming, affecting the project progress, professional labelers in certain fields are scarce, and the labeling cost increases linearly with the size of the data), the manual labeling process is prone to inconsistencies or errors (i.e. inconsistent standards between different labelers, deviations in the understanding of labeling rules, fatigue errors caused by long-term labeling, and unclear judgment standards in complex scenarios), and the data set has an unbalanced data distribution (i.e. the number of minority class samples is seriously insufficient; samples of certain feature combinations are scarce; data distribution in real scenarios changes dynamically; data distribution in different scenarios differs significantly).

[0023] In order to solve the above problems, data enhancement is mainly used to obtain high-quality training data. Currently, common data enhancement methods mainly include traditional data enhancement technology and rule-based data generation technology. Among them, traditional data enhancement technology includes: in the image field, data enhancement is carried out by rotation, flipping, cropping, scaling, adding noise, etc.; in the text field, data enhancement is carried out by synonym replacement, back translation, syntactic transformation, etc.; in the speech field, data enhancement is carried out by time stretching, pitch change, adding background noise, etc.; in the tabular data, data enhancement is carried out by random sampling, feature combination, numerical perturbation, etc. Rule-based data generation technology uses predefined template filling, business rule combination generation, parameter randomization generation, random generation under constraints, etc. to achieve data enhancement.

[0024] However, the above data enhancement methods are too mechanical and lack semantic understanding, making it difficult to ensure the semantic rationality of the generated data and to maintain complex business constraints, thus making it difficult to ensure the quality of the generated data.

[0025] Based on this, the disclosed technical solution ensures the rationality of generated data by understanding the data samples in the original data set, realizes the intelligent generation of data enhancement samples, and is conducive to reducing the cost and time of obtaining high-quality training data. At the same time, it can perform automated quality assessment, improve data reliability, and ensure the quality of data generation.

[0026] According to an embodiment of the present disclosure, a data enhancement method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0027] In this embodiment, a data enhancement method is provided, which can be used in computer devices, such as computers, servers, etc. Figure 1is a flow chart of a data enhancement method according to an embodiment of the present disclosure. Figure 1 As shown, the process includes the following steps:

[0028] Step S101 : obtaining an original data set to be enhanced, where the original data set includes data samples of various types.

[0029] The original data set is a small amount of model training data collected in advance, and the original data set includes data samples of different types to ensure that the model training data can cover multiple types and ensure the balance of data distribution of the original data set.

[0030] Specifically, the original data set can be obtained by manually annotating according to the model usage field; it can also be collected from public data sets (such as public databases, data sharing platforms, etc.); it can also be data captured online through web crawler technology. Of course, it can also be obtained by collecting user-generated content from social applications. The method for obtaining the original data set is not specifically limited here, and those skilled in the art can determine it according to actual needs.

[0031] Step S102: performing data analysis on the data samples in the original data set to generate data enhancement information.

[0032] Data enhancement information is used to expand the data samples in the original data set to obtain a large amount of model training data. The data enhancement information includes data enhancement strategies and data enhancement suggestions. The data enhancement strategies control the expansion method of data samples, and the data enhancement suggestions guide the enhancement method of data samples.

[0033] Specifically, a large language model is used to perform feature analysis on data samples in the original data set to determine the feature distribution of the original data set, and the imbalanced data and missing data in the original data set are identified according to the feature distribution. Therefore, by combining the data distribution characteristics and the identified imbalanced data and missing data, data enhancement information for the original data set can be determined to determine the data types and feature combinations that currently need to be enhanced, so as to facilitate the subsequent expansion of data samples based on the data enhancement information.

[0034] Step S103: guiding the enhancement process for different types of data samples according to the data enhancement information to generate data enhancement samples of various types.

[0035] Data enhancement samples are data samples generated through data enhancement. Since the original data set includes different types of data samples, each type of data sample is expanded according to the data enhancement information to supplement the unbalanced data samples and generate data samples with characteristic feature combinations, thereby generating different types of data enhancement samples in a targeted manner.

[0036] Step S104: perform quality assessment on the data enhancement samples, and generate target enhancement data based on the quality assessment results.

[0037] In order to ensure that the data enhancement samples meet the requirements, it is necessary to conduct quality assessment on the data enhancement samples to verify whether the data enhancement samples meet the characteristic distribution of the model training data, whether the data enhancement samples have unreasonable data or contradictory data, etc., and generate corresponding quality assessment results. Then, based on the quality assessment results, the data enhancement samples that have not passed the quality assessment are corrected to obtain data enhancement samples that meet the requirements. The data enhancement samples that have passed the quality assessment and the corrected data enhancement samples are merged to obtain the final target enhancement data.

[0038] The data enhancement method provided in this embodiment performs data analysis on the data samples in the original data set to achieve understanding of the data samples and obtain corresponding data enhancement information, thereby guiding the enhancement process of the data samples according to the data enhancement information, ensuring the rationality of the generation of the data enhancement samples, and realizing the intelligent generation of the data enhancement samples. At the same time, the quality of the data enhancement samples is evaluated, and the data reliability and data generation quality of the data enhancement samples are improved through automated quality evaluation.

[0039] In this embodiment, a data enhancement method is provided, which can be used in computer devices, such as computers, servers, etc. Figure 2 is a flow chart of a data enhancement method according to an embodiment of the present disclosure. Figure 2 As shown, the process includes the following steps:

[0040] Step S201, obtaining an original data set to be enhanced, the original data set including a plurality of different types of data samples. For details, please refer to the relevant description of the corresponding steps in the above-mentioned embodiment, which will not be repeated here.

[0041] Step S202: performing data analysis on the data samples in the original data set to generate data enhancement information.

[0042] Specifically, the above step S202 includes:

[0043] Step S2021, obtaining data analysis prompt description information for the original data set.

[0044] Data analysis prompt description information refers to the guiding information provided for data analysis, so as to effectively analyze the data samples in the original data set. Specifically, the data analysis model is trained based on the large language model architecture, and the data analysis prompt description information is the natural language description information set for analyzing the original data set. The data analysis model can analyze the data samples through the data analysis prompt description information.

[0045] The following is an example of the description information of the data analysis prompt:

[0046] You are a data scientist analyzing a dataset. Please analyze the following dataset:

[0047] [Input original data set]

[0048] Please provide a detailed analysis including the following:

[0049] 1. The distribution of features and their relationships;

[0050] 2. Class imbalance problem and severity indicators;

[0051] 3. Missing patterns and their potential causes;

[0052] 4. Data quality issues and impact assessment;

[0053] Answer in the form of a structured report including:

[0054] a) Quantitative indicators for each question;

[0055] b) The following specific recommendations:

[0056] - The target number of samples required for each type;

[0057] -Priority areas for data enhancement;

[0058] -Required edge cases and special scenarios;

[0059] - Proposed data generation method.

[0060] Step S2022: Analyze each type of data sample in the original data set according to the data analysis prompt description information to generate feature distribution information of the data samples.

[0061] Feature distribution information is used to characterize the distribution of data samples in the original data set. Specifically, the data analysis model analyzes data samples of various types in the original data set according to the data analysis prompt description information, identifies the data sample types in the original data set, such as numerical type, category type, time series type, text type, etc., and determines one or more data samples of each type. The data samples of each type are analyzed to determine the feature distribution information of data samples of different types.

[0062] For example, descriptive statistics, outlier detection, and distribution chart analysis are performed on numerical data samples to determine statistical values ​​such as the mean and median of the data samples, distribution information such as normal distribution or skewed distribution, the number of outliers, etc. For categorical data samples, the number and frequency of each type of data samples, as well as the proportion of each type of data samples in the data set, are calculated to determine the type imbalance of the data samples.

[0063] Step S2023: Determine missing data and unbalanced data in the original data set based on the feature distribution information.

[0064] Missing data refers to the lack of data samples under a certain type; unbalanced data refers to the imbalanced distribution of data samples under each type. Specifically, the feature distribution of data samples of different types in the original data set can be determined based on the feature distribution information. By analyzing the feature distribution, it is possible to identify whether data samples under each type are missing, and combined with the missing identification results, the missing data in the original data set can be determined; at the same time, by calculating the number of data samples under different types, it is determined whether data samples under each type are distributed unbalancedly, and combined with the balanced calculation results, the unbalanced data in the original data set can be determined.

[0065] Step S2024: determine data enhancement information according to missing data and unbalanced data.

[0066] Based on the missing data and unbalanced data in the original data set, we can determine the type of data, feature combination, and data enhancement ratio that currently need to be enhanced, and combine the data enhancement information for the original data set according to the type of data, feature combination, and data enhancement ratio that need to be enhanced.

[0067] In some optional implementations, the data enhancement information includes a data enhancement ratio and a target enhancement quantity, wherein the data enhancement ratio indicates the ratio of data samples used for enhancement to the total data samples; and the target enhancement quantity indicates the quantity of data samples used for enhancement.

[0068] Accordingly, the above step S2024 includes: determining the data enhancement ratio and the target enhancement quantity based on the missing data and the unbalanced data.

[0069] For missing data, the goal of data augmentation is to generate new data samples containing missing values. Specifically, missing values ​​can be interpolated by statistical methods such as mean, median, mode, etc., and missing values ​​can be predicted using regression models or other prediction methods. The data augmentation ratio needs to match the missing value ratio, so the data augmentation ratio can be set based on the missing value ratio. For example, if 10% of the data is missing, the augmentation ratio may need to be set to 10%. The target augmentation number corresponds to the number of missing values. For example, if there are 100 missing values, 100 new data samples may need to be generated.

[0070] For imbalanced data, the goal of data augmentation is to increase the number of data samples in the minority class, usually without changing the number of data samples in the majority class. Specifically, oversampling is used to increase the number of data samples in the minority class, such as repetition or synthesis; undersampling is used to reduce the number of data samples in the majority class; synthetic data is generated using generative models (such as GANs), etc. The data augmentation ratio is usually determined based on the proportion of the minority class. For example, if the minority class accounts for 1%, the augmentation ratio may need to be set to at least 1%; the target augmentation number should be large enough to significantly improve the model's predictive performance for the minority class. Specifically, the balance coefficient can be used to determine the number of augmentations, or cross-validation or a separate validation set can be used to determine the number of augmentation samples required, which is not specifically limited here.

[0071] In the above implementation, the corresponding data enhancement ratio and target enhancement quantity are determined in combination with missing data and unbalanced data to ensure that data enhancement can satisfy the balanced data distribution and achieve accurate data distribution optimization.

[0072] In some optional implementations, the data enhancement information includes boundary enhancement information, and the boundary enhancement information is used to enhance data samples in the boundary area to improve the performance of the model in the boundary area.

[0073] Accordingly, the above step S2024 includes: determining boundary enhancement information based on missing data and unbalanced data.

[0074] For missing data, analyze the distribution of missing data, identify missing value boundaries, and determine feature boundaries or feature combination boundaries with missing values. According to the distribution of missing values, select appropriate interpolation strategies to interpolate data samples, such as mean interpolation, median interpolation, K-nearest neighbor interpolation, model prediction-based interpolation, etc.

[0075] When performing boundary enhancement for missing data, interpolation methods can be used to fill in missing values ​​near the boundary to obtain boundary enhancement information for data samples in the boundary area; new data samples can also be created near the boundary to simulate the possible range of missing values ​​to obtain boundary enhancement information for data samples in the boundary area.

[0076] For imbalanced data, analyze the distribution of data samples of various types in the original data set, identify the boundaries of imbalanced data, and determine the types or type combinations with fewer data samples. According to the boundaries of imbalanced data, oversampling can be used to add data samples or synthesize data samples near the boundaries of minority classes to obtain boundary enhancement information for data enhancement of data samples in the boundary area; new data samples can also be created to simulate the possible distribution of minority classes to obtain boundary enhancement information for data enhancement of data samples in the boundary area.

[0077] In the above implementation, missing data and unbalanced data are combined to determine corresponding boundary enhancement information, which effectively solves the problem of insufficient feature coverage and ensures the balanced distribution of data samples under different types.

[0078] Step S203: guiding the enhancement process for different types of data samples according to the data enhancement information to generate data enhancement samples of various types.

[0079] Specifically, the above step S203 includes:

[0080] Step S2031, analyzing the data enhancement information, and determining data generation prompt description information corresponding to the data enhancement information.

[0081] Data generation prompt description information refers to the guiding information provided for the enhancement of data samples, so as to generate targeted data enhancement samples of different types. Specifically, based on the training of the data generation model based on the large language model architecture, the data generation prompt description information is the natural language description information set for data sample enhancement, and the data generation model can enhance the data sample through the data generation prompt description information. Among them, the examples of data generation prompt description information are as follows:

[0082] 1. Abstract:

[0083] - Type distribution: Category A (20%), Category B (70%), Category C (10%);

[0084] - Key missing feature combination: [feature X=1, feature Y>10];

[0085] -Border cases required: transaction amount exceeds USD 1 million and risk score is high;

[0086] 2. Generation requirements:

[0087] - Generate 500 samples for class A to achieve 30% representation;

[0088] -Focus on samples with feature X=1 and feature Y>10;

[0089] - Includes 50 edge cases that meet specified criteria;

[0090] Reference example:

[0091] [3-5 high-quality examples with detailed annotations]

[0092] Let's generate the data step by step:

[0093] 1. Analyze the patterns and constraints in the reference examples;

[0094] 2. Extract the main features that make these examples effective;

[0095] 3. Generate new samples by:

[0096] a) Keep the core pattern in the example;

[0097] b) Varying non-critical attributes within valid ranges;

[0098] c) Ensure target category and feature requirements;

[0099] d) Ensure compliance with business rules;

[0100] 4. Verify that each generated sample complies with:

[0101] -Format and structure requirements;

[0102] -Business logic constraints;

[0103] -Target distribution target;

[0104] Please generate a sample according to the above specifications, starting with Category A.

[0105] Step S2032: perform data enhancement on different types of data samples according to the data generation prompt description information to obtain data enhancement samples corresponding to each type.

[0106] When performing data enhancement, the type of data sample is identified and the data enhancement method of each type is determined. Then, when enhancing data samples of different types according to the data generation prompt description information, the data sample is enhanced using the data enhancement method corresponding to each type according to the type of data sample and the enhancement target of the data sample to obtain the corresponding data enhancement sample.

[0107] For example, for numerical data samples, you can add random noise or use mathematical functions (such as normal distribution, logarithmic transformation, etc.) to expand the data range; you can also perform data interpolation in areas where the data is sparse; you can also use machine learning models (such as generative adversarial networks (GANs)) to generate new numerical samples, etc. to enhance numerical data samples.

[0108] For example, for classified data samples, SMOTE technology can be used to copy data samples of the minority class to increase its sample quantity; undersampling can also be used to reduce the number of data samples of the majority class to balance the distribution of data samples under each type; clustering algorithms can also be used to synthesize new data samples under the minority class to enhance classified data samples.

[0109] In some optional implementations, the above step S203 further includes:

[0110] Step a1: Determine the data enhancement strategy and data enhancement priority based on the data enhancement information.

[0111] Step a2: According to the data enhancement strategy and data enhancement priority, data enhancement is performed on the imbalanced data and the missing data, and / or data enhancement is performed on the boundary data.

[0112] The data enhancement strategy is a strategy for expanding data samples; the data enhancement priority is the order of enhancing data samples of different types. The data enhancement information contains relevant information for enhancing data samples. By parsing the data enhancement information, the corresponding data enhancement strategy and data enhancement priority are extracted from it. Then, according to the data enhancement strategy and data enhancement priority, data samples of different types are generated in a targeted manner to supplement the data samples of the unbalanced type and generate data samples with specific feature combinations. At the same time, according to the data enhancement strategy and data enhancement priority, data samples in the boundary area can be constructed to achieve data enhancement for the data samples in the boundary area.

[0113] In the above implementation, by parsing the data enhancement strategy and the data enhancement priority in the data enhancement information and performing data enhancement according to the data enhancement strategy and the data enhancement priority, the rationality of the generation of data enhancement samples can be effectively improved.

[0114] Step S204: perform quality assessment on the data enhancement sample, and generate target enhancement data based on the quality assessment result. For details, please refer to the relevant description of the corresponding steps in the above-mentioned embodiment, which will not be repeated here.

[0115] The data enhancement method provided in this embodiment generates feature distribution information of the data samples by analyzing the data samples according to the data analysis prompt description information, thereby ensuring the rationality of the generation of feature distribution information through semantic understanding ability, thereby ensuring the intelligent extraction of missing data and unbalanced data, and then improving the rationality of the generation of data enhancement information, which is conducive to reducing the cost and time of obtaining high-quality training data. By guiding the data enhancement of data samples of different types according to the data generation prompt description information, the data enhancement samples corresponding to each type are obtained, thereby combining semantic understanding to perform data enhancement, ensuring the semantic rationality of the data enhancement process, and further ensuring the coherence of the data enhancement samples in the data context. At the same time, it is possible to quickly adapt to data enhancement in different business scenarios by adjusting the corresponding prompt description information, thereby improving the adaptability of data enhancement in various fields, and then improving the scalability of data enhancement.

[0116] In this embodiment, a data enhancement method is provided, which can be used in computer devices, such as computers, servers, etc. Figure 3 is a flow chart of a data enhancement method according to an embodiment of the present disclosure. Figure 3 As shown, the process includes the following steps:

[0117] Step S301, obtaining an original data set to be enhanced, the original data set including data samples of different types. For details, please refer to the relevant description of the corresponding steps in the above-mentioned embodiment, which will not be repeated here.

[0118] Step S302: Perform data analysis on the data samples in the original data set to generate data enhancement information. For details, please refer to the description of the corresponding steps in the above embodiment, which will not be repeated here.

[0119] Step S303: guide the enhancement process for different types of data samples according to the data enhancement information to generate data enhancement samples of various types. For details, please refer to the relevant description of the corresponding steps in the above-mentioned embodiment, which will not be repeated here.

[0120] Step S304: perform quality assessment on the data enhancement samples, and generate target enhancement data based on the quality assessment results.

[0121] Specifically, the above step S304 includes:

[0122] Step S3041, obtaining quality assessment prompt description information.

[0123] The quality assessment hint description information refers to the guiding information provided for the evaluation of data augmentation samples, so as to effectively evaluate the generation quality of data augmentation samples. Specifically, the quality assessment model is trained based on the large language model architecture, and the quality assessment hint description information is the natural language description information set for the evaluation of data augmentation samples. The quality assessment model can use the quality assessment hint description information to perform quality assessment on data augmentation samples. Examples of quality assessment hint description information are as follows:

[0124] You are a data quality expert. Please review the following generated data sample:

[0125] [Generated data augmentation samples]

[0126] For each data augmentation example, the following are evaluated:

[0127] 1. Logical consistency;

[0128] 2. Format compliance;

[0129] 3. Range validity;

[0130] 4. Semantic coherence;

[0131] If you find any issues, please provide specific suggestions for corrections.

[0132] Provide an evaluation (pass / fail) for each data augmentation example and explain your rationale.

[0133] Step S3042: Perform quality assessment on the data enhancement sample according to the quality assessment prompt description information.

[0134] The quality assessment model analyzes the data enhancement samples generated under each type according to the quality assessment prompt description information to determine whether the data enhancement samples generated under each type pass the quality assessment, including whether they meet the target distribution requirements, whether they meet the key features required for data samples, whether they meet business rules and constraints, and whether there are unreasonable or contradictory data samples.

[0135] If the data enhancement sample fails the quality assessment, steps S3043 to S3044 are executed; if the data enhancement sample passes the quality assessment, the data enhancement sample is merged into the original data set.

[0136] Step S3043: If the data enhancement sample fails the quality assessment, a data correction suggestion is generated.

[0137] Data correction suggestions are measures to correct unqualified data augmentation samples. If the data augmentation sample fails the quality assessment, it means that the generated data augmentation sample is unqualified. At this time, corresponding data correction suggestions can be given for the unqualified data augmentation sample.

[0138] Step S3044: correct the data enhancement sample according to the data correction suggestion to obtain corrected target enhancement data.

[0139] Correct the unqualified data enhancement samples according to the data correction suggestions until the data enhancement samples can pass the quality assessment, and determine the corrected data enhancement samples that pass the quality assessment as the target enhancement data.

[0140] In some optional implementations, the above method further includes:

[0141] Step b1: fuse the data enhancement samples that have passed the quality assessment and the target enhancement data with the original data set to generate an enhanced data set.

[0142] Step b2: dynamically adjust the sampling strategy of the enhanced data set to generate the feature distribution result of the enhanced data set.

[0143] The data enhancement samples that have passed the quality assessment and the target enhancement data are merged to obtain a data enhancement sample set, and then the data enhancement sample set is fused with the original data set according to a predetermined mixing ratio to generate an enhanced data set. The enhanced data set achieves a balanced distribution of data samples of various types and ensures the feature coverage of the data samples.

[0144] The sampling strategy is a strategy for sampling the enhanced dataset. The sampling strategy of the enhanced dataset is adjusted in real time according to the distribution of the dataset to obtain the dynamically adjusted feature distribution result until the feature distribution result can maintain the proportion of data samples in the enhanced dataset in each type or feature consistent or nearly consistent with the overall dataset.

[0145] In a specific example, the enhanced dataset is stratified according to key features (such as categories, labels) according to stratified sampling, ensuring that each layer maintains the same proportion as the overall dataset during the sampling process. If the amount of data in some layers is very small, resampling techniques (such as oversampling or undersampling) can be used to balance these layers, and the sampling strategy is adjusted according to the changes in the distribution of data samples under each type in each iteration by dynamically adjusting the sampling rate.

[0146] Of course, the sampling weights can also be adjusted according to the distribution of the current data samples in each iteration or batch processing. For data samples with unbalanced distribution, higher sampling weights can be assigned to data samples in the minority class. The sampling strategy is not limited here, and those skilled in the art can determine it according to actual needs.

[0147] By dynamically adjusting the sampling strategy, the feature distribution balance of data samples in the enhanced dataset is dynamically maintained, thereby improving the efficiency and accuracy of model training while maintaining the diversity of data samples.

[0148] In some optional implementations, the above method further includes:

[0149] Step c1: perform fusion evaluation on the enhanced data set according to a preset data evaluation cycle to generate a fusion evaluation result.

[0150] Step c2: adjusting the data samples in the enhanced data set based on the fusion evaluation results.

[0151] The preset data evaluation period is a preset period for triggering fusion evaluation, such as 1 hour, 6 hours, 12 hours, 24 hours, etc., which is not specifically limited here. Fusion evaluation refers to the evaluation of the fusion effect of the data enhancement sample and the original data set.

[0152] Specifically, a fusion evaluation model is trained based on a large language model architecture, and corresponding fusion evaluation prompt description information is set for the fusion evaluation process. The fusion evaluation model can then perform fusion evaluation of data samples for the enhanced dataset according to the preset data evaluation cycle through the fusion evaluation prompt description information, generate corresponding fusion evaluation results, and adjust the data samples in the enhanced dataset based on the fusion suggestions given in the fusion evaluation results so that it can maintain a balanced feature distribution of data samples of each type.

[0153] An example of the description information of the fusion evaluation prompt is as follows:

[0154] The following augmented datasets are analyzed:

[0155] Original dataset: [Statistical information of the original dataset]

[0156] Augmented dataset: [Statistics of the augmented dataset]

[0157] Evaluate:

[0158] 1. Distribution consistency between original data and enhanced data;

[0159] 2. Overall diversity indicators;

[0160] 3. Coverage of feature space;

[0161] Provides suggestions on:

[0162] -Optimal mixing ratio;

[0163] - Sampling strategy;

[0164] - Provide balance adjustment strategies if adjustments are needed.

[0165] By performing fusion evaluation on the enhanced data set according to a preset data evaluation cycle, the data samples in the enhanced data set can be adjusted according to the fusion evaluation results, thereby supporting incremental data optimization and updating.

[0166] The data enhancement method provided in this embodiment implements automated quality assessment of data enhancement samples by performing quality assessment on data enhancement samples according to the quality assessment prompt description information. When a data enhancement sample fails the quality assessment, a corresponding data correction suggestion is generated, and the unqualified data enhancement sample is corrected according to the data correction suggestion, thereby improving the reliability of the data enhancement sample through the correction mechanism and ensuring the generation quality of the data enhancement sample.

[0167] In this embodiment, a data enhancement device is also provided, which is used to implement the above embodiments and preferred implementation modes, and will not be repeated hereafter. As used below, the term "module" may be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0168] This embodiment provides a data enhancement device, such as Figure 4 As shown, including:

[0169] The acquisition module 401 is used to acquire an original data set to be enhanced, where the original data set includes data samples of different types.

[0170] The analysis module 402 is used to perform data analysis on the data samples in the original data set to generate data enhancement information.

[0171] The enhancement module 403 is used to guide the enhancement process for different types of data samples according to the data enhancement information to generate data enhancement samples of various types.

[0172] The evaluation module 404 is used to perform quality evaluation on the data enhancement samples and generate target enhancement data based on the quality evaluation results.

[0173] In some optional implementations, the analysis module 402 includes:

[0174] The first prompt description unit is used to obtain data analysis prompt description information for the original data set.

[0175] The sample analysis unit is used to analyze data samples of various types in the original data set according to the data analysis prompt description information to generate feature distribution information of the data samples.

[0176] The abnormal data determination unit is used to determine missing data and unbalanced data in the original data set based on feature distribution information.

[0177] The enhancement information generating unit is used to determine data enhancement information according to missing data and unbalanced data.

[0178] In some optional implementations, the data enhancement information includes a data enhancement ratio and a target enhancement quantity. Accordingly, the enhancement information generation unit includes:

[0179] The first enhancement information generating subunit is used to determine the data enhancement ratio and the target enhancement quantity based on the missing data and the unbalanced data.

[0180] In some optional implementations, the data enhancement information includes boundary enhancement information. Accordingly, the enhancement information generation unit includes:

[0181] The second enhancement information generating subunit is used to determine the boundary enhancement information based on the missing data and the unbalanced data.

[0182] In some optional implementations, the enhancement module 403 includes:

[0183] The first prompt description unit is used to analyze the data enhancement information and determine the data corresponding to the data enhancement information to generate prompt description information.

[0184] The first data enhancement unit is used to perform data enhancement on different types of data samples according to the data generation prompt description information to obtain data enhancement samples corresponding to each type.

[0185] In some optional implementations, the enhancement module 403 further includes:

[0186] The enhancement information determination unit is used to determine the data enhancement strategy and the data enhancement priority based on the data enhancement information.

[0187] The second data enhancement unit is used to perform data enhancement on imbalanced data and missing data and / or perform data enhancement on boundary data according to the data enhancement strategy and data enhancement priority.

[0188] In some optional implementations, the evaluation module 404 includes:

[0189] The third prompt description unit is used to obtain quality assessment prompt description information.

[0190] The quality assessment unit is used to perform quality assessment on the data augmentation samples according to the quality assessment prompt description information.

[0191] The correction suggestion generating unit is used to generate a data correction suggestion if the data augmentation sample fails the quality assessment.

[0192] The correction unit is used to correct the data enhancement sample according to the data correction suggestion to obtain the corrected target enhancement data.

[0193] In some optional embodiments, the above device further includes:

[0194] The fusion module is used to fuse the data enhancement samples that have passed the quality assessment and the target enhancement data with the original data set to generate an enhanced data set.

[0195] The sampling adjustment module is used to dynamically adjust the sampling strategy of the enhanced data set and generate the feature distribution results of the enhanced data set.

[0196] In some optional embodiments, the above device further includes:

[0197] The fusion evaluation module is used to perform fusion evaluation on the enhanced data set according to the preset data evaluation cycle and generate a fusion evaluation result.

[0198] The adjustment module is used to adjust the data samples in the enhanced data set based on the fusion evaluation results.

[0199] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0200] The data enhancement device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0201] The data enhancement device provided in this embodiment performs data analysis on the data samples in the original data set to achieve understanding of the data samples and obtain corresponding data enhancement information, thereby guiding the enhancement process of the data samples according to the data enhancement information, ensuring the rationality of the generation of the data enhancement samples, and realizing the intelligent generation of the data enhancement samples. At the same time, the quality of the data enhancement samples is evaluated, and the data reliability and data generation quality of the data enhancement samples are improved through automated quality evaluation.

[0202] The present disclosure also provides a computer device having the above Figure 4 The data enhancement device shown.

[0203] See also Figure 5 , Figure 5 is a schematic diagram of a computer device provided by an optional embodiment of the present disclosure, such as Figure 5 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.

[0204] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0205] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0206] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0207] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0208] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0209] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0210] A part of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0211] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A data enhancement method, characterized in that: The method comprises: Acquire an original data set to be enhanced, where the original data set includes data samples of multiple different types; Performing data analysis on the data samples in the original data set to generate data enhancement information; Guide the enhancement process for different types of the data samples according to the data enhancement information to generate data enhancement samples of various types; A quality assessment is performed on the data enhancement samples, and target enhancement data is generated based on the quality assessment result.

2. The method according to claim 1, characterized in that The performing data analysis on the data samples in the original data set to generate data enhancement information includes: Obtaining data analysis prompt description information for the original data set; Analyze the data samples of each type in the original data set according to the data analysis prompt description information to generate feature distribution information of the data samples; Based on the feature distribution information, determining missing data and unbalanced data in the original data set; Data augmentation information is determined according to the missing data and the imbalanced data.

3. The method according to claim 2, characterized in that The determining of data enhancement information according to the missing data and the unbalanced data includes: Based on the missing data and the unbalanced data, determine a data enhancement ratio and a target enhancement quantity, wherein the data enhancement information includes the data enhancement ratio and the target enhancement quantity; and / or, Based on the missing data and the unbalanced data, boundary enhancement information is determined, and the data enhancement information includes boundary enhancement information.

4. The method according to claim 1, characterized in that: The step of guiding the enhancement process for different types of data samples according to the data enhancement information to generate data enhancement samples of various types includes: Analyze the data enhancement information to determine data generation prompt description information corresponding to the data enhancement information; Data enhancement is performed on the data samples of different types according to the data generation prompt description information to obtain the data enhanced samples corresponding to each type.

5. The method according to claim 4, characterized in that Also includes: Based on the data enhancement information, determine a data enhancement strategy and a data enhancement priority; According to the data enhancement strategy and the data enhancement priority, data enhancement is performed on the imbalanced data and the missing data, and / or data enhancement is performed on the boundary data.

6. The method according to claim 1, characterized in that The step of performing quality assessment on the data enhancement sample and generating target enhancement data based on the quality assessment result includes: Get quality assessment prompt description information; Performing quality assessment on the data augmentation sample according to the quality assessment prompt description information; If the data augmentation sample fails the quality assessment, a data correction suggestion is generated; The data enhancement sample is corrected according to the data correction suggestion to obtain the corrected target enhancement data.

7. The method according to claim 6, characterized in that Also includes: Performing data fusion on the data enhancement samples that pass the quality assessment and the target enhancement data with the original data set to generate an enhanced data set; The sampling strategy of the enhanced data set is dynamically adjusted to generate a feature distribution result of the enhanced data set.

8. The method according to claim 7, characterized in that Also includes: According to a preset data evaluation cycle, a fusion evaluation is performed on the enhanced data set to generate a fusion evaluation result; The data samples in the enhanced data set are adjusted based on the fusion evaluation result.

9. A data enhancement device, characterized in that: The device comprises: An acquisition module is used to acquire an original data set to be enhanced, wherein the original data set includes data samples of multiple different types; An analysis module, used to perform data analysis on the data samples in the original data set to generate data enhancement information; An enhancement module, used to guide the enhancement process for different types of data samples according to the data enhancement information, and generate data enhancement samples of various types; An evaluation module is used to perform quality evaluation on the data enhancement samples and generate target enhancement data based on the quality evaluation results.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data enhancement method according to any one of claims 1 to 8 by executing the computer instructions.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data enhancement method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the data enhancement method according to any one of claims 1 to 8.