Text data enhancement method and electronic equipment

By dividing the text dataset into professional domains, representational diversity, and custom augmentation subsets, and optimizing with large models and scoring models, the problem of inconsistent generation quality and data bias in text data augmentation techniques is solved, achieving high-quality text data augmentation.

CN121145815APending Publication Date: 2025-12-16INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511410434.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing text data augmentation techniques suffer from inconsistent quality of the generated augmented text data, data bias, and insufficient coverage of low-frequency words, resulting in poor data augmentation effects.

Method used

The text dataset to be enhanced is divided into three subsets: domain-specific enhancement, expression diversity enhancement, and custom enhancement. Each subset is processed using a large model, and the generated text data is optimized through a scoring model to ensure semantic consistency and diversity.

Benefits of technology

It improves the semantic consistency and diversity of the generated augmented text data, avoids the limitations of human skills and professional background, ensures the professionalism, diversity and customization of the text data, solves the problems of data bias and insufficient coverage of low-frequency words, and enhances the data augmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121145815A_ABST
    Figure CN121145815A_ABST
Patent Text Reader

Abstract

The invention discloses a text data enhancement method and electronic equipment, and relates to the technical field of data processing, and the method comprises the steps: dividing a to-be-enhanced text data set into a plurality of sub-data sets based on the character number of a text data sample in the to-be-enhanced text data set; obtaining a first model used for text data sample enhancement and a second model used for enhancing a text data score; wherein the semantic consistency between the enhanced text data generated by the first model and the text data sample corresponding to the enhanced text data is greater than a preset semantic consistency threshold value; performing professional field enhancement on the first sub-data set by using the first model, performing expression diversity enhancement on the second sub-data set, and performing user-defined enhancement on the third sub-data set to obtain an enhanced sub-data set; and scoring the enhanced text data in the enhanced sub-dataset by using a second model to determine a target enhanced text dataset. The technical problem that the data enhancement effect is poor in the prior art is solved, and the technical effect of improving the data enhancement effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a text data enhancement method and an electronic device. BACKGROUND

[0002] With the rapid development of artificial intelligence, text data enhancement technology as a key means to enrich training data is becoming more and more important. The text data enhancement technology in the related art includes a rule-based method, a mask language model-based method and a back-translation method.

[0003] Among them, the rule-based method is limited by the skill, inspection level and professional background of artificial, resulting in uneven quality of the generated enhanced text data; the mask language model-based method is easy to amplify data bias, resulting in insufficient coverage of low-frequency words; and the back-translation method is easy to cause the generated enhanced text data to have style distortion or semantic deviation compared with the original text data. It can be seen that the data enhancement effect of the text data enhancement technology in the related art is poor. SUMMARY

[0004] The present application provides a text data enhancement method and an electronic device to at least solve the problem of poor data enhancement effect of the text data enhancement technology in the related art.

[0005] The present application provides a text data enhancement method, comprising: obtaining a text data set to be enhanced; dividing the text data set to be enhanced into a first sub-data set, a second sub-data set and a third sub-data set based on the number of characters of a text data sample in the text data set to be enhanced; obtaining a first model for text data sample enhancement and a second model for enhanced text data scoring; wherein the semantic consistency of the enhanced text data generated by the first model and the corresponding text data sample is greater than a preset semantic consistency threshold; performing professional field enhancement on the first sub-data set, expression diversity enhancement on the second sub-data set and self-defined enhancement on the third sub-data set by using the first model to obtain an enhanced sub-data set; scoring the enhanced text data in the enhanced sub-data set by using the second model; determining a target enhanced text data set based on the scoring result.

[0006] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above text data enhancement methods.

[0007] The application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program.

[0008] The application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the steps of any of the text data enhancement methods.

[0009] According to the application, the text data to be enhanced is obtained, the text data to be enhanced is divided into a first sub-data set, a second sub-data set and a third sub-data set based on the number of characters of the text data samples in the text data to be enhanced, a first model for text data sample enhancement and a second model for enhanced text data scoring are obtained, the semantic consistency of the enhanced text data generated by the first model with the corresponding text data sample is greater than a preset semantic consistency threshold, the first model is used to perform professional field enhancement on the first sub-data set, expression diversity enhancement on the second sub-data set and self-defined enhancement on the third sub-data set to obtain an enhanced sub-data set, the second model is used to score the enhanced text data in the enhanced sub-data set, and a target enhanced text data set is determined based on the scoring result. By using the first model to perform professional field enhancement on the first sub-data set, expression diversity enhancement on the second sub-data set and self-defined enhancement on the third sub-data set, the semantic consistency of the generated enhanced text data with the corresponding text data sample is high, the problem that the quality of the generated enhanced text data is uneven due to the skill, test level and professional background of artificial restriction is avoided, data bias is avoided, low-frequency words are sufficiently covered, the professionalism, expression diversity and self-defined nature of the text data are also taken into account, and therefore, the technical problem that the data enhancement effect of the text data enhancement technology in the related art is poor can be solved, and the technical effect of improving the data enhancement effect is achieved. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the application, the drawings required in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Figure 1 A structural schematic diagram of a text data enhancement system provided for the embodiments of the application; Figure 2 A flowchart of a text data enhancement method provided for the embodiments of the application; Figure 3 Another flowchart of a text data enhancement method provided for the embodiments of the application; Figure 4 A schematic diagram of the distribution histogram provided in an embodiment of this application; Figure 5 This is a schematic diagram of a sample interval in the distribution histogram provided in an embodiment of this application; Figure 6 A flowchart illustrating another text data enhancement method provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0013] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0014] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0015] With the rapid development of artificial intelligence, text data augmentation technology, as a key means of enriching training data, is becoming increasingly important. Text data augmentation refers to the process of transforming, expanding, or improving the original text in natural language processing tasks to generate new training samples, thereby enhancing the robustness and generalization ability of the model. Its main purpose is to reduce overfitting and improve the model's performance in practical applications by increasing the diversity of the dataset.

[0016] Text data augmentation techniques in related technologies include rule-based methods, masked language model-based methods, and reverse translation methods. Rule-based methods augment text data by altering its surface content, such as random insertion, random deletion, rearranging text order (e.g., changing words or sentences), or using regular expressions to replace keywords with synonyms. However, this method is limited by human skill, verification level, and professional background, resulting in inconsistent quality of the augmented text data. It is prone to generating invalid data or fixed-type data, thus limiting the effectiveness of the augmented text data.

[0017] The masked language model approach augments text by randomly masking a subset of words in the input text and using bidirectional contextual information to predict the masked words. This method maintains semantic consistency, enhances lexical diversity, and requires no additional labeled data. However, when the data sample distribution is uneven, it can amplify data bias, causing the augmented text to be dominated by the category with the larger sample size, resulting in insufficient coverage of low-frequency words. The masked language model can be a Bidirectional Encoder Representations from Transformers (BERT) model.

[0018] Reverse translation methods augment text data by translating the original text into an intermediate language and then back into the original language, generating new samples that are semantically consistent but express different meanings. While this method improves expressive diversity and semantic fidelity, the augmented text data still carries the risk of style distortion or semantic shift compared to the original text data, and its domain adaptability is poor.

[0019] It is evident that text data augmentation techniques in related technologies have poor data augmentation effects.

[0020] To address the aforementioned technical problems, embodiments of this application provide a text data augmentation method and an electronic device. The text data augmentation method includes: acquiring a text dataset to be augmented; dividing the text dataset into a first subset, a second subset, and a third subset based on the number of characters in the text data samples; acquiring a first model for text data sample augmentation and a second model for augmenting text data scoring; wherein the semantic consistency between the augmented text data generated by the first model and its corresponding text data samples is greater than a preset semantic consistency threshold; using the first model to perform domain-specific augmentation on the first subset, expressive diversity augmentation on the second subset, and custom augmentation on the third subset, thereby obtaining augmented subsets; using the second model to score the augmented text data in the augmented subsets; and determining a target augmented text dataset based on the scoring results. The method provided by the above solution enhances the first subset of text data with professional domain-specific enhancements using a first model, enhances the second subset with expressive diversity enhancements, and enhances the third subset with custom enhancements. This results in high semantic consistency between the generated enhanced text data and its corresponding text data samples, avoiding the problem of inconsistent quality of the generated enhanced text data due to limitations in human skills, testing levels, and professional backgrounds. It also avoids data bias, provides sufficient coverage of low-frequency words, and simultaneously considers the professionalism, expressive diversity, and customizability of the text data, enhancing it from multiple dimensions. Therefore, it can solve the technical problem of poor data enhancement effects in related text data enhancement techniques, achieving a significant improvement in data enhancement effectiveness.

[0021] The specific application environment architecture or specific hardware architecture on which the execution of text data augmentation methods depends is described here.

[0022] The text data enhancement method and electronic device provided in this application are applicable to enhancing text data. Figure 1The diagram shows the structure of the text data augmentation system upon which this application is based. The system includes a client and a text data augmentation device. The client sends the text dataset to be augmented to the text data augmentation device, which receives the dataset from the client and, based on the number of characters in the text data samples, divides the dataset into a first subset, a second subset, and a third subset. It then obtains a first model for text data sample augmentation and a second model for scoring the augmented text data. The semantic consistency between the augmented text data generated by the first model and its corresponding text data samples is greater than a preset semantic consistency threshold. The first model is used to augment the first subset using a specific domain, the second subset using expressive diversity enhancement, and the third subset using custom enhancement, resulting in augmented subsets. The second model is used to score the augmented text data in the augmented subsets. Based on the scoring results, a target augmented text dataset is determined.

[0023] Embodiments of this application provide a text data enhancement method. Figure 2 A flowchart illustrating the text data enhancement method provided in this application embodiment is shown below. Figure 2 As shown, this text data augmentation method includes the following steps: Step S201: Obtain the text dataset to be enhanced.

[0024] For example, in applications such as financial news sentiment analysis and market risk prediction, the text dataset to be augmented can be a massive amount of real-time financial news, financial information from social media, and financial research reports. The augmented text data can be used to identify potential market stress points and signs of systemic risk, providing early warnings for financial traders and assisting automated trading systems in making more informed decisions.

[0025] It should be noted that the text dataset to be enhanced can be collected and provided by users, and the dataset format is not limited. It can be a question-and-answer pair of prompts and responses, or it can be plain text.

[0026] Step S202: Based on the number of characters in the text data samples in the text dataset to be enhanced, divide the text dataset to be enhanced into a first subset, a second subset, and a third subset.

[0027] It is understood that the text dataset to be enhanced includes at least one text data sample. By analyzing the number of characters in the text data samples in the text dataset to be enhanced, the text dataset to be enhanced is divided into multiple subsets, namely the first subset, the second subset, and the third subset.

[0028] Step S203: Obtain a first model for text data sample enhancement and a second model for text data scoring enhancement; wherein the semantic consistency between the enhanced text data generated by the first model and its corresponding text data sample is greater than a preset semantic consistency threshold.

[0029] It should be noted that both the first and second models are large models. Large models typically refer to large-scale pre-trained models, which are deep learning models trained on massive amounts of data and have a large number of parameters.

[0030] The first model needs to balance generation quality, efficiency, and security. Specifically, the first model needs to meet the following requirements: Semantic fidelity: The semantic consistency between the enhanced text data generated by the first model and its corresponding text data sample is greater than a preset semantic consistency threshold. For example, the preset semantic consistency threshold can be 90%.

[0031] Diversity control: Supports adjusting the similarity between the augmented text data generated by the first model and its corresponding text data samples.

[0032] Domain adaptability: The terminology accuracy in the professional domain corresponding to the text dataset to be enhanced is greater than 85%.

[0033] Low-bias generation: The bias rate of gender and race-sensitive words in the enhanced text data generated by the first model is less than 5%.

[0034] The second model needs to balance scoring accuracy, interpretability, and robustness. Specifically, the second model needs to meet the following requirements: Quantitative scoring capability: Supports continuous / discrete value scoring.

[0035] Multimodal understanding: capable of handling mixed data including tables, text, and images.

[0036] Causal inference: Identifying causal relationships between variables.

[0037] Interpretability: Provides scoring criteria, such as SHAP scores. SHAP is a method for interpreting the predictions of machine learning models. It answers the question, "Why did the model make this judgment?" by calculating the contribution of each input feature to the final prediction.

[0038] It should be further noted that the first model is determined based on the application scenario of the text dataset to be augmented. For example, in general scenarios (such as customer service scenarios), the first model can be the LLaMA-2-7B model. In professional domain scenarios (such as medical / legal domain scenarios), the first model can be the GPT-3.5 model.

[0039] The second model is determined based on the application scenario of the text dataset to be augmented. For example, in a financial risk control scenario, the second model could be the ChatGLM3 model. In a medical diagnosis scenario, the second model could be the Med-PaLM2 professional large model.

[0040] Step S204: Use the first model to perform domain-specific enhancement on the first subset, enhance the representation diversity of the second subset, and perform custom enhancement on the third subset to obtain the enhanced subsets.

[0041] It should be noted that during the process of using the first model to perform domain-specific augmentation on the first subset, to enhance the expressive diversity of the second subset, and to perform custom augmentation on the third subset, the augmented subsets are obtained in real time. This means the augmented subsets are updated in real-time.

[0042] Step S205: Use the second model to score the enhanced text data in the enhanced subset.

[0043] The second model scores the enhanced text data in the enhanced subset in real time to obtain the score results.

[0044] Step S206: Based on the scoring results, determine the target augmented text dataset.

[0045] The text data augmentation method provided in this application enhances the first subset of text data by using a first model to enhance its professional domain, enhances the second subset of text data by enhancing its expressive diversity, and enhances the third subset of text data by customizing it. This results in high semantic consistency between the generated augmented text data and its corresponding text data samples, avoiding the problem of inconsistent quality of the generated augmented text data due to limitations in human skills, verification levels, and professional backgrounds. It also avoids data bias, provides sufficient coverage of low-frequency words, and simultaneously considers the professionalism, expressive diversity, and customizability of the text data. Therefore, it can solve the technical problem of poor data augmentation effect in related technologies and achieve the technical effect of improving data augmentation performance.

[0046] Embodiments of this application provide a text data enhancement method. Figure 3 A flowchart illustrating the text data enhancement method provided in this application embodiment is shown below. Figure 3 As shown, this text data augmentation method includes the following steps: Step S301: Obtain the text dataset to be enhanced. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0047] Step S302: Based on the number of characters in the text data samples in the text dataset to be enhanced, divide the text dataset to be enhanced into a first subset, a second subset, and a third subset.

[0048] Specifically, step S302 includes: Step S3021: Based on the number of characters in the text data samples in the text dataset to be enhanced, determine the maximum and minimum number of characters in the text data samples.

[0049] By analyzing the number of characters in the text data samples in the text dataset to be enhanced, the maximum and minimum number of characters corresponding to all text data samples in the text dataset to be enhanced were determined.

[0050] Step S3022: Determine the sample interval width of the text data sample based on the maximum and minimum number of characters in the text data sample.

[0051] Step S3023: Generate a distribution histogram of the text data samples based on the maximum number of characters, the minimum number of characters, and the sample interval width of the text data samples.

[0052] Understandably, after determining the sample interval width, the first sample interval is [minimum number of characters, minimum number of characters + sample interval width), and so on, generating a distribution histogram of text data samples.

[0053] For example, Figure 4 This is a schematic diagram of the distribution histogram provided in the embodiments of this application, such as... Figure 4 As shown, the histogram of the text data sample distribution includes 7 sample intervals, each with a different number of text data samples. The maximum number of characters is 526, the minimum number of characters is 198, and the total number of text data samples is 50.

[0054] Figure 5 This is a schematic diagram of a sample interval in the distribution histogram provided in the embodiments of this application, such as... Figure 5 As shown, the dashed box represents a sample interval in the distribution histogram, which is [245, 292). The number of text data samples in this sample interval is 12, and the proportion of text data samples in this sample interval is 24%.

[0055] Step S3024: Based on the distribution histogram of the text data samples, divide the text dataset to be enhanced into a first subset, a second subset, and a third subset.

[0056] Step S303: Obtain a first model for text data sample augmentation and a second model for enhancing text data scoring; wherein the semantic consistency between the augmented text data generated by the first model and its corresponding text data sample is greater than a preset semantic consistency threshold. For details, please refer to [link to relevant documentation]. Figure 2 Step S203 of the illustrated embodiment will not be described again here.

[0057] Step S304: Use the first model to perform domain-specific enhancement on the first subset, enhance the representation diversity of the second subset, and perform custom enhancement on the third subset to obtain the enhanced subsets.

[0058] Specifically, step S304 includes: S3041, Obtain a first prompt template for enhancing professional domains, a second prompt template for enhancing diversity of expression, and a third prompt template for custom enhancement.

[0059] The first and second prompt word templates are pre-set by technical personnel; the third prompt word template is a user-defined template, i.e., a template customized by the user according to business requirements.

[0060] For example, in the financial sector, the first prompt template could be: I. Task Objectives You are a seasoned financial expert responsible for enriching basic financial texts with professional depth. Based on the core information of the original text, please expand upon it from multiple professional dimensions, adding professional details, relevant concepts, and industry background to enhance the text's professional depth and richness.

[0061] II. Original Text [Insert the basic financial text that needs to be expanded here] III. Dimensions of Professional Field Expansion 1. Theoretical Framework Expansion Supplement relevant financial economics theories Add a background in econometrics or financial engineering Introduce relevant mathematical models or formulas for explanation. 2. Deepening of market mechanisms Expanding market microstructure details Supplementary Transaction Mechanism and Rules Increase market quality indicators such as liquidity and volatility 3. Risk Management Extension Add risk measurement and management framework Supplement the calculation logic of risk indicators such as VaR and ES Expanding stress testing and scenario analysis elements 4. Deepening Regulatory Compliance Supplementing relevant regulatory requirements and compliance framework Add specifications for Basel protocol, IFRS, and other standards. Expanding Compliance Risks and Controls 5. Refine the product structure Deepen the description of financial product features Supplementary pricing models and valuation methods Add description of cash flow structure and maturity IV. Requirements for Professional Expansion Depth requirement: Add 2-3 professional levels to each expansion dimension. Maintaining academic rigor and professional accuracy Citing mainstream financial theories and practices Breadth requirement: Covering at least 3 related professional subfields Includes horizontal comparison and vertical in-depth analysis Add international comparisons and industry best practices Quality requirements: All professional concepts must be accurately defined. Data computation and methodology need to be explained in detail. Maintain logical rigor and professional expression V. Constraints Must comply with: All extended content must be based on authoritative financial theories. The use of technical terms must comply with industry standards. Numerical computation requires a methodological explanation. Maintain objectivity and neutrality, and avoid subjective judgment. Prohibited items: False or unverified information must not be added. Financial regulatory provisions must not be violated. Unprofessional or vague expressions are not allowed. The original facts and data must not be altered. For example, the second prompt word template could be: I. Task Objectives You are a financial language expert tasked with expanding the expressive diversity of a given financial text. While keeping the core financial information and data absolutely unchanged, generate text variations with diverse expressions but consistent semantics through methods such as sentence structure transformation, synonym substitution, and perspective shifting.

[0062] II. Original Text [Insert the financial text to be expanded here] III. Expanding Dimensions and Requirements 1. Sentence structure diversity (generating 3-5 variations) Active and passive sentence conversion Sentence structure adjustment Word order changes between main and subordinate clauses Transformation of interrogative, declarative, and exclamatory sentences 2. Diversity of professional expression (generating 3-5 variations) Synonyms for technical terms (e.g., "stock price rises" → "stock price climbs" / "stock price increases") Adjustment of formality (academic → commercial → popular) Changes in data presentation (percentage ↔ absolute value ↔ multiple relationship) 3. Variety of perspectives and focal points (generating 2-3 variations) Investor perspective vs. analyst perspective vs. company perspective Macro perspective vs. micro perspective Risk-oriented vs. return-oriented statements 4. Diversity of emotional tone (generating 2-3 variations) Positive expression vs. cautious expression vs. neutral expression Emphasis level adjustment (mild ↔ moderate ↔ heavy emphasis) IV. Constraints Must be strictly observed: All key data and facts must be 100% accurate. The use of financial terminology must comply with industry standards. The core meaning and logical relationships of the original text must not be altered. Maintain professionalism and objectivity, and avoid excessive emotionalism. Flexible adjustment: Sentence structure and order of expression Synonyms and expressions The emphasis and hierarchy of information Text length and level of detail For example, the third prompt word template could be: 1. Roles and Tasks You are a professional financial data scientist with expertise in financial market microstructure, risk management, and econometrics. Your task is to generate high-quality, diverse synthetic data that reflects the statistical characteristics of the real financial world to enhance our existing datasets and address issues such as data scarcity, class imbalance, or model overfitting. The generated data must maintain plausibility, internal consistency, and statistical accuracy.

[0063] 2. Background and Context Original data description: [Describe your original data here].

[0064] Core issue: This dataset has [specific issues explained here].

[0065] The purpose of the enhancement is to train a model [e.g., a high-frequency trading volatility prediction model / a credit risk assessment machine learning model / a news sentiment analysis model] to improve its generalization ability under rare events and different market regimes.

[0066] 3. Instructions and Constraints (Data Generation Rules) a. Core requirements: Reasonableness: All generated data points must be reasonable in terms of financial logic.

[0067] Diversity: The generated data must cover multiple market scenarios, including: Different trends: strong upward, moderate upward, sideways movement, moderate decline, panic selling.

[0068] o Different volatility regimes: low volatility (calm market), high volatility (earnings release, macroeconomic data release), extreme volatility (market crisis, flash crash).

[0069] o Different asset types (if applicable): Generate data features for companies in different industries (such as technology, finance, energy) and with different market capitalizations.

[0070] Fidelity: The generated time series should retain the key statistical characteristics of real financial data, such as volatility clustering, leptokurtosis, autocorrelation, and the positive correlation between trading volume and volatility.

[0071] b. Specific constraints: Data format: The output must be in strict CSV or JSON format.

[0072] Time series continuity: If time series data is generated, the timestamps must be continuous and equally spaced (e.g., one data point per minute).

[0073] Event simulation: Randomly insert the following specific financial events into the generated data, and ensure that the data reflects the typical impact of the events: Earnings Announcement: Prices move significantly in one direction after the announcement, and trading volume increases dramatically.

[0074] Macroeconomic news (GDP, CPI, interest rate decisions): The entire sector or market moves in tandem, and volatility increases instantly.

[0075] Liquidity crisis: The bid-ask spread widens sharply, the price depth becomes shallower, and a "stampede" decline may occur.

[0076] Flash crash: A price drop that occurs within a very short period of time, followed by a rapid rebound.

[0077] Data volume: Please generate [e.g., 5000] data samples.

[0078] S3042, For any first text data sample in the first subset of the dataset, generate a first prompt word based on the first text data sample and the first prompt word template.

[0079] S3043, input the first prompt word into the first model to obtain the first enhanced text data corresponding to the first text data sample, so as to obtain the enhanced first subset of data.

[0080] S3044, For any second text data sample in the second subset of the dataset, generate a second prompt word based on the second text data sample and the second prompt word template.

[0081] S3045, input the second prompt word into the first model to obtain the second enhanced text data corresponding to the second text data sample, so as to obtain the enhanced second subset.

[0082] S3046, For any third text data sample in the third subset, generate a third prompt word based on the third text data sample and the third prompt word template.

[0083] S3047, input the third prompt word into the first model to obtain the third enhanced text data corresponding to the third text data sample, so as to obtain the enhanced third subset.

[0084] The augmented subsets include the augmented first subset, the augmented second subset, and the augmented third subset.

[0085] Step S305: Use the second model to score the enhanced text data in the enhanced subset.

[0086] The second model uses rating prompts to rate the augmented text data in the augmented subset. The rating rules in the rating prompts can be general or custom.

[0087] For example, rating prompts could be: I. Characters and Background You are a senior financial data analyst specializing in data quality assessment. You are now required to conduct a comprehensive quality assessment and score of a financial data sample. The assessment results will determine whether the data sample can be used for important financial modeling and decision analysis.

[0088] II. Evaluation of Financial Data Samples Data sample name: [Insert data sample name here] Data time range: [Start Time] to [End Time] Data categories: [such as market data, trading data, fundamental data, risk data, etc.] Intended uses: [such as risk modeling, investment decision-making, regulatory reporting, etc.] III. Evaluation Dimensions and Weights Please rate the data sample according to the following dimensions (each dimension is worth 10 points): 1. Completeness (weight 20%) o Check the proportion and distribution of missing values o Assessing the continuity of time series o Check the comprehensiveness of field coverage 2. Accuracy (weight 25%) cross-validation with authoritative data sources o Logical consistency checks (such as the relationship between price and trading volume) outlier detection and assessment 3. Timeliness (weight 15%) o Data update frequency o Data latency Historical data coverage depth 4. Consistency (weight 20%) o Internal consistency (cross-field logical relationships) o Time consistency (comparison of data from different periods) Cross-source consistency (compared to external data sources) 5. Availability (weight 10%) o Data document integrity o Data format standardization o Metadata richness 6. Compliance (weight 10%) o Data source legality privacy protection compliance o Regulatory compliance IV. Scoring Criteria 9-10 points: Excellent - Fully meets requirements, no improvement needed. 7-8 points: Good - Basically meets requirements, minor improvements needed. 5-6 points: Pass - requires moderate improvement 3-4 points: Insufficient - requires significant improvement 1-2 points: Critical defect - unsuitable for use Step S306: Based on the scoring results, determine the target augmented text dataset. See details below. Figure 2 Step S206 of the illustrated embodiment will not be described again here.

[0089] The text data augmentation method provided in this application generates a distribution histogram of text data samples based on the maximum number of characters, the minimum number of characters, and the sample interval width of the text data samples. This divides the text dataset to be augmented into a first subset, a second subset, and a third subset, preserving the distribution of long and short texts in the original dataset, avoiding sample bias, and thus avoiding data bias in the generated augmented text data, while ensuring sufficient coverage of low-frequency words.

[0090] The text data enhancement method provided in this application enhances text data from multiple dimensions by performing professional domain enhancement, expression diversity enhancement, and custom enhancement on three character datasets respectively, thereby improving the text data enhancement effect.

[0091] In some optional implementations, step S3022 above includes: Step a1: Based on the first formula, determine the sample interval width of the text data sample. The first formula is: BinWidth=(Max-Min) / (1+ ) Where BinWidth is the width of the sample interval of the text data sample, Max is the maximum number of characters in the text data sample, and Min is the minimum number of characters in the text data sample. The number of text data samples in the text dataset to be augmented.

[0092] The text data augmentation method provided in this application ensures the accuracy of sample interval division by determining the sample interval width of the text data sample based on a first formula, thereby ensuring the accuracy of subset division.

[0093] In some optional implementations, step S3024 above includes: Step b1: Based on the distribution histogram of the text data samples, sort the multiple sample intervals in the distribution histogram in descending order of the number of characters.

[0094] Step b2: Determine that the first subset of data includes the first text data sample located in the first sample interval. The first sample interval is the first preset number of sample intervals with the highest number of characters.

[0095] For example, the first preset quantity can be 3, and no specific limit is imposed here.

[0096] Step b3: Determine that the second subset of data includes the second text data samples located in the second sample interval. The second sample interval is the first preset number of sample intervals with the last character count.

[0097] Step b4: Determine that the third subset of data includes the third text data sample located in the third sample interval. The third sample interval is the other sample intervals in the distribution histogram besides the first and second sample intervals.

[0098] The text data augmentation method provided in this application sorts multiple sample intervals in the distribution histogram according to the number of characters from largest to smallest, thereby determining the first subset, the second subset, and the third subset, accurately realizing the differentiated division of the subsets and providing clear boundaries for subsequent targeted augmentation.

[0099] In some alternative implementations, step S306 includes: Step c1: Based on the scoring results of the first augmented text data in the augmented first subset, sort the first augmented text data in descending order of score.

[0100] Step c2 involves filtering the second preset number of first enhanced text data points that are ranked lower in the enhanced first subset of data to obtain the first enhanced text data subset.

[0101] The second preset quantity can be set by technical personnel, and no specific restrictions are imposed here.

[0102] Step c3: Perform diversity filtering and deduplication filtering on the first enhanced text data subset to obtain the second enhanced text data subset.

[0103] Diversity filtering refers to the process of analyzing the language structure, vocabulary distribution, sentence style, or multidimensional features of the generated enhanced text data to select a subset with significant differences in expression and a wider range of variation patterns, in order to ensure the expressive diversity of the enhanced text data.

[0104] Deduplication filtering refers to the process of identifying and removing redundant enhanced text data with highly repetitive or identical content by comparing the semantic similarity or string similarity between generated enhanced text data.

[0105] Step c4: Determine whether the second augmented text data subset satisfies the first derivation multiple. If it does, stop using the first model to augment the first subset of data in the professional domain and determine the second augmented text data subset as the first target augmented text data subset.

[0106] The first derivation factor is set by technical personnel and is not specifically limited here. Determining whether the second augmented text data subset satisfies the first derivation factor means determining whether the second augmented text data subset includes the first derivation factor of augmented text data corresponding to each first text data sample.

[0107] If the second augmented text data subset does not meet the first derivation multiple, then the first text data sample to be augmented in the first subset is determined, and the first model is used to perform professional domain augmentation on the first text data sample until the second augmented text data subset meets the first derivation multiple.

[0108] It should be noted that, for any first text data sample, if the number of enhanced text data corresponding to the first text data sample in the second enhanced text data subset does not reach the first derivative multiple, then the first text data sample is determined to be the first text data sample to be enhanced.

[0109] It is understandable that the number of first-target enhanced text data N1 included in the first-target enhanced text data subset is equal to the number of first text data samples M1 included in the first subset × the first derivation multiple X1.

[0110] Step c5: Based on the scoring results of the second enhanced text data in the enhanced second subset, determine the second target enhanced text data subset.

[0111] Step c6: Based on the scoring results of the third enhanced text data in the enhanced third subset, determine the third target enhanced text data subset.

[0112] The target-enhanced text dataset includes a first target-enhanced text data subset, a second target-enhanced text data subset, and a third target-enhanced text data subset.

[0113] The text data augmentation method provided in this application provides a method that filters low-quality data by scoring and sorting, and then removes redundant data by diversity filtering and deduplication filtering, ensuring that the retained augmented text data not only meets quality standards but also avoids duplication or homogenization, thereby improving the data augmentation effect.

[0114] In some alternative implementations, step c5 above includes: Step c51: Based on the scoring results of the second enhanced text data in the enhanced second subset, sort the second enhanced text data in descending order of score.

[0115] Step c52: Filter the second preset number of second enhanced text data that are ranked last in the enhanced second subset of data to obtain the third enhanced text data subset.

[0116] Step c53: Perform diversity filtering and deduplication filtering on the third enhanced text data subset to obtain the fourth enhanced text data subset.

[0117] Step c54: Determine whether the fourth augmented text data subset satisfies the second derivative multiple. If it does, stop using the first model to enhance the expressive diversity of the second subset and determine the fourth augmented text data subset as the second target augmented text data subset.

[0118] The second derivation multiple is set by technical personnel and is not specifically limited here. Determining whether the fourth augmented text data subset satisfies the second derivation multiple means determining whether the fourth augmented text data subset includes the second derivation multiple of augmented text data corresponding to each second text data sample.

[0119] If the fourth augmented text data subset does not meet the second derivation multiple, then the second text data sample to be augmented in the second subset is determined, and the first model is used to perform professional domain augmentation on the second text data sample until the fourth augmented text data subset meets the second derivation multiple.

[0120] It should be noted that, for any second text data sample, if the number of enhanced text data corresponding to the second text data sample in the fourth enhanced text data subset does not reach the second derivative multiple, then the second text data sample is determined to be the second text data sample to be enhanced.

[0121] It is understandable that the number of second-target enhanced text data included in the second-target enhanced text data subset N2 = the number of second text data samples included in the second subset M2 × the second derivation multiple X2.

[0122] In some alternative implementations, step c6 above includes: Step c61: Based on the scoring results of the third enhanced text data in the enhanced third subset, sort the third enhanced text data in descending order of score.

[0123] Step c62: Filter the second preset number of third enhanced text data that are ranked last in the enhanced third subset of data to obtain the fifth enhanced text data subset.

[0124] Step c63: Perform diversity filtering and deduplication filtering on the fifth enhanced text data subset to obtain the sixth enhanced text data subset.

[0125] Step c64: Determine whether the sixth enhanced text data subset satisfies the third derivative multiple. If it does, stop using the first model to perform custom enhancement on the third subset and determine the sixth enhanced text data subset as the third target enhanced text data subset.

[0126] The third derivative multiple is set by technical personnel and is not specifically limited here. Determining whether the sixth enhanced text data subset satisfies the third derivative multiple means determining whether the sixth enhanced text data subset includes the third derivative multiple of enhanced text data corresponding to each third text data sample.

[0127] If the sixth augmented text data subset does not meet the third derivation multiple, then the third text data sample to be augmented in the third subset is determined, and the first model is used to perform professional domain augmentation on the third text data sample until the sixth augmented text data subset meets the third derivation multiple.

[0128] It should be noted that, for any third text data sample, if the number of enhanced text data corresponding to the third text data sample in the sixth enhanced text data subset does not reach the third derivative multiple, then the third text data sample is determined to be the third text data sample to be enhanced.

[0129] It is understandable that the number of third-target enhanced text data included in the third-target enhanced text data subset N3 = the number of third text data samples included in the third subset M3 × the third derivation multiple X3.

[0130] In some optional implementations, before dividing the text dataset to be augmented into a first subset, a second subset, and a third subset based on the number of characters in the text data samples in the text dataset to be augmented, the above text data augmentation method further includes: Step d1 involves cleaning the text dataset to be enhanced. Data cleaning includes at least one of deduplication and privacy removal.

[0131] To ensure data quality for subsequent data augmentation, the text dataset to be augmented undergoes data cleaning before the augmentation process. Deduplication involves removing duplicate or blank content. Privacy removal removes private information present in the original data.

[0132] It should be noted that data cleaning can also include filtering, which means filtering text data samples with fewer than a preset threshold of characters, filtering noisy text data samples that are irrelevant to the application scenario, etc.

[0133] The text data augmentation method provided in this application improves the data augmentation effect by cleaning the dataset of text to be augmented.

[0134] In some alternative implementations, step S306 includes: Step e1: Based on the scoring results obtained by scoring the enhanced text data in the enhanced subset using the second model, determine the first target enhanced text dataset.

[0135] Step e2: Obtain the third model for enhancing the scoring of text data.

[0136] Step e3: Use the third model to score the enhanced text data in the enhanced subset. Based on the scoring results obtained by using the third model to score the enhanced text data in the enhanced subset, determine the second target enhanced text dataset.

[0137] Step e3, in response to the selection operation of the first target augmented text dataset and the second target augmented text dataset, determines the target augmented text dataset.

[0138] Users select from the first and second target augmented text datasets based on their own needs to determine the target augmented text dataset.

[0139] The text data augmentation method provided in this application uses a second model and a third model to score the augmented text data in the augmented subset, thereby obtaining a first target augmented text dataset and a second target augmented text dataset. This allows users to choose the target augmented text dataset according to their own needs, strengthening the user's control over the data and improving the user experience.

[0140] Embodiments of this application provide a text data enhancement method. Figure 6 A flowchart illustrating the text data enhancement method provided in this application embodiment is shown below. Figure 6 As shown, this text data augmentation method includes the following steps: The first step is to obtain the initial dataset. See the description of step S201 above for details, which will not be repeated here.

[0141] The second step is to clean the initial dataset. See step d1 above for details; it will not be repeated here.

[0142] The third step is to analyze the data distribution, i.e., generate a distribution histogram. See steps S3021 to S3023 above for details, which will not be repeated here.

[0143] The fourth step is to divide the dataset into three sub-datasets based on the data distribution. For details, please refer to the description of step S3024 above, which will not be repeated here.

[0144] The fifth step is to select the large model, that is, to obtain the first model for text data sample augmentation. The first model is used to perform domain-specific augmentation on the first subset, expression diversity augmentation on the second subset, and custom augmentation on the third subset, resulting in the augmented first subset, augmented second subset, and augmented third subset.

[0145] The sixth step involves using the second model to score the enhanced text data in the enhanced subset. See step S305 above for details, which will not be repeated here.

[0146] Step 7: Filter the results based on the scoring, including deduplication and diversity filtering. See steps c1 to c6 above for details; they will not be repeated here.

[0147] Step 8: Obtain the target augmented text dataset.

[0148] The text data augmentation method provided in this application determines three subsets with different sample proportions based on the data distribution of the initial dataset. Then, it performs domain-specific derivation, expression diversity derivation, and custom derivation on each subset to enhance and expand the text data. This method avoids the aimlessness of manually adjusting data augmentation methods and reduces manpower consumption. The augmented text data not only retains the sample distribution of the initial dataset but also takes into account domain-specificity, expression diversity, and business flexibility requirements. It solves the data imbalance problem and sample bias caused by long and short texts, ensuring semantic diversity and fidelity of the data from multiple dimensions.

[0149] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0150] Embodiments of this application also provide an electronic device, such as... Figure 7 As shown, it includes a processor 701 and a memory 702, in which a computer program is stored. The processor 701 is configured to run the computer program to perform the steps in any of the above-described text data enhancement method embodiments.

[0151] Embodiments of this application also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described text data enhancement method embodiments at runtime.

[0152] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0153] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described text data enhancement method embodiments.

[0154] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above-described text data enhancement method embodiments.

[0155] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0156] The foregoing has provided a detailed description of a text data enhancement method and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A text data augmentation method, characterized in that, include: Obtain the dataset of text to be enhanced; Based on the number of characters in the text data samples in the text dataset to be enhanced, the text dataset to be enhanced is divided into a first subset, a second subset, and a third subset; A first model for enhancing the text data samples and a second model for enhancing the text data scoring are obtained; wherein the semantic consistency between the enhanced text data generated by the first model and its corresponding text data samples is greater than a preset semantic consistency threshold. The first model is used to enhance the first subset of data by a specific domain, the second subset of data is used to enhance the diversity of expression, and the third subset of data is used to enhance the data in a custom way, so as to obtain the enhanced subset of data. The second model is used to score the enhanced text data in the enhanced subset; Based on the scoring results, the target augmented text dataset was determined.

2. The method according to claim 1, characterized in that, The text dataset to be enhanced is divided into a first subset, a second subset, and a third subset based on the number of characters in the text data samples, including: Based on the number of characters in the text data samples in the text dataset to be enhanced, determine the maximum and minimum number of characters in the text data samples; The sample interval width of the text data sample is determined based on the maximum and minimum number of characters in the text data sample. Based on the maximum number of characters, the minimum number of characters, and the width of the sample interval of the text data sample, a distribution histogram of the text data sample is generated; Based on the distribution histogram of the text data samples, the text dataset to be enhanced is divided into a first subset, a second subset, and a third subset.

3. The method according to claim 2, characterized in that, Determining the sample interval width of the text data sample based on the maximum and minimum number of characters in the text data sample includes: Based on the first formula, the sample interval width of the text data sample is determined. The first formula is: BinWidth=(Max-Min) / (1+ ) Wherein, BinWidth is the sample interval width of the text data sample, Max is the maximum number of characters in the text data sample, and Min is the minimum number of characters in the text data sample. The number of text data samples in the text dataset to be enhanced.

4. The method according to claim 2, characterized in that, The distribution histogram based on the text data samples divides the text dataset to be enhanced into a first subset, a second subset, and a third subset, including: Based on the distribution histogram of the text data samples, the multiple sample intervals in the distribution histogram are sorted in descending order of the number of characters. The first subset of data is determined to include a first text data sample located in a first sample interval, wherein the first sample interval is a first preset number of sample intervals sorted by the number of characters. The second subset of data is determined to include second text data samples located in the second sample interval, where the second sample interval is a first preset number of sample intervals sorted by the number of characters. The third subset of data is determined to include a third text data sample located in a third sample interval, wherein the third sample interval is the other sample intervals in the distribution histogram besides the first sample interval and the second sample interval.

5. The method according to claim 1, characterized in that, The process involves using the first model to perform domain-specific enhancement on the first subset of the dataset, enhancing the expressive diversity of the second subset of the dataset, and performing custom enhancement on the third subset of the dataset to obtain enhanced subsets, including: Get first cue word templates for enhancing professional domains, second cue word templates for enhancing diversity of expression, and third cue word templates for custom enhancement; For any first text data sample in the first subset of the dataset, a first prompt word is generated based on the first text data sample and the first prompt word template; Input the first prompt word into the first model to obtain the first enhanced text data corresponding to the first text data sample, so as to obtain the enhanced first subset of data. For any second text data sample in the second subset of the dataset, a second prompt word is generated based on the second text data sample and the second prompt word template; Input the second prompt word into the first model to obtain the second enhanced text data corresponding to the second text data sample, so as to obtain the enhanced second subset of data. For any third text data sample in the third subset of the dataset, a third prompt word is generated based on the third text data sample and the third prompt word template; Input the third prompt word into the first model to obtain the third enhanced text data corresponding to the third text data sample, so as to obtain the enhanced third subset of data. The enhanced subset includes an enhanced first subset, an enhanced second subset, and an enhanced third subset.

6. The method according to claim 1, characterized in that, The determination of the target augmented text dataset based on the scoring results includes: Based on the scoring results of the first enhanced text data in the enhanced first subset, the first enhanced text data is sorted in descending order of score; Filter the second preset number of first enhanced text data that are ranked last in the enhanced first subset of data to obtain the first enhanced text data subset. The first enhanced text data subset is subjected to diversity filtering and deduplication filtering to obtain the second enhanced text data subset; Determine whether the second enhanced text data subset satisfies the first derivative multiple. If it does, stop using the first model to enhance the first subset of data in the professional field and determine the second enhanced text data subset as the first target enhanced text data subset. Based on the scoring results of the second enhanced text data in the enhanced second subset, the second target enhanced text data subset is determined; Based on the scoring results of the third enhanced text data in the enhanced third subset, the third target enhanced text data subset is determined; The target augmented text dataset includes a first target augmented text data subset, a second target augmented text data subset, and a third target augmented text data subset.

7. The method according to claim 6, characterized in that, The determination of the second target enhanced text data subset based on the scoring results of the second enhanced text data in the enhanced second subset includes: Based on the scoring results of the second enhanced text data in the enhanced second subset, the second enhanced text data is sorted in descending order of score; Filter the second preset number of second enhanced text data that are ranked last in the enhanced second subset of data to obtain a third enhanced text data subset. The third enhanced text data subset is subjected to diversity filtering and deduplication filtering to obtain the fourth enhanced text data subset; Determine whether the fourth enhanced text data subset satisfies the second derivation multiple. If it does, stop using the first model to enhance the expressive diversity of the second subset and determine the fourth enhanced text data subset as the second target enhanced text data subset.

8. The method according to claim 6, characterized in that, The determination of the third target enhanced text data subset based on the scoring results of the third enhanced text data in the enhanced third subset includes: Based on the scoring results of the third enhanced text data in the enhanced third subset, the third enhanced text data is sorted in descending order of score; Filter the second preset number of third enhanced text data that are ranked last in the enhanced third subset of data to obtain the fifth enhanced text data subset. The fifth enhanced text data subset is subjected to diversity filtering and deduplication filtering to obtain the sixth enhanced text data subset; Determine whether the sixth enhanced text data subset satisfies the third derivative multiple. If it does, stop using the first model to perform custom enhancement on the third subset and determine the sixth enhanced text data subset as the third target enhanced text data subset.

9. The method according to claim 1, characterized in that, Before dividing the text dataset to be enhanced into a first subset, a second subset, and a third subset based on the number of characters in the text data samples in the dataset to be enhanced, the method further includes: The data cleaning process involves cleaning the text dataset to be enhanced, which includes at least one of deduplication and privacy removal.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the text data enhancement method as claimed in any one of claims 1 to 9.