Training data generation method, equipment, device, medium and product
By pre-training the initial model and analyzing the output results, and dynamically screening the training data, the problem of strong subjectivity in training data generation is solved, and high-quality training data can be efficiently generated, thereby improving the performance and reliability of the model.
Patent Information
- Application Number
- CN202510652078.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-05
AI Technical Summary
In existing technologies, training data generation relies on manual labeling and expert review, which leads to high subjectivity of training data and makes it difficult to ensure the quality of generation.
By using the training data set to pre-train the initial model, a pre-trained model is generated, and the output results are generated using the pre-trained model. Based on the difference between the output results and the input instructions and the degree of correlation between the training data, the quality of the training data is determined, and a new training data set is dynamically screened and generated.
It achieves automatic quantification of training data quality, efficiently generates high-quality training data, and improves the generalization ability and reliability of the model.
Smart Images

Figure CN120597965A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, apparatus, medium, and product for generating training data. Background Art
[0002] With the rapid development of artificial intelligence (AI), the quality of training data plays a decisive role in the performance of machine learning models. High-quality training data can improve the model's generalization ability, reduce bias, and ensure its reliability in real-world application scenarios.
[0003] Currently, the generation of training data usually relies on manual labeling and expert review. The standards of different labelers may vary, resulting in high subjectivity in the training data and difficulty in ensuring the quality of the generated training data. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present application provides a training data generation method, equipment, device, medium and product.
[0005] According to a first aspect of any embodiment of the present application, a method for generating training data is provided, the method comprising:
[0006] Pre-train the initial model using the training data set to obtain a pre-trained model;
[0007] Using the pre-trained model, generating an output result for the input instruction;
[0008] In a case where the output result differs from the target result corresponding to the input instruction, determining the data quality of each training data based on the degree of association between the output result and each training data in the training data set;
[0009] The various training data are screened according to the data quality to generate a new training data set.
[0010] According to a second aspect of any embodiment of the present application, a training data generating apparatus is provided, the apparatus comprising:
[0011] A training module is used to pre-train the initial model using a training data set to obtain a pre-trained model;
[0012] A generation module, configured to generate an output result for an input instruction using the pre-trained model;
[0013] a determination module, configured to determine the data quality of each training data based on the degree of association between the output result and each training data in the training data set when there is a difference between the output result and the target result corresponding to the input instruction;
[0014] The screening module is used to screen the various training data according to the data quality to generate a new training data set.
[0015] According to a third aspect of any embodiment of the present application, an electronic device is provided, including:
[0016] processor;
[0017] a memory for storing processor-executable instructions;
[0018] The processor implements the method described in any embodiment of the present application by running the executable instructions.
[0019] According to a fourth aspect of any embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the method described in any embodiment of the present application is implemented.
[0020] According to a fifth aspect of any embodiment of the present application, a computer program product is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the method described in any embodiment of the present application is implemented.
[0021] The technical solution provided by this application may have the following beneficial effects:
[0022] According to the above embodiments, the initial model is pre-trained by using the training data set to obtain a pre-trained model, and the pre-trained model is used to generate an output result for the input instruction. When there is a difference between the output result and the target result corresponding to the input instruction, the data quality of each training data is determined based on the degree of correlation between the output result and each training data in the training data set. Each training data is screened according to the data quality to generate a new training data set. The output result is generated by the pre-trained model and its degree of correlation with the training data is evaluated. The quality of the training data can be automatically quantified, and a new training data set can be generated based on dynamic screening of the data quality, thereby efficiently generating high-quality training data.
[0023] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0025] Figure 1 is a flowchart of a method for generating training data according to an exemplary embodiment of the present application;
[0026] Figure 2 is a schematic diagram of a data screening interface according to an exemplary embodiment of the present application;
[0027] Figure 3 is a flowchart of another method for generating training data according to an exemplary embodiment of the present application;
[0028] Figure 4 is a structural diagram of an electronic device according to an exemplary embodiment of the present application;
[0029] Figure 5 It is a block diagram of a training data generating device according to an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0031] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0032] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0033] Currently, the generation of training data usually relies on manual labeling and expert review, which makes the training data highly subjective and makes it difficult to ensure the quality of the generated training data.
[0034] In order to solve the above problems, this application proposes a training data generation method. To further illustrate this application, the following embodiments are provided:
[0035] See also Figure 1 , Figure 1 This is a flowchart illustrating a training data generation method according to an exemplary embodiment of the present application. This training data generation method can be applied to a data training system. The data training system can be applied to electronic devices such as mobile phones, computers, digital broadcast terminals, messaging devices, tablet devices, personal digital assistants, and the like. It can also be applied to server-side devices such as single servers, cluster servers, and cloud servers. This generation method can also be executed by other systems or devices in different application scenarios, and this embodiment of the present application does not limit this.
[0036] like Figure 1 As shown, the training data generation method may include the following steps:
[0037] Step 101: Pre-train the initial model using the training data set to obtain a pre-trained model.
[0038] In this step, the data training system collects and organizes the training dataset for pre-training, ensuring the diversity and quality of the data in the training dataset to cover different scenarios and characteristics. This training dataset is then fed into the initial model for multiple rounds of iterative training, updating the initial model's parameters so that the initial model gradually learns the patterns and regularities in the data. Ultimately, a pre-trained model is obtained, preparing for subsequent processing.
[0039] The initial model is the original model used at the beginning of the model training process. It has a basic architecture and parameter settings and can be a large language model, a computer vision model, etc. The pretrained model is the model obtained by pretraining the initial model using the training dataset and has certain knowledge and feature understanding capabilities.
[0040] Step 102: Generate output results for input instructions using the pre-trained model.
[0041] In this step, the preset input instructions can be input into the pre-trained model. The pre-trained model performs calculations and reasoning based on the knowledge and patterns learned internally to generate output results related to the input instructions.
[0042] The input instruction is the guidance information or problem description provided to the pre-trained model, and may include text data, image data, and / or audio data. The output result is the response or answer to the input instruction, which can reflect the model's understanding and processing ability of the input instruction. The output data may include text data, image data, and / or audio data.
[0043] Step 103: When there is a difference between the output result and the target result corresponding to the input instruction, the data quality of each training data is determined based on the degree of association between the output result and each training data in the training data set.
[0044] In this step, you can pre-define metrics to measure the difference between the output result and the target result corresponding to the input instruction. If there is a difference between the output result and the target result, the degree of correlation between the output result and each training data in the training dataset is calculated by calculating text similarity, feature matching, etc.
[0045] The data quality of each training data set is determined based on the degree of correlation between the output result and each training data set. For example, the degree of correlation and data quality may be directly proportional. Training data with a low degree of correlation may cause the model to generate inaccurate results and thus have relatively low data quality. Training data with a high degree of correlation may help the model learn to generate accurate results and thus have relatively high data quality.
[0046] Among them, the target result refers to the preset output expected from the input instruction, which can be pre-defined or determined based on specific task requirements and domain knowledge, providing a benchmark for evaluating the accuracy of the output results of the pre-trained model.
[0047] The degree of association is the correlation or similarity between the output result and each training data in the training data set. It is used to evaluate the quality and relevance of the training data and can measure the contribution and impact of the training data on the output result.
[0048] By analyzing the degree of association, we can determine the training data that has a positive effect on the model generating accurate output results and the training data that interferes with or has a negative impact on the model. This helps to screen and optimize the training data set and improve the performance and generalization ability of the model.
[0049] Step 104: Screen each training data according to data quality to generate a new training data set.
[0050] In this step, based on the data quality assessment results, you can set screening criteria to distinguish between training data of different qualities. You can then categorize or sort the training data based on data quality, removing training data below the preset quality standard and retaining high-quality training data. This filtered training data is then integrated into a new training dataset for subsequent optimization and training of pre-trained models or other models.
[0051] By comparing the output results with the target results and evaluating and filtering the training data based on the degree of correlation, it helps to remove low-quality or irrelevant data and retain high-quality data, thereby improving the overall quality of the training dataset. High-quality training data can better reflect the actual situation and task requirements, providing a more reliable foundation for model training.
[0052] The training data generation method of this embodiment pre-trains an initial model using a training data set to obtain a pre-trained model, and uses the pre-trained model to generate an output result for an input instruction. When there is a difference between the output result and the target result corresponding to the input instruction, the data quality of each training data is determined based on the degree of correlation between the output result and each training data in the training data set. Each training data is screened according to the data quality to generate a new training data set. The output result is generated by the pre-trained model and its degree of correlation with the training data is evaluated. This can automatically quantify the quality of the training data, and dynamically screen and generate a new training data set based on the data quality, thereby efficiently generating high-quality training data.
[0053] In the aforementioned embodiments, we described how to pre-train an initial model using a training dataset to generate a pre-trained model, then use this model to generate an output result. When the output result differs from the target result, the quality of the training data is determined based on the degree of correlation, thereby selecting high-quality training data. The following embodiments provide a more detailed description of the training dataset construction process, which can be applied to any of the above embodiments.
[0054] In one embodiment, a multimodal data set from multiple data sources is displayed, and a training data set is generated in response to a filtering operation on multimodal data in the multimodal data set.
[0055] The multiple data sources may include at least one of the following: a business database, a distributed storage service, or a local file. A business database is used to store business-related data, such as a question-and-answer database, a customer service database, or a troubleshooting database. Business databases may include structured or unstructured data related to business processes, customer service, product development, and so on. A distributed storage service is a storage architecture that distributes data across multiple storage nodes and is used to store and manage large amounts of data.
[0056] Multimodal data sets are comprehensive datasets containing multiple data types (such as text, images, audio, and video). They provide models with more comprehensive and rich information, helping them understand and capture complex features and relationships within the data. Multimodal data can include database tables, fields, images, audio and video, text files, and Portable Document Format (PDF) files.
[0057] See also Figure 2 , Figure 2A schematic diagram of a data filtering interface is shown. For example, the system can display the data sources of business databases, distributed storage services, and imported local files in the data filtering interface 20. After the user selects a business database, the tables and fields in the business database can be displayed in the data filtering interface.
[0058] The user can select database table 1 from database tables 1, 2 and 3 in the data filtering interface, select fields 1, 2, 3, 4 and 5 in database table 1, and click the filter button 201. The system extracts the filtered multimodal data from the multimodal data set based on the user's filtering operation and generates a training data set.
[0059] The user can also click the import button 202 in the data filtering interface to import local files into the system. The system receives local files in formats such as Excel and TXT imported by the user. During the import process, the system can display the multimodal data set in the local file. In response to the multimodal data customized by the user, the system can perform pre-processing operations such as deduplication and format checking on the multimodal data, and use the pre-processed multimodal data as training data in the training dataset.
[0060] It is understandable that Figure 2 The data screening interface shown is only an example and is not limited to this embodiment of the present application.
[0061] As mentioned above, by visually displaying multi-source and multi-modal data and supporting interactive screening, users can intuitively select high-value data to generate training sets, avoiding invalid data input while retaining multimodal characteristics, providing a richer feature learning foundation for subsequent model training. In addition, it can improve the simplicity of cross-source data integration and significantly enhance the efficiency and flexibility of the data preparation stage.
[0062] In one embodiment, the data generation instruction can be input into the data generation model to obtain the risk content data output by the data generation model. The risk content data is used as a negative sample and added to the training data set.
[0063] Data generation instructions are commands or parameter settings used to guide a data generation model in generating data of a specific type or topic. These instructions can include key information such as specific topics, keywords, or context descriptions. A data generation model, such as a large language model, generates new data based on given instructions or conditions.
[0064] Risk content data can cover various possible risk scenarios and content, such as text containing sensitive words, images or sounds with potential risks, etc.
[0065] For example, a user can enter data generation instructions in the data training system interface. The system receives the instructions and inputs them into the data generation model. The data generation model then generates corresponding risk content data based on the instructions. This risk content data is then added to the training dataset as negative samples to improve the model's ability to identify and process risk content.
[0066] As mentioned above, by actively generating risk content data and adding it to the training dataset, potential violation scenarios can be proactively covered, thereby enhancing the model's robustness in identifying risky content.
[0067] In one embodiment, data generation templates corresponding to different risk dimensions may be displayed. In response to an edit operation on the data generation template, a data generation instruction is generated.
[0068] Data generation templates provide the format and content framework for data generation instructions. Risk dimensions and data generation templates can be customized based on model training requirements. Data generation templates can include specific prompts, placeholders, and structures to guide users or the system in generating data generation instructions that meet requirements. Standardized instruction formats can be provided, such as: "Please provide some {filler content} vocabulary," "Any additional sensitive words regarding {filler content}?", etc.
[0069] For example, the system can display multiple data generation templates related to different risk dimensions in the interface, such as templates related to risk dimensions such as bias and vulgarity. Each data generation template includes a specific format and prompts to guide the user to generate data generation instructions related to that risk dimension.
[0070] When users edit these data generation templates, such as filling in keywords, adjusting parameters, or modifying the template structure, the system can generate corresponding data generation instructions in real time based on the edited content by the user, and then use the data generation model to obtain more targeted risk content data, enriching the diversity and coverage of the training data set.
[0071] As mentioned above, by visually editing data generation templates for different risk dimensions, abstract risk control requirements are converted into actionable instructions, allowing users to customize negative data generation strategies to ensure the diversity of training data and the comprehensiveness of risk coverage.
[0072] In one embodiment, the training data set may be cleaned using data cleaning rules to obtain a cleaned training data set, and a cleaning comparison result between the training data sets before and after cleaning may be displayed.
[0073] Data cleaning rules are the rules and algorithms used to clean and preprocess training datasets, identifying and addressing noise, errors, duplicates, and other issues in the data. Cleaning comparison results are the results of a comparative analysis of the training dataset before and after cleaning, demonstrating changes and improvements in the training data during the cleaning process.
[0074] For example, the system can provide multiple data cleaning rules. Users can flexibly select one or more rules and apply them in combination based on the specific circumstances of the training dataset and the needs of model training to achieve in-depth cleaning of the training dataset. The system can use the user-selected rules to comprehensively scan and process the training dataset, identify and correct training data that does not comply with the rules, and ultimately obtain a cleaned training dataset, providing a higher-quality data foundation for subsequent model training.
[0075] The system can also support manual data cleaning. Users can directly observe the content of the training data in the interface and modify the training data through manual judgment. It is suitable for handling some complex or special data problems.
[0076] To visually demonstrate the effectiveness of data cleaning, the system displays a comparison of the training dataset before and after cleaning. This interface allows users to clearly see the differences between the training dataset before and after applying the data cleaning rules. For example, the unified data format, the elimination of invalid characters, and the reduction of duplicate data after cleaning help users evaluate the cleaning results, ensure the effectiveness of the cleaning operation, and gain a more intuitive understanding of the quality of the cleaned data.
[0077] Cleaning comparison results can also include evaluation metrics for the training dataset before and after cleaning. These metrics can help users understand the positive impact of data cleaning on the entire training dataset from a macro perspective. For example, by statistically analyzing the improvement in quality metrics such as completeness, consistency, and accuracy before and after cleaning, as well as the changes in the rationality of data distribution, users can comprehensively assess the effectiveness of the cleaning operation.
[0078] As mentioned above, through interactive data cleaning and cleaning comparison results display, users can instantly verify the effectiveness of cleaning rules, avoid information loss caused by incorrect cleaning, and realize the visualization and traceability of the data cleaning process.
[0079] In one embodiment, the data cleaning rules may include at least one of the following: a data format verification rule, an invalid character processing rule, and a data deduplication rule. The data format verification rule is used to check whether the data conforms to a predefined format standard.
[0080] The system can perform format validation on training data based on predefined format standards. For example, it can check whether the date format is "YYYY-MM-DD", whether the phone number conforms to a specific regional format, and whether the email address contains "@" and the domain name suffix.
[0081] Taking the JSON format as an example, if the training data is {"name":"Alice","age":30,"city":}, the system can recognize that the JSON object is incomplete because it lacks a value for the "city" field. In response, the system can employ a processing strategy: for example, automatically marking such noncompliant data to facilitate targeted corrections, such as completing field values; or automatically correcting it based on pre-set repair rules, such as referencing the format of other complete data items to fill in missing parts or remove incomplete fields, thus ensuring the integrity and consistency of the data structure.
[0082] Invalid character processing rules are used to identify and process invalid characters in training data, such as garbled characters, special control characters (such as form feed and backspace), empty strings, unknown characters, and meaningless repeated characters. The system can use advanced text processing technology to identify and process invalid characters during the cleaning process.
[0083] For example, special control characters such as page feed and backspace can be uniformly replaced with valid characters or directly removed; for garbled characters caused by encoding errors, such as The system can use encoding detection and conversion technology to first identify the original encoding and then convert it to the target encoding. If restoration is not possible, the garbled part can be replaced with a placeholder or directly removed.
[0084] By combining pattern matching and semantic analysis, character combinations without actual semantics can be identified, accurately located, and removed; by adjusting text alignment, redundant spaces and line breaks can be replaced with standard formats while retaining text readability.
[0085] For empty strings in training data, such as {"id":123,"content":null}, the system can directly delete records containing null values. This is suitable for situations where the data volume is large and the proportion of null value records is small, and can quickly improve data quality. It can also adopt data filling strategies, such as using default values, means, or reasonable values generated by predictive models for filling. This is suitable for scenarios where null values follow a certain pattern or deleting null values will significantly affect the data scale.
[0086] For unknown characters in training data, such as "This is an example containing some placeholders: [PLACEHOLDER]" or "TODO: Needs to be filled in," the system combines keyword matching and contextual semantic analysis to identify these placeholders or unfilled content. For placeholders with clear candidates, the system provides a drop-down menu or auto-complete suggestions. For parts that require manual judgment, the system can mark them and prompt the user to complete them.
[0087] Data deduplication rules are used to remove duplicate data items from training datasets to prevent the model from over-learning the same information during training, thereby improving the model's generalization ability. The system can identify duplicate data by calculating the hash value of the data item or using unique identifiers. Duplicate data is then processed according to a pre-defined deduplication strategy (such as retaining the first-appearing item or merging duplicates).
[0088] In addition to identifying and processing duplicate data items within the training dataset, data deduplication rules can further target duplicate content within a single training data set. For example, for redundant expressions in text data, the system uses natural language processing technology to accurately locate repeated words or phrases.
[0089] During the cleaning process, redundant and repetitive parts are removed based on pre-set semantic retention logic to ensure semantic integrity and concise expression. For example, the repeated "especially" in "I especially, especially like playing ball" is removed to obtain the training data for "I especially like playing ball." This not only reduces data redundancy but also improves data readability and quality, enabling the model to more efficiently learn core information during training and avoiding overfitting caused by redundant content.
[0090] For example, semantic preservation logic can use methods based on word frequency statistics and text similarity calculations to identify recurring words or phrases. After identifying duplicate content, semantic analysis is performed to retain one instance while removing all other duplicates. The system can perform grammatical and semantic checks on the cleaned training data to ensure that the cleaning operation has not altered the original meaning of the data. Users can intuitively see the changes in the text before and after cleaning, as well as the improvements in data quality and distribution, in the cleaning comparison results.
[0091] For training data containing irrelevant information, meaningless content, or confusing formatting, the system can also use text classification and information extraction technologies to separate core information from irrelevant content, retaining key components and removing distracting items. For private data containing user personal information and potentially violating privacy regulations, the system can employ privacy-preserving algorithms such as data desensitization and anonymization to ensure data availability while adhering to privacy regulations.
[0092] For biased training data, the system can apply fairness constraints and bias detection algorithms to identify and correct biased representations, ensuring the data's objectivity and neutrality. This training data can also be added to the training dataset as negative examples. When training data lacks key information or fields, the system can use data augmentation techniques, such as generating synthetic data or supplementing information from other data sources, to ensure that the training data is as complete as possible and meets specific training requirements.
[0093] As mentioned above, by combining the application of format verification, invalid character processing and data deduplication rules, a multi-level cleaning system is formed. Format verification ensures data structure compliance, invalid character processing eliminates noise data, and deduplication rules increase data density, which can effectively improve the cleaning efficiency and accuracy of training data.
[0094] In one embodiment, the target application domain of the initial model is obtained and the initial model is pre-trained using training data in the training dataset that matches the target application domain.
[0095] Among them, the target application field is the specific field or scenario where the initial model is expected to be applied, which can provide a clear direction and goal for model development and training, such as: autonomous driving field, financial field, general field, etc.
[0096] For example, the system can obtain the target application domain of the initial model through user input, system presets, or analysis of the model's expected application scenarios. Different application domains have different requirements for model training data. For example, models used in the field of autonomous driving require training data such as road image data and sensor data; models used in the field of finance require training data such as financial transaction data and market analysis data; and models used in general fields require training data covering knowledge and information from multiple fields.
[0097] The system can use a variety of methods, such as keyword matching, data label recognition, and data content analysis, to perform detailed classification and annotation of the training dataset, identifying training data within the training dataset that matches the target application domain. For example, if the target application domain is autonomous driving, data related to road images and traffic sign recognition can be selected from the training dataset.
[0098] Pre-training begins by feeding the initial model with selected training data that matches the target application domain. During pre-training, the model learns the features, patterns, and relationships within this data, gradually adjusting its parameters to suit the specific characteristics of the target application domain. Pre-training can employ a variety of algorithms and strategies, including unsupervised and supervised learning, depending on the model type and data characteristics.
[0099] For example, in the field of autonomous driving, supervised learning can be used to feed labeled road image data into an initial model. The initial model then learns features such as traffic signs and lane markings in the images and their corresponding labels, building its ability to understand and recognize various elements in autonomous driving scenarios. As training progresses, the model's performance in the target application domain will gradually improve, laying a solid foundation for subsequent practical applications.
[0100] As mentioned above, by using the training data in the training dataset that matches the target application field, the initial model of the target application field is pre-trained, so that the model can more accurately understand and process data in subsequent tasks targeting the target application field, thereby improving the model's professional performance and reducing training resource consumption.
[0101] To further illustrate the training data generation process, Figure 3 A flow chart of another method for generating training data is shown. The method for generating training data may include the following steps:
[0102] Step 301: Display multimodal data sets from multiple data sources.
[0103] In this step, the data training system can display the table names and field names in the business database in a tabular form in the interface, display the data in the distributed storage service in a hierarchical structure of folders and files, and display the file names and file types of local files.
[0104] Step 302: In response to a screening operation on multimodal data in a multimodal data set, a training data set is generated.
[0105] In this step, users can set filtering criteria using filter components within the interface, such as drop-down menus, checkboxes, and text input boxes. Based on the received filtering criteria, the system extracts matching multimodal data from the multimodal data set to generate a training dataset. Users can also annotate and evaluate the training data within the training dataset within the interface.
[0106] Step 303: In response to the editing operation on the data generation template, generate a data generation instruction.
[0107] In this step, the system can display multiple data generation templates for different risk dimensions on the interface. Users edit the data generation templates on the interface. After the user completes the editing, the system analyzes the user's input content and parameters in real time and generates data generation instructions.
[0108] Step 304: Input the data generation instruction into the data generation model to obtain the risk content data output by the data generation model.
[0109] In this step, the system sends the user-confirmed data generation instructions to the data generation model via an API or other communication protocol. The data generation model then performs data generation processing based on the received instructions. For example, using the large language model, the model uses the keywords and topic descriptions in the instructions to invoke its internal language generation algorithm to generate risk content data containing sensitive terms and risk scenario descriptions.
[0110] Step 305: Add the risky content data as negative samples to the training dataset.
[0111] In this step, specific labels can be added to the risk content data generated by the model and manually added, marking them as negative samples and adding them to the training dataset.
[0112] Step 306: Clean the training data set using data cleaning rules to obtain a cleaned training data set.
[0113] In this step, the system provides a visual interface displaying various data cleaning rules, including data format verification, invalid character handling, and data deduplication. Users select appropriate data cleaning rules based on the characteristics and quality requirements of the training dataset. The system then cleans the training dataset according to the selected data cleaning rules and displays a comparison of the cleansing results before and after the cleansing.
[0114] Step 307: Use the training data in the training dataset that matches the target application domain to pre-train the initial model.
[0115] In this step, the system determines the target application domain of the initial model and obtains training data from the training dataset that matches the target application domain. The matching training data is imported into the disk address extracted for initial model training, and the initial model is pre-trained to obtain a pre-trained model.
[0116] Step 308: Input the input instruction to the pre-trained model to obtain the output result.
[0117] In this step, the system inputs the preset input instructions into the pre-trained model to obtain the output results.
[0118] Step 309: When there is a difference between the output result and the target result corresponding to the input instruction, the data quality of each training data is determined based on the degree of association between the output result and each training data in the training data set.
[0119] In this step, the system predefines various metrics to measure the difference between the output result and the target result, such as precision, recall, and F1 score for classification tasks; mean squared error and mean absolute error for regression tasks. The output result can be compared with the target result to calculate the difference metric. If the output result differs from the target result corresponding to the input instruction, the data quality of each training data point is determined based on the degree of correlation between the output result and the training data point.
[0120] Step 310: Screen each training data according to data quality to generate a new training data set.
[0121] In this step, the system sets filtering criteria based on data quality, such as retaining only training data with a data quality score higher than 80 (out of 100). The system classifies or sorts each training data in the training dataset, arranging it from high to low according to its data quality score, removes training data with a score below the screening score threshold, and retains training data with a score higher than or equal to the score threshold to generate a new training dataset.
[0122] The system can display the training data in the new training data set in the interface, repeat steps 303-310, iteratively optimize the training data set, improve the performance and accuracy of the model, and make it better meet the actual application needs.
[0123] Figure 4 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present application. The electronic device may be, for example, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a personal digital assistant, a server, a smart home appliance, a car computer, etc. Figure 4 At the hardware level, the electronic device includes a processor 401, an internal bus 402, a network interface 403, a memory 404, and a non-volatile memory 405. Of course, it may also include hardware required for other services. The processor 401 reads the corresponding computer program from the non-volatile memory 405 into the memory 404 and then runs it, forming a training data generation device at the logical level. Of course, in addition to software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0124] Figure 5 This is a block diagram of a training data generation device according to an exemplary embodiment of the present application. Figure 5 The apparatus may include: a training module 501, a generation module 502, a determination module 503 and a screening module 504, wherein:
[0125] The training module 501 is used to pre-train the initial model using the training data set to obtain a pre-trained model;
[0126] The generating module 502 is configured to generate an output result for an input instruction using the pre-trained model;
[0127] The determining module 503 is configured to determine the data quality of each training data based on the degree of association between the output result and each training data in the training data set when there is a difference between the output result and the target result corresponding to the input instruction;
[0128] The screening module 504 is configured to screen the various training data according to the data quality to generate a new training data set.
[0129] In one example, the training module 501, before being used to pre-train the initial model using the training data set to obtain the pre-trained model, further includes: displaying a multimodal data set from multiple data sources, the multiple data sources including at least one of the following: a business database, a distributed storage service, and a local file; generating the training data set in response to a screening operation on the multimodal data in the multimodal data set.
[0130] In one example, the training module 501, before being used to pre-train the initial model using the training data set to obtain the pre-trained model, also includes: inputting data generation instructions into the data generation model to obtain risk content data output by the data generation model; and adding the risk content data as a negative sample to the training data set.
[0131] In one example, the training module 501, before being used to input data generation instructions into a data generation model to obtain risk content data output by the data generation model, further includes: displaying data generation templates corresponding to different risk dimensions; and generating the data generation instructions in response to an editing operation on the data generation template.
[0132] In one example, the screening module 504 is further configured to clean the training data set using data cleaning rules to obtain a cleaned training data set; and display a cleaning comparison result between the training data sets before and after cleaning.
[0133] In one example, the data cleaning rules include at least one of the following: a data format verification rule, an invalid character processing rule, and a data deduplication rule.
[0134] In one example, the training module 501, when used to pre-train the initial model using a training data set, includes: obtaining the target application field of the initial model; and pre-training the initial model using training data in the training data set that matches the target application field.
[0135] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0136] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application scheme. Those of ordinary skill in the art can understand and implement it without paying any creative work.
[0137] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory including instructions. The above instructions can be executed by a processor of a training data generating device to implement any method as described in the above embodiments.
[0138] The non-temporary computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc., and this application does not limit this.
[0139] In an exemplary embodiment, a computer program product including a computer program / instruction is further provided. The computer program / instruction can be executed by a processor of a training data generating device to implement any of the methods described in the above embodiments.
[0140] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0141] Other embodiments of the present invention will readily occur to those skilled in the art after consideration of the specification and practice of the invention claimed herein. The present application is not limited to the precise construction described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present application is limited solely by the appended claims.
[0142] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A training data generation method, characterized in that: The method comprises: Pre-train the initial model using the training data set to obtain a pre-trained model; Using the pre-trained model, generating an output result for the input instruction; In a case where the output result differs from the target result corresponding to the input instruction, determining the data quality of each training data based on the degree of association between the output result and each training data in the training data set; The various training data are screened according to the data quality to generate a new training data set.
2. The method according to claim 1, characterized in that Before pre-training the initial model using the training data set to obtain the pre-trained model, the method further includes: Displaying a multimodal data set from multiple data sources, wherein the multiple data sources include at least one of the following: a business database, a distributed storage service, and a local file; The training data set is generated in response to a screening operation on multimodal data in a multimodal data set.
3. The method according to claim 1, characterized in that Before pre-training the initial model using the training data set to obtain the pre-trained model, the method further includes: Inputting the data generation instruction into the data generation model to obtain the risk content data output by the data generation model; The risky content data is used as a negative sample and added to the training data set.
4. The method according to claim 3, characterized in that Before inputting the data generation instruction into the data generation model to obtain the risk content data output by the data generation model, the method further includes: Display data generation templates corresponding to different risk dimensions; The data generation instruction is generated in response to an editing operation on the data generation template.
5. The method according to claim 1, wherein The method further comprises: Cleaning the training data set using a data cleaning rule to obtain a cleaned training data set; Shows the cleaning comparison results between the training dataset before and after cleaning.
6. The method according to claim 5, characterized in that The data cleaning rules include at least one of the following: data format verification rules, invalid character processing rules, and data deduplication rules.
7. The method according to claim 1, characterized in that The pre-training of the initial model using the training data set includes: Obtaining a target application field of the initial model; The initial model is pre-trained using training data in the training data set that matches the target application field.
8. A training data generating device, characterized in that: The device comprises: A training module is used to pre-train the initial model using a training data set to obtain a pre-trained model; A generation module, configured to generate an output result for an input instruction using the pre-trained model; a determination module, configured to determine the data quality of each training data based on the degree of association between the output result and each training data in the training data set when there is a difference between the output result and the target result corresponding to the input instruction; The screening module is used to screen the various training data according to the data quality to generate a new training data set.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 7 by running the executable instructions.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product having a computer program / instructions stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.