Method and apparatus for sample generation, and device and storage medium

By introducing automated data generation methods into machine learning models and using evaluation criteria to build prompt word input, the problems of high cost of traditional data annotation process and data bias are solved, and high-quality and efficient data generation is achieved.

WO2025091983A1PCT designated stage expired Publication Date: 2025-05-08DOUYIN VISION CO LTD

Patent Information

Application Number
PCT/CN2024/102266
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-06-28
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The traditional data annotation process is expensive and time-consuming, and the generated data may be biased and cannot cover the comprehensive information required by the task.

Method used

Through an automated data generation method based on machine learning models, the existing data samples and evaluation criteria are used to construct prompt word input, and the model is guided to generate data samples that meet the evaluation criteria.

Benefits of technology

It realizes the quality and efficiency of data generation while reducing the burden of manual labeling, and ensures that the generated data samples meet the quality and reliability of the requirements of specific tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024102266_08052025_PF_FP_ABST
    Figure CN2024102266_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a method and apparatus for sample generation, and a device and a storage medium. The method comprises: determining at least one data sample, wherein the at least one data sample is classified into a first category; on the basis of feature information of the at least one data sample, generating a first assessment criterion for the first category; on the basis of the at least one data sample and the first assessment criterion, constructing a first prompt input, wherein the first prompt input is at least used for guiding a first machine learning model to generate a data sample which conforms to the first assessment criterion; and by means of providing the first prompt input to the first machine learning model, obtaining at least one further data sample which is output by the first machine learning model, wherein the at least one further data sample belongs to the first category.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for sample generation

[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, equipment and storage media for sample generation” and application number 2023114243506, filed on October 30, 2023, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for sample generation. Background Art

[0003] With advances in machine learning and deep learning technologies, machine learning models have been widely applied in many fields, including natural language processing, speech processing, and video / image processing. Model performance and output quality depend on the comprehensiveness, diversity, and accuracy of training samples. For models designed for specific tasks, it is desirable to obtain more qualified training samples at a lower cost for model training or adjustment.

[0004] Summary of the Invention

[0005] In a first aspect of the present disclosure, a method for sample generation is provided. The method includes: determining at least one data sample, the at least one data sample being classified into a first category; generating a first evaluation criterion for the first category based on feature information of the at least one data sample; constructing a first prompt word input based on the at least one data sample and the first evaluation criterion, the first prompt word input being used to at least guide a first machine learning model to generate a data sample that meets the first evaluation criterion; and obtaining at least one additional data sample output by the first machine learning model by providing the first prompt word input to the first machine learning model, the at least one additional data sample belonging to the first category.

[0006] In a second aspect of the present disclosure, a device for sample generation is provided. The device includes: a sample determination module, configured to determine at least one data sample, at least one data sample being classified into a first category; a criterion generation module, configured to generate a first evaluation criterion for the first category based on feature information of at least one data sample; a prompt word construction module, configured to construct a first prompt word input based on at least one data sample and the first evaluation criterion, the first prompt word input being at least used to guide a first machine learning model to generate a data sample that meets the first evaluation criterion; and an extended sample acquisition module, configured to obtain at least one additional data sample output by the first machine learning model by providing the first prompt word input to the first machine learning model, the at least one additional data sample belonging to the first category.

[0007] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0009] In a fifth aspect of the present disclosure, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions, which, when executed by a device, cause the device to perform the method of the first aspect.

[0010] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0012] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0013] FIG2 illustrates a block diagram of an example architecture for sample generation according to some embodiments of the present disclosure;

[0014] FIG3 shows a schematic diagram of a processing flow for sample generation according to some embodiments of the present disclosure;

[0015] FIG4 shows a flowchart of a process for sample generation according to some embodiments of the present disclosure;

[0016] FIG5 shows a block diagram of an apparatus for sample generation according to some embodiments of the present disclosure; and

[0017] FIG6 illustrates an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0020] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0021] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0022] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0025] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0026] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.

[0027] Generally speaking, machine learning can be roughly divided into three stages, namely the training stage, the testing stage and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application stage, the trained model can be used to process the actual model input based on the parameter values ​​obtained through training to determine the corresponding model output.

[0028] As mentioned earlier, in model-based task processing, the performance and output quality of the model depend on the comprehensiveness, diversity, and accuracy of the training samples.

[0029] FIG1 illustrates a schematic diagram of a model training and application environment 100 in which embodiments of the present disclosure can be implemented. The environment 100 in FIG1 illustrates three distinct stages of a model, including a pre-training stage 102, a fine-tuning stage 104, and an application stage 106. Although not shown in the figure, a testing / validation stage of the model may also occur after the pre-training or fine-tuning stage is complete.

[0030] In the pre-training phase 102, the model pre-training system 110 is configured to perform pre-training of the machine learning model 105 using a training dataset 112. The machine learning model 105 may be configured with a corresponding model structure based on the task to be processed.

[0031] At the beginning of pre-training, the machine learning model 105 may have initial parameter values. The pre-training process is to update the parameter values ​​of the machine learning model 105 to the desired values ​​based on the training data. During the pre-training process, one or more pre-training tasks 107-1, 107-2, etc. may be designed. The pre-training tasks are used to help update the parameters of the machine learning model 105. Some pre-training tasks may require connecting the machine learning model 105 to the output layer associated with the pre-training task.

[0032] During the pre-training phase 102, large-scale training data can be used to enable the machine learning model 105 to acquire strong generalization capabilities. After pre-training is complete, the parameters of the machine learning model 105 have been updated to the pre-trained values. The pre-trained machine learning model 105 can more accurately extract feature representations.

[0033] The pre-trained machine learning model 105 can be provided to the fine-tuning stage 104, where it is fine-tuned by the model fine-tuning system 120 for different downstream tasks. Downstream tasks can involve various visual tasks, such as text generation, image classification, object detection, semantic segmentation, etc. In some embodiments, depending on the specific downstream task, the pre-trained machine learning model 105 can be connected to the output layer 127 required by the downstream task to construct the downstream task model 125. This is because different downstream tasks may require different outputs.

[0034] In the fine-tuning phase 104, the training dataset 122 is further utilized to adjust the parameter values ​​of the machine learning model 105. If necessary, the parameters of the output layer 127 may also be adjusted. The machine learning model 105 may extract feature representations from the model input and provide them to the output layer 127 to provide outputs corresponding to the task.

[0035] During fine-tuning, the corresponding training algorithm is also used to update and adjust the parameters of the overall model. Since the machine learning model 105 has learned a lot of knowledge from the training data in the pre-training phase, a small amount of training data can be used in the fine-tuning phase 104 to obtain a model that meets the desired downstream task.

[0036] In the application phase 106, the obtained downstream task model 125 has trained parameter values ​​and can be provided to the model application system 130 for use. In the application phase 106, the downstream task model 125 can be used to process corresponding inputs in actual scenarios and provide corresponding outputs.

[0037] In Figure 1 , the model pre-training system 110 , model fine-tuning system 120 , and model application system 130 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may include any type of mobile, fixed, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframe computers, edge computing nodes, computing devices in cloud environments, and the like.

[0038] It should be understood that the components and arrangements of environment 100 shown in FIG1 are merely examples, and that a computing system suitable for implementing the exemplary implementations described herein may include one or more different components, other components, and / or different arrangements. For example, although shown as separate, model pre-training system 110, model fine-tuning system 120, and model application system 130 may be integrated into the same system or device, or distributed across a cloud computing environment. Implementations of the present disclosure are not limited in this respect.

[0039] In some embodiments, the training phase of the machine learning model 105 may not be divided into the pre-training phase and the fine-tuning phase shown in FIG1 , but a specific downstream task model may be directly configured and the model may be trained using a large amount of training data.

[0040] As can be seen from FIG1 , before obtaining a usable model, it is necessary to use training data (eg, training data sets 112 and 122 ) to complete model training.

[0041] Traditional data sample collection and labeling processes are costly and time-consuming. This is primarily because they typically require significant human effort to collect, process, and label data. The data labeling process can involve experts across multiple fields or individuals without relevant backgrounds, compromising sample quality and accuracy. Furthermore, the data labeling process may require multiple validations to ensure the accuracy of the labeled results. All of these factors make traditional data labeling difficult to adapt to the demands of large-scale, efficient data generation.

[0042] Currently, several generative models have been developed and are available for generating new data. Generative models can automatically generate meaningful and coherent content based on input content, including natural language text, images, audio, video, etc. Generative models can include language models to understand natural language input or other input from users. Generative models can also include other types of models to analyze and understand input content in other modalities (e.g., images, videos, audio).

[0043] During model-based data generation, the data output by the model may be limited by its training dataset, resulting in biased data or a failure to fully capture the information required for the task. Therefore, when using models to generate data, effective measures are needed to ensure the quality and reliability of the generated data, which is one of the current challenges.

[0044] Taking into account the problems of cumbersome traditional data annotation processes, high costs, and insufficient data accuracy, an improved sample generation solution is provided in an embodiment of the present disclosure. This solution utilizes the data generation capability of the model to automatically expand more data samples based on available data samples. In order to ensure the quality of the data samples output by the model, in the solution of the present disclosure, data quality evaluation criteria are generated based on the characteristics of the data samples under the corresponding category. A prompt word input is constructed based on the available data samples and the generated evaluation criteria, and the first prompt word input is at least used to guide the model to generate data samples that meet the first evaluation criteria. Then, based on the available data samples and the prompt word input, more data samples are generated with the help of the model.

[0045] This solution leverages the model's data generation capabilities and introduces evaluation criteria for data samples, enabling the automatic generation of large quantities of high-quality data samples with minimal or even no human intervention. This data can then be used for subsequent model training, fine-tuning, monitoring, and more. This automated and model-based approach can reduce the burden of manual sample annotation while improving the quality and efficiency of data generation.

[0046] Some example embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0047] Figure 2 shows a block diagram of an example architecture 200 for sample generation according to some embodiments of the present disclosure. In Figure 2, an electronic device 215 is configured to implement automatic augmentation of data samples.

[0048] The electronic device 215 receives a data sample set 210 to be expanded, which includes one or more data samples 212-1, 212-2, ... 212-N (collectively or individually referred to as data samples 212). The number N of data samples 212 can be an integer greater than or equal to 1. The data samples 212 can be considered as high-quality data samples suitable for model training. In embodiments of the present disclosure, it is possible to support the expansion of more data samples based on a small number of existing data samples. Therefore, there is no limit on the number of data samples 212.

[0049] In some embodiments, data samples 212 may be provided or input by user 202. For example, user 202 may obtain data samples labeled as meeting model training requirements through various methods. In other embodiments, data samples 212 may also be obtained through any other appropriate method. Data samples 212 may also sometimes be referred to as data samples to be augmented.

[0050] The electronic device 215 is further configured to expand more data samples with quality requirements based on the available one or more data samples 212. In particular, the electronic device 215 expands the data samples by category according to the one or more categories to which the data samples 212 belong.

[0051] Assume that the data samples 212 in the data sample set 210 include K categories (K is an integer greater than or equal to 1). The electronic device 215 generates an expanded data sample set 230-1 based on the data samples 212 belonging to the first category (category 1) in the data sample set 210, which includes one or more data samples 232-1, 232-2, ... 232-L (collectively or individually referred to as data samples 232). The electronic device 215 generates an expanded data sample set 230-2 based on the data samples 212 belonging to the second category (category 2) in the data sample set 210, which includes one or more data samples 234-1, 234-2 , ...234-K (collectively or individually referred to as data sample 234. Similarly, the electronic device 215 generates an expanded data sample set 230-M based on the data sample 212 belonging to the Mth category (category M) in the data sample set 210, which includes one or more data samples 236-1, 236-2, ...236-K (collectively or individually referred to as data sample 236. The expanded data sample sets 230-1, 230-2, ...230-M can be collectively or individually referred to as expanded data sample sets 230.

[0052] In an embodiment of the present disclosure, the expansion of data samples is achieved with the help of the data generation capability of the model. The electronic device 215 uses the machine learning model 220 to generate an expanded data sample set 230 based on the existing data sample 212. The machine learning model 220 can be any model with data generation capabilities, which can generate data in one or more modalities in response to model input, such as generating text, images, voice and other data. In some embodiments, the machine learning model 220 may include at least a language model to support the understanding of model input in the form of natural language. In this way, by conveniently inputting input in the form of natural language, the requirements for data generation can be described, and the model can generate corresponding content accordingly.

[0053] Generally, in the process of model-based data generation, the construction of prompts is one of the important steps to ensure the quality and reliability of the generated data. Prompts refer to the information used to interact with the model, with the purpose of guiding or triggering the model to produce corresponding responses or actions. Prompts, also known as prompt inputs, can be input into the model for use. For example, in a chatbot scenario built based on a generative model, the chatbot will analyze the messages input by the user and construct prompt inputs to guide the model to generate appropriate responses, such as answers, questions or suggestions. This depends on the training and design of the model. The prompt input can be in any type of text form, as long as it can effectively guide the model to generate the corresponding reply.

[0054] Prompt word engineering focuses on designing and optimizing prompt words to guide the model to generate the desired response. Currently, there are some implementation solutions for prompt word engineering.

[0055] However, existing prompt word generation methods cannot be directly used to meet the needs of data sample expansion, and cannot generate targeted data samples that meet data quality requirements.

[0056] In an embodiment of the present disclosure, by automatically analyzing the features of the data samples to be expanded that meet the task requirements and effectively utilizing the prompt word technology, the machine learning model 220 is guided to generate data samples that meet the data quality assessment criteria. Specifically, the electronic device 215 generates an assessment criterion for the category based on the feature information of at least one data sample 212 under a specific category. Then, a prompt word input is constructed based on at least one data sample 212 and the generated assessment criteria, and the prompt word input is at least used to guide the machine learning model 220 to generate data samples that meet the assessment criteria. By providing the model with the assessment criteria obtained based on the data features, the model can be guided to generate high-quality data samples that meet the requirements. Then, the generated prompt word is input and provided to the machine learning model 220 to obtain at least one additional data sample output by the machine learning model 220 (i.e., a data sample expanded based on the model). The data sample obtained at this time belongs to the category of the data sample 212 used. For each category of data sample 212, the corresponding data sample can be expanded according to the above process.

[0057] According to this solution, while using the model's automated data generation, data that meets specific task requirements can be generated, ensuring the quality of the generated data samples.

[0058] The data samples in one or more expanded data sample sets 230 (in combination with or without the data samples 232 under each category) can be provided for subsequent use. As shown in Figure 2, the data samples in one or more expanded data sample sets 230 can be provided to the electronic device 240 for training or fine-tuning the target machine learning model 250. The target machine learning model 250 can be any model that is expected to be trained. Generally, the type of data sample 212 can depend on the model input requirements of the target machine learning model 250 to which the data sample is to be applied. The types of the expanded data samples 232, 234, 236, etc. are consistent with the data samples 212.

[0059] In some embodiments, the training data required to train or fine-tune the target machine learning model 250 is in text mode, so data samples 212, data samples 232, 234, 236, etc. are all data samples in text mode. For example, the target machine learning model 250 can be a text generation model based on a language model. It is expected that the target machine learning model 250 will be trained or fine-tuned to be able to achieve text generation tasks in specific scenarios, industries, and fields. Accordingly, it is necessary to utilize a large number of text samples belonging to specific scenarios, industries, and fields to perform model updates.

[0060] In some embodiments, the training data required to train or fine-tune the target machine learning model 250 is in image modality, so data samples 212, data samples 232, 234, 236, etc. are all data samples in image modality. For example, the target machine learning model 250 can be an image generation model based on a language model and an image processing model. It is expected that the target machine learning model 250 will be trained or fine-tuned to be able to achieve image generation tasks within specific scenarios, industries, and fields. Accordingly, it is necessary to utilize a large number of image samples belonging to specific scenarios, industries, and fields to perform model updates.

[0061] Of course, the embodiments of the present disclosure do not limit the subsequent uses of the data samples obtained after expansion.

[0062] In the architecture 200 of FIG2 , electronic device 215 and / or electronic device 240 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Although shown as separate, electronic device 215 and electronic device 240 may be integrated into the same system or device, or distributed in a cloud computing environment. Implementations of the present disclosure are not limited in this respect.

[0063] The data sample generation process at the electronic device 215 will be described in detail below in conjunction with Figure 3. Figure 3 shows a schematic diagram of a processing flow 300 for sample generation according to some embodiments of the present disclosure. The processing flow 300 can be implemented by the electronic device 215, for example.

[0064] As shown in Figure 3, during the data classification 310 phase of processing flow 300, data classification is performed on data samples 212 in data sample set 210. Different categories of data samples allow for more targeted data sample expansion and subsequent quality assessment. This is because different model tasks may require different training data categories. By performing subsequent data generation based on category, it is ensured that the resulting data samples meet the task requirements.

[0065] During the data classification 310 stage, the electronic device 215 may classify the data samples 212 based on the feature information of each data sample 212 in the data sample set 210. In some embodiments, the electronic device 215 may extract features of each data sample 212 and cluster all data samples 212 based on the extracted features. Based on the clustering results, the data samples 212 in the data sample set 210 are divided into one or more categories. In this way, data samples 212 in each category can be obtained, including at least one data sample 212 in the first category, at least one data sample 212 in the second category, ..., and at least one data sample 212 in the Kth category. By using clustering rather than a fixed classification method, data features can be extracted more flexibly and the quality of data samples subsequently expanded by category can be effectively improved.

[0066] In some embodiments, the classification of the data sample 212 can be achieved with the help of a machine learning model 220. The machine learning model 220 can be deployed at the electronic device 215, or can be remotely deployed and callable by the electronic device 215. For each data sample 212 in the data sample set 210, the electronic device 215 can construct a second prompt word input based on the data sample 212, and the second prompt word input is at least used to guide the machine learning model 220 to determine the category of the data sample by analyzing the feature information of the data sample 212. By providing the second prompt word input to the machine learning model 220, the electronic device 215 can obtain a classification result output by the machine learning model 220, and the classification result indicates the category of the data sample. In this way, automatic classification can be performed based on the input data sample with the help of the model. By guiding the model to classify the data sample input by the user, data classification can be achieved more flexibly and accurately, and the subsequent high-quality data generation can be better guaranteed.

[0067] In some embodiments, in addition to the machine learning model 220, any other suitable classifier model can be used to perform data classification 310. A classifier model refers to a model that is specifically trained to perform data classification tasks. The selection of a classifier model can be based on the type of data sample 212. For example, a classifier model suitable for text classification or a classifier model suitable for image classification can be selected.

[0068] In some embodiments, data classification 310 may be performed optionally. For example, user 202 may provide data samples 212 to be expanded by category, and then all data samples 212 in data sample set 210 may be considered to belong to the same category. For another example, user 202 may specify or label the category of data samples 212 through other means.

[0069] After determining the data samples 212 under each category, in the stage of criterion generation 320, the electronic device 215 can generate an evaluation criterion for the category based on the feature information of at least one data sample 212 under the category. For example, for the first category (category 1), a first evaluation criterion for the category can be generated; for the second category (category 2), a second evaluation criterion for the category can be generated, and so on. In the embodiment of the present disclosure, the data samples 212 to be expanded are considered to be high-quality data samples that meet the requirements of the model training task. Therefore, by extracting the features of such data samples, the evaluation criteria for the data samples under the category can be analyzed and determined.

[0070] In some embodiments, if the number of data samples 212 under a category exceeds one, the electronic device 215 can extract feature information (also called key feature points) of each data sample 212 under the category and generate evaluation criteria for the category based on the extracted feature information.

[0071] In some embodiments, the criterion generation 320 stage can also be implemented with the aid of the machine learning model 220. For a specific category, the electronic device 215 can construct a third prompt word input based on at least one data sample 212 of that category. The third prompt word input is used to at least guide the machine learning model 220 to determine the evaluation criteria for that category by analyzing the feature information of the at least one data sample 212. The electronic device 215 can then provide the third prompt word input to the machine learning model 220 and obtain the evaluation criteria output by the machine learning model 220.

[0072] For ease of understanding, an example of generating evaluation criteria based on data samples is given below. Table 1 below shows a data sample (single data sample) under one category. Table 2 shows the output of the machine learning model 220, which indicates the evaluation criteria.

[0073] Table 1

[0074] Table 2

[0075] In some embodiments, in addition to the machine learning model 220, any other appropriate method may be used to generate the evaluation criteria for each category. For example, the evaluation criteria may be obtained by extracting key features of the data samples 212 under each category and aggregating the key features of the data samples 212 under the category.

[0076] After determining the evaluation criteria for each category, the electronic device 215 can perform prompt word construction 330 by category for sample generation. For a specific category, the electronic device 215 constructs a first prompt word input based on at least one data sample 212 under the category and the evaluation criteria corresponding to the category. The first prompt word input is at least used to guide the machine learning model 220 to generate data samples that meet the corresponding evaluation criteria. For example, the first prompt word model can guide the machine learning model 220 to perform text imitation or image generation based on the data sample 212 under the category, and requires that the generated text or image meet the corresponding evaluation criteria.

[0077] In some embodiments, the first prompt word input instructs the machine learning model 220 to use at least one data sample 212 as an example of data sample generation. That is, at least one data sample 212 under each category can be included in the first prompt word input to provide an example of expanding the sample of the machine learning model 220. If each category includes multiple data samples 212, one or more of the data samples 212 can be selected to be included in the prompt word input, or all of the multiple data samples 212 can be included in the prompt word input. In the prompt word input, different data samples 212 can be separated by specific symbols. In some embodiments, the first prompt word input can also indicate the category of the data sample provided. In some embodiments, the first prompt word input can also indicate the number of data samples output by the first machine learning model 220.

[0078] For ease of understanding, the following Table 3 gives an example of generating evaluation criteria based on data samples.

[0079] Table 3

[0080] In some embodiments, for each category, a plurality of different prompt word inputs may be constructed, and the forms of these prompt word inputs may be different from each other. The prompt word inputs constructed for each category may be stored in the prompt word library 332 .

[0081] After determining the evaluation criteria for each category and constructing the prompt word input, the electronic device 215 can perform category-based sample expansion 340. By providing the first prompt word input to the machine learning model 220, at least one additional data sample output by the machine learning model 220 is obtained, and the at least one additional data sample belongs to the category of the used data sample 212. For example, based on the data sample 212 under the first category, a data sample 232 of the first category (category 1) can be generated; based on the data sample 212 under the first category, a data sample 234 of the second category (category 2) can be generated; and so on.

[0082] Table 4 shows the data samples expanded based on the prompt word input in Table 3. Note that although a single data sample is given here, the machine learning model 220 can be required to generate more data samples as needed.

[0083] Table 4

[0084] In some embodiments, the quality of the generated data samples can also be evaluated. For the expanded data samples under each category, the electronic device 215 can evaluate the quality level of each expanded data sample based on the evaluation criteria corresponding to the category. In some embodiments, if the evaluation criteria indicate one or more aspects that a high-quality data sample must meet, the expanded data samples can be scored for each of these aspects, and the quality level of the data sample can be determined based on the score of each aspect.

[0085] In some embodiments, sample evaluation 350 may also be implemented using target machine learning 250. The electronic device 215 may generate a fourth prompt word input based on the evaluation criteria corresponding to each category, and the fourth prompt word input is used to guide the machine learning model 220 to evaluate the quality of the expanded data samples under the category according to the evaluation criteria. The electronic device 215 may provide the fourth prompt word input to the machine learning model 220. In some examples, the fourth prompt word input may also indicate a grade classification of the quality of the data sample, for example, it may indicate that each data sample is to be quality evaluated according to a predetermined quality score. The electronic device 215 may obtain the quality grade for each data sample output by the machine learning model 220.

[0086] Furthermore, the electronic device 215 can filter out data samples from each category whose quality level meets the quality requirements and provide them to the user. The quality requirement can indicate a quality level threshold, and data samples that exceed or equal to the quality threshold can be provided for subsequent use. This can further ensure the quality of the expanded data samples.

[0087] In some embodiments, the evaluation results of the data samples can also be used to filter out unqualified prompt word inputs from the prompt word library 332 generated by the sample, thereby achieving iterative optimization of the prompt words. The processing flow 300 can also include a sample evaluation 350 stage to evaluate the quality of the data samples generated in the sample expansion 340 stage.

[0088] During the sample evaluation 350 phase, the electronic device 215 may adjust the corresponding prompt input based on the quality level of at least one additional data sample. For example, if the quality levels of all or most of the data samples generated based on a prompt input do not meet the predetermined quality requirements, the electronic device 215 may determine that the prompt input requires further optimization. The electronic device 215 may adjust the prompt input in the prompt word library 332. The electronic device 215 may use the adjusted prompt input to perform sample expansion 340 again until the expanded data samples that meet the quality requirements meet the predetermined data generation target, such as, for example, the number of qualified data samples meets a predetermined threshold.

[0089] Note that although in the embodiments of Figures 2 and 3, a single machine learning model 220 is shown to be used to implement the processing of various stages such as data classification, criterion generation, sample expansion and sample evaluation, in other embodiments, the same or different models can be selected to implement data classification, criterion generation, sample expansion and sample evaluation in all stages or multiple stages.

[0090] According to the embodiments of the present disclosure, for the data samples to be expanded input by the user, evaluation criteria close to the data characteristics can be automatically generated from multiple aspects such as data format, task classification, and writing style, so that the data evaluation results are more reliable, thereby obtaining data samples with more reliable quality generated by the model. In this way, features can be extracted and data samples can be expanded based on a small number of data samples to be expanded that meet the task requirements, which can reduce the cost of manual labeling while ensuring the quality and reliability of the samples. In addition, in the process of generating data samples, data samples are expanded by category according to the task requirements, so that the data samples under each category can meet the task requirements.

[0091] FIG4 shows a flow chart of a process 400 for sample generation according to some embodiments of the present disclosure. The process 400 may be implemented at the electronic device 215 and / or the electronic device 240 of FIG2.

[0092] At block 410, electronic device 215 determines at least one data sample, which is classified into a first category. At block 420, electronic device 215 generates a first data quality assessment criterion based on the feature information of the at least one data sample. At block 430, electronic device 215 constructs a first prompt word input based on the at least one data sample and the first assessment criterion. The first prompt word input is used to at least guide the first machine learning model to generate a data sample that meets the first assessment criterion.

[0093] In box 440, the electronic device 215 obtains at least one additional data sample output by the first machine learning model by providing the first prompt word input to the first machine learning model, and the at least one additional data sample belongs to the first category.

[0094] In some embodiments, determining at least one data sample includes: obtaining a data sample set, the data sample set including multiple data samples; dividing each data sample in the data sample set into at least one category based on feature information of each data sample in the data sample set, the at least one category including a first category; and obtaining at least one data sample in the data sample set that is divided into the first category.

[0095] In some embodiments, classifying each data sample in the data sample set into at least one category includes: constructing a second prompt word input for the data sample in the data sample set based on the data sample, the second prompt word input being used to at least guide a second machine learning model to determine the category of the data sample by analyzing feature information of the data sample; and providing the second prompt word input to the second machine learning model to obtain a classification result output by the second machine learning model, the classification result indicating the category of the data sample. The second machine learning model can be the same as the first machine learning model, or different from the first machine learning model.

[0096] In some embodiments, generating a first evaluation criterion for a first category includes: constructing a third prompt word input based on at least one data sample, the third prompt word input being used to at least guide a third machine learning model to determine the first evaluation criterion for the first category by analyzing feature information of the at least one data sample; and providing the third prompt word input to the third machine learning model to obtain the first evaluation criterion output by the third machine learning model. The third machine learning model can be the same as the first machine learning model or the second machine learning model, or different from the first machine learning model and the second machine learning model.

[0097] In some embodiments, the first prompt word input further indicates at least one of the following: a first category, with at least one data sample as an example of data sample generation.

[0098] In some embodiments, process 400 further includes: determining a quality level of each of at least one additional data sample based on the first evaluation criterion; and providing a data sample of the at least one additional data sample whose quality level meets the quality requirement to the user.

[0099] In some embodiments, process 400 further includes adjusting the first prompt word input based on a respective quality level of at least one additional data sample.

[0100] In some embodiments, determining the quality level of at least one additional data sample includes: generating a fourth prompt word input based on the first evaluation criteria, the fourth prompt word input being used to guide a fourth machine learning model to evaluate the quality of the at least one additional data sample according to the first evaluation criteria; and obtaining the quality level of the at least one additional data sample output by the fourth machine learning model by providing the fourth prompt word input to the fourth machine learning model. The fourth machine learning model can be the same as the first machine learning model, the second machine learning model, or the third machine learning model, or different from the first machine learning model, the second machine learning model, and the third machine learning model.

[0101] In some embodiments, process 400 also includes: the electronic device 240 uses at least one additional data sample to train or fine-tune the target machine learning model.

[0102] In some embodiments, the at least one data sample and the at least one further data sample are data samples of a text modality or data samples of an image modality.

[0103] 5 shows a block diagram of an apparatus 500 for sample generation according to some embodiments of the present disclosure. The apparatus 500 may be implemented as or included in the electronic device 215 and / or the electronic device 240. Each module / component in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0104] As shown in the figure, the device 500 includes a sample determination module 510, which is configured to determine at least one data sample, and the at least one data sample is classified into a first category. The device 500 also includes a criterion generation module 520, which is configured to generate a first evaluation criterion for the first category based on feature information of the at least one data sample. The device 500 also includes a first prompt word construction module 530, which is configured to construct a first prompt word input based on at least one data sample and the first evaluation criterion, and the first prompt word input is at least used to guide the first machine learning model to generate a data sample that meets the first evaluation criterion. The device 500 also includes an extended sample acquisition module 540, which is configured to obtain at least one additional data sample output by the first machine learning model by providing the first prompt word input to the first machine learning model, and the at least one additional data sample belongs to the first category.

[0105] In some embodiments, the sample determination module 510 includes: a sample set acquisition module, configured to obtain a data sample set, the data sample set including multiple data samples; a category division module, configured to divide each data sample in the data sample set into at least one category based on feature information of each data sample in the data sample set, the at least one category including a first category; and a category sample acquisition module, configured to obtain at least one data sample in the data sample set that is divided into the first category.

[0106] In some embodiments, the category classification module includes: a second prompt word construction module, configured to construct a second prompt word input based on the data sample in the data sample set, the second prompt word input is at least used to guide the second machine learning model to determine the category of the data sample by analyzing the feature information of the data sample; and a classification result acquisition module, configured to obtain a classification result output by the second machine learning model by providing the second prompt word input to the second machine learning model, the classification result indicating the category of the data sample.

[0107] In some embodiments, the criterion generation module 520 includes: a third prompt word construction module, configured to construct a third prompt word input based on at least one data sample, the third prompt word input being at least used to guide the third machine learning model to determine a first evaluation criterion for the first category by analyzing feature information of at least one data sample; and a criterion acquisition module, configured to obtain the first evaluation criterion output by the third machine learning model by providing the third prompt word input to the third machine learning model.

[0108] In some embodiments, the first prompt word input further indicates at least one of the following: a first category, with at least one data sample as an example of data sample generation.

[0109] In some embodiments, the device 500 also includes: a quality determination module, configured to determine the quality level of each of at least one additional data sample based on the first evaluation criterion; and a sample providing module, configured to provide a data sample whose quality level meets the quality requirements in at least one additional data sample to the user.

[0110] In some embodiments, the apparatus 500 further includes a prompt word adjustment module configured to adjust the first prompt word input based on a quality level of each of the at least one additional data samples.

[0111] In some embodiments, the quality determination module includes: a fourth prompt word construction module, configured to generate a fourth prompt word input based on the first evaluation criteria, and the fourth prompt word input is used to guide the fourth machine learning model to evaluate the quality of at least one additional data sample according to the first evaluation criteria; and a quality acquisition module, configured to obtain the quality level of each of the at least one additional data samples output by the fourth machine learning model by providing the fourth prompt word input to the fourth machine learning model.

[0112] In some embodiments, the apparatus 500 further includes a sample utilization module configured to utilize at least one additional data sample to train or fine-tune the target machine learning model.

[0113] In some embodiments, the at least one data sample and the at least one further data sample are data samples of a text modality or data samples of an image modality.

[0114] FIG6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in FIG6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 600 shown in FIG6 can be used to implement the model pre-training system 110, the model fine-tuning system 120, and / or the model application system 130 of FIG1, the electronic device 215 and / or the electronic device 240 of FIG2. The electronic device 600 shown in FIG6 can be used to implement the apparatus 500 of FIG5.

[0115] As shown in FIG6 , electronic device 600 is in the form of a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 600.

[0116] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.

[0117] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG6 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0118] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0119] The input device 650 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 660 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 600, or with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0120] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0121] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0122] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0123] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0124] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0125] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for sample generation, comprising: determining at least one data sample, the at least one data sample being classified into a first category; Based on the characteristic information of the at least one data sample, generating a first evaluation criterion for data quality; Constructing a first prompt word input based on the at least one data sample and the first evaluation criterion, wherein the first prompt word input is at least used to guide the first machine learning model to generate a data sample that meets the first evaluation criterion; as well as By providing the first prompt word input to the first machine learning model, at least one additional data sample output by the first machine learning model is obtained, and the at least one additional data sample belongs to the first category.

2. The method of claim 1 , wherein determining at least one data sample comprises: Obtaining a data sample set, wherein the data sample set includes a plurality of data samples; Based on feature information of each data sample in the data sample set, classify each data sample in the data sample set into at least one category, wherein the at least one category includes the first category; as well as At least one data sample in the data sample set classified into the first category is obtained.

3. The method according to claim 2, wherein dividing each data sample in the data sample set into at least one category comprises: For the data samples in the data sample set, constructing a second prompt word input based on the data sample, wherein the second prompt word input is at least used to guide a second machine learning model to determine a category of the data sample by analyzing feature information of the data sample; as well as By providing the second prompt word input to the second machine learning model, a classification result output by the second machine learning model is obtained, wherein the classification result indicates the Describe the category of the data sample.

4. The method of claim 1 , wherein generating a first evaluation criterion for the first category comprises: constructing a third prompt word input based on the at least one data sample, wherein the third prompt word input is at least used to guide a third machine learning model to determine a first evaluation criterion of the first category by analyzing feature information of the at least one data sample; as well as By providing the third prompt word input to the third machine learning model, the first evaluation criterion output by the third machine learning model is obtained.

5. The method according to claim 1, wherein the first prompt word input further indicates at least one of the following: The first category, The at least one data sample is taken as an example of data sample generation.

6. The method according to claim 1, further comprising: determining a quality level of each of the at least one further data sample based on the first evaluation criterion; as well as The data sample whose quality level meets the quality requirement among the at least one other data sample is provided to the user.

7. The method according to claim 6, further comprising: The first cue word input is adjusted based on a respective quality level of the at least one additional data sample.

8. The method of claim 6, wherein determining the quality level of each of the at least one additional data samples comprises: generating a fourth prompt word input based on the first evaluation criterion, the fourth prompt word input being used to guide a fourth machine learning model to evaluate the quality of the at least one additional data sample according to the first evaluation criterion; as well as By providing the fourth prompt word input to the fourth machine learning model, a quality level of each of the at least one additional data samples output by the fourth machine learning model is obtained.

9. The method according to claim 1, further comprising: The target machine learning model is trained or fine-tuned using at least the at least one additional data sample.

10. The method of claim 1, wherein the at least one data sample and the at least one additional data sample are data samples in a text modality or data samples in an image modality.

11. A device for sample generation, comprising: a sample determination module, configured to determine at least one data sample, wherein the at least one data sample is classified into a first category; a criterion generating module, configured to generate a first evaluation criterion for the first category based on the feature information of the at least one data sample; A first prompt word construction module is configured to construct a first prompt word input based on the at least one data sample and the first evaluation criterion, wherein the first prompt word input is at least used to guide the first machine learning model to generate a data sample that meets the first evaluation criterion; as well as The extended sample acquisition module is configured to obtain at least one additional data sample output by the first machine learning model by providing the first prompt word input to the first machine learning model, and the at least one additional data sample belongs to the first category.

12. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by a processor.

14. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Machine translation model training method and device, machine translation method and device and computing equipment

    CN115130534A

  • Model training method and language model training method and device

    CN116401551A

  • Method, device and equipment for sample generation and storage medium

    CN117473316A

  • Training sample set generation from imbalanced data in view of user goals

    US20230136125A1

Cited By

  • Prompt information optimization method of generative model, quality detection method and medium

    CN120996110A