Data evaluation using reinforcement learning
By dynamically selecting a subset of training samples and updating model parameters using a reinforcement learning framework, the problem of high training costs for large-scale datasets is solved, and efficient data value evaluation and model performance improvement are achieved.
Patent Information
- Application Number
- CN202080065876.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2020-09-19
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2040-09-19
AI Technical Summary
Existing technologies face significant challenges in training machine learning models due to the high cost and difficulty in collecting large-scale, high-quality datasets, as well as the difficulty in effectively quantifying data values, which limits the improvement of model performance.
By employing a reinforcement learning (DVRL) framework, a subset of training samples is dynamically selected through a data value estimator and a predictor model, and the model parameters are updated using reinforcement signals, thereby achieving efficient data value evaluation.
It improves the computational efficiency and quality ranking of the dataset, saves time, optimizes the model training process, and enhances model performance.
Smart Images

Figure CN114424204B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to data valuation using reinforcement learning. BACKGROUND
[0002] Machine learning models receive inputs and generate outputs, e.g., prediction outputs, based on the received inputs. Machine learning models are trained from data. However, quantifying the value of data is a fundamental problem in machine learning. Machine learning models generally improve when trained from large-scale and high-quality datasets. However, collecting such large-scale and high-quality datasets can be expensive and challenging. Moreover, there is additional complexity in determining the samples in a large-scale dataset that are most useful for training and labeling. Real-world training datasets often contain incorrect labels, or input samples differ in relevance, sample quality, or usefulness for the target task.
[0003] Accurately quantifying the value of data improves model performance of training datasets. Instead of treating all data samples equally, data can be assigned lower priority when its value is low to obtain a higher performance model. In general, quantifying data valuation performance requires removing samples individually to compute the performance loss, which is then assigned as the data of that sample. However, these methods scale linearly with the number of training samples, making them prohibitively expensive for large-scale datasets and complex models. Besides establishing insights about the problem, data valuation has different use cases such as in domain adaptation, corrupted sample discovery, and robust learning. SUMMARY
[0004] One aspect of the present disclosure provides a method for evaluating training samples. The method includes obtaining, at data processing hardware, a set of training samples. During each of a plurality of training iterations, the method further includes sampling, by the data processing hardware, a batch of training samples from the set of training samples. The method includes, for each training sample in the batch of training samples, determining, by the data processing hardware, a selection probability using a data value estimator. The selection probability of the training sample is based on estimator parameter values of the data value estimator. The method further includes selecting, by the data processing hardware, a subset of training samples from the batch of training samples based on the selection probability of each training sample, and determining, by the data processing hardware, a performance measure using a predictor model with the subset of training samples. The method further includes adjusting, by the data processing hardware, model parameter values of the predictor model based on the performance measure, and updating, by the data processing hardware, the estimator parameter values of the data value estimator based on the performance measure.
[0005] Implementations of the present disclosure can include one or more of the following optional features. In some implementations, determining the performance measure using the predictor model includes determining loss data by a loss function. In these implementations, adjusting the model parameter values of the predictor model based on the performance measure includes adjusting the model parameter values of the predictor model based on the loss data. Additionally, in some implementations, updating the estimator parameter values of the data value estimator based on the performance measure includes determining an augmentation signal from the loss data and updating the estimator parameter values of the data value estimator based on the augmentation signal. Updating the estimator parameter values of the data value estimator based on the augmentation signal further includes determining a reward value based on the loss data and updating the estimator parameter values of the data value estimator based on the reward value. In these implementations, determining the reward value based on the loss data includes determining a moving average of loss data based on the last N training iterations of the predictor model, determining a difference between the loss data of the last training iteration and the moving average of the loss data, and determining the reward value based on the difference between the loss data of the last training iteration and the moving average of the loss data.
[0006] In some examples, the data value estimator includes a neural network, and updating the estimator parameter values of the data value estimator includes updating layer parameter values of the neural network of the data value estimator. In some examples, the predictor model is trained using stochastic gradient descent. In some implementations, selecting the subset of training samples from the batch of training samples based on the selection probability for each training sample includes, for each training sample in the batch of training samples, determining a corresponding selection value indicating selection or non-selection. When the corresponding selection value indicates selection, the method includes adding the training sample to the subset of training samples, and when the corresponding selection value indicates non-selection, the method further includes discarding the training sample. In some examples, sampling the batch of training samples includes, for each training iteration in the plurality of training iterations, sampling a different batch of training samples from the set of training samples.
[0007] Another aspect of the present disclosure provides a system for evaluating training samples. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations including obtaining a set of training samples. During each training iteration of a plurality of training iterations, the operations further include sampling a batch of training samples from the set of training samples. The operations further include, for each training sample in the batch of training samples, determining a selection probability using a data value estimator. The selection probability for the training sample is based on estimator parameter values of the data value estimator. The operations further include selecting a subset of training samples from the batch of training samples based on the selection probability for each training sample and determining a performance measure using a predictor model with the subset of training samples. The operations further include adjusting model parameter values of the predictor model based on the performance measure and updating the estimator parameter values of the data value estimator based on the performance measure.
[0008] This aspect can include one or more of the following optional features. In some implementations, determining the performance measure using the predictor model includes determining loss data by a loss function. In these implementations, adjusting the model parameter values of the predictor model based on the performance measure includes adjusting the model parameter values of the predictor model based on the loss data. Additionally, in some implementations, updating the estimator parameter values of the data value estimator based on the performance measure includes determining a boost signal from the loss data and updating estimator parameter values of the data value estimator based on the boost signal. Updating the estimator parameter values of the data value estimator based on the boost signal further includes determining a reward value based on the loss data and updating the estimator parameter values of the data value estimator based on the reward value. In these implementations, determining the reward value based on the loss data includes determining a moving average of loss data based on the last N training iterations of the predictor model, determining a difference between the loss data of the last training iteration and the moving average of the loss data, and determining the reward value based on the difference between the loss data of the last training iteration and the moving average of the loss data.
[0009] In some examples, the data value estimator comprises a neural network, and updating estimator parameter values of the data value estimator comprises updating layer parameter values of the neural network of the data value estimator. In some examples, the predictor model is trained using stochastic gradient descent. In some implementations, selecting the subset of training samples from the batch of training samples based on the selection probability for each training sample comprises, for each training sample in the batch of training samples, determining a corresponding selection value indicating selection or non-selection. When the corresponding selection value indicates selection, the operations further comprise adding the training sample to the subset of training samples, and when the corresponding selection value indicates non-selection, the operations further comprise discarding the training sample. In some examples, sampling the batch of training samples comprises, for each training iteration in the plurality of training iterations, sampling a different batch of training samples from the set of training samples.
[0010] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is a schematic diagram of an example system for conducting data evaluations.
[0012] Figure 2 is Figure 1 is a schematic diagram of example components of the system of
[0013] Figure 3 is Figure 1 is a schematic diagram of additional example components of the system of
[0014] Figure 4 is a schematic diagram of an algorithm for training a model for data evaluations.
[0015] Figure 5 is a flowchart of an example arrangement of operations of a method for data evaluations using reinforcement learning.
[0016] Figure 6 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0017] Like reference numbers in different drawings indicate like elements. DETAILED DESCRIPTION
[0018] Training deep neural networks to be highly accurate in predictions often requires large amounts of training data. However, collecting large-scale and high-quality real-world datasets is expensive and challenging. Additionally, accurately training neural networks can take a large amount of time and computational expense. Accurately quantifying the value of training data has significant potential to improve model performance on real-world training datasets, which often contain incorrect labels or vary in quality and usefulness. Rather than treating all data samples in a training dataset equally, a lower priority can be assigned to samples with lower quality to obtain a higher performance model. In addition to improving performance, data valuation can also help develop better data collection practices. However, historical data valuation has been limited by computational cost, as the approach scales linearly with the number of training samples in the dataset.
[0019] Embodiments herein point to data valuation using deep reinforcement learning (DVRL), which is a meta-learning framework to adaptively learn data values in conjunction with training of a predictor model. A data value estimator function modeled by a deep neural network outputs a likelihood that a training sample will be used in training of the predictor model. Training of the data value estimator is based on reinforcement signals using rewards obtained directly from ongoing target tasks. With a small validation set, DVRL can provide computationally efficient and high-quality ranking of data values for a training dataset, which saves time and outperforms other approaches. DVRL can be used in various applications across multiple types of datasets.
[0020] Reference Figure 1 In some embodiments, an example system 100 includes a processing system 10. The processing system 10 can be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) with fixed or scalable / elastic computing resources 12 (e.g., data processing hardware) and / or storage resources 14 (e.g., memory hardware). The processing system 10 executes a meta-learning framework 110 (also referred to herein as a DVLR framework or simply DVLR). The DVLR framework 110 obtains a set of training samples 102. Each training sample includes training data and a label for the training data. The label includes an annotation or other indication of a correct result based on a prediction for the training data. In contrast, unlabeled training samples include only training data without a corresponding label.
[0021] For example, the training samples 102 can include tabular datasets, audio datasets (e.g., for transcription or speech recognition, etc.), image datasets (e.g., for object detection or classification, etc.), and / or text datasets (e.g., for natural language classification, text translation, etc.). The set of training samples 112G can be stored in the processing system 10 (e.g., within the memory hardware 14) or received from another entity via a network or other communication channel. The data value estimator 120 can select training samples 102 in batches from the set of training samples 102 (i.e., a selected or random portion of the set of training samples 102). In some examples, the data value estimator 120 samples the batches of training samples 102 (i.e., different batches for each iteration of training).
[0022] The DVLR framework 110 includes a data value estimator model 120 (e.g., a machine learning model). In some implementations, the data value estimator model 120 is a neural network. For each training sample 102 in a batch of training samples 102, the data value estimator model 120 determines a selection probability 106 based on estimator parameter values 122 of the data value estimator model 120. The selection probability 106 represents a value of the training sample 102 to a prediction of a value of the predictor model 142 for each training sample 102 in the batch of training samples 102. In some examples, the data value estimator model 120 determines a value of an input training sample 102 by quantifying a relevance of the input training sample 102 to the predictor model 142.
[0023] The DVLR framework 110 includes a sampler 130. The sampler 130 receives as input the selection probabilities 106 determined by the data value estimator model 120 for each training sample 102 in a batch. The sampler 130 selects a subset of the training samples 102 based on the selection probability 106 of each training sample 102 to provide to the predictor model 142. As discussed in more detail below, the sampler 130 can discard remaining training samples 102 in the batch of training samples 102 based on the selection probabilities 106. In some implementations, the selection probabilities 106 provided as input to the sampler 130 are based on a multinomial distribution.
[0024] The predictor model 142 (e.g., a machine learning model) receives the subset of training samples 102 sampled by the sampler 130. The predictor model 142 determines a performance measure based on the subset of training samples 102 sampled from the batch of input training samples 102 selected for a current training iteration. The predictor model 142 is trained only with the subset of training samples 102 sampled by the sampler 130. That is, in some implementations, the predictor model 142 is not trained on training samples 102 that are not selected or sampled by the sampler 130.
[0025] The predictor model 142 includes model parameter values 143 that control the predictive capabilities of the predictor model 142. The predictor model 142 makes predictions 145 based on the input training samples 102. The performance evaluator 150 receives the predictions 145 and determines a performance measure (e.g., accuracy of the predictions 145) based on the predictions 145 and the training samples 102 (i.e., the labels associated with the training samples 102). In some implementations, the performance measure includes loss data (e.g., cross-entropy loss data). In these implementations, the DVLR framework 110 determines the augmentation signal based on the loss data. Optionally, the DVLR framework 110 can generate a reward value 230 based on the performance measure. Figure 2
[0026] The DVLR framework 110 adjusts and / or updates the model parameter values 143 of the predictor model 142 and the estimator parameter values 122 of the data value estimator model 120 based on the performance measure. During each training iteration of the plurality of training iterations, the DVLR 110 can adjust the model parameter values 143 of the predictor model 142 based on the performance measure of the training iteration using a feedback loop 148 (e.g., backpropagation). The DVLR 110 can adjust the estimator parameter values 122 of the data value estimator model 120 based on the performance measure of the training iteration using the same or a different feedback loop 148. In some implementations, the DVLR framework 110 updates the estimator parameter values 122 of the data value estimator model 120 by updating the layer parameter values of the neural network of the data value estimator 120.
[0027] Referring now to Figure 2 , the schematic diagram 200 includes the DVLR 110 with the augmentation signal 260 and the feedback loop 148. The performance measure can include loss data. The DVRL framework 110 can determine the loss data 144 based on a subset of the training samples 102 input to the predictor model 142 using a loss function. In some examples, the DVRL framework 110 trains the predictor model 142 using a stochastic gradient descent optimization algorithm with a loss function (e.g., mean squared error (MSE) for regression or cross-entropy for classification). When the performance evaluator 150 determines the loss data 144 based on the loss function, the DVLR 110 uses the feedback loop 148 to update the model parameter values 143 of the predictor model 142 with the performance measure (e.g., the loss data 144).
[0028] After the DVRL framework 110 determines the loss data 144 for a training iteration, the DVLR 110 can generate an augmentation signal 260. In some implementations, the DVRL framework 110 updates the estimator parameter values 122 of the data value estimator model 120 based on the augmentation signal 260. The augmentation signal 260 can also include reward data 220. The performance evaluator 150 can determine the reward data 220 by quantifying a performance measure. For example, when the performance measure indicates low loss data 144 (i.e., minimal error or accurate prediction) from a subset of the training samples 102 received by the predictor model 142, the reward data 220 can augment the estimator parameter values 122 of the data value estimator model 120. Conversely, when the performance measure indicates high loss data 144 (i.e., high error) from a subset of the training samples 102 received by the predictor model 142, the reward data 220 can indicate that the estimator parameter values 122 of the data value estimator model 120 need further updating.
[0029] In some implementations, the performance evaluator 150 computes the reward data 220 based on historical loss data. For example, the performance evaluator 150 uses a moving average calculator 146 to determine a moving average of the loss data based on the last N training iterations of the predictor model 142. In other words, for each training iteration, the moving average calculator 146 can obtain the loss data 144 and determine a difference between the current training iteration loss data 144 and an average of the last N training iterations of the loss data. The DVLR 110 can generate a reward value 230 based on the moving average of the loss data determined by the moving average calculator 146. The reward value 230 can be based on the difference between the current training iteration loss data 144 and the average of the last N training iterations of the loss data. In some implementations, the DVRL framework 110 adds the reward value 230 to the reward data 220 of the augmentation signal 260. In other implementations, the DVRL framework 110 only uses the reward value 230 to influence the reward data 220 by increasing or decreasing the reward data 220 of the augmentation signal 260.
[0030] Reference is now made to Figure 3The schematic diagram 300 includes the DVLR 110 selecting a subset of the training samples 102. In some implementations, the DVLR 110 selects the training samples 102 in the batch of training samples 102 by determining a selection value 132 for each training sample 102. The selection value 132 can indicate selection or non-selection for the corresponding training sample 102. After the data value estimator model 120 generates the selection probability 106 for each training sample 102 in the batch of training samples 102, the sampler 130 determines the corresponding selection value 132 indicating selection 310 or non-selection 320. Optionally, the selection probability 106 generated by the data value estimator model 120 conforms to a multinomial distribution. The sampler 130 obtains the distribution of selection probabilities 106 and the corresponding training sample 102 in the batch of training samples 102 and determines the selection value 132 by determining the likelihood of the training predictor model 142 for each training sample 102 in the batch of training samples 102.
[0031] When the sampler 130 determines that the selection value 132 for the training sample 102 indicates selection 310, the sampler 130 adds the training sample 102 to the subset of training samples 102. Conversely, when the sampler 130 determines that the selection value for the training sample 102 indicates non-selection 320, the sampler 130 can discard the training sample 102 (e.g., to discarded training samples 340). In some implementations, the DVLR framework 110 returns the discarded training samples 340 to the set of training samples 102 for future training iterations. In other implementations, the DVRL framework 110 quarantines the discarded training samples 340 (i.e., removes from the set of training samples 102) to prevent inclusion in future training iterations.
[0032] Reference is now made to Figure 4In some implementations, the DVLR 110 implements the algorithm 400 to train the data value estimator 120 and the predictor model 142. Here, the DVLR 110 takes in a set of training samples 102 (i.e., D), and initializes the estimator parameter values of the data value estimator model 120, the model parameter values of the predictor model 142, and resets the moving average loss in the moving average loss calculator 146. For each training iteration until convergence, the DVLR 110 samples a batch of training samples 102 (i.e., mini-batch B) from the set of training samples 102, and updates the estimator parameter values 122 of the data value estimator model 120 and the model parameter values 143 of the predictor model 142. Using the algorithm 400, for each training sample 102 (i.e., j) in the batch of training samples 102, the data value estimator model 120 uses the sampler 130 to select the value 132 to compute the selection probability 106 and the sample. For each training iteration (i.e., t), the DVLR 110 samples the batch of training samples 102 with the respective selection probabilities 106 and the selection values 132 indicating the selections 310, and determines a performance measure (i.e., loss data). In the next step, the DVLR 110 updates the model parameter values 143 of the predictor model 142 based on the performance measure for the training iteration. Then, the DVLR 110 updates the estimator parameter values 122 of the data value estimator model 120 based on the performance measure for the training iteration, including the moving average loss from the moving average loss calculator 146. In the final step, the DVLR updates the moving average loss in the moving average loss calculator 146.
[0033] Figure 5 is a flowchart of an example arrangement of operations of the method 500 for data evaluation using reinforcement learning. At operation 502, the method 500 includes obtaining, at the data processing hardware 12, a set of training samples 102. At operation 504, the method 500 includes, during each training iteration of a plurality of training iterations, for each training sample 102 in a batch of training samples 102, determining, by the data processing hardware 12 using the data value estimator 120, a selection probability 106 for the training sample 102 based on the estimator parameter values of the data value estimator 120.
[0034] At operation 506, the method 500 includes selecting, by the data processing hardware 12, a subset of the training samples 102 from the batch of training samples 102 based on the selection probabilities 106 for each training sample 102. At operation 508, the method 500 includes determining, by the data processing hardware 12, a performance measure using the predictor model 142 with the subset of training samples 102. At operation 510, the method 500 further includes adjusting, by the data processing hardware 12, the model parameter values 143 of the predictor model 142 based on the performance measure. At operation 512, the method includes updating, by the data processing hardware 12, the estimator parameter values 122 of the data value estimator 120 based on the performance measure.
[0035] Figure 6 FIG. 6 is a schematic diagram of an example computing device 600 that can be used to implement the systems and methods described in this document. The computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions, are meant to be examples for the purposes of this document only, and are not meant to limit implementations of the applications described and / or claimed in this document.
[0036] The computing device 600 includes a processor 610, memory 620, a storage device 630, a high-speed interface / controller 640 connecting the memory 620 and the processor 610 to each other
[0037] Memory 620 non-transitorily stores information within computing device 600. Memory 620 can be a computer-readable medium, a volatile memory unit(s) or non-volatile memory unit(s). The non-transitory memory 620 can be physical devices used to store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as BIOS). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM) as well as disks or tapes.
[0038] Storage device 630 can provide mass storage for computing device 600. In some implementations, storage device 630 is a computer-readable medium. In various implementations, storage device 630 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine- readable medium, such as the memory 620, the storage device 630, or memory on processor 610.
[0039] High-speed controller 640 manages bandwidth-intensive operations for computing device 600, while low-speed controller 660 manages lower bandwidth-intensive operations. Such allocation of functions is exemplary only. In some implementations, high-speed controller 640 is coupled to memory 620, display 680 (e.g., through a graphics processor or accelerator), and to high-speed expansion ports 650, which can accept various expansion cards (not shown). In some implementations, low-speed controller 660 is coupled to storage device 630 and low-speed expansion port 690. The low-speed expansion port 690, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.
[0040] The computing device 600 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a standard server 600a or multiple times in a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.
[0041] Various implementations of the systems and techniques described herein can be realized in digital electronic and / or optical circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0042] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform a task. In some examples, a software application can be referred to as an “application,” an “app,” or a “program.” Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0043] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer readable medium, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0044] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks, or a computer program product comprising a computer program for performing the processes described herein, and a computer program product comprising a computer program for performing the processes described herein. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0045] To provide for interaction with a user, one or more aspects of the disclosure can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display), or touch screen for displaying information to the user and optionally a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device of the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0046] A number of implementations have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the following claims.
Claims
1. A method (500) for evaluating a training sample (102), characterized in that, The method (500) includes: obtaining, at data processing hardware (12), a set of training samples (102); and during each training iteration of a plurality of training iterations: sampling, by the data processing hardware (12), a batch of training samples (102) from the set of training samples (102); for each training sample (102) in the batch of training samples (102), determining, by the data processing hardware (12), using a data value estimator (120), a selection probability (106) for the training sample (102) based on estimator parameter values (122) of the data value estimator (120), wherein the selection probability (106) for each training sample (102) is determined based on a priority value assigned to each training sample (102) based on a sample quality of each training sample (102); selecting, by the data processing hardware (12), a subset of training samples (102) from the batch of training samples (102) based on the selection probability (106) for each training sample (102); determining, by the data processing hardware (12), a performance measure using a predictor model (142) with the subset of training samples (102), wherein determining the performance measure using the predictor model (142) comprises determining loss data (144) by a loss function; adjusting, by the data processing hardware (12), model parameter values (143) of the predictor model (142) based on the performance measure; and updating, by the data processing hardware (12), the estimator parameter values (122) of the data value estimator (120) based on the performance measure, wherein updating the estimator parameter values (122) of the data value estimator (120) based on the performance measure comprises: determining an augmentation signal (260) from the loss data (144); and updating the estimator parameter values (122) of the data value estimator (120) based on the augmentation signal (260).
2. The method (500) according to claim 1, wherein Adjusting the model parameter values (143) of the predictor model (142) based on the performance measure comprises adjusting the model parameter values (143) of the predictor model (142) based on the loss data (144).
3. The method (500) of claim 1, wherein Updating the estimator parameter values (122) of the data value estimator (120) based on the augmentation signal (260) further comprises: determining a reward value (230) based on the loss data (144); and updating the estimator parameter values (122) of the data value estimator (120) based on the reward value (230).
4. The method (500) according to claim 3, wherein Determining the reward value (230) based on the loss data (144) comprises: determining a moving average of loss data based on the last N training iterations of the predictor model (142); determining a difference between the loss data (144) of the last training iteration and the moving average of loss data; and determining the reward value (230) based on the difference between the loss data (144) for the most recent training iteration and a moving average of the loss data.
5. The method (500) according to any one of claims 1-4, wherein, The data value estimator (120) comprises a neural network, and updating estimator parameter values (122) of the data value estimator (120) comprises updating layer parameter values of the neural network of the data value estimator (120).
6. The method (500) according to any one of claims 1-4, wherein, selecting the subset of training samples (102) from the batch of training samples (102) based on the selection probability (106) for each training sample (102) comprises, for each training sample (102) in the batch of training samples (102): determining a corresponding selection value (132) that indicates selection (310) or non-selection (320); when the corresponding selection value (132) indicates selection (310), adding the training sample (102) to the subset of training samples (102); and when the corresponding selection value (132) indicates non-selection (320), discarding the training sample (102).
7. The method (500) according to any one of claims 1-4, wherein, training the predictor model (142) using stochastic gradient descent.
8. The method (500) according to any one of claims 1-4, wherein, sampling the batch of training samples (102) comprises, for each training iteration in the plurality of training iterations, sampling a different batch of training samples (102) from the set of training samples (102).
9. A system (100) for evaluating a training sample (102), characterized in that comprises: data processing hardware (12); and memory hardware (14) in communication with the data processing hardware (12), the memory hardware (14) storing instructions that, when executed on the data processing hardware (12), cause the data processing hardware (12) to perform operations comprising: obtaining a set of training samples (102); and during each training iteration of a plurality of training iterations: sampling a batch of training samples (102) from the set of training samples (102); for each training sample (102) in the batch of training samples (102), determining, using a data value estimator (120), a selection probability (106) for the training sample (102) based on estimator parameter values (122) of the data value estimator (120), wherein the selection probability (106) for each training sample (102) is determined based on a priority value assigned to each training sample (102) based on a sample quality of each training sample (102); selecting a subset of training samples (102) from the batch of training samples (102) based on the selection probability (106) for each training sample (102); determining a performance measure using a predictor model (142) with the subset of training samples (102), wherein determining the performance measure using the predictor model (142) comprises determining loss data (144) by a loss function; adjusting model parameter values (143) of the predictor model (142) based on the performance measure; and determining the reward value (230) based on the difference between the loss data (144) for the most recent training iteration and a moving average of the loss data. updating the estimator parameter values (122) of the data value estimator (120) based on the performance measure, wherein updating the estimator parameter values (122) of the data value estimator (120) based on the performance measure comprises: determining an augmented signal (260) from the loss data (144); and updating estimator parameter values (122) of the data value estimator (120) based on the augmented signal (260).
10. The system (100) according to claim 9, characterized in that updating the model parameter values (143) of the predictor model (142) based on the performance measure comprises updating the model parameter values (143) of the predictor model (142) based on the loss data (144).
11. The system (100) according to claim 9, characterized in that updating the estimator parameter values (122) of the data value estimator (120) based on the augmented signal (260) comprises: determining a reward value (230) from the loss data (144); and updating the estimator parameter values (122) of the data value estimator (120) based on the reward value (230).
12. The system (100) according to claim 11, characterized by determining the reward value (230) from the loss data (144) comprises: determining a moving average of loss data based on the last N training iterations of the predictor model (142); determining a difference between the loss data (144) of the last training iteration and the moving average of loss data; and determining the reward value (230) based on the difference between the loss data (144) of the last training iteration and the moving average of loss data.
13. The system (100) according to any one of claims 9-12, characterized by, the data value estimator (120) comprises a neural network, and updating estimator parameter values (122) of the data value estimator (120) comprises updating layer parameter values of the neural network of the data value estimator (120).
14. The system (100) according to any one of claims 9-12, characterized by, selecting a subset of the training samples (102) from a batch of the training samples (102) based on the selection probabilities (106) of each training sample (102) further comprises, for each training sample (102) in the batch of training samples (102): determining a corresponding selection value (132) indicating selection (310) or non-selection (320); when the corresponding selection value (132) indicates selection (310), adding the training sample (102) to the subset of training samples (102); and when the corresponding selection value (132) indicates non-selection (320), discarding the training sample (102).
15. The system (100) according to any one of claims 9-12, characterized by, training the predictor model (142) using stochastic gradient descent.
16. The system (100) according to any one of claims 9-12, characterized by sampling the batch of training samples (102) comprises, for each training iteration in the plurality of training iterations, sampling a different batch of training samples (102) from the set of training samples (102).
Citation Information
Patent Citations
Target customer screening method and device
CN107730286A
Training method and device for depth learning model
CN109034365A