A model training method and system, a storage medium and a terminal device
By classifying and adjusting the sample weights in the training sample set, the problem of uneven recognition accuracy of machine learning models on long-tailed distributed datasets is solved, the recognition accuracy of the model for each category of samples is improved, and the overall performance is enhanced.
Patent Information
- Application Number
- CN202110952438.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-19
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-11-05
AI Technical Summary
When existing machine learning models are trained on long-tailed datasets, some categories have a large amount of data while others have a small amount. This results in the model performing well on categories with more data but poorly on categories with less data, leading to an imbalance in recognition accuracy.
The training sample set is divided into multiple types of sample subsets, and training samples are selected from each sample subset according to the sample weight value to form the current batch of training subsets. The sample weight values are adjusted to balance the probability of training samples being selected, and a multi-task convolutional neural network is used to train the object recognition model.
It improved the model's accuracy in recognizing various types of training samples, alleviated the long-tail distribution problem, and improved the model's overall recognition performance by about 1-2%.
Smart Images

Figure CN114332550B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information processing technology, and in particular to a model training method, system, storage medium, and terminal device. Background Technology
[0002] With the development of machine learning models, computer vision technology is being applied in various scenarios, such as autonomous driving, smart retail, and visual content understanding, all relying on machine learning models to complete specific tasks. Currently, the integration of deep learning with industry is a major trend, improving recognition accuracy and reducing manual labor costs. The foundation of deep learning-based computer vision tasks is the machine learning model. A good machine learning model can be directly transferred to other related tasks, thereby upgrading a range of tasks. Generally, an excellent machine learning model not only depends on a good basic network architecture, but also on training methods and data sampling methods that can stably improve model performance, which are currently hot research topics. Furthermore, excellent data sampling methods and model training strategies can also be transferred to many downstream tasks, such as image classification, video classification, and video content understanding.
[0003] Specifically, in news reading apps, content creators generate a large amount of news content of various categories every day. This leads to a long-tail distribution problem, where the amount of content in some common categories far exceeds the amount in rare categories, resulting in an imbalance. For example, in a news app, video content about TV drama trailers accounts for 7% of the total, while science-related content accounts for only 1%. Due to this phenomenon, the performance of deep learning-based machine learning models improves significantly with the increase in data volume. Therefore, if there is a large amount of data in certain categories, the machine learning model will perform better during training in those categories, but perform worse in other categories. Summary of the Invention
[0004] This invention provides a model training method, system, storage medium, and terminal device, enabling more accurate training of object recognition models.
[0005] One embodiment of the present invention provides a model training method, comprising:
[0006] Obtain a training sample set, which includes sample subsets of multiple types;
[0007] Determine the sample weight values for each training sample in each sample subset of each type;
[0008] Based on the sample weight values of the training samples in each type of sample subset, at least one training sample is selected from the corresponding type of sample subset, and the at least one training sample selected from each of the multiple types of sample subsets forms the training subset of the current batch.
[0009] The object recognition model is trained based on the training subset of the current batch, and the object recognition model is used to identify the type of the target object.
[0010] Another aspect of this invention provides a model training system, comprising:
[0011] A sample acquisition unit is used to acquire a training sample set, which includes multiple types of sample subsets.
[0012] The sample weight unit is used to determine the sample weight value of each training sample in each type of sample subset;
[0013] The sample selection unit is used to select at least one training sample from the corresponding type of sample subset according to the sample weight value of the training samples in each type of sample subset, and to form the training subset of the current batch by selecting at least one training sample from the multiple types of sample subsets respectively.
[0014] The training unit is used to train an object recognition model based on the training subset of the current batch, the object recognition model being used to identify the type of the target object.
[0015] Another aspect of the present invention provides a computer-readable storage medium storing a plurality of computer programs adapted for loading by a processor and executing the model training method as described in one aspect of the present invention.
[0016] Another embodiment of the present invention provides a terminal device, including a processor and a memory;
[0017] The memory is used to store multiple computer programs, which are loaded and executed by a processor as described in one aspect of the model training method of the present invention; the processor is used to implement each of the multiple computer programs.
[0018] As can be seen, in the method of this embodiment, the model training system determines the sample weight value of each training sample in each type of sample subset of the training sample set, and selects at least one training sample from each sample subset according to the sample weight value. The at least one training sample selected from multiple sample subsets is then combined to form the current batch of training subsets to train the object recognition model. This divides the training sample set into multiple types of sample subsets, and uses the sample weight value to measure the probability of a training sample being selected in each sample subset to form the current batch of training subsets. This allows the probability of training samples being selected to be balanced by adjusting the sample weight value, thereby improving the accuracy of the trained object recognition model in recognizing various types of training samples. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of a model training method provided in an embodiment of the present invention;
[0021] Figure 2 This is a flowchart of a model training method provided in one embodiment of the present invention;
[0022] Figure 3 This is a flowchart of a method for training an object recognition model in one embodiment of the present invention;
[0023] Figure 4 This is a flowchart of a model training method provided in an application embodiment of the present invention;
[0024] Figure 5 This is a schematic diagram illustrating the selection of training samples from a subset of samples of one type in an application embodiment of the present invention;
[0025] Figure 6 This is a schematic diagram of the training subsets that make up each batch in one application embodiment of the present invention;
[0026] Figure 7 This is a schematic diagram of a distributed system to which the model training method is applied in another application embodiment of the present invention;
[0027] Figure 8 This is a schematic diagram of the block structure in another application embodiment of the present invention;
[0028] Figure 9This is a schematic diagram of the logical structure of a model training system provided in an embodiment of the present invention;
[0029] Figure 10 This is a schematic diagram of the logical structure of a terminal device provided in an embodiment of the present invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] This invention provides a model training method, which mainly involves processing a training sample set and then training an object recognition model based on the processed training sample set. Specifically, for example... Figure 1 As shown, the model training system can train the object recognition model according to the following steps:
[0033] Obtain a training sample set, which includes sample subsets of multiple types (n types are used as an example in the figure); determine the sample weight value of each training sample in the sample subset of each type; select at least one training sample from the corresponding sample subset according to the sample weight value of the training sample in the sample subset of each type, and form the at least one training sample selected from the sample subsets of the multiple types into the current batch of training subsets; train an object recognition model according to the current batch of training subsets, which is used to identify the type of target object.
[0034] This allows the training sample set to be divided into multiple types of sample subsets, and the probability of a training sample being selected from each sample subset is measured by the sample weight value to form the current batch of training subsets. This enables the probability of a training sample being selected to be balanced by adjusting the sample weight value, thereby improving the accuracy of the trained object recognition model in recognizing various types of training samples.
[0035] In practical applications, the method of this embodiment can be applied not only to video / image / text classification and recognition products, but also to the training of a series of models such as video / text tag recognition, high-quality content recognition, and low-quality content recognition.
[0036] The aforementioned object recognition model is a machine learning model based on artificial intelligence, specifically a multi-task convolutional neural network (MTCNN), such as the output network (ONet) of MTCNN. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0037] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0038] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.
[0039] One embodiment of the present invention provides a model training method, mainly a method executed by a model training system, the flowchart of which is shown below. Figure 2 As shown, it includes:
[0040] Step 101: Obtain the training sample set, which includes sample subsets of multiple types.
[0041] It is understandable that when training any machine learning model, it is necessary to learn the parameter values of any structure of the machine learning model based on a large number of training samples. These training samples form a training sample set. Each training sample includes the sample object and the annotation information of the sample object. The specific type of information in the annotation information of the sample object needs to be determined according to the function of the machine learning model to be trained. For example, if the machine learning model is mainly used to identify a certain type of object in an image, then the annotation information of the sample object is mainly the position information of the corresponding type of object in the sample object.
[0042] In this embodiment, the machine learning model to be trained is mainly an object recognition model, used to identify the attribute information of target objects, such as whether any image is a face image, or the user emotion represented by any piece of voice data. Essentially, it identifies the type of the target object, which can be multimedia information such as images, videos, or text. Correspondingly, the training samples in the training sample set also need to include multiple types of sample subsets. The specific types of sample subsets included must be consistent with the types that the object recognition model can identify. For example, if the object recognition model can identify n types of target objects, then the training sample set needs to include the corresponding n types of sample subsets.
[0043] To improve the accuracy of the trained object recognition model in identifying target objects as any type of object, it is necessary to ensure that the number of training samples in each type of sample subset is sufficient, and that the difference in the number of training samples between any two type of sample subsets is not significant; that is, the sample subsets of each type need to be balanced. However, in practical applications, some types of data in the existing datasets that can be used as training samples may be relatively scarce. For example, in a news video dataset, there are many news videos about TV drama trailers, but relatively few news videos related to science, resulting in an imbalance in the obtained sample subsets of each type. To reduce the impact of sample subset imbalance on the training of the object recognition model, in this embodiment of the invention, after obtaining the training sample set, the following steps 102 to 104 need to be performed.
[0044] Step 102: Determine the sample weight value of each training sample in each type of sample subset. The sample weight value is a weight value corresponding to any training sample.
[0045] Specifically, when determining the sample weights of training samples in a sample subset of a certain type, the model training system may only consider all training samples in that sample subset, without considering training samples in other sample subsets.
[0046] In a specific scenario, the model training system can first identify the hard and easy samples among all training samples in a subset of samples of a certain type. It then determines that the sample weights of the hard samples are greater than one threshold, while the sample weights of the easy samples are less than another threshold. This increases the probability of hard samples being selected to form the training subset, thereby improving the accuracy of the object recognition model in learning from hard samples. Here, one threshold can be greater than the other.
[0047] In this context, "difficult samples" refer to training samples that are easily predicted incorrectly by the object recognition model, or training samples with low prediction confidence. "Simple samples" refer to training samples that are easily learned by the object recognition model. Specifically, when identifying difficult samples, cosine clustering or Euclidean clustering distances can be calculated among the training samples to filter out difficult samples.
[0048] Step 103: Based on the sample weight values of the training samples in each type of sample subset, select at least one training sample from the corresponding type of sample subset, and combine the at least one training sample selected from the multiple types of sample subsets to form the training subset of the current batch.
[0049] Specifically, the model training system needs to select at least one training sample from each sample subset and combine these training samples into a batch of training subsets. The number of training samples selected from each sample subset can be the same or different, for example, selecting 3 training samples from one sample subset and 4 training samples from another sample subset, etc.
[0050] When selecting at least one training sample from a subset of samples, one can choose the at least one training sample with the largest sample weight value from that subset, or choose at least one training sample with a sample weight value greater than a preset value. Specifically, at least one training sample from any subset can be selected in the following ways:
[0051] (1) Select at least one training sample based on binary search.
[0052] Specifically, the sample weights of all training samples in a sample subset of a certain type are first used as a vector. Then, a binary search is used to find at least one training sample with the largest sample weight.
[0053] The basic idea of binary search is as follows: Assuming the elements in the table are sorted in ascending order, compare the key of the record in the middle position with the search key. If they are equal, the search is successful; otherwise, divide the table into two sub-tables using the middle position record. If the key of the middle position record is greater than the search key, continue searching the first sub-table; otherwise, continue searching the second sub-table. Repeat this process until a record that meets the conditions is found, making the search successful, or until the sub-table no longer exists, at which point the search is unsuccessful. This method offers fast search speed and good average performance when selecting training samples.
[0054] (2) Select at least one training sample based on the roulette wheel rotation method.
[0055] Specifically, the sample weight values of each training sample in a subset of samples of a certain type can be mapped to a unit component of a preset wheel. Then, the preset wheel is rotated, and the training sample corresponding to the sample weight value pointed to by the pointer in the preset wheel is selected.
[0056] The preset wheel mainly consists of multiple unit components and pointers. During the rotation of the wheel, the pointers will cycle and point to each unit component in sequence. The unit components can be of any shape, such as a sector or a square.
[0057] When assigning the sample weights of each training sample in a subset of samples of a certain type to a unit component of a pre-set roulette wheel, it is necessary to consider the relationship between the number of training samples 'a' in the subset and the number of unit components 'b' in the roulette wheel. If 'a' and 'b' are the same, the sample weights can be assigned one-to-one with the unit components. If 'a' is less than 'b', 'a' unit components can be selected from 'b' unit components, and the sample weights can be assigned one-to-one with the selected unit components.
[0058] If a is greater than b, first, assign x*b sample weight values to b unit components one-to-one. Then, select ax*b unit components from the b unit components and assign ax*b sample weight values to the selected ax*b unit components one-to-one. In this case, one unit component may correspond to one or more sample weight values. Here, x is a natural number greater than or equal to 1.
[0059] (3) Select at least one training sample based on a preset number of samplings.
[0060] Specifically, based on the sample weight values of each training sample in a sample subset, a corresponding number of samplings is assigned to each training sample; based on the number of samplings of each training sample, the training sample with the largest sample weight value is selected from a sample subset, and the number of samplings of the selected training sample is updated; in this way, the steps of selecting training samples and updating the number of samplings are executed cyclically, so as to obtain at least one training sample corresponding to a sample subset.
[0061] Specifically, when assigning sampling counts to each training sample, training samples with larger weight values are sampled more times, while training samples with smaller weight values are sampled less times. When updating the sampling count of selected training samples, the current sampling count is decremented by 1. If the sampling count of a training sample is 0, that training sample will not be selected.
[0062] Furthermore, the number of iterations for the steps of selecting training samples and updating the number of samples can be determined by the number of iterations preset in the model training system, or by the number of training samples in the current batch of training subsets. For example, if the number of training samples in the current batch of training subsets is a1, and there are a2 types of sample subsets, then a1 / a2 training samples can be selected from each sample subset. For a sample subset, a3 training samples can be selected in each iteration. Therefore, the number of iterations for the steps of selecting training samples and updating the number of samples can be obtained as (a1 / a2) / a3.
[0063] Step 104: Train an object recognition model based on the training subset of the current batch. This object recognition model is used to identify the type of the target object.
[0064] Specifically, each of the above training samples includes a sample object and its annotation information. Therefore, the sample weight value of a training sample is the same as the sample weight value of the sample object, such as... Figure 3 As shown, the model training system can train an object recognition model based on the training subset of the current batch according to the following steps:
[0065] Step 201: Determine the initial model for object recognition.
[0066] It is understandable that when determining the initial model for object recognition, the model training system mainly determines the initial values of the parameters in the multi-layered structure and mechanism of each layer of the initial model for object recognition.
[0067] The initial model for object recognition may include a feature extraction module and a recognition module. The feature extraction module is used to extract feature information of sample objects in the training samples, and the recognition module is used to determine the type of sample object based on the feature information extracted by the feature extraction module. Specifically, the recognition module will output the probability information of the sample object belonging to a certain type.
[0068] The parameters of the initial model for object recognition refer to the fixed parameters used in the calculation of each layer of the initial model for object recognition, which do not need to be assigned values at any time, such as parameter size, number of network layers, user vector length, etc.
[0069] Step 202: Identify the type of each sample object in the training subset of the current batch using the object recognition initial model.
[0070] Specifically, in the initial model of object recognition, the feature extraction module first extracts the feature information of each sample object, and then the recognition module determines the type of each sample object based on the feature information extracted by the feature extraction module.
[0071] Step 203: Based on the type of each sample object obtained from the initial object recognition model, and the annotation information and sample weight values of the corresponding sample objects in the training subset of the current batch, adjust the initial object recognition model to obtain the final object recognition model.
[0072] Specifically, the model training system device first calculates a loss function related to the initial object recognition model based on the types of each sample object obtained from the initial object recognition model in step 202 above, and the annotation information and sample weight values of the corresponding sample objects in the current batch of training subsets. Specifically, the loss function may include: the error between the type of each sample object recognized by the initial object recognition model and the actual type of the corresponding sample object (obtained from the annotation information of the sample object), and the product of the sample weight value of the corresponding sample object, such as the cross-entropy loss function; and then adjusts the parameter values in the initial object recognition model according to the calculated loss function.
[0073] It should be noted that the process of training the object recognition model aims to minimize the aforementioned error values. This training process involves continuously optimizing the parameter values in the initial object recognition model determined in step 201 using a series of mathematical optimization techniques, such as backpropagation differentiation and gradient descent, to minimize the calculated value of the loss function. Specifically, when the calculated loss function value is large, such as exceeding a preset value, the parameter values need to be changed, for example, by reducing the weight value connected to a neuron, so that the loss function value calculated based on the adjusted parameter values is reduced.
[0074] In one specific embodiment, the model training system will consider the type weight values of each sample object during the process of calculating the loss function. Specifically:
[0075] The model training system determines a type weight value for each type of sample subset. This type weight value corresponds to a weight value for any sample subset of a given type, and all training samples within a subset correspond to a single type weight value. Thus, when calculating the loss function, the system can specifically calculate the loss function related to the initial object recognition model based on the types of each sample object obtained from the initial model, the annotation information of the corresponding sample objects in the current batch of training subsets, and their sample weight values and type weight values. For example, the loss function could include: the error between the type of each sample object identified by the initial model and the actual type of the corresponding sample object (obtained from the annotation information), multiplied by the sample weight value and type weight value of the corresponding sample object.
[0076] When determining the type weight value for each type of sample subset, it can be achieved in the following ways, but not limited to:
[0077] (1) Determine the corresponding type weight value by the number of training samples included in a sample subset of a type.
[0078] Specifically, the reciprocal of the number of training samples included in a subset of a type can be taken to obtain a reciprocal value. This reciprocal value can then be used as the type weight value for a subset of samples of a type. This allows subsets with fewer training samples to have higher type weight values, while subsets with more training samples have lower type weight values. This increases the importance of smaller subsets and alleviates the long-tail distribution phenomenon of training samples.
[0079] (2) Determine the type weight value through a smoothing strategy.
[0080] Specifically, the ratio of the reciprocal of the number of training samples included in a subset of a type to the sum of the reciprocals of the number of training samples in each subset is calculated, and the function value of this ratio is used as a type weight value for a subset of samples of a type. Here, the number of training samples included in a subset of a type is Count. i reciprocal of ratio i It can be expressed by the following formula 1, where the type weight value of a subset of samples of a certain type is weight. i This can be expressed by the following formula 2:
[0081]
[0082]
[0083] Where M is the number of sample subsets, j represents the sample subset of each type, and i represents the sample subset of one type.
[0084] The type weight values determined by this smoothing strategy ensure that the type weight values of sample subsets with a large number of training samples are not too low.
[0085] Additionally, it should be noted that steps 202 to 203 above are adjustments to the parameter values in the initial object recognition model based on the types of each sample object determined by the initial object recognition model. In practical applications, steps 202 to 203 need to be executed repeatedly until the adjustment of the parameter values meets certain stopping conditions.
[0086] Therefore, after executing steps 201 to 203 of the above embodiment, the model training system also needs to determine whether the current adjustment of the parameter value meets the preset stopping conditions. If it does, the parameter value obtained in step 203 is used as the parameter value of the finally trained object recognition model, and the process ends. If it does not meet the conditions, steps 202 to 203 are returned to be executed for the initial object recognition model after the parameter value adjustment. The preset stopping conditions include, but are not limited to, any one of the following conditions: the difference between the currently adjusted parameter value and the previously adjusted parameter value is less than a threshold, i.e., the adjusted parameter value has converged; and the number of parameter value adjustments is equal to the preset number, etc.
[0087] It should be further noted that, since the number of training samples in the training sample set obtained in step 101 is relatively large, in the actual training process, these training samples need to be divided into multiple batches of training subsets. There can be overlap between any two batches of training subsets. Then, for each batch of training subsets, the object recognition model can be trained according to the methods in steps 202 to 203 above. When the object recognition model has been trained for all batches of training subsets, one round of training of the object recognition model is completed. Furthermore, the model training system can perform multiple rounds of training on the object recognition model to make the trained object recognition model more accurate. Specifically, when training the object recognition model for a batch of training subsets, steps 202 and 203 above can also be executed repeatedly.
[0088] In this process, the sample weights of each sample object can be continuously updated. Since there is a one-to-one correspondence between training samples and sample objects, the sample weights of the sample objects are the same as the sample weights of the training samples. This achieves the continuous updating of the sample weights of each training sample. Specifically, this can be done in ways including but not limited to the following two:
[0089] (1) After training the object recognition model for the current batch of training subsets, the sample weight values of the corresponding sample objects can be updated according to the type of each sample object obtained from the initial object recognition model and the annotation information of the corresponding sample objects in the current batch of training subsets, so as to obtain the updated sample weight values of each sample object in the current batch of training subsets; then, for the updated sample weight values, return to execute the above steps 103 and 104, that is, execute the steps of forming another batch of training subsets and training the object recognition model. In this way, the above steps 103 and 104 and the steps of updating sample weight values can be executed repeatedly until all batches of training subsets have been processed.
[0090] Specifically, when updating the sample weight value of a sample object, the update can be based on the difference between the type of the sample object obtained from the initial object recognition model and the annotation information of the corresponding sample objects in the current batch of training subsets. If the difference is greater than a threshold, the sample weight value of the sample object can be increased; if the difference is less than another threshold, the sample weight value of the sample object can be decreased. Thus, when the initial object training model's prediction of the type of a certain sample object is inaccurate, the sample weight value of that sample object needs to be increased. This ensures that when adjusting the parameter values in the initial object training model, the consideration given to that sample object is increased, thereby making the training of the object training model more accurate.
[0091] (2) After training the object recognition model for all batches of training subsets corresponding to the training sample set, update the sample weight values of all training samples in the training sample set to obtain the updated sample weight values of all training samples; then start the next round of training of the object recognition model based on the updated sample weight values, that is, return to the steps in steps 103 and 104 above to form the current batch of training subsets and train the object recognition model.
[0092] In this case, updating the sample weight value of a training sample specifically means updating the sample weight value of the sample object in that training sample. Specifically, we can first count the difference between the type of the sample object obtained by the initial object recognition model obtained when training the object recognition model based on the sample object in at least one batch of training subsets in one round of training, and the labeling information of the corresponding sample object in a batch of training subsets. Then, we calculate the average value of the difference obtained for a batch of training subsets. If the average value is greater than a threshold, the sample weight value of the sample object can be increased; if the average value is less than another threshold, the sample weight value of the sample object can be decreased.
[0093] As can be seen, in the method of this embodiment, the model training system determines the sample weight value of each training sample in each type of sample subset of the training sample set, and selects at least one training sample from each sample subset according to the sample weight value. The at least one training sample selected from multiple sample subsets is then combined to form the current batch of training subsets to train the object recognition model. This divides the training sample set into multiple types of sample subsets, and uses the sample weight value to measure the probability of a training sample being selected in each sample subset to form the current batch of training subsets. This allows the probability of training samples being selected to be balanced by adjusting the sample weight value, thereby improving the accuracy of the trained object recognition model in recognizing various types of training samples.
[0094] The following uses a specific application example to illustrate the model training method of the present invention, such as... Figure 4 As shown, the model training method in this embodiment may include the following steps:
[0095] Step 301: Obtain the training sample set, which includes multiple types of sample subsets. Each type of sample subset includes multiple training samples, and each training sample includes a sample object and its annotation information. The annotation information of a sample object is specifically the type information of the sample object.
[0096] Any subset of samples of any type can be denoted as D = {(x i ,y i )}, i∈(1,2,...,N), where N is the number of training samples in the sample subset, where x i For the sample object, y i The annotation information for the sample objects.
[0097] Step 302: Determine the type weight values for each type of sample subset.
[0098] Specifically, the number of training samples included in each type of sample subset is first counted, and the type weight value of the corresponding type of sample subset is determined based on the counted number. The specific method for determining the type weight value is described in the above embodiments and will not be repeated here.
[0099] Step 303: Determine the sample weight value of each training sample in the sample subset of each type. In this way, each training sample in the training sample set corresponds to a sample weight value and a type weight value, that is, each sample object corresponds to a sample weight value and a type weight value.
[0100] It should be noted that object recognition models will perform differently when predicting different types of training samples. Some training samples can quickly learn the corresponding characteristics, and these training samples are simple samples that are easy for object recognition models to learn. However, some training samples are either predicted as the wrong type or the prediction confidence is very low when the object recognition model is used. These training samples are underfitting for the object recognition model and therefore need to be focused on in subsequent learning stages. These training samples are considered to be difficult samples.
[0101] In this context, when determining the sample weights of each training sample in a sample subset, the model training system can first identify the difficult samples in the training samples included in that subset and assign higher sample weights to these difficult samples. This increases the probability that these difficult samples will be selected into the training subset, while also ensuring the probability that simple samples will be selected. In other words, it improves the accuracy of the object recognition model in recognizing both difficult and simple samples.
[0102] Step 304: Select at least one training sample from the corresponding sample subset based on the sample weight values of each training sample in each type of sample subset. The specific selection method is described in the above embodiments and will not be repeated here.
[0103] For example Figure 5 As shown, this represents all training sample items included in a subset of samples of a certain type. 0,i Select the training sample item with the largest sample weight from the i = 1, 2, ..., k, ..., n. 0,k They are then placed into a training subset of a batch.
[0104] Step 305: Select at least one training sample from each type of sample subset according to step 304 above to form the training subset of the current batch.
[0105] Specifically, during the training process of the object recognition model, due to hardware limitations, it is necessary to divide all training samples in the entire training sample set into multiple batches of training subsets. In order to ensure that the object recognition model attaches the same importance to all types of training samples, when forming a batch of training subsets, the training subsets of each type can be numbered i∈(0,1,2,...,M), and at least one training sample selected from each training subset in step 304 above can be filled into a batch of training subsets according to the number. This can ensure that each type of training sample has the same probability of being learned by the object recognition model.
[0106] For example Figure 6As shown, the sample subsets of each type are numbered 0, 1, 2, ..., 9. At least one training sample obtained from each sample subset according to its number needs to be added to the training subset of each batch. Thus, the training sample item in any batch of the training subset... i,j This represents the training samples selected from a subset i and added to the training subset of batch j.
[0107] It should be noted that, in this embodiment, the model training system can obtain a training subset corresponding to the training sample set as the training subset of the current batch, and perform the following steps to train the object recognition model. After adjusting the sample weight values, it can obtain another training subset corresponding to the training sample set as the training subset of the current batch to train the object recognition model. This cyclic operation ensures that all batches of training subsets corresponding to the training sample set are processed.
[0108] Alternatively, in other embodiments, the model training system can acquire all batches of training subsets corresponding to the training sample set at once, and sequentially select one training subset as the training subset of the current batch, and execute the following training steps of the object recognition model respectively. Then, after completing one round of training, the sample weight values of all training samples are adjusted.
[0109] Step 306: Train the object recognition model based on the training subset of the current batch. When calculating the loss function during training, it can be calculated based on the error of the object recognition initial model in predicting the type of each sample object in the training subset of the current batch, as well as the type weight value and sample weight value of the corresponding sample object.
[0110] Step 307: Determine whether the above steps have been performed for all batches of training subsets corresponding to the training sample set. If all have been performed, it means that one round of training of the object recognition model has been completed for the training sample set, and you can continue to step 309. If there are still some batches of training subsets that have not been trained, continue to step 308.
[0111] Step 308: Based on the error of the object recognition initial model in predicting the type of each sample object in the current batch of training subset, update the sample weight value of each sample object in the current batch of training subset, obtain the updated sample weight value of each sample object, and return to execute step 304 above for the updated sample weight value.
[0112] Step 309: Determine whether another round of training is needed for the training sample set. If so, return to step 303; otherwise, the process can end.
[0113] It should be noted that in this embodiment, after training the object recognition model for each batch of training subsets, the sample weight values of each sample object in the batch of training subsets will be adjusted. In another specific embodiment, the sample weight values of all training samples can be adjusted after one round of training of the object recognition model, that is, after training the object recognition model for all batches of training subsets corresponding to the training set.
[0114] As can be seen, the method in this embodiment is mainly based on the dynamic sampling method of training samples to train the object recognition model. It can be directly applied to the training of various machine learning models and can alleviate the long-tail distribution problem in existing datasets. At the same time, in each round of training iteration, the sample weight value of each training sample is dynamically updated according to the latest trained object recognition model to realize dynamic sampling of training samples. This allows for targeted training of training samples with poor prediction performance of the object recognition model, thereby enabling the object recognition model to focus on learning certain samples, such as difficult samples, which can significantly improve the performance of the object recognition model, generally by about 1-2%.
[0115] The following uses another specific application example to illustrate the model training method in this invention. The model training system in this embodiment is mainly a distributed system 100. The distributed system may include a client 300 and multiple nodes 200 (any form of computing device in the network, such as a server or user terminal). The client 300 and the nodes 200 are connected through network communication.
[0116] Taking a distributed system as an example, see blockchain system. Figure 7 This is an optional structural diagram of the distributed system 100 provided in this embodiment of the invention applied to a blockchain system. It consists of multiple nodes 200 (any form of computing device connected to the network, such as servers or user terminals) and clients 300. The nodes form a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In the distributed system, any machine, such as a server or terminal, can join and become a node. A node includes a hardware layer, a middleware layer, an operating system layer, and an application layer.
[0117] See Figure 7 The functions of each node in the blockchain system shown include:
[0118] 1) Routing: A basic function of nodes used to support communication between nodes.
[0119] In addition to routing capabilities, nodes can also have the following functions:
[0120] 2) Applications are deployed in the blockchain to implement specific business needs. They record data related to the implementation of functions to form record data, carry digital signatures in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes successfully verify the source and integrity of the record data, they add the record data to the temporary block.
[0121] For example, the business logic implemented by the application includes: code that implements the model training function, which mainly includes:
[0122] A training sample set is obtained, which includes multiple types of sample subsets; the sample weight value of each training sample in each type of sample subset is determined; based on the sample weight value of the training samples in each type of sample subset, at least one training sample is selected from the corresponding type of sample subset, and the at least one training sample selected from each of the multiple types of sample subsets forms the current batch of training subsets; an object recognition model is trained based on the current batch of training subsets, and the object recognition model is used to identify the type of target object.
[0123] 3) A blockchain consists of a series of blocks that are sequentially generated. Once a new block is added to the blockchain, it will not be removed. The blocks contain the data submitted by the nodes in the blockchain system.
[0124] See Figure 8 This is an optional schematic diagram of the block structure provided in an embodiment of the present invention. Each block includes the hash value of the transaction records stored in this block (the hash value of this block) and the hash value of the previous block. The blocks are connected through their hash values to form a blockchain. Additionally, the block may include information such as a timestamp when it was generated. A blockchain is essentially a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains relevant information used to verify the validity of the information (anti-counterfeiting) and to generate the next block.
[0125] This invention also provides a model training system, the structural schematic of which is shown below. Figure 9 As shown, it can specifically include:
[0126] The sample acquisition unit 10 is used to acquire a training sample set, which includes multiple types of sample subsets.
[0127] The sample weight unit 11 is used to determine the sample weight value of each training sample in each type of sample subset acquired by the sample acquisition unit 10.
[0128] The sample weight unit 11 is specifically used to determine the hard and easy samples among all training samples of a sample subset of a certain type; to determine that the sample weight value of the hard sample is greater than a threshold, and to determine that the sample weight value of the easy sample is less than another threshold.
[0129] The sample weight unit 11 is specifically used to assign the sample weight values of each training sample in a sample subset to a unit component of a preset wheel, rotate the preset wheel; select the training sample corresponding to the sample weight value pointed to by the pointer in the preset wheel; or, take the sample weight values of all training samples in a sample subset as a vector, and obtain at least one training sample with the largest sample weight value through a binary search method.
[0130] Alternatively, the sample weight unit 11 is specifically used to assign a corresponding number of samplings to each training sample based on the sample weight value of each training sample in a sample subset; select the training sample with the largest sample weight value from the sample subset based on the number of samplings of each training sample; update the number of samplings of the selected training sample; and repeatedly execute the steps of selecting training samples and updating the number of samplings to obtain at least one training sample corresponding to the sample subset.
[0131] The sample selection unit 12 is used to select at least one training sample from the corresponding type of sample subset according to the sample weight value of the training sample in each type of sample subset determined by the sample weight unit 11, and to form the training subset of the current batch by selecting at least one training sample from the multiple types of sample subsets respectively.
[0132] Training unit 13 is used to train an object recognition model based on the training subset of the current batch obtained by sample selection unit 12. The object recognition model is used to identify the type of the target object.
[0133] The training unit 13 is specifically used to determine the initial model for object recognition; to identify the type of each sample object in the current batch of training subsets using the initial model for object recognition; and to adjust the initial model for object recognition based on the type of each sample object obtained from the initial model for object recognition, as well as the annotation information and sample weight values of the corresponding sample objects in the current batch of training subsets, so as to obtain the final object recognition model.
[0134] Specifically, when the training unit 13 adjusts the initial object recognition model based on the types of each sample object obtained from the initial object recognition model and the annotation information and sample weight values of the corresponding sample objects in the current batch of training subsets, it calculates a loss function related to the initial object recognition model based on the types of each sample object obtained from the initial object recognition model and the annotation information and sample weight values of the corresponding sample objects in the current batch of training subsets; and adjusts the parameter values in the initial object recognition model based on the loss function.
[0135] Furthermore, the model training system in this embodiment also includes: a type weight unit 14 and a weight adjustment unit 15, wherein:
[0136] Type weight unit 14 is used to determine the type weight value of each type of sample subset; then training unit 13 is specifically used to calculate the loss function related to the object recognition initial model based on the type of each sample object obtained by the object recognition initial model, the annotation information of the corresponding sample objects in the current batch of training subsets and their sample weight values and the type weight value determined by type weight unit 14.
[0137] Specifically, the type weight unit 14 is used to calculate the reciprocal of the number of training samples included in a sample subset of a type, obtain a reciprocal value, and use the reciprocal value as the type weight value of the sample subset of the type; or, it calculates the ratio of the reciprocal of the number of training samples included in a sample subset of a type to the sum of the reciprocals of the number of training samples in the sample subset of each type, and uses the function value of the ratio as the type weight value of the sample subset of the type.
[0138] The weight adjustment unit 15 is used to update the sample weight value of the corresponding sample object according to the type of each sample object obtained by the initial model of object recognition and the annotation information of the corresponding sample object in the training subset of the current batch, so as to obtain the updated sample weight value of each sample object; and to notify the sample selection unit 12 to execute the step of training the object recognition model by the training subset that makes up the current batch and the training unit 13.
[0139] Alternatively, the weight adjustment unit 15 is used to perform one round of training of the object recognition model for all batches of training subsets corresponding to the training sample set, update the sample weight values of all training samples in the training sample set, and obtain the updated sample weight values of all training samples; and for the updated sample weight values, notify the sample selection unit 12 to perform the steps of training the object recognition model for the training subsets that make up the current batch and training unit 13.
[0140] As can be seen, in the model training system of this embodiment, the sample weight unit 11 determines the sample weight value of each training sample in each type of sample subset of the training sample set. The sample selection unit 12 selects at least one training sample from each sample subset according to the sample weight value. The training unit 13 forms the current batch of training subsets from the at least one training sample selected from multiple sample subsets to train the object recognition model. In this way, the training sample set can be divided into multiple types of sample subsets, and the probability of training samples being selected in each sample subset can be measured by the sample weight value to form the current batch of training subsets. This allows the probability of training samples being selected to be balanced by adjusting the sample weight value, thereby improving the accuracy of the trained object recognition model in recognizing various types of training samples.
[0141] This invention also provides a terminal device, the structural schematic of which is shown below. Figure 10 As shown, the terminal device can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) 20 (e.g., one or more processors) and memory 21, and one or more storage media 22 (e.g., one or more mass storage devices) for storing application programs 221 or data 222. The memory 21 and storage media 22 can be temporary or persistent storage. The program stored in the storage media 22 may include one or more modules (not shown in the figure), each module including a series of instruction operations on the terminal device. Furthermore, the CPU 20 may be configured to communicate with the storage media 22 and execute the series of instruction operations in the storage media 22 on the terminal device.
[0142] Specifically, the application program 221 stored in the storage medium 22 includes a model training application program, which may include the sample acquisition unit 10, sample weight unit 11, sample selection unit 12, training unit 13, type weight unit 14, and weight adjustment unit 15 in the aforementioned model training system, which will not be elaborated here. Furthermore, the central processing unit 20 may be configured to communicate with the storage medium 22 and execute a series of operations corresponding to the model training application program stored in the storage medium 22 on the terminal device.
[0143] The terminal device may also include one or more power supplies 23, one or more wired or wireless network interfaces 24, one or more input / output interfaces 25, and / or one or more operating systems 223, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0144] The steps performed by the model training system in the above method embodiments can be based on this. Figure 10 The structure of the terminal device shown is illustrated.
[0145] Another aspect of the present invention provides a computer-readable storage medium storing a plurality of computer programs adapted for loading by a processor and executing a model training method as described above by the model training system.
[0146] Another embodiment of the present invention provides a terminal device, including a processor and a memory;
[0147] The memory is used to store multiple computer programs, which are loaded by the processor and executed as in the model training method performed by the model training system described above; the processor is used to implement each of the multiple computer programs.
[0148] Additionally, according to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the model training methods provided in the various alternative implementations described above.
[0149] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0150] The foregoing has provided a detailed description of a model training method, system, storage medium, and terminal device provided by embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A model training method, characterized in that, The method comprises the following steps: obtaining a training sample set, wherein the training sample set comprises a plurality of sample subsets of different types; determining a sample weight value of each training sample in each sample subset of different types respectively; allocating a corresponding sampling frequency to each training sample in a sample subset according to the sample weight value of each training sample in the sample subset; selecting a training sample with the largest sample weight value from the sample subset according to the sampling frequency of each training sample in the sample subset; updating the sampling frequency of the selected training sample; repeating the steps of selecting a training sample and updating the sampling frequency to obtain at least one training sample corresponding to the sample subset, and grouping at least one training sample selected from the plurality of sample subsets of different types to form a current batch of training subsets, wherein any training sample comprises a sample object and its label information, and the number of training samples in the current batch of training subsets determines the number of cycles of repeating the steps of selecting a training sample and updating the sampling frequency; determining an initial object recognition model; recognizing the type of each sample object in the current batch of training subsets by using the initial object recognition model; adjusting the initial object recognition model according to the type of each sample object obtained by using the initial object recognition model, the label information of the corresponding sample object in the current batch of training subsets, and the sample weight value of the corresponding sample object to obtain a final object recognition model, wherein the object recognition model is used to recognize the type of a target object and output probability information of the target object belonging to a certain type, and the target object is multimedia information, which comprises image, video, voice data or text; updating the sample weight value of each sample object according to the type of each sample object obtained by using the initial object recognition model and the label information of the corresponding sample object in the current batch of training subsets to obtain an updated sample weight value of each sample object; returning to the steps of grouping the current batch of training subsets and training the object recognition model according to the updated sample weight value.
2. The method of claim 1, wherein, The method of determining the sample weight value of each training sample in each sample subset of different types comprises the following steps: determining difficult samples and simple samples in all training samples of a sample subset of a certain type; determining that the sample weight value of the difficult sample is greater than a threshold value and the sample weight value of the simple sample is less than another threshold value.
3. The method of claim 1, wherein, The method of selecting at least one training sample from a sample subset of a certain type according to the sample weight value of each training sample in the sample subset comprises the following steps: corresponding to a unit part of a preset roulette, rotating the preset roulette, and selecting a training sample corresponding to the sample weight value pointed by a pointer in the preset roulette; or, regarding the sample weight values of all training samples in a sample subset as a vector, and obtaining at least one training sample with the largest sample weight value by using a binary search method.
4. The method of claim 1, wherein, The type of each sample object obtained according to the object recognition initial model and the annotation information and the sample weight value of the corresponding sample object in the current batch of training subsets are used to adjust the object recognition initial model, and specifically include: The type of each sample object obtained according to the object recognition initial model and the annotation information and the sample weight value of the corresponding sample object in the current batch of training subsets are used to calculate a loss function related to the object recognition initial model; The parameter value in the object recognition initial model is adjusted according to the loss function.
5. The method of claim 4, wherein, Before the object recognition model is trained according to the current batch of training subsets, the method further includes: The type weight value of each type of sample subset is determined respectively; The type of each sample object obtained according to the object recognition initial model and the annotation information and the sample weight value of the corresponding sample object in the current batch of training subsets are used to calculate a loss function related to the object recognition initial model, and specifically include: The type of each sample object obtained according to the object recognition initial model and the annotation information and the sample weight value and the type weight value of the corresponding sample object in the current batch of training subsets are used to calculate a loss function related to the object recognition initial model.
6. The method of claim 5, wherein, The type weight value of each type of sample subset is determined respectively, and specifically includes: The number of training samples included in a type of sample subset is inverted to obtain an inverse value, and the inverse value is used as the type weight value of the type of sample subset; Alternatively, the ratio of the inverse of the number of training samples included in a type of sample subset to the sum of the inverses of the number of training samples in each type of sample subset is calculated, and the function calculation value of the ratio is used as the type weight value of the type of sample subset.
7. The method of claim 5, wherein, The method further includes: After one round of training of the object recognition model is performed on all batch training subsets corresponding to the training sample set, the sample weight values of all training samples in the training sample set are updated to obtain updated sample weight values of all training samples; For the updated sample weight values, the steps of forming a training subset of a current batch and training an object recognition model are returned.
8. A model training system, comprising: It includes: A sample acquisition unit is configured to acquire a training sample set, and the training sample set includes multiple types of sample subsets; A sample weight unit is configured to determine the sample weight value of each training sample in each type of sample subset respectively; A sample selection unit is configured to assign a corresponding sampling frequency to each training sample in a sample subset according to the sample weight value of each training sample in the sample subset; and select a training sample with the largest sample weight value from the sample subset according to the sampling frequency of each training sample. updating the sampling times of the selected training samples; and performing the steps of selecting training samples and updating the sampling times cyclically to obtain at least one training sample corresponding to the one sample subset, and grouping at least one training sample selected from the plurality of types of sample subsets into a current batch of training subsets, any training sample including a sample object and its annotation information, wherein the number of training samples in the current batch of training subsets determines the number of cycles of performing the steps of selecting training samples and updating the sampling times cyclically; a training unit configured to determine an initial object recognition model, recognize the type of each sample object in the current batch of training subsets by using the initial object recognition model, and adjust the initial object recognition model according to the type of each sample object obtained by using the initial object recognition model and the annotation information and sample weight value of the corresponding sample object in the current batch of training subsets, to obtain a final object recognition model, wherein the object recognition model is used to recognize the type of a target object and output probability information of the target object belonging to a certain type, the target object is multimedia information, and the multimedia information includes images, videos, voice data, or text; updating the sample weight value of each sample object according to the type of each sample object obtained by using the initial object recognition model and the annotation information of the corresponding sample object in the current batch of training subsets, to obtain an updated sample weight value of each sample object; returning to the steps of grouping the current batch of training subsets and training the object recognition model according to the updated sample weight value.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a plurality of computer programs, and the computer programs are adapted to be loaded and executed by the processor to implement the model training method according to any one of claims 1 to 7.
10. A terminal device, comprising: The computer readable storage medium stores a plurality of computer programs, and the computer programs are adapted to be loaded and executed by the processor to implement the model training method according to any one of claims 1 to 7. The computer readable storage medium stores a plurality of computer programs, and the computer programs are adapted to be loaded and executed by the processor to implement the model training method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Target detection model training method and device, electronic equipment and storage medium
CN111832614A
Data sample balanced distribution method and device and storage medium
CN111860568A