A method and device for model training
By extracting the relevant terms from the basic loss function as derivative functions, and determining the initial function based on the adjustment gradient required during the model training process, the actual loss function used to train the model is solved, and the problem of the impact of redundant terms in model training is improved.
Patent Information
- Application Number
- CN202111283429.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2041-11-01
AI Technical Summary
In the prior art, there are redundant terms in the loss function during model training, resulting in a decrease in model training efficiency.
By extracting the relevant terms used to converge the model parameters from the basic loss function as derivative function, and determining the initial function based on the adjustment gradient required during the model training process, the actual loss function is constructed for training the model.
Effectively remove the influence of redundant terms in the basic loss function, and improve the efficiency of model training.
Smart Images

Figure CN114118218B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and device for model training. Background Art
[0002] With the continuous development of computer technology, artificial intelligence has made great progress and has been applied to various fields, such as image recognition, face verification, unmanned driving, etc. Among them, deep neural networks have developed rapidly in recent years and have been favored for their good results. Therefore, they have become one of the most important branches of artificial intelligence.
[0003] Before a model built based on a deep neural network can be used in practice, it needs to be trained with a certain number of samples. During the training process, it is often necessary to introduce a suitable loss function to measure the degree of model training. However, in practical applications, there are often certain redundant terms in the loss function, which reduces the efficiency of model training under the influence of these redundant terms.
[0004] Therefore, how to obtain a reasonable loss function during model training to improve the efficiency of model training is an urgent problem to be solved. Summary of the invention
[0005] This specification provides a model training method and device to partially solve the above-mentioned problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This manual provides a model training method, including:
[0008] The terminal obtains the basic loss function required by the model to be trained;
[0009] Determining relevant items for converging model parameters in the model to be trained from the basic loss function;
[0010] Extract the related term from the basic loss function as a derivative function;
[0011] Determine an initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determine an actual loss function for training the model to be trained based on the initial function;
[0012] The model to be trained is trained using the actual loss function and the acquired training samples.
[0013] Optionally, obtain the basic loss function required for the model to be trained, including:
[0014] If it is determined that the model to be trained is used for image matching, determining a basic loss function consistent with the image matching;
[0015] Determining relevant items for converging model parameters in the model to be trained from the basic loss function specifically includes:
[0016] If it is determined that minimizing the number of images irrelevant to the target image identified by the model to be trained is the optimization goal and the model to be trained is trained, then a related term representing the number of images irrelevant to the target image is determined from the basic loss function as a related term for converging the model parameters in the model to be trained, and the target image is the image input into the model to be trained.
[0017] Optionally, according to the adjustment gradient required in the training process of the model to be trained, determining the initial function obtained based on the derivative function specifically includes:
[0018] The initial function is determined by taking the method of solving the original function of the derivative function as a constraint condition and according to the adjustment gradient required in the training process of the model to be trained.
[0019] Optionally, the initial function is determined by taking the original function of the derivative function as a constraint condition and adjusting the gradient required in the training process of the model to be trained, specifically including:
[0020] Determine the model parameter adjustment rate during the training process of the model to be trained according to the adjustment gradient required during the training process of the model to be trained, and determine the original function obtained by solving the derivative function;
[0021] If it is determined that the numerical change rate of the original function matches the model parameter adjustment rate, the original function is used as the initial function. If it is determined that the numerical change rate of the original function does not match the model parameter adjustment rate, the derivative function is adjusted to obtain an adjusted derivative function in which the numerical change rate of the solved original function matches the model parameter adjustment rate. The initial function is obtained by solving the original function of the adjusted derivative function.
[0022] Optionally, the initial function is determined by taking the original function of the derivative function as a constraint condition and adjusting the gradient required in the training process of the model to be trained, specifically including:
[0023] Determine the model parameter adjustment rate during the training process of the model to be trained according to the adjustment gradient required during the training process of the model to be trained, and determine the original function obtained by solving the derivative function;
[0024] If it is determined that the numerical change rate of the original function matches the model parameter adjustment rate, the original function is used as the initial function; if it is determined that the numerical change rate of the original function does not match the model parameter adjustment rate, the original function is adjusted to obtain an adjusted original function whose numerical change rate matches the model parameter adjustment rate, and the adjusted original function is used as the initial function.
[0025] Optionally, before determining an initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determining an actual loss function for training the model to be trained based on the initial function, the method further includes:
[0026] The model to be trained is trained according to the basic loss function to obtain a numerical correspondence between a function value of the basic loss function and the related items when training the model to be trained;
[0027] According to the corresponding relationship, the adjustment gradient required in the training process of the model to be trained is determined.
[0028] Optionally, the basic loss function includes: a loss function of average accuracy AP.
[0029] This specification provides a model training device, including:
[0030] The acquisition module is used to obtain the basic loss function required by the model to be trained;
[0031] A determination module, used to determine relevant items for converging model parameters in the model to be trained from the basic loss function;
[0032] An extraction module, used for extracting the related term from the basic loss function as a derivative function;
[0033] A solution module, used to determine an initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determine an actual loss function for training the model to be trained based on the initial function;
[0034] The training module is used to train the model to be trained through the actual loss function and the acquired training samples.
[0035] This specification provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned model training method is implemented.
[0036] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned model training method when executing the program.
[0037] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0038] In the model training method provided in this specification, the terminal can obtain the basic loss function required for the model to be trained, and determine the relevant items used to converge the model parameters in the model to be trained from the basic loss function. Then, the terminal can extract the relevant items from the basic loss function as the derivative function. The terminal can determine the initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and based on the initial function, determine the actual loss function for training the model to be trained, and then perform model training on the model to be trained through the actual loss function and the acquired training samples.
[0039] It can be seen from the above method that since the terminal can use the relevant terms used to converge the model parameters in the model to be trained from the basic loss function required by the model to be trained as the derivative function, and obtain the actual loss function used to train the model to be trained by solving the derivative function, this can effectively remove the adverse effects of redundant terms in the basic loss function on the model training process, thereby effectively improving the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0041] Figure 1 A flowchart of a model training method provided in this specification;
[0042] Figure 2A and 2B A function graph of a related item provided in this specification, and a schematic diagram of the change rate of model parameter adjustment required by a model to be trained during the training process;
[0043] Figure 3 A schematic diagram of the change of the model parameter adjustment rate required for a model to be trained during the training process provided in this specification;
[0044] Figure 4 A schematic diagram of a model training device provided in this manual;
[0045] Figure 5A method corresponding to the Figure 1 Schematic diagram of an electronic device. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0047] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0048] Figure 1 A flow chart of a model training method provided in this specification includes the following steps:
[0049] S101: The terminal obtains the basic loss function required for the model to be trained.
[0050] In practical applications, before training the model, it is necessary to construct a corresponding loss function to determine the deviation between the model output value and the label value during the model training process, and to adjust the model parameters contained in the model. Therefore, in the process of model training, the terminal needs to first obtain the basic loss function required by the model to be trained.
[0051] The basic loss function required by the model to be trained obtained by the terminal is the original loss function used to train the model to be trained. The basic loss function can be a loss function that matches the model to be trained and pre-queried from a preset loss function library during the model training process by the terminal, or a loss function input by the user into the terminal for training the model to be trained.
[0052] It should be noted that the execution entity used to perform model training can be a terminal such as a desktop computer, a laptop computer, or a server. For the sake of ease of description, the model training method provided in this specification is described in detail using the terminal as an example below.
[0053] S102: Determine relevant items for converging model parameters in the model to be trained from the basic loss function.
[0054] After obtaining the above basic loss function, the terminal can further identify each item in the basic loss function to determine the relevant items used to converge the model parameters in the model to be trained. Specifically, the basic loss function contains multiple function items, some of which can play a role in the convergence of model parameters during model training, while some function items do not play an effective role in the convergence of model parameters during model training. Furthermore, the models used in different business scenarios are different, and correspondingly, the relevant items corresponding to the models in different scenarios are also different.
[0055] For example, a user sends a picture to the server through the client installed on the smartphone he holds. The server can use the trained image matching model to match the picture with the product pictures corresponding to each product, obtain the product pictures that match the picture, and then return the product information corresponding to these product pictures to the client to display it to the user.
[0056] As can be seen from the above examples, in the business scenario of image matching, the image matching model may match multiple images that match the image to be matched (hereinafter referred to as the target image for ease of explanation). In the training stage of the image matching model, the multiple images matched by the image matching model may not all match the target image, that is, some images may not match the target image. From the perspective of the training goal of the image matching model, it is hoped that the images matched by the image matching model will match the target image as much as possible. Therefore, for the basic loss function used by the image matching model, the function term included in the basic loss function for indicating the number of images matched by the image matching model that do not match the target image is the related term used to converge the model parameters in the image matching model during the model training stage. The function term included in the basic loss function for indicating the number of images matched by the image matching model that match the target image is a redundant term that does not play an effective role in converging the model parameters in the image matching model. It is shown in the following formula:
[0057]
[0058] The above formula is the loss function of the average accuracy (AP). In this loss function formula, S p It is used to represent the image preceding image i that matches the target image, R(i, S p ) is used to represent the number of images preceding image i that match the target image, S n It is used to represent the image before image i that does not match the target image, R(i, S n) is used to represent the number of images that do not match the target image before image i. The so-called image i can be understood as, if the multiple images matched by the image matching model that match the target image are regarded as an image sequence (some of the images in this image sequence actually match the target image, and some do not actually match the target image), then the image at the i-th position in the image sequence is image i.
[0059] Because in the process of training the image matching model, it is hoped that the number of images matched by the image matching model that do not match the target image is as small as possible, so, it can be seen from the above formula that R(i, S p ) is actually a redundant term in the image matching model training process. n ) is a related item in the image matching model training process. In this way, in the subsequent process, the terminal can extract the related item from the basic loss function to reconstruct the loss function for training the image matching model around the related item.
[0060] S103: Extract the related term from the basic loss function as a derivative function.
[0061] S104: Determine an initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determine an actual loss function for training the model to be trained based on the initial function.
[0062] After the terminal extracts the relevant items from the above basic loss function, it can use the relevant items as derivative functions and solve the original function to generate the actual loss function for model training based on the actual training requirements of the model to be trained.
[0063] Since the original function is solved based on the related terms, the original function does not contain redundant terms in the previous basic loss function that have no effective effect on the convergence of model parameters during model training. Therefore, the actual loss function constructed for model training under this circumstance can effectively reduce the amount of calculation in the model training process compared to the original basic loss function, thereby significantly improving the efficiency of model training.
[0064] In this specification, the terminal can solve the original function of the derivative function as a constraint condition, and determine the initial function according to the adjustment gradient required during the training process of the model to be trained. Among them, the initial function can be obtained based on the following different methods:
[0065] One is to adjust the derivative first, and then solve the original function of the adjusted derivative to get the corresponding initial function; one is to solve the original function of the derivative first, and then adjust the original function to get the initial function; one is to adjust the derivative first, and then solve the original function of the adjusted derivative, and finally adjust the solved original function to get the initial function. The above methods will be introduced below.
[0066] Specifically, corresponding to the first method, the terminal can first determine the model parameter adjustment rate of the model to be trained during the training process according to the adjustment gradient required during the training process of the model to be trained, and determine the original function obtained by solving the above derivative function. If the terminal determines that the numerical change rate of the original function matches the determined model parameter adjustment rate, the original function can be directly used as the initial function mentioned above, and in the subsequent process, based on the initial function, the actual loss function used to train the model to be trained is obtained.
[0067] If the terminal determines that the numerical change rate of the original function does not match the determined model parameter adjustment rate, the above-mentioned derivative function can be adjusted to obtain an adjusted derivative function in which the solved original function matches the above-mentioned model parameter adjustment rate in terms of numerical change rate, and then the above-mentioned initial function is obtained by solving the original function of the adjusted derivative function.
[0068] The following examples will be used to explain in detail: Figure 2A and Figure 2B shown.
[0069] Figure 2A and 2B A function graph of a related item provided in this specification, and a schematic diagram of the change in the model parameter adjustment rate required for a model to be trained during the training process.
[0070] Assume that the relevant items required in the training process of the model to be trained determined from the AP loss function are:
[0071]
[0072] This correlation term can be understood as the redundant term R(i, S p ) is removed, and judging from this correlation term alone, the correlation term is monotonically decreasing as a whole, and judging from the function graph of this correlation term (such as Figure 2A (As shown in the function graph), the change range of the function value in the early stage is relatively large, while the change range of the function value in the later stage is relatively small.
[0073] If it is determined that the model parameter adjustment rate required by the model to be trained during the training process is, the related term R(i, S n) gradually increases, the growth rate of the loss value of the loss function required for the training model will gradually decrease, and the model parameter adjustment rate will also gradually decrease, such as Figure 2B As shown, from Figure 2B It can be seen from the schematic diagram of the change of the model parameter adjustment rate shown in that the numerical change rate of the original function of the related item should match the model parameter adjustment rate, so the terminal can directly solve the original function of the related item (i.e., the derivative function) and use the obtained original function as the obtained initial function, as shown in the following formula:
[0074]
[0075] It should be pointed out that the above formula is not the initial function obtained, but the actual loss function for training the training model finally determined on the basis of the initial function, log(1+R(i, S n ))This part can be understood as the initial function obtained.
[0076] Furthermore, if it is determined that the model to be trained needs the relevant item R(i, S n ) gradually increases, the model parameter adjustment rate further accelerates and slows down, then it is determined that the value change rate of the original function corresponding to the above-mentioned related item does not match the model parameter adjustment rate required in the training process of the model to be trained. In this case, the terminal can adjust this related item, that is, the derivative function, to obtain the adjusted derivative function, and solve the original function of the adjusted derivative function, such as Figure 3 shown.
[0077] Figure 3 A schematic diagram of the change in the model parameter adjustment rate required for a model to be trained during the training process provided in this specification.
[0078] from Figure 3 It can be seen that compared with Figure 2B , related terms R(i, S n ) gradually increases, the growth rate of the loss value of the loss function required for the training model will further accelerate the reduction, and the model parameter adjustment rate will also gradually decrease. In this case, the required actual loss function can be obtained by adjusting the derivative function.
[0079] The specific formula is as follows:
[0080]
[0081] It can be seen from this formula that by adding an exponential term to the denominator, the numerical change rate of the original function can be significantly improved. Therefore, the above formula can be used to produce an original function that matches the model parameter adjustment rate, namely:
[0082]
[0083] This formula is not the initial function obtained, but the actual loss function that is ultimately used to train the model to be trained, determined on the basis of the initial function. In this formula, This term can be understood as the original function obtained by solving the adjusted derivative function.
[0084] For the second method mentioned above, the terminal can also first determine the model parameter adjustment rate during the training process of the model to be trained according to the adjustment gradient required during the training process of the model to be trained, and determine the original function obtained by solving the above derivative function. Then, if the terminal determines that the numerical change rate of the original function matches the model parameter adjustment rate, the original function can be used as the initial function. If it is determined that the numerical change rate of the original function does not match the model parameter adjustment rate, the original function can be adjusted to obtain an adjusted original function whose numerical change rate matches the model parameter adjustment rate, and the adjusted original function is used as the initial function.
[0085] The matching situation is the same as the matching situation in the first situation, so it will not be described in detail here. As for the non-matching situation, an example will be used below to further illustrate it.
[0086] Assume that the original function The rate of change of the value of is opposite to the rate of adjustment of the model parameters required during the training process of the model to be trained. According to actual needs, a minus sign can be added in front of this original function, and S can be added before the minus sign. p This item, that is, This adjusted original function is used as the above initial function. Furthermore, the actual loss function that can be finally obtained through this initial function is:
[0087]
[0088] As for the third method, it is obtained by combining the above two methods, that is, the above derivative function is adjusted according to the adjustment gradient required in the training process of the model to be trained to obtain the adjusted derivative function, and the original function of the adjusted derivative function is solved, and then based on the actual model training requirements, the original function is further adjusted to obtain the above initial function. This process can be combined with the examples listed above, and no detailed examples will be given here.
[0089] It should be noted that before the terminal determines to obtain the actual loss function, it needs to first train the model to be trained according to the basic loss function to obtain the corresponding relationship between the function value of the basic loss function and the above-mentioned related items in numerical value when training the model to be trained. This corresponding relationship can characterize the model parameter adjustment rate required for the model to be trained to a certain extent. For example, the terminal can train the model to be trained based on the basic loss function according to a preset sample set to fit a function curve, which is used to represent the corresponding relationship between the function value of the basic loss function and the above-mentioned related items in numerical value. In this way, the terminal can determine the adjustment gradient required in the training process of the model to be trained based on this corresponding relationship.
[0090] S105: Train the model to be trained using the actual loss function and the acquired training samples.
[0091] After constructing the actual loss function in the above manner, the terminal can obtain training samples for training the model to be trained, and train the model to be trained through the obtained training samples. For different actual scenarios, the model to be trained is different in function, and accordingly, the obtained training samples are also different.
[0092] For example, in an image recognition scenario, the model to be trained is used to identify targets from images, so the acquired training samples may be sample images with label information, where the label information is used to indicate which targets are specifically contained in the sample images. For another example, in the control process of an unmanned driving device, the model to be trained may be used to output driving decisions for controlling the unmanned driving device based on the sensor data used by the unmanned driving device itself and the surrounding environmental data. In this case, the acquired training samples may be historical sensor data and historical environmental data with label information of actual driving decisions.
[0093] It can be seen from the above content that the model training method provided in this specification can be applied to model training in a variety of scenarios, except that in different scenarios, the models differ in function, training samples, and model training requirements. And from the above method, it can be seen that since the terminal can use the relevant terms used to converge the model parameters in the model to be trained from the basic loss function required by the model to be trained as the derivative function, and by solving the derivative function, the actual loss function used to train the model to be trained can be obtained, which can effectively remove the adverse effects of redundant terms in the basic loss function on the model training process, thereby effectively improving the efficiency of model training.
[0094] From the above example, we can further see that the AP loss function has a redundant term, which is used to represent the number of images R(i, S i) that match the target image before image i. p ), if the value of the redundant term is too large (especially in the later stage of model training), the AP loss function almost becomes a constant function, which leads to the failure of the AP loss function to achieve good optimization. As for the actual loss function determined by the model training method provided in this specification, more attention is paid to the number of images R(i, S) used to represent the image i that does not match the target image. n ), which means that even if there is a slight difference in the number of images that do not match the target image in adjacent rounds of training, the actual loss function obtained can effectively reflect the gradient changes required in the model training process, thereby achieving better optimization of the training model.
[0095] In this specification, in combination with the actual model training requirements, a variety of actual loss functions for training the model to be trained can be constructed through the above method. Several actual loss functions obtained through the above method are listed below.
[0096]
[0097]
[0098] The above two formulas are the actual loss functions constructed according to the actual training requirements, which can be collectively referred to as L I Loss function( and All can be classified as L I ), it can be seen from these two formulas that L I The loss function is constructed based on the AP loss function. I The loss function only contains the relevant items in the AP loss function (basic loss function). In other words, the redundant items that originally appeared in the basic loss function are not in the L loss function. I This is reflected in the loss function.
[0099] Furthermore, L I The loss function is an increasing function, so using L I When the loss function is used for model training, samples with large loss function values will be given larger gradients. Therefore, L I The loss function actually optimizes the aggregation of all relevant samples.
[0100] Compared with these two actual loss functions, the L mentioned above D ( and All can be classified as L D ) can give smaller gradients to samples with large loss function values, so L D The loss function only quickly updates samples that are sufficiently confident, and slowly updates samples that may belong to another center or noise. In other words, if images that do not match the target image are called counterexamples, the larger the number of counterexamples, the smaller the increase in the corresponding loss value, and the smaller the number of counterexamples, the greater the decrease in the corresponding loss value. Therefore, when there are multiple subcategories of data that need to be distinguished in the actual application scenario, use L D By training the model to be trained, the trained model can accurately distinguish these subcategories.
[0101] The above is a method for model training provided in one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding model training device, such as Figure 4 shown.
[0102] Figure 4 A schematic diagram of a model training device provided in this specification includes:
[0103] An acquisition module 401 is used to acquire a basic loss function required by the model to be trained;
[0104] A determination module 402 is used to determine, from the basic loss function, relevant items for converging model parameters in the model to be trained;
[0105] An extraction module 403, used to extract the related term from the basic loss function as a derivative function;
[0106] A solution module 404 is used to determine an initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determine an actual loss function for training the model to be trained based on the initial function;
[0107] The training module 405 is used to train the model to be trained by using the actual loss function and the acquired training samples.
[0108] Optionally, the acquisition module 401 is specifically used to, if it is determined that the model to be trained is used for image matching, determine a basic loss function consistent with the image matching;
[0109] The determination module 402 is specifically used to, if it is determined that minimizing the number of images irrelevant to the target image identified by the model to be trained is the optimization goal, train the model to be trained, then determine the relevant terms representing the number of images irrelevant to the target image from the basic loss function as the relevant terms for converging the model parameters in the model to be trained, and the target image is the image input into the model to be trained.
[0110] Optionally, the solving module 404 is specifically used to determine the initial function according to the adjustment gradient required in the training process of the model to be trained, taking the method of solving the original function of the derivative function as a constraint condition.
[0111] Optionally, the solution module 404 is specifically used to determine the model parameter adjustment rate during the training process of the model to be trained, and determine the original function obtained by solving the derivative function according to the adjustment gradient required during the training process of the model to be trained; if it is determined that the numerical change rate of the original function matches the model parameter adjustment rate, the original function is used as the initial function; if it is determined that the numerical change rate of the original function does not match the model parameter adjustment rate, the derivative function is adjusted to obtain an adjusted derivative function in which the numerical change rate of the solved original function matches the model parameter adjustment rate; the initial function is obtained by solving the original function of the adjusted derivative function.
[0112] Optionally, the solution module 404 is specifically used to determine the model parameter adjustment rate during the training process of the model to be trained, and determine the original function obtained by solving the derivative function according to the adjustment gradient required during the training process of the model to be trained; if it is determined that the numerical change rate of the original function matches the model parameter adjustment rate, the original function is used as the initial function; if it is determined that the numerical change rate of the original function does not match the model parameter adjustment rate, the original function is adjusted to obtain an adjusted original function whose numerical change rate matches the model parameter adjustment rate, and the adjusted original function is used as the initial function.
[0113] Optionally, the solution module 404 is used to train the model to be trained according to the basic loss function before determining the initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determining the actual loss function for training the model to be trained based on the initial function, so as to obtain the numerical correspondence between the function value of the basic loss function and the related items when training the model to be trained; and determine the adjustment gradient required in the training process of the model to be trained according to the corresponding relationship.
[0114] Optionally, the basic loss function includes: a loss function of average accuracy AP.
[0115] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A model training method is provided.
[0116] This manual also provides Figure 5 The one shown corresponds to Figure 1 A schematic diagram of the electronic device. Figure 5 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The model training method described. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0117] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0118] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0119] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0120] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0121] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0122] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0123] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0125] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0126] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0127] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0128] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0129] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0130] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0131] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0132] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A model training method, characterized in that: include: The terminal obtains a basic loss function required by a model to be trained, where the model to be trained is used to identify a target object from an image; Determining relevant items for converging model parameters in the model to be trained from the basic loss function; Extract the related term from the basic loss function as a derivative function; Determine an initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determine an actual loss function for training the model to be trained based on the initial function; The model to be trained is trained by using the actual loss function and the acquired training samples, wherein the acquired training samples are sample images with label information.
2. The method according to claim 1, characterized in that Get the basic loss function required for the model to be trained, including: The model to be trained is used for image matching, and a basic loss function consistent with the image matching is determined; Determining relevant items for converging model parameters in the model to be trained from the basic loss function specifically includes: if it is determined that minimizing the number of images irrelevant to the target image identified by the model to be trained is the optimization goal and the model to be trained is trained, then determining relevant items for representing the number of images irrelevant to the target image from the basic loss function as relevant items for converging model parameters in the model to be trained, and the target image is the image input into the model to be trained.
3. The method according to claim 1, characterized in that According to the adjustment gradient required in the training process of the model to be trained, determining the initial function obtained based on the derivative function specifically includes: The initial function is determined by taking the method of solving the original function of the derivative function as a constraint condition and according to the adjustment gradient required in the training process of the model to be trained.
4. The method according to claim 3, characterized in that Taking the method of solving the original function of the derivative function as a constraint condition, the initial function is determined according to the adjustment gradient required in the training process of the model to be trained, specifically including: Determine the model parameter adjustment rate during the training process of the model to be trained according to the adjustment gradient required during the training process of the model to be trained, and determine the original function obtained by solving the derivative function; If it is determined that the numerical change rate of the original function matches the model parameter adjustment rate, the original function is used as the initial function. If it is determined that the numerical change rate of the original function does not match the model parameter adjustment rate, the derivative function is adjusted to obtain an adjusted derivative function in which the numerical change rate of the solved original function matches the model parameter adjustment rate. The initial function is obtained by solving the original function of the adjusted derivative function.
5. The method according to claim 3, characterized in that Taking the method of solving the original function of the derivative function as a constraint condition, the initial function is determined according to the adjustment gradient required in the training process of the model to be trained, specifically including: Determine the model parameter adjustment rate during the training process of the model to be trained according to the adjustment gradient required during the training process of the model to be trained, and determine the original function obtained by solving the derivative function; If it is determined that the numerical change rate of the original function matches the model parameter adjustment rate, the original function is used as the initial function. If it is determined that the numerical change rate of the original function does not match the model parameter adjustment rate, the original function is adjusted to obtain an adjusted original function whose numerical change rate matches the model parameter adjustment rate, and the adjusted original function is used as the initial function.
6. The method according to claim 1, characterized in that Before determining an initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determining an actual loss function for training the model to be trained based on the initial function, the method further includes: The model to be trained is trained according to the basic loss function to obtain a numerical correspondence between a function value of the basic loss function and the related items when training the model to be trained; According to the corresponding relationship, the adjustment gradient required in the training process of the model to be trained is determined.
7. The method according to any one of claims 1 to 6, characterized in that: The basic loss function includes: the loss function of the average accuracy AP.
8. A device for model training, characterized in that: include: An acquisition module, used to acquire a basic loss function required by a model to be trained, wherein the model to be trained is used to identify a target object from an image; A determination module, used to determine relevant items for converging model parameters in the model to be trained from the basic loss function; An extraction module, used for extracting the related term from the basic loss function as a derivative function; A solution module, used to determine an initial function obtained based on the derivative function according to the adjustment gradient required in the training process of the model to be trained, and determine an actual loss function for training the model to be trained based on the initial function; The training module is used to train the model to be trained through the actual loss function and the acquired training samples, where the acquired training samples are sample images with label information.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Data processing model training method and device
CN111639684A
Training method and system, strategy generation method and system, computer device and storage medium
CN113298329A