Image recognition model training method and device, vehicle and electronic equipment

The dynamic generation function and backpropagation algorithm adjust the convolution kernel of the image recognition model, which solves the problem of fixed convolution kernel restricting feature extraction, and achieves higher accuracy and robustness.

CN120339664APending Publication Date: 2025-07-18BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410073824.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing image recognition model has limited its feature extraction ability at different locations due to the fixed-sized convolution kernel, which affects the generalization ability and feature extraction ability of the model.

Method used

By calculating the convolution kernel based on the training image and dynamic generation functions, updating the learning parameters using the backpropagation algorithm, dynamically adjusting the size and configuration of the convolution kernel until the model training is completed.

Benefits of technology

The feature extraction ability of the image recognition model is improved, the accuracy and robustness of the model are enhanced, the overfitting phenomenon is reduced, and the model's recognition ability of new images is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339664A_ABST
    Figure CN120339664A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition model training method and device, a vehicle and electronic equipment, based on a training image, a dynamic generation function is calculated, a convolution kernel in an image recognition model is configured according to a calculation result, and the dynamic generation function is composed of a generation function and learning parameters; inputting the training image into the image recognition model for training so as to update learning parameters of the dynamic generation function; and according to the loss function of the image recognition model, judging whether the image recognition model is trained or not. Compared with the prior art, the method has the advantages that the convolution kernel of the image recognition model is configured according to the dynamic generation function, so that the convolution kernel parameters can be adaptively adjusted according to different task requirements; therefore, the image recognition model can be helped to better learn feature information in the image, the image feature extraction capability of the image recognition model is improved, and the accuracy and robustness of the image recognition model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image recognition technologies, and in particular, to a method and apparatus for training an image recognition model, a vehicle, and an electronic device. Background Art

[0002] With the development of technologies such as artificial intelligence, the application of machine vision in the field of autonomous driving has become increasingly widespread. Through machine vision algorithms, high-precision target detection and recognition can be achieved, thereby helping autonomous driving vehicles accurately judge the surrounding environment and avoid collisions and accidents. The essence of machine vision is based on the processing of input videos / images, and the level of image feature extraction ability greatly affects the quality of machine vision algorithms.

[0003] As a part of machine vision algorithms, the size of the convolution kernel affects the generalization ability and feature extraction ability of the model. A convolution kernel with a fixed size is fixed in the spatial dimension, which may cause the image recognition model to not respond equally to input features at different positions, thus affecting the generalization ability of the model; a convolution kernel with a fixed size is the same for each channel, which limits the feature extraction ability of the convolution kernel to adapt to features at different spatial positions. Therefore, optimizing the feature extraction ability of the image recognition model by changing the size of the convolution kernel is crucial for improving the performance of the image recognition model. Summary of the Invention

[0004] The present disclosure provides a method and apparatus for training an image recognition model, a vehicle, and an electronic device. Its main purpose is to improve the image feature extraction ability of the image recognition model, and further improve the accuracy and robustness of the image recognition model.

[0005] According to a first aspect of the present disclosure, there is provided a method for training an image recognition model, including:

[0006] Calculating a convolution kernel based on a training image and a dynamic generation function; the training image is used to train the image recognition model, the dynamic generation function is used to generate the convolution kernel, and the dynamic generation function consists of a generation function and learning parameters;

[0007] Inputting the training image into the image recognition model configured with the convolution kernel, and performing a convolution operation on the training image through the convolution kernel to obtain a feature image corresponding to the training image;

[0008] In the case where it is determined by a loss function in the image recognition model that the image recognition model has not completed training, performing backpropagation on the feature image, and updating the learning parameters in the dynamic generation function according to the backpropagation result;

[0009] Reconfigure the convolutional kernels of the image recognition model based on the updated dynamically generated function, and train the image recognition model with reconfigured convolutional kernels based on the training images until the training of the image recognition model is completed.

[0010] In some embodiments, calculating the convolutional kernel based on the training images and the dynamically generated function includes:

[0011] Calculate the configuration information of the convolutional kernel by calculating the dynamically generated function based on the feature data of the training images and the task requirements corresponding to the image recognition model;

[0012] Configure the convolutional kernels in the image recognition model using the configuration information.

[0013] In some embodiments, backpropagating the feature images and updating the learning parameters in the dynamically generated function according to the backpropagation results includes:

[0014] Calculate the gradient of the loss function with respect to each layer in the image recognition model based on the feature images;

[0015] Update the learning parameters of the dynamically generated function using the calculated gradient information.

[0016] In some embodiments, calculating the gradient of the loss function with respect to each layer in the image recognition model based on the feature images includes:

[0017] Perform gradient calculation on the feature images to obtain the gradients of the feature images;

[0018] Perform a transposed convolution operation on the gradients of the feature images and the convolutional kernels, and calculate the gradient of the loss function with respect to each layer in the image recognition model layer by layer according to the calculation results of the transposed convolution.

[0019] In some embodiments, updating the learning parameters of the dynamically generated function using the calculated gradient information includes:

[0020] Obtain the preset learning rate of the image recognition model;

[0021] Calculate and update the learning parameters of the dynamically generated function based on the preset learning rate and the gradient information.

[0022] According to the second aspect of the present disclosure, there is provided an apparatus for training an image recognition model, including:

[0023] A configuration unit for calculating a convolution kernel based on a training image and a dynamic generation function; the training image is used to train the image recognition model, and the dynamic generation function is used to generate a convolution kernel based on it. The dynamic generation function consists of a generation function and learning parameters;

[0024] A first training unit for inputting the training image into the image recognition model with a configured convolution kernel, performing a convolution operation on the training image through the convolution kernel, and obtaining a feature image corresponding to the training image;

[0025] An update unit for, in the case where it is determined by the loss function in the image recognition model that the image recognition model has not been trained yet, performing backpropagation on the feature image and updating the learning parameters in the dynamic generation function according to the backpropagation result;

[0026] A second training unit for reconfiguring the convolution kernel of the image recognition model based on the updated dynamic generation function and training the image recognition model with the reconfigured convolution kernel based on the training image until the image recognition model is trained completely.

[0027] In some embodiments, the configuration unit includes:

[0028] A first calculation module for calculating configuration information of the convolution kernel by calculating the dynamic generation function based on the feature data of the training image and the task requirements corresponding to the image recognition model;

[0029] A configuration module for configuring the convolution kernel in the image recognition model using the configuration information to obtain the configured image recognition model.

[0030] In some embodiments, the update unit includes:

[0031] A second calculation module for calculating the gradient of the loss function with respect to each layer in the image recognition model based on the feature image;

[0032] An update module for updating the learning parameters of the dynamic generation function using the calculated gradient information.

[0033] In some embodiments, the second calculation module is further used for:

[0034] Calculating the gradient of the feature image to obtain the gradient of the feature image;

[0035] Performing a transposed convolution operation on the gradient of the feature image and the convolution kernel, and calculating the gradient of the loss function with respect to each layer in the image recognition model layer by layer according to the calculation result of the transposed convolution.

[0036] In some embodiments, the updating module is further configured to:

[0037] Obtain a preset learning rate of the image recognition model;

[0038] Based on the preset learning rate and the gradient information, calculate and update the learning parameters of the dynamically generated function.

[0039] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0040] At least one processor; and

[0041] A memory communicatively connected to the at least one processor; wherein,

[0042] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the foregoing first aspect.

[0043] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.

[0044] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in the foregoing first aspect is implemented.

[0045] According to a sixth aspect of the present disclosure, there is provided a vehicle, and the vehicle includes the apparatus for training an image recognition model described in the foregoing second aspect, or the electronic device described in the foregoing third aspect, or the non-transitory computer-readable storage medium described in the foregoing fourth aspect.

[0046] The present disclosure provides a method and apparatus for training an image recognition model, a vehicle, and an electronic device. A convolution kernel is calculated based on a training image and a dynamic generation function. The training image is used to train the image recognition model, and the dynamic generation function is used to generate a convolution kernel. The dynamic generation function consists of a generation function and learning parameters. The training image is input into the image recognition model with the configured convolution kernel, and the training image is subjected to a convolution operation through the convolution kernel to obtain a feature image corresponding to the training image. In the case where it is determined by the loss function in the image recognition model that the image recognition model is not trained yet, the feature image is backpropagated, and the learning parameters in the dynamic generation function are updated according to the backpropagation result. The convolution kernel of the image recognition model is reconfigured based on the updated dynamic generation function, and the image recognition model with the reconfigured convolution kernel is trained based on the training image until the image recognition model is trained. Compared with the related art, the present disclosure can adaptively adjust the convolution kernel parameters according to different task requirements by configuring the convolution kernel of the image recognition model according to the dynamic generation function. Furthermore, it can help the image recognition model better learn the feature information in the image, improve the image feature extraction ability of the image recognition model, and further improve the accuracy and robustness of the image recognition model.

[0047] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0049] Figure 1 is a schematic flowchart of a method for training an image recognition model provided by an embodiment of the present disclosure;

[0050] Figure 2 is a schematic flowchart of another method for training an image recognition model provided by an embodiment of the present disclosure;

[0051] Figure 3 is a schematic structural diagram of an apparatus for training an image recognition model provided by an embodiment of the present disclosure;

[0052] Figure 4 is a schematic structural diagram of another apparatus for training an image recognition model provided by an embodiment of the present disclosure;

[0053] Figure 5 is a schematic block diagram of an exemplary electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0055] The method and apparatus for training an image recognition model, a vehicle, and an electronic device according to embodiments of the present disclosure will be described below with reference to the drawings.

[0056] Figure 1 The flowchart of a method for training an image recognition model provided by an embodiment of the present disclosure.

[0057] As Figure 1 shown, the method includes the following steps:

[0058] Step 101, calculating a convolution kernel based on a training image and a dynamic generation function; the training image is used to train the image recognition model, the dynamic generation function is used to generate a convolution kernel, and the dynamic generation function is composed of a generation function and learning parameters.

[0059] Different from traditional convolutional neural networks (CNNs), Involution convolution does not use a fixed convolution kernel (such as a 3x3 or 5x5 filter), but dynamically generates a convolution kernel that matches the size of the input image through a linear transformation function. This method can adaptively adjust the convolution kernel according to the image content, thereby capturing long-range dependencies when processing images and improving the feature expression ability to a certain extent. However, since a fixed linear transformation function is used when generating the convolution kernel, the fixed function expression limits the ability of the convolution kernel to be dynamically generated, and thus limits the feature extraction ability of the image recognition model.

[0060] In the embodiments of the present disclosure, the dynamic generation function is composed of a generation function and learning parameters. The generation function can be the default generation function of Involution convolution or a generation function created according to task requirements; the present disclosure does not limit this. When training the image recognition model, the parameters are learned and updated through the backpropagation algorithm; the learning parameters can be a learnable weight matrix or a learnable bias term; the specific form of the learning parameters needs to be determined according to specific task requirements and the input image, and the present disclosure does not limit this.

[0061] The function for generating the convolution kernel is a dynamic generation function, which can dynamically generate the convolution kernel according to the requirements of the task and the input images (including training images and images input during the model inference stage), so as to better capture image information and improve the feature extraction ability of the image recognition model.

[0062] Step 102: Input the training image into the image recognition model with the configured convolution kernel, and perform convolution operation on the training image through the convolution kernel to obtain the feature image corresponding to the training image.

[0063] In the embodiments of the present disclosure, after determining the size of the convolution kernel, training the image recognition model enables the model to better learn and understand the features in the training image, so as to more accurately identify and classify; the image recognition model can further optimize its parameters and the dynamic generation function, thereby improving the calculation efficiency and accuracy. By adjusting the parameters and structure of the model to make it more suitable for specific image recognition tasks, the performance and response speed of the model can be further improved; the training process enables the image recognition model to extract more generalizable features from the training image, thereby improving the model's recognition ability for new images.

[0064] Step 103: In the case that it is determined by the loss function in the image recognition model that the image recognition model has not completed training, perform backpropagation on the feature image, and update the learning parameters in the dynamic generation function according to the backpropagation result.

[0065] In the embodiments of the present disclosure, in the case that the model training fails to meet the standard, the backpropagation algorithm is used to process the feature image, and then the parameters in the dynamic generation function are adjusted. The backpropagation algorithm can accurately guide the image recognition model on how to adjust the learning parameters according to the gradient information of the loss function to reduce the loss and improve the accuracy; to improve the model's feature extraction ability for images. The learning parameters of the dynamic generation function can be dynamically adjusted according to the features of the training image, making the image recognition model more adaptable. This means that the image recognition model can better adapt to different types and features of images, thereby improving the recognition ability for images. The dynamic generation function can calculate and adjust the convolution kernel configuration in real time as the training data changes, so that the model can adapt to the continuously changing image features and task requirement parameters in real time. By continuously adjusting and optimizing the learning parameters of the dynamic generation function, the model can better generalize to new unseen image data. This helps to reduce the overfitting phenomenon and improve the generalization ability of the model.

[0066] Step 104: Reconfigure the convolution kernel of the image recognition model based on the updated dynamic generation function, and train the image recognition model with the reconfigured convolution kernel based on the training image until the image recognition model is trained.

[0067] In an embodiment of the present disclosure, a loss function is used to quantify the difference between the predicted value and the actual value of the model. In order to enable the learning ability of the image recognition model to meet the expectations, it is necessary to determine whether the image recognition model is trained completely according to the loss function. Through continuous training of the image recognition model, normally, the difference between the predicted value and the actual value of the image recognition model becomes smaller and smaller. According to the loss function of the image recognition model, the training situation of the model can be judged in a timely manner, so as to make adjustments and optimizations. This helps to improve the performance of the model and accelerate the training process.

[0068] During the training process of the model, the configuration of the convolutional kernel is continuously adjusted according to the current performance of the model and the information of the loss function, and the training is continued until the model reaches the preset accuracy rate or the number of training rounds. By reconfiguring the convolutional kernel, the model can continuously optimize its feature extraction ability according to the progress and requirements of the training. Continuously adjusting the configuration of the convolutional kernel during the training process can help the model avoid overfitting the training data, thereby improving the generalization ability of the model.

[0069] The present disclosure provides a method for training an image recognition model, calculating a convolutional kernel based on training images and a dynamic generation function; the training images are used to train the image recognition model, the dynamic generation function is used to generate a convolutional kernel, and the dynamic generation function consists of a generation function and learning parameters; inputting the training images into the image recognition model with the convolutional kernel configured, performing a convolutional operation on the training images through the convolutional kernel to obtain a feature image corresponding to the training images; in the case that it is judged by the loss function in the image recognition model that the image recognition model is not trained completely, performing backpropagation on the feature image, and updating the learning parameters in the dynamic generation function according to the backpropagation result; reconfiguring the convolutional kernel of the image recognition model based on the updated dynamic generation function, and training the image recognition model with the reconfigured convolutional kernel based on the training images until the image recognition model is trained completely. Compared with the related technology, the present disclosure can adaptively adjust the convolutional kernel parameters according to different task requirements by configuring the convolutional kernel of the image recognition model according to the dynamic generation function; furthermore, it can help the image recognition model better learn the feature information in the images, improve the image feature extraction ability of the image recognition model, and further improve the accuracy and robustness of the image recognition model.

[0070] To clearly illustrate the embodiments of the present disclosure, the present embodiment provides a schematic flowchart of another method for training an image recognition model.

[0071] As Figure 2 shown, the method includes the following steps:

[0072] Step 201: Calculate the configuration information of the convolutional kernel for the dynamic generation function based on the feature data of the training image and the task requirements corresponding to the image recognition model.

[0073] Step 202: Configure the convolutional kernels in the image recognition model using the configuration information to obtain the configured image recognition model.

[0074] Specifically, in Steps 201 to 202, by parameterizing the task requirements, task requirement parameters are obtained; the task requirement parameters define the output requirements and performance metrics of the model. For example, for a face recognition task, the task requirements may include recognizing faces at different angles, lighting conditions, and expressions. These parameters will guide the training and optimization of the model.

[0075] Configure the convolutional kernels in the image recognition model using the calculated configuration information. This process may involve adjusting parameters such as the size, stride, and padding of the convolutional kernels to more effectively extract image features during the convolution operation. Through efficient convolutional kernel configuration, the model can achieve better performance with limited computing resources. This helps reduce the time and resource consumption during training and inference.

[0076] Step 203: Input the training image into the image recognition model with configured convolutional kernels, and perform a convolution operation on the training image through the convolutional kernels to obtain the feature image corresponding to the training image.

[0077] Specifically, in Step 203, after determining the size of the convolutional kernel, train the image recognition model. The image recognition model can better learn and understand the features in the training image, thus more accurately recognizing and classifying; the image recognition model can further optimize its parameters and dynamic generation function, thereby improving the computing efficiency and accuracy. By adjusting the parameters and structure of the model to make it more suitable for a specific image recognition task, the performance and response speed of the model can be further improved; the training process enables the image recognition model to extract more generalizable features from the training image, thereby improving the model's recognition ability for new images.

[0078] Step 204: Calculate the gradient of the loss function with respect to each layer in the image recognition model based on the feature image.

[0079] As an implementable manner of the embodiments of the present disclosure, calculating the gradient of the loss function with respect to each layer in the image recognition model based on the feature image may include but is not limited to the following methods:

[0080] Step 2041: Calculate the gradient of the feature image to obtain the gradient of the feature image.

[0081] Step 2042: Perform a transposed convolution operation on the gradient of the feature image and the convolution kernel, and layer by layer calculate the gradient of the loss function with respect to each layer in the image recognition model according to the calculation result of the transposed convolution.

[0082] Specifically in step 204, based on the chain rule, calculate the gradient according to the output of the loss function and the dynamic generation function. For the convolutional layer, the backpropagation of the gradient requires the use of transposed convolution (or deconvolution) operations. This step passes the gradient of the output feature map back to the feature map of the training image, realizing the backpropagation of the gradient. The backpropagation algorithm utilizes the chain rule to calculate the gradient layer by layer. It starts from the output layer and propagates the gradient backward to each layer until reaching the input layer. In this way, we can calculate the gradient of the loss function with respect to the parameters of each layer.

[0083] By calculating the gradient and performing the transposed convolution operation, the backpropagation of the gradient is realized, enabling the error information to be transmitted from the output layer to the input layer. This helps to guide the weight update and optimization direction of the model during training. The use of the gradient descent algorithm enables the learning parameters to be updated according to the error gradient. This helps to adjust parameters such as the convolution kernel weights and biases in the model, enabling the model to gradually adapt to the task requirements during training.

[0084] Step 205: Update the learning parameters of the dynamic generation function by using the calculated gradient information.

[0085] As an implementable way of the embodiment of the present disclosure, the updating of the learning parameters of the dynamic generation function by using the calculated gradient information may include but is not limited to the following ways:

[0086] Step 2051: Obtain the preset learning rate of the image recognition model.

[0087] Step 2052: Calculate and update the learning parameters of the dynamic generation function based on the preset learning rate and the gradient information.

[0088] Specifically, in step 205, the parameters are updated according to the direction of the gradient, and the preset learning rate is used to control the step size of each update. The larger the preset learning rate, the larger the step size of each update, but it may cause the algorithm to not converge; the smaller the preset learning rate, the smaller the step size of each update, but it may cause the algorithm to converge too slowly. The gradient descent algorithm is called to process the calculated gradient. Gradient descent is an optimization algorithm that gradually reduces the loss function by updating the weights in the opposite direction of the gradient. According to the processing results of the gradient descent algorithm, the learning parameters of the dynamic generation function are updated. These parameters are adjusted by adjusting the weights and biases of the convolution kernel to optimize the performance of the model. By continuously updating the learning parameters, the model can converge to the optimal solution faster. This shortens the training time and improves the training efficiency. By optimizing the learning parameters, the model can better generalize to unseen image data. This helps to improve the generalization ability of the model, making it more stable and reliable in practical applications. The model that has undergone gradient back propagation and parameter optimization has strong robustness to noise and abnormal input. This helps to improve the robustness of the model, making it more adaptable and stable in practical applications.

[0089] Step 206, reconfigure the convolution kernel of the image recognition model based on the updated dynamic generation function, and train the image recognition model with the reconfigured convolution kernel based on the training image until the image recognition model training is completed.

[0090] Specifically in step 206, the loss function is used to quantify the difference between the predicted value and the actual value of the model. In order to make the learning ability of the image recognition model meet expectations, it is necessary to judge whether the image recognition model is trained based on the loss function. Through continuous training of the image recognition model, under normal circumstances, the difference between the predicted value and the actual value of the image recognition model is getting smaller and smaller. The loss function can be a cross entropy loss, a mean square error loss, etc., and the specific selection should be determined according to the specific task requirements. During the training process of the model, the configuration of the convolution kernel is continuously adjusted according to the current performance of the model and the information of the loss function, and the training is continued until the model reaches a preset accuracy or number of training rounds. By reconfiguring the convolution kernel, the model can continuously optimize its feature extraction capability according to the progress and needs of the training. Continuously adjusting the configuration of the convolution kernel during the training process can help the model avoid overfitting the training data, thereby improving the generalization ability of the model.

[0091] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers do not limit the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.

[0092] Corresponding to the above method for training an image recognition model, the present invention also provides an apparatus for training an image recognition model. Since the apparatus embodiments of the present invention correspond to the above method embodiments, details not disclosed in the apparatus embodiments may be referred to the above method embodiments and will not be elaborated herein.

[0093] Figure 3 It is a schematic structural diagram of an apparatus for training an image recognition model provided by an embodiment of the present disclosure. As Figure 3 shown, it includes:

[0094] A configuration unit 31, configured to calculate a convolution kernel based on a training image and a dynamic generation function; the training image is used to train the image recognition model, the dynamic generation function is used to generate a convolution kernel, and the dynamic generation function is composed of a generation function and learning parameters;

[0095] A first training unit 32, configured to input the training image into the image recognition model with the configured convolution kernel, and perform a convolution operation on the training image through the convolution kernel to obtain a feature image corresponding to the training image;

[0096] An update unit 33, configured to, when it is determined by a loss function in the image recognition model that the image recognition model is not trained yet, perform backpropagation on the feature image, and update the learning parameters in the dynamic generation function according to the backpropagation result;

[0097] A second training unit 34, configured to reconfigure the convolution kernel of the image recognition model based on the updated dynamic generation function, and train the image recognition model with the reconfigured convolution kernel based on the training image until the image recognition model is trained.

[0098] The present disclosure provides an apparatus for training an image recognition model, which calculates a convolution kernel based on a training image and a dynamic generation function; the training image is used to train the image recognition model, and the dynamic generation function is used to generate the convolution kernel, and the dynamic generation function is composed of a generation function and learning parameters; input the training image into the image recognition model with the configured convolution kernel, perform a convolution operation on the training image through the convolution kernel to obtain a feature image corresponding to the training image; in the case where it is determined by the loss function in the image recognition model that the image recognition model has not been trained, perform backpropagation on the feature image, and update the learning parameters in the dynamic generation function according to the backpropagation result; reconfigure the convolution kernel of the image recognition model based on the updated dynamic generation function, and train the image recognition model with the reconfigured convolution kernel based on the training image until the image recognition model is trained. Compared with the related art, the present disclosure can adaptively adjust the convolution kernel parameters according to different task requirements by configuring the convolution kernel of the image recognition model according to the dynamic generation function; furthermore, it can help the image recognition model better learn the feature information in the image, improve the image feature extraction ability of the image recognition model, and further improve the accuracy and robustness of the image recognition model.

[0099] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the configuration unit 31 includes:

[0100] A first calculation module 311, configured to calculate configuration information of a convolution kernel by calculating a dynamic generation function based on feature data of the training image and task requirements corresponding to the image recognition model;

[0101] A configuration module 312, configured to configure the convolution kernel in the image recognition model by using the configuration information to obtain the configured image recognition model.

[0102] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the update unit 33 includes:

[0103] A second calculation module 331, configured to calculate a gradient of the loss function with respect to each layer in the image recognition model based on the feature image;

[0104] An update module 332, configured to update the learning parameters of the dynamic generation function by using the calculated gradient information.

[0105] Further, in a possible implementation manner of this embodiment, the second calculation module 331 is further configured to:

[0106] Perform gradient calculation on the feature image to obtain the gradient of the feature image;

[0107] Perform transposed convolution operation on the gradient of the feature image and the convolution kernel, and calculate the gradient of the loss function with respect to each layer in the image recognition model layer by layer according to the calculation result of the transposed convolution.

[0108] Further, in a possible implementation manner of this embodiment, the update module 332 is further configured to:

[0109] Obtain the preset learning rate of the image recognition model;

[0110] Based on the preset learning rate and the gradient information, calculate and update the learning parameters of the dynamically generated function.

[0111] It should be noted that the foregoing explanations of the method embodiments also apply to the apparatus in this embodiment. The principles are the same and will not be limited in this embodiment.

[0112] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0113] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 400 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0114] As Figure 5 shown, the device 400 includes a computing unit 401, which can execute various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded from a storage unit 408 into a RAM (Random Access Memory) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.

[0115] Multiple components in device 400 are connected to I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0116] The computing unit 401 can be various general-purpose and / or dedicated processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the method for training an image recognition model. For example, in some embodiments, the method for training an image recognition model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the aforementioned method for training an image recognition model in any other suitable manner (e.g., by means of firmware).

[0117] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0118] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0119] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0120] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0121] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0122] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server may also be a server of a distributed system or a server combined with a blockchain.

[0123] Here, it should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0124] Furthermore, the present disclosure also provides a vehicle. The vehicle in this embodiment further includes the device for training the image recognition model as described above, or the electronic device as described above, or the readable storage medium as described above. For specific descriptions, please refer to the above embodiments and will not be elaborated here one by one.

[0125] The various digital numbers such as the first and the second involved in the present disclosure are only for the convenience of description for distinction and do not limit the scope of the embodiments of the present disclosure, nor do they represent the order of precedence.

[0126] At least one in the present disclosure may also be described as one or more. The plurality may be two, three, four, or more, and the present disclosure does not make any limitations. In the embodiments of the present disclosure, for a technical feature, the technical features in this technical feature are distinguished by "the first", "the second", "the third", "A", "B", "C", and "D", etc. There is no order of precedence or size order among the technical features described by "the first", "the second", "the third", "A", "B", "C", and "D".

[0127] It should be understood that various forms of processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recorded in the present disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are made herein.

[0128] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for training an image recognition model, characterized in that, Including: Calculating a convolution kernel based on a training image and a dynamically generated function; The training image is used to train the image recognition model, the dynamically generated function is used to generate a convolution kernel based on it, and the dynamically generated function consists of a generation function and learning parameters; Inputting the training image into the image recognition model with the configured convolution kernel, and performing a convolution operation on the training image through the convolution kernel to obtain a feature image corresponding to the training image; In the case where it is determined by the loss function in the image recognition model that the image recognition model is not trained yet, performing backpropagation on the feature image and updating the learning parameters in the dynamically generated function according to the backpropagation result; Reconfiguring the convolution kernel of the image recognition model based on the updated dynamically generated function, and training the image recognition model with the reconfigured convolution kernel based on the training image until the image recognition model is trained.

2. The method according to claim 1, wherein The calculating the convolution kernel based on the training image and the dynamically generated function includes: Calculating the configuration information of the convolution kernel by calculating the dynamically generated function based on the feature data of the training image and the task requirements corresponding to the image recognition model; Configuring the convolution kernel in the image recognition model by using the configuration information.

3. The method according to claim 1, wherein The performing backpropagation on the feature image and updating the learning parameters in the dynamically generated function according to the backpropagation result includes: Calculating the gradient of the loss function with respect to each layer in the image recognition model based on the feature image; Updating the learning parameters of the dynamically generated function by using the calculated gradient information.

4. The method according to claim 3, wherein The calculating the gradient of the loss function with respect to each layer in the image recognition model based on the feature image includes: Performing gradient calculation on the feature image to obtain the gradient of the feature image; Performing a transposed convolution operation on the gradient of the feature image and the convolution kernel, and calculating the gradient of the loss function with respect to each layer in the image recognition model layer by layer according to the calculation result of the transposed convolution.

5. The method according to claim 3, characterized in that, The updating the learning parameters of the dynamically generated function by using the calculated gradient information includes: Obtaining the preset learning rate of the image recognition model; Calculating and updating the learning parameters of the dynamically generated function based on the preset learning rate and the gradient information.

6. An apparatus for training an image recognition model, characterized in that, Including: A configuration unit for calculating a convolution kernel based on a training image and a dynamically generated function; The training image is used to train the image recognition model, the dynamically generated function is used to generate a convolution kernel based on it, and the dynamically generated function consists of a generation function and learning parameters; A first training unit for inputting the training image into the image recognition model with the configured convolution kernel, and performing a convolution operation on the training image through the convolution kernel to obtain a feature image corresponding to the training image; An updating unit for, in the case where it is determined by the loss function in the image recognition model that the image recognition model is not trained yet, performing backpropagation on the feature image and updating the learning parameters in the dynamically generated function according to the backpropagation result; A second training unit, configured to reconfigure the convolutional kernels of the image recognition model based on the updated dynamic generation function, and train the image recognition model with reconfigured convolutional kernels based on the training images until the training of the image recognition model is completed.

7. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-5.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.

9. A computer program product, characterized in that, Comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.

10. A vehicle, characterized in that, The vehicle includes the apparatus for training the image recognition model according to claim 6, or the electronic device according to claim 7, or the non-transitory computer-readable storage medium according to claim 8.

Citation Information

Cited By

  • Cross-department routing gateway based on image recognition and dynamic permission configuration method

    CN121098635A

  • Cross-department routing gateway and dynamic permission configuration method based on image recognition

    CN121098635B