Training method, device, computer equipment and storage medium for neural network model
By conducting directional training on the general image recognition model and generating user-specific weight parameters, the problem of insufficient adaptability to the personalized environment in the image processing field is solved, and the recognition accuracy and robustness of the model are improved.
Patent Information
- Application Number
- CN201910101300.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-01-31
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-01-31
AI Technical Summary
The existing neural network models have limited adaptability to personalized environments in the field of image processing, and the recognition accuracy needs to be improved.
During the user's use process, the general image recognition model is directionally trained to generate weight parameters exclusive to the user, so that the model can be personalized to adapt to the user's actual use needs.
The neural network model improves the accuracy of the user's image recognition, and enhances the robustness and adaptability of the model.
Smart Images

Figure CN111507467B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of model training, and in particular to a training method, apparatus, computer equipment, and storage medium for a neural network model. Background Art
[0002] Since the advent of mathematical methods to simulate actual human neural networks, people have gradually become accustomed to calling such artificial neural networks directly neural networks. Neural networks have broad and attractive prospects in the fields of system identification, pattern recognition, intelligent control, etc., especially in intelligent control, people are particularly interested in the self-learning function of neural networks, and regard this important feature of neural networks as one of the key keys to solving the problem of controller adaptability in automatic control.
[0003] In the prior art, neural network models have good performance in the field of image processing. By repeatedly training the neural network model with a large number of pictures of the same type, the neural network model learns to recognize one or more image categories. Once the neural network model is trained to a convergence state, the weight parameters therein will be fixed and cannot be changed. Therefore, the adaptability of the neural network model in the prior art to the personalized environment is limited, and the accuracy of the neural network model recognition needs to be improved. Summary of the invention
[0004] The embodiments of the present invention provide a training method, apparatus, computer equipment and storage medium that can perform targeted training on a general image recognition model according to the actual usage needs of the user during the user's use, so as to generate a neural network model with corresponding weight parameters for the user.
[0005] In order to solve the above technical problems, a technical solution adopted by an embodiment of the present invention is: to provide a training method for a neural network model, comprising:
[0006] Get the user image of the target user;
[0007] Inputting the user image into a preset image recognition model to obtain a classification result of the user image by the image recognition model, wherein the image recognition model is a neural network model pre-trained to a convergence state and used to classify the input image;
[0008] Reading the classification result output by the image recognition model and displaying the classification result to obtain verification information made by the target user on the classification result;
[0009] The image recognition model is trained according to the verification information so that a weight parameter specific to the target user is generated in the image recognition model.
[0010] Optionally, before inputting the user image into a preset image recognition model to obtain a classification result of the image recognition model on the user image, the method further includes:
[0011] Inputting the user image into a preset environment classification model, wherein the environment classification model is a neural network model that is pre-trained to a convergence state and is used to perform environment classification on the input image;
[0012] Reading the environment classification result output by the environment classification model;
[0013] Identify whether the environmental scene represented by the environmental classification result has been trained to a convergence state;
[0014] When the environment scene has not been trained to a convergence state, confirming that the image recognition model recognizes the user image.
[0015] Optionally, the identifying whether the environment scene represented by the environment classification result is trained to a convergence state includes:
[0016] Using the environmental classification result as a search condition, searching in a preset historical classification information database to obtain training information having a mapping relationship with the environmental scene represented by the environmental classification result, wherein the training information includes the classification accuracy of the image recognition model in the historical records for the user image containing the environmental scene;
[0017] Comparing the classification accuracy with a preset accuracy threshold;
[0018] When the classification accuracy is less than the accuracy threshold, it is confirmed that the environment scene has not been trained to a convergence state; otherwise, it is confirmed that the environment scene has been trained to a convergence state.
[0019] Optionally, when the environment scene has not been trained to a convergence state, after confirming that the image recognition model recognizes the user image, the method further comprises:
[0020] Searching a preset strategy database for an image enhancement strategy having a mapping relationship with the environmental scene;
[0021] The user image is subjected to image enhancement processing according to the image enhancement strategy to generate an enhanced image derived from the user image.
[0022] Optionally, the image recognition model is a multi-channel model, and the step of inputting the user image into a preset image recognition model to obtain a classification result of the image recognition model on the user image includes:
[0023] Inputting the user image and the enhanced image into different channels of the image recognition model respectively;
[0024] Comparing whether the classification result of the user image calculated by the image recognition model is consistent with the classification result of the enhanced image;
[0025] When the classification result of the user image is consistent with the classification result of the enhanced image, confirm and output the classification result of the user image.
[0026] Optionally, after comparing whether the classification result of the user image calculated by the image recognition model is consistent with the classification result of the enhanced image, the method further includes:
[0027] When the classification result of the user image is inconsistent with the classification result of the enhanced image, calculating a first vector difference between a feature vector of the user image and a feature vector of the enhanced image;
[0028] The first vector difference is back-propagated to correct the weight parameters in the image recognition model so that the classification result of the user image is consistent with the classification result of the enhanced image.
[0029] Optionally, the verification information includes a verification classification result of the user image recognized by the target user, and the training of the image recognition model according to the verification information so as to generate a weight parameter specifically for the target user in the image recognition model includes:
[0030] When the verification result represented by the verification information is that the classification result is wrong, searching a preset feature database for a calibration feature vector having a mapping relationship with the verification classification result;
[0031] Calculating a second vector difference between the calibration feature vector and the feature vector of the user image;
[0032] The second vector difference is back-propagated to correct the weight parameters in the image recognition model so that the classification result of the user image is consistent with the verification classification result.
[0033] In order to solve the above technical problems, an embodiment of the present invention further provides a training device for a neural network model, comprising:
[0034] An acquisition module, used to acquire a user image of a target user;
[0035] A processing module, used for inputting the user image into a preset image recognition model to obtain a classification result of the user image by the image recognition model, wherein the image recognition model is a neural network model pre-trained to a convergence state and used for image classification of the input image;
[0036] A reading module, used to read the classification result output by the image recognition model and display the classification result to obtain verification information made by the target user on the classification result;
[0037] An execution module is used to train the image recognition model according to the verification information so that a weight parameter specific to the target user is generated in the image recognition model.
[0038] Optionally, the training device of the neural network model further includes:
[0039] A first processing submodule, used for inputting the user image into a preset environment classification model, wherein the environment classification model is a neural network model pre-trained to a convergence state and used for performing environment classification on the input image;
[0040] A first reading submodule is used to read the environment classification result output by the environment classification model;
[0041] A first identification submodule, used to identify whether the environment scene represented by the environment classification result is trained to a convergence state;
[0042] The first execution submodule is used to confirm that the image recognition model recognizes the user image when the environment scene has not been trained to a convergence state.
[0043] Optionally, the training device of the neural network model further includes:
[0044] A second processing submodule is used to search in a preset historical classification information database using the environmental classification result as a retrieval condition to obtain training information having a mapping relationship with the environmental scene represented by the environmental classification result, wherein the training information includes the classification accuracy of the image recognition model in the historical records for the user image containing the environmental scene;
[0045] A first comparison submodule, used for comparing the classification accuracy with a preset accuracy threshold;
[0046] The second execution submodule is used to confirm that the environment scene has not been trained to a convergence state when the classification accuracy is less than the accuracy threshold; otherwise, confirm that the environment scene has been trained to a convergence state.
[0047] Optionally, the training device of the neural network model further includes:
[0048] A third processing submodule is used to search for an image enhancement strategy having a mapping relationship with the environmental scene in a preset strategy database;
[0049] The third execution submodule is used to perform image enhancement processing on the user image according to the image enhancement strategy to generate an enhanced image derived from the user image.
[0050] Optionally, the image recognition model is a multi-channel model, and the training device of the neural network model further includes:
[0051] a fourth processing submodule, configured to input the user image and the enhanced image into different channels of the image recognition model respectively;
[0052] A second comparison submodule, used to compare whether the classification result of the user image calculated by the image recognition model is consistent with the classification result of the enhanced image;
[0053] The fourth execution submodule is used to confirm and output the classification result of the user image when the classification result of the user image is consistent with the classification result of the enhanced image.
[0054] Optionally, the training device of the neural network model further includes:
[0055] a fifth execution submodule, configured to calculate a first vector difference between a feature vector of the user image and a feature vector of the enhanced image when the classification result of the user image is inconsistent with the classification result of the enhanced image;
[0056] The fifth processing submodule is used to back-propagate the first vector difference to correct the weight parameters in the image recognition model so that the classification result of the user image is consistent with the classification result of the enhanced image.
[0057] Optionally, the verification information includes a verification classification result of the user image identified by the target user, and the training device of the neural network model further includes:
[0058] A sixth execution submodule, configured to search a preset feature database for a calibration feature vector having a mapping relationship with the verification classification result when the verification result represented by the verification information is that the classification result is wrong;
[0059] A first calculation submodule, configured to calculate a second vector difference between the calibration feature vector and the feature vector of the user image;
[0060] The sixth processing submodule is used to back-propagate the second vector difference to correct the weight parameters in the image recognition model so that the classification result of the user image is consistent with the verification classification result.
[0061] In order to solve the above technical problems, an embodiment of the present invention also provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the above-mentioned training method of the neural network model.
[0062] In order to solve the above technical problems, an embodiment of the present invention also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the above-mentioned training method of the neural network model.
[0063] The beneficial effect of the embodiment of the present invention is: after obtaining the user image of the user, the user image is input into a general image recognition model, wherein the image recognition model has been trained to convergence. The classification result output by the image recognition model is displayed, and verification information of the user's judgment on the classification result is obtained. According to the judgment result represented in the verification information, the image recognition model is trained in a targeted manner so that the image recognition model has personalized recognition capabilities. Since the targeted training can make the weight parameters in the image recognition model adaptively adjusted, the image recognition accuracy of the adjusted image recognition model for the user is improved, and the robustness of the model is more stable. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0065] Figure 1 A basic flow chart of a training method for a neural network model according to an embodiment of the present invention;
[0066] Figure 2 A schematic diagram of a process for confirming whether a user image is to be recognized according to an environment scene represented by the user image according to an embodiment of the present invention;
[0067] Figure 3 A schematic diagram of a process for determining whether an environment scene represented by a user image is involved in training through historical records according to an embodiment of the present invention;
[0068] Figure 4 A schematic diagram of a process of enhancing a user image according to an environmental scenario according to an embodiment of the present invention;
[0069] Figure 5 A schematic diagram of a process for controlling the output of user classification results by enhancing the classification results of an image according to an embodiment of the present invention;
[0070] Figure 6 A schematic diagram of a process for adjusting weight parameters of an image recognition model according to an enhanced image according to an embodiment of the present invention;
[0071] Figure 7 A schematic diagram of a process of training an image recognition model by verifying classification results according to an embodiment of the present invention;
[0072] Figure 8 A schematic diagram of the basic structure of a training device for a neural network model according to an embodiment of the present invention;
[0073] Fig. 9 The figure is a basic structural block diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0074] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0075] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0076] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0077] It will be understood by those skilled in the art that the "terminal" and "terminal device" used herein include both devices with wireless signal receivers, which are devices with only wireless signal receivers without transmission capabilities, and devices with receiving and transmitting hardware, which are devices with receiving and transmitting hardware capable of performing two-way communications on a two-way communication link. Such devices may include: cellular or other communication devices, which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which may combine voice, data processing, fax and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, a pager, Internet / intranet access, a web browser, a notepad, a calendar and / or a GPS (Global Positioning System) receiver; conventional laptop and / or palmtop computers or other devices, which have and / or include a conventional laptop and / or palmtop computer or other device with a radio frequency receiver. The "terminal" or "terminal device" used herein may be portable, transportable, installed in a vehicle (air, sea and / or land), or adapted and / or configured to operate locally, and / or in a distributed form, at any other location on the earth and / or in space. The "terminal" or "terminal device" used herein may also be a communication terminal, an Internet terminal, a music / video playing terminal, such as a PDA, a MID (Mobile Internet Device) and / or a mobile phone with a music / video playing function, or a smart TV, a set-top box and other devices.
[0078] Please refer to Figure 1 , Figure 1 Schematic diagram of the basic flow of the training method of the neural network model in this embodiment.
[0079] like Figure 1 As shown, a training method for a neural network model includes:
[0080] S1100, obtaining a user image of a target user;
[0081] In this embodiment, the user image of the target user refers to: an image collected or uploaded by the target user. The user image is not limited to the face image of the target user, any body part of the target user, or a full-body photo, but includes all images collected or uploaded by the terminal held by the target user. The target user refers to any user who uses the image recognition model in this embodiment.
[0082] The image content in the user image corresponds to the recognition type of the image recognition model. For example, if the image recognition model is used for the identity information, age, gender or appearance of a facial image, the content of the user image includes a facial image; if the image recognition model is used for the recognition of body movements, the content of the user image includes an image of the human body shape; if the image recognition model is used for the recognition of text information, the content of the user image includes a text image.
[0083] S1200, inputting the user image into a preset image recognition model to obtain a classification result of the user image by the image recognition model, wherein the image recognition model is a neural network model that is pre-trained to a convergence state and is used to classify the input image;
[0084] The acquired user image is input into a preset image recognition model, wherein the image recognition model is a neural network model pre-trained to a convergence state. The image category recognized by the image recognition model can be (but not limited to): recognizing the user's identity information, extracting the user's font, or recognizing the user's body movements, etc.
[0085] In this embodiment, the image recognition model can be a convolutional neural network model (CNN) that has been trained to a convergence state, but the image recognition model can also be: a deep neural network model (DNN), a recurrent neural network model (RNN) or a deformation model of the above three network models.
[0086] When training an image recognition model, a large number of images related to the classification categories set for it are used for training. After training to a convergence state, the model can accurately identify related images. For example, a large number of face images are used for training, so that the image recognition model has the ability to determine the user's identity information; a large number of font images are used for training, so that the image recognition model has the ability to recognize text; a large number of body movement images are used for training, so that the image recognition model has the ability to recognize body movements.
[0087] The image recognition model performs image recognition on the user image input therein to obtain the corresponding classification result. The classification result is the recognition result of the user image by the image recognition model. Different types of image recognition models output different classification results. For example, when the image recognition model recognizes the user's identity information, the classification result it outputs is the user's identity information represented by the user image; when the image recognition model extracts the user's font, the classification result it outputs is the text information recorded in the user image; when the image recognition model recognizes the user's body movements, the classification result it outputs is the information represented by the body movements in the user image.
[0088] S1300, reading the classification result output by the image recognition model and displaying the classification result to obtain verification information made by the target user on the classification result;
[0089] After the image recognition model outputs the classification result of the user image, the user terminal reads the classification result and displays it, and the target user views the classification result based on the displayed content. For example, when a user uploads a face image for face recognition, the name of the user image displayed in the classification result is consistent with the user's cognition. If so, the user confirms that the classification result is correct; otherwise, the user confirms that the classification result is wrong. In some embodiments, when the user confirms that the classification result is wrong, it is also necessary to input the user's correct cognition of the user image as a verification of the classification result.
[0090] The target user's judgment result on the classification result is recorded to generate verification information. That is, the verification information includes the target user's judgment information on the classification result.
[0091] S1400: training the image recognition model according to the verification information, so that a weight parameter specific to the target user is generated in the image recognition model.
[0092] The user terminal reads the content in the verification information and further trains the image recognition model according to the content recorded in the verification information.
[0093] When the information recorded in the verification information is "confirmed", that is, the target user believes that the classification result is consistent with his / her cognition, at this time, the weights of the convolutional layer in the image recognition model are not changed.
[0094] When the information recorded in the verification information is "error", that is, the target user believes that the classification result is different from his or her cognition, at this time, it is necessary to calculate the vector difference between the feature vector of the user image and the feature vector corresponding to the correct judgment result through the loss function in the image recognition model, and perform back propagation based on the vector difference to correct the weight parameters of the image recognition model. The feature vector of the user image is multiplied by the vector difference to obtain the gradient of the weight, and the gradient is multiplied by a gradient parameter and inverted and added to the previous weight of the image recognition model to complete the update of the weight parameter.
[0095] Through the above steps, the weight parameters in the image recognition model are corrected repeatedly and iteratively until the classification result of the user image is consistent with the user's cognition and the training is terminated. The retrained image recognition model can extract more detailed image features that distinguish the user image from other images, and then convert the above detailed identification features into unique features of the target user, thereby improving the recognition accuracy of the image recognition model for the target user. For example, repeated training can enable the image recognition model to learn the subtle differences between the facial features of the target user and other users or the subtle differences in different environmental scenes, or learn the user's writing habits, or learn the difference between the user's body movements and regular body movements.
[0096] In the above implementation, after obtaining the user image of the user, the user image is input into a general image recognition model, wherein the image recognition model has been trained to convergence. The classification result output by the image recognition model is displayed, and verification information of the user's judgment on the classification result is obtained. According to the judgment result represented in the verification information, the image recognition model is trained in a targeted manner so that the image recognition model has personalized recognition capabilities. Since the targeted training can make the weight parameters in the image recognition model adaptively adjusted, the image recognition accuracy of the adjusted image recognition model for the user is improved, and the robustness of the model is more stable.
[0097] In some embodiments, when the image recognition model is used for face recognition or person recognition, in order to enhance the accuracy of face recognition or person recognition in different environmental scenes, after acquiring the user image, the environment category of the user image is first identified to determine whether the environment scene represented by the current user image is trained. Figure 2 , Figure 2 This is a schematic diagram of a flow chart of confirming whether a user image is to be recognized according to an environment scene represented by the user image in this embodiment.
[0098] like Figure 2 As shown, Figure 1 Before the S1200 step shown, it includes:
[0099] S1111, inputting the user image into a preset environment classification model, wherein the environment classification model is a neural network model that is pre-trained to a convergence state and is used to perform environment classification on the input image;
[0100] The user image is input into a preset environment classification model, wherein the environment classification model is a neural network model that is pre-trained to a convergence state and is used to classify the environment of the input image. The environment classification model can classify the environment in which the user's face image or body image is located, for example, the environment classification result is: indoor, outdoor, rivers, lakes, seas, or mountains and rivers, etc. That is, the environmental information represented by the face image in the user image or the background image in the body image is identified.
[0101] In this embodiment, the environment classification model can be a convolutional neural network model (CNN) that has been trained to a convergence state, but the environment classification model can also be: a deep neural network model (DNN), a recurrent neural network model (RNN) or a deformation model of the above three network models.
[0102] When training the environment classification model, a large number of pictures recording the environment scenes are used for training. After training to a convergence state, the environment classification model can accurately identify the environment scenes in the input images.
[0103] S1112, reading the environment classification result output by the environment classification model;
[0104] After the environment classification model classifies the user image, the environment classification result output by the environment classification model is read. The environment classification result records the environment scene to which the user image belongs.
[0105] S1113, identifying whether the environment scene represented by the environment classification result has been trained to a convergence state;
[0106] In order to increase the training efficiency of the image recognition model, for environmental scenes with high recognition accuracy (for example, 99.5%), or when the image recognition model in the same environmental scene has accurately recognized 20 times in a row, it is judged that the image recognition model in this environmental scene mode has been trained to convergence, and the rate of return for continuing to train the model will be very low, so the training of the image recognition model in this environmental scene is stopped.
[0107] S1114: When the environment scene has not been trained to a convergence state, confirm that the image recognition model recognizes the user image.
[0108] When the judgment accuracy of the environmental scene does not reach the above standard, it is determined that the image recognition model in the environmental scene mode has not been trained to convergence, and the user image in the scene mode needs further training.
[0109] By identifying the environmental scenes in user images, identifying the recognition accuracy of the image recognition model in each environmental scene, and excluding user images in environmental scenes with accuracy higher than a certain standard from training, it is beneficial to improve the yield of image recognition model training, shorten the image recognition model training cycle, and save training resources.
[0110] In some implementations, in order to continuously train the image recognition model, the recognition information of the image recognition model in history is recorded, and the historical records are used to identify whether user images in different environmental scenes need image recognition training. Figure 3 , Figure 3 This is a flow chart of determining whether the environment scene represented by the user image participates in the training through historical records in this embodiment.
[0111] like Figure 3 As shown, Figure 2 The step S1113 shown includes:
[0112] S1121, searching in a preset historical classification information database using the environmental classification result as a search condition, and obtaining training information having a mapping relationship with the environmental scene represented by the environmental classification result, wherein the training information includes the classification accuracy of the image recognition model in the historical records for the user image containing the environmental scene;
[0113] In this embodiment, a historical classification information database is established, and the historical classification information database records the environmental classification results of the environmental scenes represented by the user images at each classification in the past history of the image recognition model. And the classification accuracy of the user images of each environmental scene is counted by the category of the environmental scene, and the accuracy is recorded in the training information corresponding to each environmental scene. Therefore, the corresponding training information can be found in the historical classification information database through the environmental classification results.
[0114] S1122, comparing the classification accuracy with a preset accuracy threshold;
[0115] The classification accuracy corresponding to the environmental scene corresponding to the current user image is compared with the preset accuracy threshold, and the accuracy threshold is a value that measures whether the recognition accuracy of the classification recognition model for each environmental scene meets the standard. For example, the accuracy threshold is 99.5%. It should be pointed out that the value of the accuracy threshold is not limited to this. According to different specific application scenarios, in some embodiments, the value of the accuracy threshold can be larger or smaller.
[0116] S1123. When the classification accuracy is less than the accuracy threshold, confirm that the environment scene has not been trained to a convergence state; otherwise, confirm that the environment scene has been trained to a convergence state.
[0117] When the classification accuracy of the comparison result is less than the accuracy threshold, it is judged that the image recognition model in the environment scene mode has not been trained to convergence, and the user images in this scene mode need further training. When the classification accuracy of the comparison result is greater than or equal to the accuracy threshold, it is judged that the image recognition model in the environment scene mode has been trained to convergence, and the yield of continuing to train the model will be very low, so the training of the image recognition model in this environment scene is stopped.
[0118] By recording historical classification results, it is convenient to count the recognition accuracy of the image recognition model in various environmental scenarios. The separate statistics of the accuracy in different environmental scenarios are conducive to analyzing the key training objects, thereby improving the recognition accuracy of the image recognition model as a whole.
[0119] In some implementations, in order to further improve the robustness of the image recognition model, when recognizing a user image, an enhanced image derived from the user image is first generated, wherein the enhancement strategy of the enhanced image needs to be selected according to the environment scene. Figure 4 , Figure 4 The figure is a schematic diagram of the process of enhancing the user image according to the environment scene in this embodiment.
[0120] See also Figure 4 ,like Figure 4 After the step S1114 shown, the following steps are included:
[0121] S1131, searching a preset strategy database for an image enhancement strategy that has a mapping relationship with the environment scene;
[0122] In this embodiment, a policy database is set up, and various enhancement strategies for enhancing user images are recorded in the policy database. The strategies recorded in the policy database are mainly classified into two categories, among which the first category is that the environment scene of the user image is relatively simple, such as the indoor environment or the background color composition of the background pixel is relatively simple, such as the background pixel value includes three or less pixel values. At this time, the image enhancement strategy is to add other background elements to the user image so that the confusion between the face image or the body image and the background image in the enhanced user image is enhanced. Among them, the increase of background elements is to randomly extract the background elements of the policy database for addition. The second category is that the environment scene of the user image is relatively complex, such as in the outdoor environment or the background color composition is relatively complex, such as the background pixel value includes three or more pixel values. At this time, the image enhancement strategy is to reduce the brightness and sharpening parameters of the user image to reduce the contrast between the face image or the body image and the background image, so that the confusion between the face image or the body image and the background image in the enhanced user image is enhanced.
[0123] S1132: Perform image enhancement processing on the user image according to the image enhancement strategy to generate an enhanced image derived from the user image.
[0124] The user image is enhanced according to the image enhancement strategy to generate an enhanced image derived from the user image. The enhanced image is derived from the user image, and the features recorded therein that are effective for the image recognition model to recognize are completely consistent with those of the user image. The difference is that due to the enhanced confusion between the face image or body image in the background image and the background image, the judgment difficulty during the training of the image recognition model is increased. The image recognition model trained with the enhanced image is more robust and has improved recognition accuracy in complex environments.
[0125] In some embodiments, the image recognition model is a multi-channel image recognition model, that is, the image recognition model includes at least two convolution channels, and the user image and the enhanced image are respectively input into the image recognition model to train the image recognition model to enhance the robustness of the image recognition model. Figure 5 , Figure 5 This is a flow chart of controlling the output of user classification results by enhancing the classification results of images in this embodiment.
[0126] like Figure 1 The step S1200 shown includes:
[0127] S1211, inputting the user image and the enhanced image into different channels of the image recognition model respectively;
[0128] In this embodiment, the image recognition model includes two convolution channels, one of which is used to extract image features of the user image, and the other is used to extract features of the enhanced image. However, the image recognition model is not limited to two convolution channels. Depending on the specific application scenario, in some embodiments, when there are N enhanced images, the number of convolution channels of the image recognition model is N+1.
[0129] The user image and the enhanced image are respectively input into different channels of the image recognition model, so that the image processing model can extract feature vectors of the user image and the enhanced image respectively.
[0130] S1212: Compare whether the classification result of the user image calculated by the image recognition model is consistent with the classification result of the enhanced image;
[0131] The classification result of the user image calculated by the comparison image recognition model is consistent with the classification result of the enhanced image. The comparison method is to calculate the Hamming distance between the classification result of the user image and the classification result of the enhanced image. When the Hamming distance between the two classification results is equal to zero, it indicates that the classification result of the user image is consistent with the classification result of the enhanced image; otherwise, it indicates that the classification result of the user image is inconsistent with the classification result of the enhanced image.
[0132] S1213: When the classification result of the user image is consistent with the classification result of the enhanced image, confirm and output the classification result of the user image.
[0133] When the comparison result shows that the classification result of the user image is consistent with the classification result of the enhanced image, it indicates that the image recognition model can avoid being affected by the interference content of the enhanced image and correctly extract the features of the face image or body image in the user image. No further training is required to confirm the classification result of the output user image.
[0134] In some embodiments, when the classification result of the user image is inconsistent with the classification result of the enhanced image, the image recognition model needs to be trained to enhance the robustness of the image recognition model. Figure 6 , Figure 6 This is a schematic diagram of the process of adjusting the weight parameters of the image recognition model according to the enhanced image in this embodiment.
[0135] like Figure 6 As shown, Figure 5 After the step S1212 shown, the following steps are included:
[0136] S1221. When the classification result of the user image is inconsistent with the classification result of the enhanced image, calculate a first vector difference between the feature vector of the user image and the feature vector of the enhanced image;
[0137] When the classification result of the user image is inconsistent with the classification result of the enhanced image, a first vector difference between the feature vector of the user image and the feature vector of the enhanced image is calculated. The calculation method is to calculate the vector difference between the two feature vectors through a loss function, and define the vector difference as the first vector difference. The value of the first vector difference indicates the gap between the feature vector extracted by the image recognition device and the correct feature vector.
[0138] It should be noted that S1213 and S1222 are different processing methods belonging to the same processing step, have the attribute of selective execution, and do not have a clear order relationship.
[0139] S1222. Back-propagate the first vector difference to correct the weight parameters in the image recognition model so that the classification result of the user image is consistent with the classification result of the enhanced image.
[0140] Back propagation is performed according to the first vector difference to correct the weight parameters of the image recognition model. The feature vector of the enhanced image is multiplied by the first vector difference to obtain the gradient of the weight, and the gradient is multiplied by a gradient parameter (for example, 0.05, but not limited to) and inverted and added to the weight before the convolution channel of the enhanced image of the image recognition model to complete the update of the weight parameters of the convolution channel of the enhanced image of the image recognition model. After the weight parameters are updated, the enhanced image feature vector extracted by the convolution channel of the enhanced image of the image recognition model tends to the feature vector extracted from the user image.
[0141] Through the above steps, the weight parameters in the image recognition model are corrected repeatedly and iteratively until the classification result of the enhanced image is consistent with the classification result of the user image, and the training is terminated.
[0142] Due to the increased confusion between the face image or body image in the background image and the background image, the judgment difficulty during image recognition model training is increased. The image recognition model trained with enhanced images is more robust and the recognition accuracy in complex environments is improved.
[0143] In some implementations, when the user confirmation result represented by the verification information is a classification error, the user is required to input the correct classification result, and the classification result is defined as the verification classification result. The image recognition model needs to be trained based on the verification classification result. Figure 7 , Figure 7 This is a flow chart of training the image recognition model by verifying the classification results in this embodiment.
[0144] like Figure 7 As shown, Figure 1 The steps S1400 shown include:
[0145] S1411, when the verification result represented by the verification information is that the classification result is wrong, searching a preset feature database for a calibration feature vector having a mapping relationship with the verification classification result;
[0146] When the information recorded in the verification information is "error", it means that the target user believes that the classification result is different from his / her cognition. It is necessary to find a calibrated feature vector with a mapping relationship to the verification classification result in the feature database. The feature database stores the calibrated feature vector of the set user image, and the calibrated feature vector is a correct feature vector with reference. For example, when the user image is a face image or a body image, the calibrated feature vector of the user's face image or body image that is manually calibrated and stored in the feature database, the verification classification result is the correct name of the user image entered by the target user, and the calibrated feature vector with a mapping relationship to the verification classification result is found in the feature database through the name.
[0147] S1412, calculating a second vector difference between the calibration feature vector and the feature vector of the user image;
[0148] Calculate the second vector difference between the feature vector of the user image and the calibration feature vector. The calculation method is to calculate the vector difference between the two feature vectors through the loss function, and define the vector difference as the second vector difference. The value of the second vector difference indicates the gap between the feature vector extracted by the image recognition device and the correct feature vector.
[0149] S1413. Back-propagate the second vector difference to correct the weight parameters in the image recognition model so that the classification result of the user image is consistent with the verification classification result.
[0150] Back propagation is performed based on the second vector difference to correct the weight parameters of the image recognition model. The feature vector of the user image is multiplied by the second vector difference to obtain the gradient of the weight, and the gradient is multiplied by a gradient parameter (for example, 0.05, but not limited to) and inverted and added to the weight before the image recognition model to complete the update of the convolution channel weight parameters of the enhanced image of the image recognition model. The feature vector extracted by the image recognition model after the weight parameters are updated tends to the calibration feature feature vector.
[0151] Through the above steps, the weight parameters in the image recognition model are corrected repeatedly and iteratively until the classification result of the user's image is consistent with the user's cognition, and the training is terminated.
[0152] By setting the calibration feature vector, the training convergence of the image recognition model has a direction and a reference, which can speed up the training of the image recognition model and improve the training efficiency.
[0153] In order to solve the above technical problems, an embodiment of the present invention also provides a training device for a neural network model.
[0154] Please refer to Figure 8 , Figure 8 Schematic diagram of the basic structure of the training device of the neural network model in this embodiment.
[0155] like Figure 8 As shown, a training device for a neural network model includes: an acquisition module 2100, a processing module 2200, a reading module 2300 and an execution module 2400. The acquisition module 2100 is used to acquire a user image of a target user; the processing module 2200 is used to input the user image into a preset image recognition model to obtain a classification result of the user image by the image recognition model, wherein the image recognition model is a neural network model pre-trained to a convergence state and used to classify the input image; the reading module 2300 is used to read the classification result output by the image recognition model and display the classification result to obtain verification information of the classification result made by the target user; the execution module 2400 is used to train the image recognition model according to the verification information so that a weight parameter exclusive to the target user is generated in the image recognition model.
[0156] After acquiring the user image of the user, the training device of the neural network model inputs the user image into a general image recognition model, wherein the image recognition model has been trained to convergence. The classification result output by the image recognition model is displayed, and verification information of the user's judgment on the classification result is obtained. According to the judgment result represented in the verification information, the image recognition model is trained in a targeted manner so that the image recognition model has personalized recognition ability. Since the targeted training can make the weight parameters in the image recognition model adaptively adjusted, the image recognition accuracy of the adjusted image recognition model for the user is improved, and the robustness of the model is more stable.
[0157] In some embodiments, the training device of the neural network model further includes: a first processing submodule, a first reading submodule and a first execution submodule. The first processing submodule is used to input the user image into a preset environment classification model, wherein the environment classification model is a neural network model that is pre-trained to a convergence state and is used to classify the input image; the first reading submodule is used to read the environment classification result output by the environment classification model; the first recognition submodule is used to identify whether the environment scene represented by the environment classification result is trained to a convergence state; the first execution submodule is used to confirm that the image recognition model recognizes the user image when the environment scene is not trained to a convergence state.
[0158] In some embodiments, the training device of the neural network model further includes: a second processing submodule, a first comparison submodule, and a second execution submodule. The second processing submodule is used to search in a preset historical classification information database with the environmental classification result as a retrieval condition, and obtain training information having a mapping relationship with the environmental scene represented by the environmental classification result, wherein the training information includes the classification accuracy of the image recognition model for user images containing the environmental scene in the historical records; the first comparison submodule is used to compare the classification accuracy with a preset accuracy threshold; the second execution submodule is used to confirm that the environmental scene has not been trained to a convergence state when the classification accuracy is less than the accuracy threshold; otherwise, it is confirmed that the environmental scene has been trained to a convergence state.
[0159] In some embodiments, the training device of the neural network model further includes: a third processing submodule and a third execution submodule. The third processing submodule is used to search for an image enhancement strategy having a mapping relationship with the environmental scene in a preset strategy database; the third execution submodule is used to perform image enhancement processing on the user image according to the image enhancement strategy to generate an enhanced image derived from the user image.
[0160] In some embodiments, the image recognition model is a multi-channel model, and the training device of the neural network model further includes: a fourth processing submodule, a second comparison submodule, and a fourth execution submodule. The fourth processing submodule is used to input the user image and the enhanced image into different channels of the image recognition model respectively; the second comparison submodule is used to compare whether the classification result of the user image calculated by the image recognition model is consistent with the classification result of the enhanced image; and the fourth execution submodule is used to confirm and output the classification result of the user image when the classification result of the user image is consistent with the classification result of the enhanced image.
[0161] In some embodiments, the training device of the neural network model further includes: a fifth execution submodule and a fifth processing submodule. The fifth execution submodule is used to calculate the first vector difference between the feature vector of the user image and the feature vector of the enhanced image when the classification result of the user image is inconsistent with the classification result of the enhanced image; the fifth processing submodule is used to back-propagate the first vector difference to correct the weight parameter in the image recognition model so that the classification result of the user image is consistent with the classification result of the enhanced image.
[0162] In some embodiments, the verification information includes the verification classification result of the user image recognized by the target user, and the training device of the neural network model further includes: a sixth execution submodule, a first calculation submodule and a sixth processing submodule. Among them, the sixth execution submodule is used to search for a calibration feature vector with a mapping relationship with the verification classification result in a preset feature database when the verification result represented by the verification information is that the classification result is wrong; the first calculation submodule is used to calculate the second vector difference between the calibration feature vector and the feature vector of the user image; the sixth processing submodule is used to back-propagate the second vector difference to correct the weight parameter in the image recognition model, so that the classification result of the user image is consistent with the verification classification result.
[0163] To solve the above technical problems, the embodiment of the present invention also provides a computer device. Fig. 9 , Fig. 9 This is a basic structural block diagram of the computer device in this embodiment.
[0164] like Fig. 9 As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a non-volatile storage medium, a memory, and a network interface connected via a system bus. Among them, the non-volatile storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database may store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a training method for a neural network model. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute a training method for a neural network model. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0165] In this embodiment, the processor is used to execute Figure 8The memory stores the program codes and data required to execute the above modules. The network interface is used to transmit data between user terminals or servers. The memory in this embodiment stores the program codes and data required to execute all submodules in the face image key point detection device, and the server can call the program codes and data of the server to execute the functions of all submodules.
[0166] The computer device does not need to mark the sample images involved in the training when training the fast model, thus saving the time and effort required for marking and improving the training speed. At the same time, the distance (Euclidean distance and / or cosine distance) between the feature vector representing the characteristics of the sample image output by the auxiliary model and the feature vector representing the characteristics of the sample image output by the fast model is directly calculated and back-propagated. This method converts the training of the neural network model into a simple regression algorithm, which can shorten the training time to the maximum extent and ensure the accuracy of the fast model output when the training is completed.
[0167] The present invention also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the training method of the neural network model in any of the above-mentioned embodiments.
[0168] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0169] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
Claims
1. A training method for a neural network model, It is characterized in that include: Get the user image of the target user; Inputting the user image into a preset image recognition model to obtain a classification result of the user image by the image recognition model, wherein the image recognition model is a neural network model pre-trained to a convergence state and used to classify the input image; Reading the classification result output by the image recognition model and displaying the classification result to obtain verification information of the classification result made by the target user, wherein when the target user determines that the classification result is wrong, the verification information includes the target user's correct cognition of the user image; The image recognition model is trained according to the verification information so that a weight parameter specific to the target user is generated in the image recognition model.
2. The training method of the neural network model according to claim 1, It is characterized in that Before inputting the user image into a preset image recognition model to obtain a classification result of the user image by the image recognition model, the method includes: Inputting the user image into a preset environment classification model, wherein the environment classification model is a neural network model that is pre-trained to a convergence state and is used to perform environment classification on the input image; Reading the environment classification result output by the environment classification model; Identify whether the environmental scene represented by the environmental classification result has been trained to a convergence state; When the environment scene has not been trained to a convergence state, confirming that the image recognition model recognizes the user image.
3. The training method of the neural network model according to claim 2, It is characterized in that The step of identifying whether the environment scene represented by the environment classification result has been trained to a convergence state includes: Using the environmental classification result as a search condition, searching in a preset historical classification information database to obtain training information having a mapping relationship with the environmental scene represented by the environmental classification result, wherein the training information includes the classification accuracy of the image recognition model in the historical records for the user image containing the environmental scene; Comparing the classification accuracy with a preset accuracy threshold; When the classification accuracy is less than the accuracy threshold, it is confirmed that the environment scene has not been trained to a convergence state; otherwise, it is confirmed that the environment scene has been trained to a convergence state.
4. The training method of the neural network model according to claim 2, It is characterized in that When the environment scene has not been trained to a converged state, after confirming that the image recognition model recognizes the user image, the method includes: Searching a preset strategy database for an image enhancement strategy having a mapping relationship with the environmental scene; The user image is subjected to image enhancement processing according to the image enhancement strategy to generate an enhanced image derived from the user image.
5. The training method of the neural network model according to claim 4, It is characterized in that The image recognition model is a multi-channel model, and the step of inputting the user image into a preset image recognition model to obtain a classification result of the user image by the image recognition model includes: Inputting the user image and the enhanced image into different channels of the image recognition model respectively; Comparing whether the classification result of the user image calculated by the image recognition model is consistent with the classification result of the enhanced image; When the classification result of the user image is consistent with the classification result of the enhanced image, confirm and output the classification result of the user image.
6. The training method of the neural network model according to claim 5, It is characterized in that After comparing whether the classification result of the user image calculated by the image recognition model is consistent with the classification result of the enhanced image, the method further comprises: When the classification result of the user image is inconsistent with the classification result of the enhanced image, calculating a first vector difference between a feature vector of the user image and a feature vector of the enhanced image; The first vector difference is back-propagated to correct the weight parameters in the image recognition model so that the classification result of the user image is consistent with the classification result of the enhanced image.
7. The training method of the neural network model according to claim 1, It is characterized in that The verification information includes a verification classification result of the user image identified by the target user, and the training of the image recognition model according to the verification information so as to generate a weight parameter specifically for the target user in the image recognition model includes: When the verification result represented by the verification information is that the classification result is wrong, searching a preset feature database for a calibration feature vector having a mapping relationship with the verification classification result; Calculating a second vector difference between the calibration feature vector and the feature vector of the user image; The second vector difference is back-propagated to correct the weight parameters in the image recognition model so that the classification result of the user image is consistent with the verification classification result.
8. A training device for a neural network model, It is characterized in that include: An acquisition module, used to acquire a user image of a target user; A processing module, used for inputting the user image into a preset image recognition model to obtain a classification result of the user image by the image recognition model, wherein the image recognition model is a neural network model pre-trained to a convergence state and used for image classification of the input image; a reading module, configured to read the classification result output by the image recognition model and display the classification result, so as to obtain verification information of the classification result made by the target user, wherein when the target user determines that the classification result is wrong, the verification information includes the target user's correct cognition of the user image; An execution module is used to train the image recognition model according to the verification information so that a weight parameter specific to the target user is generated in the image recognition model.
9. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the training method of the neural network model as claimed in any one of claims 1 to 7.
10. A storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the training method of the neural network model as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Image discrimination model training method and device and readable storage medium
CN107609598A
Face image gender judging method and device, computer equipment and storage medium
CN108446688A