Neural network model training method and electronic equipment
By performing specific image set processing and training on the gesture recognition neural network model, simulating the situation of hand jitter and small field of view angle, the problem of low robustness of existing models is solved and the accuracy of gesture recognition is improved.
Patent Information
- Application Number
- CN202311409424.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-10-27
AI Technical Summary
The existing gesture recognition neural network model has low robustness, especially when the field of view of the camera device is small, the accuracy of gesture recognition is affected.
By obtaining an image set including a plurality of first sample images, moving the bounding box of each sample image in the image set, generating a second image set, and training it based on the initial key point recognition model, improving the robustness of the model.
By simulating the situation where the hand part is removed from the image due to hand jitter and small field of view angle, the type of training sample images is enriched and the robustness and accuracy of the key point recognition model after training is improved.
Smart Images

Figure CN119942254A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and more specifically, to a training method and electronic device for a neural network model. Background Art
[0002] At present, gesture recognition, as a very important human-computer interaction method, is widely used in products such as smart phones, smart wearables, car-computer interaction, augmented reality (AR) and virtual reality (VR). Among them, gesture recognition can refer to collecting a video stream about gesture information through a camera device, recognizing the gesture in the video stream through a preset neural network model, and then performing corresponding operations based on the recognized gesture. This is equivalent to the user not having to touch the electronic device and not having to wear additional sensors, and the electronic device can perform corresponding operations according to the user's gesture.
[0003] In one possible case, due to the small field of view of the camera, the gestures in the video stream captured by the camera are prone to shaking, or the hands move out of the field of view of the camera, resulting in low accuracy of gesture recognition by the preset neural network model, which means that the robustness of the neural network model is low.
[0004] Based on this, how to improve the robustness of the neural network model for gesture recognition has become an urgent problem to be solved. Summary of the invention
[0005] The present application provides a training method for a neural network model, which can improve the robustness of the neural network model for gesture recognition.
[0006] In a first aspect, a training method for a neural network model is provided, comprising:
[0007] Acquire a first image set, the first image set including a plurality of first sample images, the first sample images including a target object, and a first bounding box obtained by detecting the target object;
[0008] Moving a first bounding box of a first sample image in the first image set to obtain a second image set;
[0009] An initial first key point recognition model is trained based on the first image set and the second image set to obtain a trained first key point recognition model, and the first key point recognition model is used to recognize key points in the target object.
[0010] The neural network model training method provided in the embodiment of the present application first obtains a first image set including multiple first sample images, wherein the first sample image includes a target object and a first bounding box obtained by detecting the target object, then moves the first bounding box of the first sample image in the first image set to obtain a second image set, and then trains the initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, wherein the first key point recognition model is used to recognize key points in the target object, and by moving the first bounding box in the first sample image to simulate the situation where the position of the hand in the image differs greatly due to hand shaking during image acquisition, and the situation where the hand partially moves out of the image due to the small field of view of the camera, the sample images for training the initialized first key point recognition model include not only the initial first sample images, but also sample images adjusted for the small field of view of the camera and the hand shaking problem, thereby enriching the types of sample images for training the initial first key point recognition model, thereby improving the robustness of the trained first key point model.
[0011] In combination with the first aspect, in certain embodiments of the first aspect, the second image set includes a second sample image and a third sample image, and moving the bounding box of the first sample image in the first image set to obtain the second image set includes: randomly moving the first bounding box in the first sample image to obtain the second sample image, the second sample image including the target object and the second bounding box obtained by randomly moving the first bounding box; parallel moving the bounding box in the first sample image to obtain the third sample image, the third sample image including the target object and the third bounding box obtained by parallel moving the first bounding box, the third bounding box being obtained by parallel moving the first bounding box along the X-axis, or the third bounding box being obtained by parallel moving the first bounding box along the Y-axis.
[0012] The second sample image obtained by randomly moving the first bounding box can simulate an image obtained by hand shaking. The third sample image obtained by parallel moving the first bounding box can simulate an image when the hand moves out of the field of view of the camera.
[0013] The neural network model training method provided in the embodiment of the present application obtains a first image set, wherein the first image set includes multiple first sample images, the first sample image includes a target object, and a first bounding box obtained by detecting the target object, randomly moves the bounding box in the first sample image to obtain a second sample image, the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box; parallel moves the first bounding box in the first sample image to obtain a third sample image, the third sample image includes the target object and a third bounding box obtained by parallel moving the first bounding box, the third bounding box is obtained by moving the first bounding box along the X-axis direction, or the third bounding box is obtained by moving the first bounding box along the Y-axis direction, based on the first image set, the second sample image and the third sample image for initial The first key point recognition model is trained to obtain the trained first key point recognition model, which is used to recognize key points in the target object. The first bounding box in the first sample image is randomly moved to simulate the situation where the position of the hand in the image is greatly different due to hand shaking during the image acquisition process. The parallel movement is used to simulate the situation where the hand partially moves out of the image due to the small field of view of the camera. The sample image of the first key point recognition model initialized in the training includes not only the initial first sample image, but also the second sample image and the third sample image adjusted for hand shaking and the hand moving out of the field of view of the camera, which enriches the types of sample images for training the initial first key point recognition model, thereby improving the robustness of the trained first key point model.
[0014] In combination with the first aspect, in some embodiments of the first aspect, parallel moving a bounding box in a first sample image to obtain a third sample image includes: parallel moving a first bounding box in the first sample image according to a first parameter to obtain the third sample image, the first parameter including a first sub-parameter and a second sub-parameter, the first sub-parameter being used to indicate a ratio of the number of third sample images obtained by parallel movement to the number of first sample images in the first image set, and the second sub-parameter being used to indicate a ratio of a first distance to a width or height of the first bounding box, the first distance being a distance from the first bounding box to the third bounding box; training an initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, including: training an initial first key point recognition model based on the first image set, the second sample image and the first image set. The initial first key point recognition model is pre-trained with three sample images to obtain a pre-trained first key point recognition model; the bounding box in the first sample image is parallel moved according to the second parameter to obtain a fourth sample image, the fourth sample image includes a target object and a fourth bounding box, the second parameter includes a third sub-parameter and a fourth sub-parameter, the third sub-parameter is used to indicate the ratio of the number of fourth sample images obtained by parallel movement to the number of first sample images in the first image set, the fourth sub-parameter is used to indicate the ratio of the second distance to the width or height of the first bounding box, the second distance refers to the distance from the first bounding box to the fourth bounding box; the pre-trained first key point recognition model is trained based on the first image set, the second sample image and the fourth sample image to obtain a trained first key point recognition model.
[0015] The training method of the neural network model provided in the embodiment of the present application obtains a first image set, wherein the first image set includes multiple first sample images, the first sample image includes a target object, and a first bounding box obtained by detecting the target object, randomly moves the bounding box in the first sample image to obtain a second sample image, the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box; uses a first parameter to parallel move the first bounding box in the first sample image to obtain a third sample image, the third sample image includes the target object and a third bounding box obtained by parallel moving the first bounding box, wherein the third bounding box is obtained by moving the first bounding box along the X-axis direction, or the third bounding box is obtained by moving the first bounding box along the Y-axis direction, and an initial first key point recognition model is trained based on the first image set, the second sample image, and the third sample image. Pre-training is performed to obtain a pre-trained first key point recognition model, and then the first bounding box in the first sample image is moved in parallel using the second parameter to obtain fourth sample data, and then the pre-trained first key point recognition model is continuously trained based on the first sample image, the second sample image and the fourth sample image to obtain a trained first key point recognition model. This is equivalent to dividing the training of the initial first key point recognition model into two steps: pre-training and training. Since the hand part moves out of the camera's field of view, the hand key point may not be within the bounding box. Therefore, if the initial first key point recognition model is trained in only one step, the model training may not converge due to the hand key point not being in the bounding box. The model training is performed in two steps of pre-training and training, which can effectively avoid the model training not convergence due to the hand key point not being in the bounding box.
[0016] In combination with the first aspect, in certain embodiments of the first aspect, the first sub-parameter is greater than the third sub-parameter, and the second sub-parameter is greater than the fourth sub-parameter.
[0017] In the training method of the neural network model provided in the embodiment of the present application, when pre-training the initial first key point recognition model, the data enhancement rate (first sub-parameter) used is greater than the data enhancement rate (third sub-parameter) used when continuing to train the pre-trained first key point recognition model, and the moving distance of the third bounding box in the third sample image used when pre-training the initial first key point recognition model is greater than the moving distance of the fourth bounding box in the fourth sample image used when continuing to train the pre-trained first key point recognition model. This is equivalent to first pre-training the initial first key point recognition model with a third sample image with a larger change, so that the next step of continuing to train the pre-trained first key point recognition model with a fourth sample image with a smaller change is easier to converge, thereby improving the accuracy of the trained first key point recognition model.
[0018] In combination with the first aspect, in certain embodiments of the first aspect, the initial first key point recognition model is pre-trained based on the first image set, the second sample image and the third sample image to obtain the pre-trained first key point recognition model, including: removing the image area outside the third bounding box in the third sample image to obtain an updated third sample image; pre-training the initial first key point recognition model based on the first image set, the second sample image and the updated third sample image to obtain the pre-trained first key point recognition model.
[0019] The training method of the neural network model provided in the embodiment of the present application, during the pre-training process of the initial first key point recognition model, removes the image area outside the third bounding box in the third sample image to obtain an updated third sample image, and then pre-trains the initial first key point recognition model based on the first image set, the second sample image and the updated third sample image to obtain the pre-trained first key point recognition model. This can avoid points outside the third bounding box from participating in model training, reduce the amount of data for pre-training the neural network model, and thereby reduce the difficulty of pre-training the initial first key point recognition model.
[0020] In combination with the first aspect, in certain embodiments of the first aspect, the method also includes: acquiring a third image set, the third image set including a fifth sample image, the fifth sample image being an image captured by a camera device; obtaining a first sample image based on the fifth sample image, depth information corresponding to the fifth sample image, and a target detection model, and the target detection model is used to mark a first bounding box of the target object.
[0021] In combination with the first aspect, in certain embodiments of the first aspect, obtaining a first sample image based on a fifth sample image, depth information corresponding to the fifth sample image, and a target detection model includes: obtaining a sixth sample image based on the fifth sample image and the target detection model, the sixth sample image including a first bounding box; and performing a three-dimensional conversion on the sixth sample image based on the depth information corresponding to the fifth sample image to obtain the first sample image.
[0022] It is understandable that some gestures are perpendicular to the screen of the electronic device, such as Fig.12 In this case, if the coordinates of the key points of the hand in the image are usually two-dimensional coordinates, gesture recognition based only on the coordinates of the key points of the hand in the image may lead to inaccurate gesture recognition results.
[0023] It can be understood that since the coordinates of each pixel point in the first sample image are three-dimensional coordinates, and the second sample image is obtained by randomly moving the first bounding box in the first sample image, the coordinates of each pixel point in the second sample image are also three-dimensional coordinates, including the depth information of each pixel point.
[0024] It can be understood that since the coordinates of each pixel point in the first sample image are three-dimensional coordinates, the third sample image is obtained by parallel moving the first boundary in the first sample image according to the first parameter, so the coordinates of each pixel point in the third sample image are also three-dimensional coordinates, including the depth information of each pixel point.
[0025] In combination with the first aspect, in some embodiments of the first aspect, performing three-dimensional conversion on the sixth sample image based on the depth information corresponding to the fifth sample image to obtain the first sample image includes:
[0026] The first formula and the depth information corresponding to the fifth sample image are used to perform three-dimensional coordinate transformation on the sixth sample image to obtain the first sample image. The first formula includes:
[0027] x = X / W;
[0028] y=Y / H;
[0029] z=(Z-Z0) / W;
[0030] Among them, x is the coordinate of the pixel point on the x-axis in the first sample image, y is the coordinate of the pixel point on the y-axis in the first sample image, z is the coordinate of the pixel point on the z-axis in the first sample image, X is the coordinate of the pixel point corresponding to x in the sixth sample image, Y is the coordinate of the pixel point corresponding to y in the sixth sample image, Z is the coordinate of the pixel point corresponding to z in the sixth sample image, Z0 represents the coordinate of the target key point on the z-axis in the sixth sample image, W represents the width of the first bounding box, and H represents the height of the first bounding box.
[0031] The training method of the neural network model provided in the embodiment of the present application obtains a third image set including a fifth sample image captured by a camera device, and then obtains a sixth sample image based on the fifth sample image and a target detection model, wherein the sixth sample image includes a first bounding box, and then performs a three-dimensional conversion on the sixth sample image based on the depth information corresponding to the fifth sample image to obtain a first sample image, which is equivalent to the first sample image being an image of three-dimensional coordinates, that is, the first sample image includes the depth information of each pixel point, and then the second sample image, the third sample image and the fourth sample image obtained based on the first sample image also include the depth information of each pixel point. In this way, in the process of training the initial first key point model, the first sample image, the second sample image, the third sample image and the fourth sample image used are all three-dimensional images and all include depth information, so that the coordinates of the hand key points obtained based on the trained first key point recognition model are three-dimensional coordinates including depth information, so that the gesture recognition based on the coordinates of the three-dimensional hand key points is more accurate, especially when recognizing gestures perpendicular to the screen of the electronic device, the gesture recognition result obtained is more accurate.
[0032] In combination with the first aspect, in some embodiments of the first aspect, the target object includes a hand in the image.
[0033] In combination with the first aspect, in certain embodiments of the first aspect, the first key point recognition model is used to recognize hand key points in an image.
[0034] In combination with the first aspect, in certain embodiments of the first aspect, after the initial first key point recognition model is trained based on the first image set and the second image set to obtain the trained first key point recognition model, the method also includes: acquiring a fourth image set, the fourth image set including a seventh sample image, the seventh sample image being an image obtained for a preset gesture of the target object; training the initial second key point recognition model based on the fourth image set to obtain a trained second key point recognition model, the second key point recognition model being used to identify key points corresponding to the preset gesture; and correcting parameters in the trained first key point recognition model based on parameters in the trained second key point recognition model to obtain an updated first key point recognition model.
[0035] The training method of the neural network model provided in the embodiment of the present application is to obtain a trained second key point recognition model by training the hand key points corresponding to special gestures, and then use the parameters of the trained second key point recognition model to correct the parameters of the trained first key point recognition model to obtain an updated first key point recognition model, so that the updated first key point recognition model can also accurately recognize the hand key points corresponding to special gestures, thereby improving the robustness of the updated first key point recognition model.
[0036] In combination with the first aspect, in some embodiments of the first aspect, the first key point recognition model includes a first network and a first regression layer, the second key point recognition model includes a first network and a second regression layer, and the dimension of the first regression layer is higher than the dimension of the second regression layer.
[0037] The training method of the neural network model provided in the embodiment of the present application, the first key point recognition model includes a first network and a first regression layer, the second key point recognition model includes a first network and a second regression layer, the dimension of the first regression layer is higher than the dimension of the second regression layer, and the second regression layer with lower dimension is used to train the hand key points corresponding to the preset gesture, and only a small number of hand key points corresponding to the preset gesture are trained. In this way, the hand key points corresponding to the preset gesture can be strengthened while avoiding the influence on other hand key points, thereby improving the recognition ability of the hand key points corresponding to the preset gesture.
[0038] In a second aspect, a training method for a neural network model is provided, comprising:
[0039] Training an initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model;
[0040] Acquire a fourth image set, where the fourth image set includes a seventh sample image, where the seventh sample image is an image obtained by a preset gesture of the target object;
[0041] Training the initial second key point recognition model based on the fourth image set to obtain a trained second key point recognition model, where the second key point recognition model is used to recognize key points corresponding to the preset gesture;
[0042] The parameters in the trained first key point recognition model are corrected based on the parameters in the trained second key point recognition model to obtain an updated first key point recognition model.
[0043] The first image set includes multiple first sample images, the first sample images include the target object and a first bounding box obtained by detecting the target object; the sample images in the second image set are obtained by moving the first bounding box of the first sample image in the first image set.
[0044] The training method of the neural network model provided in the embodiment of the present application is to obtain a trained second key point recognition model by training the hand key points corresponding to special gestures, and then use the parameters of the trained second key point recognition model to correct the parameters of the trained first key point recognition model to obtain an updated first key point recognition model, so that the updated first key point recognition model can also accurately recognize the hand key points corresponding to special gestures, thereby improving the robustness of the updated first key point recognition model.
[0045] In combination with the second aspect, in certain embodiments of the second aspect, the first key point recognition model includes a first network and a first regression layer, the second key point recognition model includes a first network and a second regression layer, and the dimension of the first regression layer is higher than the dimension of the second regression layer.
[0046] The training method of the neural network model provided in the embodiment of the present application, the first key point recognition model includes a first network and a first regression layer, the second key point recognition model includes a first network and a second regression layer, the dimension of the first regression layer is higher than the dimension of the second regression layer, and the second regression layer with lower dimension is used to train the hand key points corresponding to the preset gesture, and only a small number of hand key points corresponding to the preset gesture are trained. In this way, the hand key points corresponding to the preset gesture can be strengthened while avoiding the influence on other hand key points, thereby improving the recognition ability of the hand key points corresponding to the preset gesture.
[0047] In combination with the second aspect, in certain embodiments of the second aspect, the initial first key point recognition model is trained based on the first image set and the second image set, and before obtaining the trained first key point recognition model, the method also includes: acquiring a first image set, the first image set including multiple first sample images, the first sample images including a target object, and a first bounding box obtained by detecting the target object; moving the first bounding box of the first sample image in the first image set to obtain the second image set.
[0048] In combination with the second aspect, in certain embodiments of the second aspect, the second image set includes a second sample image and a third sample image, and moving the bounding box of the first sample image in the first image set to obtain the second image set includes: randomly moving the first bounding box in the first sample image to obtain the second sample image, the second sample image including the target object and the second bounding box obtained by randomly moving the first bounding box; parallel moving the bounding box in the first sample image to obtain the third sample image, the third sample image including the target object and the third bounding box obtained by parallel moving the first bounding box, the third bounding box being obtained by parallel moving the first bounding box along the X-axis, or the third bounding box being obtained by parallel moving the first bounding box along the Y-axis.
[0049] In combination with the second aspect, in some embodiments of the second aspect, parallel moving a bounding box in a first sample image to obtain a third sample image includes: parallel moving a first bounding box in the first sample image according to a first parameter to obtain the third sample image, the first parameter including a first sub-parameter and a second sub-parameter, the first sub-parameter being used to indicate a ratio of the number of third sample images obtained by parallel movement to the number of first sample images in the first image set, and the second sub-parameter being used to indicate a ratio of a first distance to a width or height of the first bounding box, the first distance being a distance from the first bounding box to the third bounding box; training an initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, including: training an initial first key point recognition model based on the first image set, the second sample image, and the first image set. The initial first key point recognition model is pre-trained with three sample images to obtain a pre-trained first key point recognition model; the bounding box in the first sample image is parallel moved according to the second parameter to obtain a fourth sample image, the fourth sample image includes a target object and a fourth bounding box, the second parameter includes a third sub-parameter and a fourth sub-parameter, the third sub-parameter is used to indicate the ratio of the number of fourth sample images obtained by parallel movement to the number of first sample images in the first image set, the fourth sub-parameter is used to indicate the ratio of the second distance to the width or height of the first bounding box, the second distance refers to the distance from the first bounding box to the fourth bounding box; the pre-trained first key point recognition model is trained based on the first image set, the second sample image and the fourth sample image to obtain a trained first key point recognition model.
[0050] In combination with the second aspect, in certain embodiments of the second aspect, the first sub-parameter is greater than the third sub-parameter, and the second sub-parameter is greater than the fourth sub-parameter.
[0051] In combination with the second aspect, in certain embodiments of the second aspect, the initial first key point recognition model is pre-trained based on the first image set, the second sample image and the third sample image to obtain the pre-trained first key point recognition model, including: removing the image area outside the third bounding box in the third sample image to obtain an updated third sample image; pre-training the initial first key point recognition model based on the first image set, the second sample image and the updated third sample image to obtain the pre-trained first key point recognition model.
[0052] In combination with the second aspect, in certain embodiments of the second aspect, the method also includes: acquiring a third image set, the third image set including a fifth sample image, the fifth sample image being an image captured by a camera device; obtaining a first sample image based on the fifth sample image, depth information corresponding to the fifth sample image, and a target detection model, and the target detection model is used to mark a first bounding box of the target object.
[0053] In combination with the second aspect, in certain embodiments of the second aspect, obtaining a first sample image based on a fifth sample image, depth information corresponding to the fifth sample image, and a target detection model includes: obtaining a sixth sample image based on the fifth sample image and the target detection model, the sixth sample image including a first bounding box; performing a three-dimensional conversion on the sixth sample image based on the depth information corresponding to the fifth sample image to obtain the first sample image.
[0054] In combination with the second aspect, in some embodiments of the second aspect, performing three-dimensional conversion on the sixth sample image based on the depth information corresponding to the fifth sample image to obtain the first sample image includes: performing three-dimensional coordinate conversion on the sixth sample image using the first formula and the depth information corresponding to the fifth sample image to obtain the first sample image, wherein the first formula includes:
[0055] x = X / W;
[0056] y=Y / H;
[0057] z=(Z-Z0) / W;
[0058] Among them, x is the coordinate of the pixel point on the x-axis in the first sample image, y is the coordinate of the pixel point on the y-axis in the first sample image, z is the coordinate of the pixel point on the z-axis in the first sample image, X is the coordinate of the pixel point corresponding to x in the sixth sample image, Y is the coordinate of the pixel point corresponding to y in the sixth sample image, Z is the coordinate of the pixel point corresponding to z in the sixth sample image, Z0 represents the coordinate of the target key point on the z-axis in the sixth sample image, W represents the width of the first bounding box, and H represents the height of the first bounding box.
[0059] In conjunction with the second aspect, in certain embodiments of the second aspect, the target object includes a hand in the image.
[0060] In combination with the second aspect, in certain embodiments of the second aspect, the first key point recognition model is used to recognize hand key points in an image.
[0061] In a third aspect, a training device for a neural network model is provided, comprising a unit for executing any one of the methods in the first aspect or the second aspect. The device may be a server, a terminal device, or a chip in a terminal device. The device may include an acquisition unit and a processing unit.
[0062] When the device is a terminal device, the processing unit may be a processor, and the acquisition unit may be a communication interface; the terminal device may also include a memory, which is used to store computer program code. When the processor executes the computer program code stored in the memory, the terminal device executes any one of the methods in the first aspect or the second aspect.
[0063] When the device is a chip in a terminal device, the processing unit may be a processing unit inside the chip, and the acquisition unit may be an output interface, a pin or a circuit, etc.; the chip may also include a memory, which may be a memory inside the chip (for example, a register, a cache, etc.) or a memory located outside the chip (for example, a read-only memory, a random access memory, etc.); the memory is used to store computer program code, and when the processor executes the computer program code stored in the memory, the chip executes any one of the methods in the first aspect or the second aspect.
[0064] In one possible implementation, a memory is used to store computer program code; a processor executes the computer program code stored in the memory, and when the computer program code stored in the memory is executed, the processor is used to perform: acquiring a first image set, the first image set including multiple first sample images, the first sample images including a target object, and a first bounding box obtained by detecting the target object; moving the first bounding box of the first sample image in the first image set to obtain a second image set; training an initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, the first key point recognition model being used to identify key points in the target object.
[0065] In one possible implementation, a memory is used to store computer program code; a processor executes the computer program code stored in the memory, and when the computer program code stored in the memory is executed, the processor is used to perform: training an initial first key point recognition model based on a first image set and a second image set to obtain a trained first key point recognition model; acquiring a fourth image set, the fourth image set including a seventh sample image, the seventh sample image being an image obtained for a preset gesture of a target object; training an initial second key point recognition model based on the fourth image set to obtain a trained second key point recognition model, the second key point recognition model being used to identify key points corresponding to preset gestures; and correcting parameters in the trained first key point recognition model based on parameters in the trained second key point recognition model to obtain an updated first key point recognition model.
[0066] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program code. When the computer program code is executed by a training device for a neural network model, the training device for the neural network model executes any one of the training methods for the neural network model in the first aspect or the second aspect.
[0067] In a fifth aspect, a computer program product is provided, comprising: a computer program code, which, when executed by a training device for a neural network model, enables the training device for the neural network model to execute any one of the device methods in the first aspect or the second aspect.
[0068] The neural network model training method and electronic device provided in the embodiments of the present application first obtain a first image set including multiple first sample images, wherein the first sample image includes a target object and a first bounding box obtained by detecting the target object, then move the first bounding box of the first sample image in the first image set to obtain a second image set, and then train the initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, the first key point recognition model is used to recognize key points in the target object, and the first bounding box in the first sample image is moved to simulate the situation where the position of the hand in the image is greatly different due to hand shaking during image acquisition, and the situation where the hand is partially moved out of the image due to the small field of view of the camera, so that the sample images for training the initialized first key point recognition model include not only the initial first sample images, but also sample images adjusted for the small field of view of the camera and the hand shaking problem, thereby enriching the types of sample images for training the initial first key point recognition model, thereby improving the robustness of the trained first key point model. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a flow chart of gesture recognition;
[0070] Figure 2 is a schematic diagram of a bounding box of a hand;
[0071] Figure 3 It is a schematic diagram of the key points of the hand;
[0072] Figure 4 It is a structural diagram of a convolutional neural network model;
[0073] Figure 5 It is a schematic diagram of the structure of a depth-separable convolutional layer;
[0074] Figure 6 It is a structural diagram of the system architecture provided by the embodiment of the present application;
[0075] Figure 7 It is a flowchart of a training method for a neural network model provided in an embodiment of the present application;
[0076] Figure 8 It is a flowchart of another neural network model training method provided in an embodiment of the present application;
[0077] Fig. 9 is a schematic diagram of a first bounding box and a second bounding box provided in an embodiment of the present application;
[0078] Fig.10 is a schematic diagram of a first bounding box and a third bounding box provided in an embodiment of the present application;
[0079] Fig.11 It is a flowchart of another neural network model training method provided in an embodiment of the present application;
[0080] Fig.12 is a schematic diagram of a preset gesture provided in an embodiment of the present application;
[0081] Fig.13 It is a flowchart of another neural network model training method provided in an embodiment of the present application;
[0082] Fig.14 is a schematic diagram of a first bounding box and a third bounding box provided in an embodiment of the present application;
[0083] Fig.15 It is a flowchart of another neural network model training method provided in an embodiment of the present application;
[0084] Fig.16 It is a structural schematic diagram of a first key point recognition model and a second key point recognition model provided in an embodiment of the present application;
[0085] Fig.17 It is a flowchart of another neural network model training method provided in an embodiment of the present application;
[0086] Fig.18 It is a schematic diagram of a training device for a neural network model provided by the present application;
[0087] Fig.19 It is a schematic diagram of an electronic device for training a neural network model provided in the present application. DETAILED DESCRIPTION
[0088] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0089] In the following, the terms "first", "second", and "third" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", and "third" may explicitly or implicitly include one or more of the features.
[0090] To facilitate understanding, some of the examples given are provided for reference to the description of concepts related to the embodiments of the present application.
[0091] (1) Neural network.
[0092] A neural network can be composed of neural units. A neural unit can be a s The output of the operation unit can be shown as formula (1):
[0093]
[0094] Where, s = 1, 2, ... n, n is a natural number greater than 1, W s For x s , b is the bias of the neural unit. f is the activation function of the neural unit, which is used to perform nonlinear transformation on the features in the neural network, thereby converting the input signal in the neural unit into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0095] (2) Convolutional neural network (CNN).
[0096] CNN is a deep neural network with a convolutional structure. It contains a feature extractor consisting of a convolutional layer and a subsampling layer. The feature extractor can be regarded as a filter. Among them, the convolutional layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of the convolutional neural network, a neuron is only connected to some neurons in the adjacent layers. A convolutional layer usually contains several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights here are convolution kernels. Shared weights can be understood as the way to extract image information is independent of position. The convolution kernel can be initialized in the form of a matrix of random size. During the training process of the convolutional neural network, the convolution kernel can obtain reasonable weights through learning. In addition, the direct benefit of shared weights is to reduce the connection between the layers of the convolutional neural network, while reducing the risk of overfitting.
[0097] When extracting features, it can be achieved through convolution operations, where the convolution operation can be composed of convolution operators with the same convolution kernel size, rectified linear units (ReLU) and batch normalization (BN). ReLU is used to enhance the nonlinear relationship between the input and output of the model, and the expression is RelU(x) = max(0, x). In one possible case, ReLU can also be called an activation function; BN is used to use a normalized transformation function to transform the input data into a normal distribution with zero mean and unit variance.
[0098] (3) Back propagation algorithm.
[0099] Neural networks can use the error back propagation (BP) algorithm to correct the values of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, the forward transmission of the input signal to the output will generate error loss, and the parameters in the initial neural network model are updated by back propagating the error loss information, so that the error loss converges. The back propagation algorithm is a back propagation movement dominated by error loss, which aims to obtain the optimal parameters of the neural network model, such as the weight matrix.
[0100] As a very important human-computer interaction method, gesture recognition is widely used in electronic devices such as smart phones, smart wearables, car-computer interaction, augmented reality (AR) and virtual reality (VR). Among them, gesture recognition can refer to collecting a video stream about gesture information through a camera device, recognizing the gesture in the video stream through a preset neural network model, and then performing corresponding operations based on the recognized gesture. This is equivalent to the user not having to touch the electronic device and not having to wear additional sensors, and the electronic device can perform corresponding operations according to the user's gesture.
[0101] In one possible case, the camera device is the front camera of a smartphone, which usually has a small field of view. When a user performs gesture recognition, the distance between the user's hand and the camera is usually small. A slight shake of the user's hand may cause the gesture in the video stream captured by the camera to shake greatly, resulting in inaccurate gesture recognition results due to excessive hand displacement; or the user may slightly move the hand and the hand may move out of the camera's field of view, resulting in inaccurate gesture recognition results based on incomplete hands; that is, when the camera device has a small field of view, the problem of low accuracy in gesture recognition may easily occur.
[0102] In view of this, an embodiment of the present application provides a training method for a neural network model and an electronic device. The neural network model training method provided by the embodiment of the present application first obtains a first image set including multiple first sample images, wherein the first sample image includes a target object and a first bounding box obtained by detecting the target object, and then moves the first bounding box of the first sample image in the first image set to obtain a second image set, and then trains the initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, the first key point recognition model is used to identify key points in the target object, and the first bounding box in the first sample image is moved to simulate the situation where the position of the hand in the image is greatly different due to hand shaking during image acquisition, and the situation where the hand is partially moved out of the image due to the small field of view of the camera, so that the sample images of the first key point recognition model for training the initialization include not only the initial first sample images, but also sample images adjusted for the small field of view of the camera and the hand shaking problem, enriching the types of sample images for training the initial first key point recognition model, thereby improving the robustness of the trained first key point model.
[0103] It is understandable that the process of performing gesture recognition on the video stream captured by the camera device and outputting the gesture recognition result can be divided into multiple steps. Among them, the camera device can be specifically a camera (for example, a color camera, a grayscale camera, a depth camera, etc.), through which an image containing a hand image stream (for example, a continuous frame image containing a hand image stream) can be obtained; then the image containing the hand image stream is processed, and the hand movements therein are recognized as predefined gesture categories; finally, the gesture category recognized by the dynamic gesture recognition system is responded to (for example, taking a photo, playing music, etc.), thereby realizing gesture interaction.
[0104] For example, Figure 1 As shown, in the process of performing gesture recognition on the video stream collected by the camera device, the following steps can be followed:
[0105] Step 1: Perform target detection on the image in the video stream and identify the hand area in the image.
[0106] Perform object detection on the image in the video stream, identify the hand area in the image, and set a bounding box in the hand area in the image. Figure 2 The bounding box 11 in the image 1 in the video stream. Among them, the target detection of the image in the video stream can be achieved by a target detection model.
[0107] Step 2: Perform key point recognition on the image area in the bounding box to obtain the hand key points.
[0108] Since the joints usually drive the muscles when the human body moves, the movement of the human body can be determined based on the movement of the joints. The key points of the hand can refer to the joints of the hand. For example, Figure 3 As shown, the key points of the hand may include 21 key points of the hand, namely, the wrist (0), the carpometacarpal joint (1), the metacarpophalangeal joint of the thumb (2), the interphalangeal joint of the thumb (3), the tip of the thumb (4), the metacarpophalangeal joint of the index finger (5), the proximal interphalangeal joint of the index finger (6), the distal interphalangeal joint of the index finger (7), the tip of the index finger (8), the metacarpophalangeal joint of the middle finger (9), the proximal interphalangeal joint of the middle finger (10), the distal interphalangeal joint of the middle finger (11), the tip of the middle finger (12), the metacarpophalangeal joint of the ring finger (13), the proximal interphalangeal joint of the ring finger (14), the distal interphalangeal joint of the ring finger (15), the tip of the ring finger (16), the metacarpophalangeal joint of the little finger (17), the proximal interphalangeal joint of the little finger (18), the distal interphalangeal joint of the little finger (19), and the tip of the little finger (20).
[0109] Among them, key point recognition is performed on the image area in the bounding box to obtain the hand key points, which can be achieved through a key point recognition model. The key point recognition model can be a neural network model, such as a convolutional neural network model of MobileNet v1.
[0110] The convolutional neural network model of MobileNet v1 provided in the embodiment of the present application, wherein: Figure 4 As shown, MobileNetv1 may include an input layer 210 , a convolution layer 220 , a deep classifiable convolution layer (Deep Sequential Convolutional Network, DSConv) 230 , an adaptive pooling layer 240 , a linear regression layer 250 and an output layer 260 .
[0111] The image with the hand area marked is used as the input of the input layer 210, and the hand key points are determined through the convolution layer 220, the depth classifiable convolution layer 230, the adaptive pooling layer 240, and the linear regression layer 250, and the image with the hand key points marked is output through the output layer 260.
[0112] The following will take the convolutional layer as an example to introduce the internal working principle of a convolutional layer.
[0113] The convolution layer can include many convolution operators, also known as kernels. Its role in image processing is equivalent to a filter that extracts specific information from the input image matrix. The convolution operator can essentially be a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix is usually processed horizontally on the input image one pixel after another (or two pixels after two pixels... depending on the value of the stride) to complete the work of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as the depth dimension of the input image. During the convolution operation, the weight matrix will extend to the entire depth of the input image.
[0114] Therefore, convolution with a single weight matrix will produce a convolution output with a single depth dimension, but in most cases, a single weight matrix is not used, but multiple weight matrices of the same size (rows × columns), that is, multiple isotype matrices, are applied. The output of each weight matrix is stacked to form the depth dimension of the convolution image, where the dimension can be understood as being determined by the "multiple" mentioned above. Different weight matrices can be used to extract different features in the image, for example, one weight matrix is used to extract image edge information, another weight matrix is used to extract specific colors of the image, and another weight matrix is used to blur unwanted noise in the image, etc. The multiple weight matrices have the same size (rows × columns), and the convolution feature maps extracted by the multiple weight matrices of the same size are also of the same size. The extracted multiple convolution feature maps of the same size are then merged to form the output of the convolution operation.
[0115] The weight values in these weight matrices need to be obtained through a lot of training in practical applications. The weight matrices formed by the weight values obtained through training can be used to extract information from the input image, so that the convolutional neural network can make correct predictions.
[0116] When a convolutional neural network has multiple convolutional layers, the initial convolutional layer often extracts more general features, which can also be called low-level features. As the depth of the convolutional neural network increases, the features extracted by the subsequent convolutional layers become more and more complex, such as high-level semantic features. Features with higher semantics are more suitable for the problem to be solved.
[0117] Adaptive pooling layer (referred to as pooling layer):
[0118] Since it is often necessary to reduce the number of training parameters, it is often necessary to periodically introduce a pooling layer after the convolution layer. It can be a convolution layer followed by a pooling layer, or multiple convolution layers followed by one or more pooling layers. In the image processing process, the only purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a maximum pooling operator to sample the input image to obtain an image of smaller size. The average pooling operator can calculate the pixel values in the image within a specific range to produce an average value as the result of average pooling. The maximum pooling operator can take the pixel with the largest value in the range as the result of maximum pooling within a specific range. In addition, just as the size of the weight matrix used in the convolution layer should be related to the image size, the operator in the pooling layer should also be related to the image size. The size of the image output after processing by the pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel in the image output by the pooling layer represents the average value or maximum value of the corresponding sub-region of the image input to the pooling layer.
[0119] Linear regression layer (or fully connected layer):
[0120] After being processed by the convolution layer / pooling layer, the convolutional neural network is not sufficient to output the required output information. Because as mentioned above, the convolution layer / pooling layer will only extract features and reduce the parameters brought by the input image. However, in order to generate the final output information (the required class information or other related information), the convolutional neural network needs to use a fully connected layer to generate one or a group of outputs of the required number of classes. Therefore, the fully connected layer may include multiple hidden layers and an output layer, and the parameters contained in the multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type may include image recognition, image classification, image super-resolution reconstruction, etc.
[0121] After the multiple hidden layers in the fully connected layer, the last layer of the entire convolutional neural network is the output layer 260, which has a loss function similar to the classification cross entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, the back propagation will begin to update the weight values and biases of the aforementioned layers to reduce the loss of the convolutional neural network and the error between the result output by the convolutional neural network through the output layer and the ideal result.
[0122] Step 3: Based on the hand key points, the gesture recognition result is obtained.
[0123] After obtaining the key points of the hand, the result of gesture recognition can be determined based on the displacement between the key points of the hand in N frames of images. The result of gesture recognition may include flip up, flip down, flip left, flip right, screen capture, pinch, etc.
[0124] For example, the coordinate origin of the mobile phone display screen is in the upper left corner. In the first frame, the coordinate of the distal interphalangeal joint (7) of the index finger is (X1, Y1), in the second frame, the coordinate of the distal interphalangeal joint (7) of the index finger is (X2, Y1), in the third frame, the coordinate of the distal interphalangeal joint (7) of the index finger is (X3, Y1) ..., in the Nth frame, the coordinate of the distal interphalangeal joint (7) of the index finger is (XN, Y1), wherein X1>X2>X3...>XN, based on which, the result of gesture recognition can be determined as a flip-up gesture.
[0125] It should be noted that if Figure 4 The neural network model shown is only an example of a convolutional neural network. In specific applications, the convolutional neural network may also exist in the form of other network models.
[0126] Among them, DSConv230 can be Figure 5 As shown, it includes a 3*3 depth-separable convolution layer 231, a normalization layer BN232, an activation function ReLU233, a 1*1 convolution layer 234, a normalization layer BN235 and an activation function ReLU236.
[0127] DSConv consists of a multi-layer convolutional neural network, each layer consists of a convolutional layer, a pooling layer (normalization layer BN) and an activation function.
[0128] The convolution layer is used to extract local features of the input sequence. It performs convolution operations on the input sequence by sliding a learnable convolution kernel (filter) to obtain a series of feature maps. Each feature map corresponds to a convolution kernel, and different convolution kernels can extract different features.
[0129] The pooling layer is used to reduce the size of the feature map and retain important features. Commonly used pooling operations include maximum pooling and average pooling.
[0130] The activation function is a nonlinear transformation in DSConv, which is used to introduce nonlinear capabilities. Common activation functions include ReLU, Sigmoid, and Tanh. The activation function can enhance the expressive power of the model and enable it to learn more complex features and patterns. Figure 5 The activation function in the DSConv shown is ReLU.
[0131] Below through Figure 6 The system architecture provided by the embodiment of the present application is described. Figure 6 , Figure 6 Schematic diagram of the system architecture of the embodiment of the present application. Figure 6 As shown, the system architecture 100 includes an execution device, a training device 120, a database 130, a client device 140, a data storage system 150, and a data acquisition system 160. In addition, the execution device 110 includes a computing module 111, an I / O interface 112, a preprocessing module 113, and a preprocessing module 114. The computing module 111 may include a target model / rule 101, and the preprocessing module 113 and the preprocessing module 114 are optional.
[0132] The data acquisition device 160 is used to collect training data. After acquiring the training data, the target model / rule 101 can be trained according to the training data. Next, the trained target model / rule 101 can be used to perform the relevant processes of gesture recognition in the embodiment of the present application.
[0133] The target model / rule 101 may include a plurality of neural network sub-models, each of which is used to perform a corresponding recognition process.
[0134] Specifically, the above-mentioned target model / rule 101 can be composed of a first neural network model, a second neural network model, a third neural network model and a fourth neural network model. The functions of these three neural network models are described below.
[0135] The first neural network model (e.g., target detection model) is used to detect the image stream to determine the hand frame of each frame in the image stream, that is, the bounding box of the hand;
[0136] A second neural network model (e.g., a key point detection model) is used to determine the key points of the hand based on the hand box of each frame image, that is, the bounding box of the hand;
[0137] The third neural network model (gesture recognition model): is used to perform gesture recognition on the image stream based on the key points of the hand to determine the user's gesture action;
[0138] The fourth neural network model is used to recognize a frame of image to determine the user's hand gesture type.
[0139] The above four neural network sub-models can be obtained through separate training.
[0140] Among them, the first neural network sub-model can be trained by the first type of training data, and the first type of training data includes multiple frames of hand images and label data of the multiple frames of hand images, wherein the label data of the multiple frames of hand images includes the enclosing box where the hand is located in each frame of the hand image, that is, the bounding box of the hand.
[0141] The second neural network sub-model can be trained by the second type of training data, and the second type of training data includes multiple image streams and label data of the multiple image streams, wherein each of the multiple image streams is composed of multiple frames of continuous hand images, and the label data of the multiple image streams includes hand key points corresponding to each image stream.
[0142] The third neural network sub-model can be trained by the third type of training data, and the third type of training data includes multiple image streams and label data of the multiple image streams, wherein each of the multiple image streams is composed of multiple frames of continuous hand images, and the label data of the multiple image streams includes gesture actions corresponding to each image stream.
[0143] The fourth neural network sub-model can be trained by the fourth type of training data, and the fourth type of training data includes multiple frames of hand images and label data of the multiple frames of images, wherein the label data of the multiple frames of images includes the gesture posture corresponding to each frame of the image.
[0144] The target model / rule 101 trained by the training device 120 can be applied to different systems or devices, such as Figure 6 The execution device 110 shown in the figure may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR), a vehicle terminal, etc., or a server or a cloud. Figure 6 In the embodiment, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with an external device. A user can input data to the I / O interface 112 through a client device 140. The input data can include: an image to be processed input by the client device. The client device 140 here can specifically be a terminal device.
[0145] The preprocessing module 113 and the preprocessing module 114 are used to perform preprocessing according to the input data (such as the image to be processed) received by the I / O interface 112. In the embodiment of the present application, there may be no preprocessing module 113 and the preprocessing module 114 or only one preprocessing module. When there is no preprocessing module 113 and the preprocessing module 114, the computing module 111 may be directly used to process the input data.
[0146] When the execution device 110 preprocesses the input data, or when the computing module 111 of the execution device 110 performs calculations and other related processing, the execution device 110 can call the data, code, etc. in the data storage system 150 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.
[0147] It is worth noting that Figure 6 This is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 6 In the embodiment, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 may also be placed in the execution device 110.
[0148] like Figure 6 As shown, the target model / rule 101 obtained by training with the training device 120 may be a neural network in an embodiment of the present application. Specifically, the neural network provided in the embodiment of the present application may be a CNN and a deep convolutional neural network (DCNN), etc.
[0149] The following is a brief introduction to the application scenarios of the neural network model training method provided in the embodiments of the present application.
[0150] The neural network model training method provided by the embodiment of the present application can be applied in the scene of gesture recognition. For example, in the gesture interaction scene of a smartphone, simple, natural and convenient operation of the smartphone can be achieved through gesture recognition. For example, a smartphone can use a camera or other peripheral camera as an image sensor to obtain image information containing a hand image stream, and then process the image information containing the hand image stream to obtain gesture recognition information, and then report the gesture recognition information to the operating system for response. Through gesture recognition, functions such as turning pages up and down, audio and video playback, volume control, reading and browsing can be realized, which greatly improves the technological sense and interactive convenience of smartphones.
[0151] It should be understood that the above is an example of an application scenario and does not limit the application scenario of the present application.
[0152] Combine the following Figures 7 to 17 The training method of the neural network model provided in the embodiment of the present application is described in detail.
[0153] Figure 7 A flowchart of a training method for a neural network model provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the method includes:
[0154] S101: Acquire a first image set, where the first image set includes a plurality of first sample images, the first sample images include a target object, and a first bounding box obtained by detecting the target object.
[0155] For example, the target detection model can be used to obtain the image captured by the camera for target detection to obtain a first sample image, that is, an image including the target object and a first bounding box annotated on the target object. The first image set may include multiple first sample images.
[0156] Optionally, the target object may refer to a hand of the user, and the first bounding box annotating the target object may refer to a bounding box annotating the hand.
[0157] For example, Figure 2 As shown, the first bounding box may be a bounding box 11 of the hand.
[0158] S102 . Move a first bounding box of a first sample image in the first image set to obtain a second image set.
[0159] The first bounding boxes of all the first sample images in the first image set may be moved to obtain the second image set, or the first bounding boxes of some of the first sample images in the first image set may be moved to obtain the second image set, which is not limited in this embodiment of the present application.
[0160] Exemplarily, there are 1000 first sample images in the first image set, and the first bounding boxes of 30% of the first sample images in the first image set are moved, that is, 300 first sample images are moved to obtain 300 moved sample images, and the 300 moved sample images are used as the second image set.
[0161] The first bounding box in the first sample image may be moved randomly or in a preset direction, which is not limited in this embodiment of the present application.
[0162] S103 . Train the initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model.
[0163] Among them, the first key point recognition model is used to recognize the key points in the target object. The target object may refer to the user's hand, so the first key point model can be used to recognize the hand key points. The initial first key point recognition model may refer to the parameters in the neural network as the initialized parameters, and the hand key points cannot be accurately recognized from the image through the initial first key point recognition model. By training the initialized first key point model and adjusting the initialized parameters, the trained first key point recognition model is obtained, so that the trained first key point recognition model can accurately identify the hand key points in the image.
[0164] When training the initial first key point model, the first sample image in the first image set and the sample image obtained by moving the first bounding box in the first sample image are used. That is to say, the sample image simulates the situation where the hand partially moves out of the camera's field of view when the field of view is small, and the sample image simulates the situation where the hand position in the sample image differs greatly due to hand shaking. Therefore, the trained first key point recognition model obtained by training the initial first key point recognition model is trained for the situation where the hand partially moves out of the camera's field of view due to the small field of view, and the situation where the hand position in the sample image differs greatly due to hand shaking.
[0165] The neural network model training method provided in the embodiment of the present application, the neural network model training method provided in the embodiment of the present application, first obtains a first image set including multiple first sample images, wherein the first sample image includes a target object and a first bounding box obtained by detecting the target object, then moves the first bounding box of the first sample image in the first image set to obtain a second image set, and then trains the initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, the first key point recognition model is used to identify the key points in the target object, and the first bounding box in the first sample image is moved to simulate the situation where the position of the hand in the image is greatly different due to hand shaking during image acquisition, and the situation where the hand is partially moved out of the image due to the small field of view of the camera, so that the sample images for training the initialized first key point recognition model include not only the initial first sample images, but also sample images adjusted for the small field of view of the camera and the hand shaking problem, enriching the types of sample images for training the initial first key point recognition model, thereby improving the robustness of the trained first key point model.
[0166] Since inaccurate gesture recognition results may be caused by hand shaking or the hand moving out of the camera's field of view, we can simulate the hand shaking scene by randomly moving the first bounding box and simulate the hand moving out of the camera's field of view scene by parallel movement, so that the adjusted sample image is more in line with the actual situation and the accuracy of gesture recognition results is improved. Figure 8 The illustrated embodiment is described in detail.
[0167] Figure 8 A flowchart of another neural network model training method provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, the method includes:
[0168] S201: Acquire a first image set, where the first image set includes a plurality of first sample images, the first sample images include a target object, and a first bounding box obtained by detecting the target object.
[0169] S202 : Randomly move a first bounding box in the first sample image to obtain a second sample image, where the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box.
[0170] The following description takes the target object as a hand. Randomly moving the first bounding box in the first sample image can be as follows Fig. 9 (a) and Fig. 9 As shown in (b) in FIG. 1 , the first sample image can be Fig. 9 As shown in (a), including the hand 11 and the first bounding box 12 of the hand, the first bounding box 12 is randomly moved to obtain the following Fig. 9 The second sample image shown in (b) in FIG. 1 , wherein the second sample image also includes the hand 11 and the second bounding box 22. It can be seen that the position of the hand 11 in the second sample image remains unchanged, and is the same as the position of the hand 11 in the first sample image, however, the positions of the first bounding box 12 and the second bounding box 21 are different. In other words, the second sample image is a sample image obtained by moving the first bounding box 12 in the first sample image.
[0171] That is, the second sample image is a sample image obtained by randomly moving the first boundary box in the first sample image. The second sample image obtained by randomly moving the first boundary box can simulate an image obtained by hand shaking.
[0172] The second sample image obtained by randomly moving the first bounding box can simulate an image obtained by hand shaking.
[0173] It should be understood that the number of sample images when training the neural network model is usually large. Therefore, when randomly moving the first bounding box, it can be moved according to preset parameters. For example, the preset parameters include setting the data enhancement probability parameter prob, and the ratio of the height and width of the first bounding box to move_ratio. Among them, prob represents the proportion of the first sample image that needs to move the first bounding box in the first image set; move_ratio represents the ratio of the first distance to the width of the first bounding box, and the ratio of the first distance to the height of the first bounding box; wherein the first distance refers to the maximum moving distance of the first bounding box to the second bounding box. When moving the first bounding box according to move_ratio, a value is randomly selected between 0-move_ratio, and the first bounding box is moved according to the random value. For example, if move_ratio is 0.3, the first bounding box is moved according to move_ratio, and a value is randomly selected from 0-0.3 multiple times to move the first bounding box to obtain multiple second sample images, so that the obtained second sample images can be enriched.
[0174] Fig. 9 (c) in FIG. 1 is a schematic diagram of placing the first bounding box 12 and the second bounding box 22 in the same image, as shown in FIG. Fig. 9 As shown in (c) in FIG. 1 , all sides of the first bounding box 12 and the second bounding box 22 overlap. The first distance may refer to the distance from point A1 in the first bounding box 12 to point A2 in the second bounding box 22.
[0175] S203, parallel moving the first bounding box in the first sample image to obtain a third sample image, where the third sample image includes the target object and a third bounding box obtained by parallel moving the first bounding box, where the third bounding box is obtained by moving the first bounding box along the X-axis direction, or the third bounding box is obtained by moving the first bounding box along the Y-axis direction.
[0176] Different from randomly moving the first bounding box in the first sample image, parallel moving the first bounding box in the first sample image means that when the first bounding box is moved to obtain the third bounding box, the coordinates of the first bounding box and the third bounding box in one direction are the same. For example, the coordinates of the first bounding box on the X-axis are the same as the coordinates of the third bounding box on the X-axis, that is, the first bounding box is translated along the Y-axis direction to obtain the third bounding box; or, the coordinates of the first bounding box on the Y-axis are the same as the coordinates of the third bounding box on the Y-axis, that is, the first bounding box is translated along the X-axis direction to obtain the third bounding box.
[0177] For example, Fig.10 (a) to Fig.10 As shown in (c) in FIG. 1 , the first sample image can be Fig.10As shown in (a) of FIG. 1 , including the hand 11 and the first bounding box 12 of the hand, the first bounding box 12 is moved parallel to the X axis to obtain the following Fig.10 The third sample image shown in (b) in FIG. 1 , wherein the third sample image also includes the hand 11 and the third bounding box 32. It can be seen that the position of the hand 11 in the third sample image remains unchanged, which is the same as the position of the hand 11 in the first sample image, however, the positions of the first bounding box 12 and the third bounding box 32 are different, but the coordinate of the horizontal side of the third bounding box 32 on the Y axis is the same as the coordinate of the horizontal side of the first bounding box 12 on the Y axis. Fig.10 (c) in FIG. 1 is a schematic diagram of placing the first bounding box 12 and the third bounding box 32 in the same image, as shown in FIG. Fig.10 As shown in (c), the horizontal side of the first bounding box 12 and the horizontal side of the third bounding box 32 completely overlap, but the vertical side on the left side of the third bounding box 32 does not overlap with the vertical side on the left side of the first bounding box 12. The vertical side on the left side of the third bounding box 32 is obtained by moving the vertical side on the left side of the first bounding box 12 to the right along the X-axis direction. It can be understood that the bounding box is used to mark the hand and is usually set along the edge of the hand. Therefore, although the vertical side on the left side of the third bounding box 32 and the vertical side on the left side of the first bounding box 12 do not overlap, the vertical side on the right side of the third bounding box 32 overlaps with the vertical side on the right side of the first bounding box 12, and both are vertical sides set along the right edge of the hand. At the same time, it can be seen that in the third sample image, the key points of the little finger tip and the key points of the ring finger tip in the hand key points are outside the third bounding box 32.
[0178] For example, Fig.10 (a) Fig.10 (d) and Fig.10 As shown in (e) in FIG. 1 , the first sample image can be Fig.10 As shown in (a) of FIG. 1 , including the hand 11 and the first bounding box 12 of the hand, the first bounding box 12 is moved parallel to the Y axis to obtain the following Fig.10 The third sample image shown in (d) in FIG. 4 , wherein the third sample image also includes the hand 11 and the third bounding box 42. It can be seen that the position of the hand 11 in the third sample image remains unchanged, and is the same as the position of the hand 11 in the first sample image. However, the positions of the first bounding box 12 and the third bounding box 42 are different, but the coordinates of the vertical side of the third bounding box 42 on the X-axis are the same as the coordinates of the vertical side of the first bounding box 12 on the X-axis. Fig.10 (e) in FIG. 1 is a schematic diagram of placing the first bounding box 12 and the third bounding box 42 in the same image, as shown in FIG. Fig.10As shown in (e), the vertical side of the first bounding box 12 and the horizontal side of the third bounding box 42 completely overlap, but the horizontal side below the third bounding box 4 does not overlap with the horizontal side below the first bounding box 12, and is obtained by moving upward along the Y-axis direction. It can be understood that the bounding box is used to mark the hand, and is usually set along the edge of the hand. Therefore, although the horizontal side below the third bounding box 42 is the horizontal side below the first bounding box 12 moved upward along the Y-axis, the horizontal side above the third bounding box 42 overlaps with the horizontal side above the first bounding box 12, and both are horizontal sides set along the upper edge of the hand.
[0179] That is, the third sample image is a sample image obtained by parallel moving the first boundary frame 12 in the first sample image. The third sample image obtained by parallel moving the first boundary frame can simulate an image when the hand moves out of the field of view of the camera.
[0180] It should be understood that the number of sample images when training a neural network model is usually large. Therefore, when moving the first bounding box in parallel, it can also be moved according to preset parameters. For example, the preset parameters include setting the data enhancement probability parameter prob, and the ratio of the height and width of the first bounding box to the parallel movement shift_ratio. Among them, prob represents the proportion of the first sample images in the first image set that need to move the first bounding box; shift_ratio represents the ratio of the first distance to the width of the first bounding box, and the ratio of the first distance to the height of the first bounding box; wherein the first distance refers to the maximum moving distance of the first bounding box to the third bounding box. When moving the first bounding box according to shift_ratio, a value is randomly selected between 0-shift_ratio, and the first bounding box is moved according to the random value. For example, if shift_ratio is 0.2, the first bounding box is moved according to shift_ratio, and a value is randomly selected from 0-0.2 multiple times to move the first bounding box to obtain multiple third sample images, so that the obtained third sample images can be enriched. Since parallel movement moves along the X-axis direction or the Y-axis direction. Therefore, the first distance can refer to Fig.10 The distance from point A1 in the first bounding box 12 to point A3 in the third bounding box 32 shown in (c) of FIG. 1 ; or, the first distance may refer to Fig.10 The distance from point A1 in the first bounding box 12 to point A4 in the third bounding box 43 shown in (e) in FIG.
[0181] S204: Train the initial first key point recognition model based on the first image set, the second sample image and the third sample image to obtain a trained first key point recognition model, where the first key point recognition model is used to recognize key points in the target object.
[0182] The neural network model training method provided in the embodiment of the present application obtains a first image set, wherein the first image set includes multiple first sample images, the first sample image includes a target object, and a first bounding box obtained by detecting the target object, randomly moves the bounding box in the first sample image to obtain a second sample image, the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box; parallel moves the first bounding box in the first sample image to obtain a third sample image, the third sample image includes the target object and a third bounding box obtained by parallel moving the first bounding box, the third bounding box is obtained by moving the first bounding box along the X-axis direction, or the third bounding box is obtained by moving the first bounding box along the Y-axis direction, based on the first image set, the second sample image and the third sample image for initial The first key point recognition model is trained to obtain the trained first key point recognition model, which is used to recognize key points in the target object. The first bounding box in the first sample image is randomly moved to simulate the situation where the position of the hand in the image is greatly different due to hand shaking during the image acquisition process. The parallel movement is used to simulate the situation where the hand partially moves out of the image due to the small field of view of the camera. The sample image of the first key point recognition model initialized in the training includes not only the initial first sample image, but also the second sample image and the third sample image adjusted for hand shaking and the hand moving out of the field of view of the camera, which enriches the types of sample images for training the initial first key point recognition model, thereby improving the robustness of the trained first key point model.
[0183] Since the camera's field of view is small and the distance the hand moves out of the camera's field of view is large, when the third sample image is obtained by parallel moving the first bounding box in the first sample image to simulate the above situation, the first bounding box is usually moved a large distance, which makes the gap between the first sample image and the third sample image large. If the initial first key point recognition model is trained directly with the third sample image, it may cause the model to not converge. Therefore, the initial first key point recognition model can be pre-trained with sample images with a small first bounding box movement distance to obtain a pre-trained first key point recognition model, and then the pre-trained first key point recognition model can be trained with sample images with a large first bounding box movement distance to obtain a trained first key point recognition model. Fig.11 The illustrated embodiment is described in detail.
[0184] Fig.11 A flowchart of another neural network model training method provided in an embodiment of the present application is shown in FIG. Fig.11 As shown, the method includes:
[0185] S301: Acquire a first image set, where the first image set includes a plurality of first sample images, the first sample images include a target object, and a first bounding box obtained by detecting the target object.
[0186] S302 : randomly moving a bounding box in the first sample image to obtain a second sample image, where the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box.
[0187] S303 : parallel move the first boundary box in the first sample image according to the first parameter to obtain a third sample image.
[0188] The first parameter includes a first sub-parameter and a second sub-parameter, the first sub-parameter is used to indicate the ratio of the number of third sample images obtained by parallel movement to the number of first sample images in the first image set, and the second sub-parameter is used to indicate the ratio of the first distance to the width or height of the first bounding box, and the first distance refers to the distance from the first bounding box to the third bounding box.
[0189] For example, the first sub-parameter may refer to setting the data enhancement probability parameter prob; the second sub-parameter may refer to the ratio shift_ratio of the parallel shift relative to the height and width of the first bounding box.
[0190] S304 , pre-training the initial first key point recognition model based on the first image set, the second sample image, and the third sample image to obtain a pre-trained first key point recognition model.
[0191] It is understandable that the third sample image is an image simulating a hand partially moving out of the camera's field of view, in which some key points of the hand may be missing. Therefore, if a smaller data enhancement probability parameter and a smaller ratio of the height and width of the parallel movement relative to the first bounding box are directly used to obtain the third sample image, the parameters of the neural network model may not converge during the training process, and the first key point recognition model after training cannot be obtained. Therefore, a larger data enhancement probability parameter and a larger ratio of the height and width of the parallel movement relative to the first bounding box can be used to obtain the third sample image to reduce the difficulty of training.
[0192] For example, the third sample image can be obtained by first using the data enhancement probability parameter prob=0.3 and the parameter shift_ratio=0.4 of the ratio of the height and width of the first bounding box to parallel movement, and then the initial first key point recognition model is pre-trained using the first image set, the second sample image and the third sample image to obtain the pre-trained first key point recognition model.
[0193] In order to further reduce the difficulty of training the neural network model, the points outside the bounding box can be shielded first. During the training of the neural network model, the points outside the bounding box do not participate in the model training when the loss function is returned, which can further reduce the difficulty of training the neural network model.
[0194] Optionally, the image area outside the third bounding box in the third sample image is removed to obtain an updated third sample image; and the initial first key point recognition model is pre-trained based on the first image set, the second sample image and the updated third sample image to obtain a pre-trained first key point recognition model.
[0195] The training method of the neural network model provided in the embodiment of the present application, during the pre-training process of the initial first key point recognition model, removes the image area outside the third bounding box in the third sample image to obtain an updated third sample image, and then pre-trains the initial first key point recognition model based on the first image set, the second sample image and the updated third sample image to obtain the pre-trained first key point recognition model. This can avoid points outside the third bounding box from participating in model training, reduce the amount of data for pre-training the neural network model, and thereby reduce the difficulty of pre-training the initial first key point recognition model.
[0196] S305 . Move the bounding box in the first sample image in parallel according to the second parameter to obtain a fourth sample image, where the fourth sample image includes the target object and a fourth bounding box.
[0197] The second parameter includes a third sub-parameter and a fourth sub-parameter, the third sub-parameter is used to indicate the ratio of the number of fourth sample images obtained by parallel movement to the number of first sample images in the first image set, and the fourth sub-parameter is used to indicate the ratio of the second distance to the width or height of the first bounding box, and the second distance refers to the distance from the first bounding box to the fourth bounding box.
[0198] For example, the third sub-parameter may be a data enhancement probability parameter, and the fourth sub-parameter may be a ratio of the parallel movement to the height and width of the first bounding box.
[0199] Optionally, the first sub-parameter is greater than the third sub-parameter, and the second sub-parameter is greater than the fourth sub-parameter.
[0200] Exemplarily, the third sub-parameter prob=0.1, and the fourth sub-parameter shift_ratio=0.3.
[0201] If a smaller second sub-parameter is used to obtain the fourth sample image, the difference between the fourth sample image and the first sample image is smaller. Therefore, the neural network model can be trained based on the first sample image and the fourth sample image with a smaller difference from the first sample image, which can improve the accuracy of the trained neural network model, that is, improve the accuracy of the first key point recognition model after training.
[0202] S306 , training the pre-trained first key point recognition model based on the first image set, the second sample image, and the fourth sample image to obtain a trained first key point recognition model.
[0203] The training method of the neural network model provided in the embodiment of the present application obtains a first image set, wherein the first image set includes multiple first sample images, the first sample image includes a target object, and a first bounding box obtained by detecting the target object, randomly moves the bounding box in the first sample image to obtain a second sample image, the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box; uses a first parameter to parallel move the first bounding box in the first sample image to obtain a third sample image, the third sample image includes the target object and a third bounding box obtained by parallel moving the first bounding box, which is obtained by moving the first bounding box along the X-axis direction, or the third bounding box is obtained by moving the first bounding box along the Y-axis direction, and pre-trains an initial first key point recognition model based on the first image set, the second sample image, and the third sample image to obtain To the pre-trained first key point recognition model, and then use the second parameter to parallel move the first bounding box in the first sample image to obtain the fourth sample data, and then continue to train the pre-trained first key point recognition model based on the first sample image, the second sample image and the fourth sample image to obtain the trained first key point recognition model. This is equivalent to dividing the training of the initial first key point recognition model into two steps: pre-training and training. Since the hand part moves out of the field of view of the camera, the hand key point may not be within the bounding box. Therefore, if the initial first key point recognition model is trained in only one step, it may lead to the model training not converging due to the hand key point not being in the bounding box. The use of pre-training and training in two steps for model training can effectively avoid the model training not converging due to the hand key point not being in the bounding box.
[0204] It is understandable that some gestures are perpendicular to the screen of the electronic device, such as Fig.12In this case, if the coordinates of the hand key points in the image are usually two-dimensional coordinates, gesture recognition based only on the coordinates of the hand key points in the image may lead to inaccurate gesture recognition results. Therefore, the electronic device converts the hand key points into three-dimensional coordinates based on the depth information of the image to improve the accuracy of gesture recognition based on the hand key points. Fig.13 Let me explain in detail.
[0205] Fig.13 A flowchart of another neural network model training method provided in an embodiment of the present application is shown in FIG. Fig.13 As shown, the method includes:
[0206] S401: Acquire a third image set, where the third image set includes a fifth sample image, and the fifth sample image is an image captured by a camera device.
[0207] It is understandable that the camera device of the electronic device generally includes a color camera and a depth camera. The fifth sample image may be an image captured by the color camera and the depth camera, which includes not only the color information of each pixel in the image, but also the depth information of each pixel.
[0208] S402: Obtain a sixth sample image based on the fifth sample image and the object detection model, where the sixth sample image includes a first bounding box.
[0209] The target detection model can be used to annotate the target object in the image to obtain the bounding box of the target object. By annotating the target object in the fifth sample image through the target detection model, a sixth sample image including the fifth sample image and the bounding box of the target object in the fifth sample image (that is, the first bounding box) can be obtained. This is equivalent to the sixth sample including the fifth sample image and the first bounding box.
[0210] S403 . Perform three-dimensional conversion on the sixth sample image based on depth information corresponding to the fifth sample image to obtain a first sample image.
[0211] Among them, the electronic device can perform a three-dimensional conversion on the sixth sample image including the first bounding box based on the depth information in the fifth sample image to obtain the first sample image represented by three-dimensional coordinates, that is, the coordinates of each pixel point in the first sample image also include depth information.
[0212] Optionally, the third sample image is transformed into a third sample image by using the first formula and the depth information corresponding to the fifth sample image to obtain the first sample image. The first formula includes:
[0213] x = X / W;
[0214] y=Y / H;
[0215] z=(Z-Z0) / W;
[0216] Among them, x is the coordinate of the pixel point on the x-axis in the first sample image, y is the coordinate of the pixel point on the y-axis in the first sample image, z is the coordinate of the pixel point on the z-axis in the first sample image, X is the coordinate of the pixel point corresponding to x in the sixth sample image, Y is the coordinate of the pixel point corresponding to y in the sixth sample image, Z is the coordinate of the pixel point corresponding to z in the sixth sample image, Z0 represents the coordinate of the target key point on the z-axis in the sixth sample image, W represents the width of the first bounding box, and H represents the height of the first bounding box.
[0217] The target key point can be Figure 3 The wrist (0) is a key point of the hand. The coordinates of each pixel point on the z-axis can be represented by the depth information of each pixel point.
[0218] S404: Acquire a first image set, where the first image set includes a plurality of first sample images, the first sample images include a target object, and a first bounding box obtained by detecting the target object.
[0219] S405 . Randomly move the bounding box in the first sample image to obtain a second sample image, where the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box.
[0220] It can be understood that since the coordinates of each pixel point in the first sample image are three-dimensional coordinates, and the second sample image is obtained by randomly moving the first bounding box in the first sample image, the coordinates of each pixel point in the second sample image are also three-dimensional coordinates, including the depth information of each pixel point.
[0221] S406 : Move the boundary box in the first sample image in parallel according to the first parameter to obtain a third sample image.
[0222] The first parameter includes a first sub-parameter and a second sub-parameter, the first sub-parameter is used to indicate the ratio of the number of third sample images obtained by parallel movement to the number of first sample images in the first image set, and the second sub-parameter is used to indicate the ratio of the first distance to the width or height of the first bounding box, and the first distance refers to the distance from the first bounding box to the third bounding box.
[0223] It can be understood that since the coordinates of each pixel point in the first sample image are three-dimensional coordinates, the third sample image is obtained by parallel moving the first boundary in the first sample image according to the first parameter, so the coordinates of each pixel point in the third sample image are also three-dimensional coordinates, including the depth information of each pixel point.
[0224] S407 , pre-training the initial first key point recognition model based on the first image set, the second sample image, and the third sample image to obtain a pre-trained first key point recognition model.
[0225] When pre-training the initial first key point recognition model, the image area outside the third bounding box in the third sample image can be removed, which can further reduce the amount of data for pre-training the initial first key point recognition model, thereby reducing the difficulty of training the initial first key point recognition model.
[0226] Optionally, the image area outside the third bounding box in the third sample image is removed to obtain an updated third sample image, and the initial first key point recognition model is pre-trained based on the first image set, the second sample image and the updated third sample image to obtain a pre-trained first key point recognition model.
[0227] The training method of the neural network model provided in the embodiment of the present application, when pre-training the initial first key point recognition model, removes the image area outside the third bounding box in the third sample image to obtain an updated third sample image, and then pre-trains the initial first key point recognition model based on the first image set, the second sample image and the updated third sample image to obtain the pre-trained first key point recognition model, which can reduce the amount of data for pre-training the initial first key point recognition model, thereby reducing the difficulty of training the initial first key point recognition model.
[0228] S408 . Move the boundary box in the first sample image in parallel according to the second parameter to obtain a fourth sample image.
[0229] The fourth sample image includes the target object and the fourth bounding box, the second parameter includes a third sub-parameter and a fourth sub-parameter, the third sub-parameter is used to indicate the ratio of the number of fourth sample images obtained by parallel movement to the number of first sample images in the first image set, and the fourth sub-parameter is used to indicate the ratio of the second distance to the width or height of the first bounding box, and the second distance refers to the distance from the first bounding box to the fourth bounding box.
[0230] Optionally, the first sub-parameter is greater than the third sub-parameter, and the second sub-parameter is greater than the fourth sub-parameter.
[0231] It should be understood that the first sub-parameter is used to indicate the ratio of the number of third sample images obtained by parallel movement to the number of first sample images in the first image set, and the third sub-parameter is used to indicate the ratio of the number of fourth sample images obtained by parallel movement to the number of first sample images in the first image set. Since the third sample image obtained by moving the first bounding box is to simulate the hand moving out of the field of view of the camera, the first sub-parameter can also be regarded as the data enhancement rate, that is, the proportion of images simulating the hand moving out of the field of view of the camera. Similarly, the third sub-parameter is used to indicate the ratio of the number of fourth sample images obtained by parallel movement to the number of first sample images in the first image set, which can also be regarded as the data enhancement rate, that is, the proportion of images simulating the hand moving out of the field of view of the camera. The first sub-parameter being greater than the third sub-parameter may mean that the ratio of the number of third sample images to the number of first sample images is greater than the ratio of the number of fourth sample images to the number of first sample images.
[0232] The second sub-parameter is used to indicate the ratio of the first distance to the width or height of the first bounding box, and the first distance refers to the distance from the first bounding box to the third bounding box; the fourth sub-parameter is used to indicate the ratio of the second distance to the width or height of the first bounding box, and the second distance refers to the distance from the first bounding box to the fourth bounding box. The third sample image is obtained by moving the first bounding box in the first sample image parallel to the X axis, or along the Y axis, and the fourth sample image is obtained by moving the first bounding box in the first sample image parallel to the X axis, or along the Y axis. Therefore, the second sub-parameter is greater than the fourth sub-parameter, which means that the moving distance of the first bounding box corresponding to the third sample image is greater than the moving distance of the first bounding box corresponding to the fourth sample image, that is, the coordinate difference between the third bounding box and the first bounding box in the third sample image is greater than the coordinate difference between the fourth bounding box and the first bounding box in the fourth sample image.
[0233] For example, Fig.14 (a) is a schematic diagram of setting the first bounding box 22 and the third bounding box 32 in the same image. Fig.14 (b) is a schematic diagram of setting the first bounding box 22 and the fourth bounding box 42 in the same image. The second sub-parameter being greater than the fourth sub-parameter may mean that the distance between the left vertex A3 of the third bounding box and the left vertex A1 of the first bounding box is greater than the distance between the left vertex A4 of the fourth bounding box and the left vertex A1 of the first bounding box. This is equivalent to that the movement distance of the third bounding box in the third sample image is greater than the movement distance of the fourth bounding box in the fourth sample image.
[0234] In the training method of the neural network model provided in the embodiment of the present application, when pre-training the initial first key point recognition model, the data enhancement rate (first sub-parameter) used is greater than the data enhancement rate (third sub-parameter) used when continuing to train the pre-trained first key point recognition model, and the moving distance of the third bounding box in the third sample image used when pre-training the initial first key point recognition model is greater than the moving distance of the fourth bounding box in the fourth sample image used when continuing to train the pre-trained first key point recognition model. This is equivalent to first pre-training the initial first key point recognition model with the third sample image that shields the image area outside the bounding box until the model converges, so that the next step is to continue to train the pre-trained first key point recognition model with the fourth sample image including the image area outside the bounding box, so that when the model can converge, the prediction ability of the trained first key point recognition model for the image area outside the bounding box is improved, and the stability of the trained first key point recognition model is improved.
[0235] S409: Train the pre-trained first key point recognition model based on the first image set, the second sample image, and the fourth sample image to obtain a trained first key point recognition model.
[0236] It should be noted that the first sample image is an image of three-dimensional coordinates, that is, the first sample image includes depth information. Correspondingly, the fourth sample image is also an image of three-dimensional coordinates, and the fourth sample image also includes depth information.
[0237] The training method of the neural network model provided in the embodiment of the present application obtains a third image set including a fifth sample image captured by a camera device, and then obtains a sixth sample image based on the fifth sample image and a target detection model, wherein the sixth sample image includes a first bounding box, and then performs a three-dimensional conversion on the sixth sample image based on the depth information corresponding to the fifth sample image to obtain a first sample image, which is equivalent to the first sample image being an image of three-dimensional coordinates, that is, the first sample image includes the depth information of each pixel point, and then the second sample image, the third sample image and the fourth sample image obtained based on the first sample image also include the depth information of each pixel point. In this way, in the process of training the initial first key point model, the first sample image, the second sample image, the third sample image and the fourth sample image used are all three-dimensional images and all include depth information, so that the coordinates of the hand key points obtained based on the trained first key point recognition model are three-dimensional coordinates including depth information, so that the gesture recognition based on the coordinates of the three-dimensional hand key points is more accurate, especially when recognizing gestures perpendicular to the screen of the electronic device, the gesture recognition result obtained is more accurate.
[0238] It is understandable that during the recognition of some special gestures, the coordinate positions of some hand key points change greatly, while the positions of other hand key points change less. For example, in the gesture of pinching the thumb and index finger, the changes to the two hand key points, the thumb tip (4) and the index finger tip (8), are relatively large. If a neural network model is trained based on all the hand key points, it will occupy a large amount of electronic device resources, resulting in unnecessary waste of resources. Therefore, the key point recognition model can be trained for some hand key points. Fig.15 Let me explain in detail.
[0239] Fig.15 A flowchart of another neural network model training method provided in an embodiment of the present application is shown in FIG. Fig.15 As shown, the method includes:
[0240] S501: Acquire a first image set, where the first image set includes a plurality of first sample images, the first sample images include a target object, and a first bounding box obtained by detecting the target object.
[0241] S502 : randomly moving a bounding box in the first sample image to obtain a second sample image, where the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box.
[0242] S503 : parallel-move the first boundary box in the first sample image to obtain a third sample image.
[0243] The third sample image includes the target object and a third bounding box obtained by parallel movement of the first bounding box, the third bounding box has the same coordinates as the first bounding box on the X-axis and different coordinates on the Y-axis, or the third bounding box has the same coordinates as the first bounding box on the Y-axis and different coordinates on the X-axis.
[0244] S504: Train the initial first key point recognition model based on the first image set, the second sample image and the third sample image to obtain a trained first key point recognition model, where the first key point recognition model is used to recognize key points in the target object.
[0245] S505: Acquire a fourth image set, where the fourth image set includes a seventh sample image, and the seventh sample image is an image obtained by a preset gesture for the target object.
[0246] Taking the recognition of the thumb and index finger pinching gesture as an example, the seventh sample image may be a sample image in which the two hand key points of the thumb tip (4) and the index finger tip (8) are annotated. Compared with the first sample image, the number of hand key points in the seventh sample image is less than the number of hand key points in the first sample image.
[0247] S506: Train the initial second key point recognition model based on the fourth image set to obtain a trained second key point recognition model, where the second key point recognition model is used to recognize key points corresponding to the preset gesture.
[0248] The second key point recognition model may be a neural network model with the same structure as the first key point recognition model. The electronic device may train the initial second key point recognition model using the fourth image set, that is, a plurality of seventh sample images annotated for the preset gestures, to obtain a trained second key point recognition model. Since the second key point recognition model may be a neural network model with the same structure as the first key point recognition model, the trained second key point recognition model has the same structure as the trained first key point recognition model, but has different parameters.
[0249] In one possible case, the first key point recognition model and the second key point recognition model reuse part of the network structure.
[0250] Optionally, the first key point recognition model includes a first network and a first regression layer, the second key point recognition model includes a first network and a second regression layer, and the dimension of the first regression layer is higher than the dimension of the second regression layer.
[0251] For example, Fig.16 As shown, the first key point recognition model and the second key point recognition model reuse the first network. Since the number of conventional hand key points is 21, in order to recognize the 21 hand key points, the first regression layer in the first key point recognition model can be a regression layer with an output dimension of 3*21, that is, a regression layer with an output dimension of 63. When recognizing the pinch gesture, the focus is on recognizing the two hand key points of the thumb tip (4) and the index finger tip (8), so a regression layer with an input dimension of n*3 can be used, where n is the number of hand key points that need to be strengthened, that is, the second regression layer can be a regression layer with an output dimension of 2*3, that is, a regression layer with an output dimension of 6.
[0252] After the initial first key point recognition model is trained by the first image set and the second image set to obtain the trained first key point recognition model, the electronic device can freeze the parameters of the first network and use the fourth image set to train the second key point recognition model composed of the first network and the second regression layer until the model converges.
[0253] The training method of the neural network model provided in the embodiment of the present application, the first key point recognition model includes a first network and a first regression layer, the second key point recognition model includes a first network and a second regression layer, the dimension of the first regression layer is higher than the dimension of the second regression layer, and the second regression layer with lower dimension is used to train the hand key points corresponding to the preset gesture, and only a small number of hand key points corresponding to the preset gesture are trained. In this way, the hand key points corresponding to the preset gesture can be strengthened while avoiding the influence on other hand key points, thereby improving the recognition ability of the hand key points corresponding to the preset gesture.
[0254] S507: Correct the parameters in the trained first key point recognition model based on the parameters in the trained second key point recognition model to obtain an updated first key point recognition model.
[0255] As can be seen from the above description of the neural network, the output of the computing unit of the neural unit in the neural network can be shown as formula (1), W s Input x to the neural unit s The weight of x in the second key point recognition model after training can be expressed as W′=W1*W2*…*W i It means that the bias b′ of the neural unit can be obtained by b′=f(W i , b i-1 , b i ) represents, where f(W i , b i-1 , b i )=W i *b i-1 +b i .
[0256] The weight W0 of x input to the neural unit in the trained first key point recognition model and the bias b0 of the neural unit. The weight W′ of x input to the neural unit in the trained second key point recognition model and the bias b′ of the neural unit are used to correct the trained first key point recognition model, and the updated first key point recognition model can be obtained. The weight W=W0+W′ of x input to the neural unit and the bias b=b0+b′ of the neural unit in the updated first key point recognition model.
[0257] The training method of the neural network model provided in the embodiment of the present application is to obtain a trained second key point recognition model by training the hand key points corresponding to special gestures, and then use the parameters of the trained second key point recognition model to correct the parameters of the trained first key point recognition model to obtain an updated first key point recognition model, so that the updated first key point recognition model can also accurately recognize the hand key points corresponding to special gestures, thereby improving the robustness of the updated first key point recognition model.
[0258] Fig.17 A flowchart of another neural network model training method provided in an embodiment of the present application is shown in FIG. Fig.17 As shown, the method includes:
[0259] S601: Acquire a third image set, where the third image set includes a fifth sample image, and the fifth sample image is an image captured by a camera device.
[0260] S602: Obtain a sixth sample image based on the fifth sample image and the target detection model, where the sixth sample image includes a first bounding box.
[0261] S603 . Perform three-dimensional conversion on the sixth sample image based on depth information corresponding to the fifth sample image to obtain a first sample image.
[0262] Optionally, the third sample image is transformed into a third sample image by using the first formula and the depth information corresponding to the fifth sample image to obtain the first sample image. The first formula includes:
[0263] x = X / W;
[0264] y=Y / H;
[0265] z=(Z-Z0) / W;
[0266] Among them, x is the coordinate of the pixel point on the x-axis in the first sample image, y is the coordinate of the pixel point on the y-axis in the first sample image, z is the coordinate of the pixel point on the z-axis in the first sample image, X is the coordinate of the pixel point corresponding to x in the sixth sample image, Y is the coordinate of the pixel point corresponding to y in the sixth sample image, Z is the coordinate of the pixel point corresponding to z in the sixth sample image, Z0 represents the coordinate of the target key point on the z-axis in the sixth sample image, W represents the width of the first bounding box, and H represents the height of the first bounding box.
[0267] S604: Acquire a first image set, where the first image set includes a plurality of first sample images, the first sample images include a target object, and a first bounding box obtained by detecting the target object.
[0268] S605 : Randomly move the boundary box in the first sample image to obtain a second sample image.
[0269] The second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box.
[0270] S606 : Move the boundary box in the first sample image in parallel according to the first parameter to obtain a third sample image.
[0271] The first parameter includes a first sub-parameter and a second sub-parameter, the first sub-parameter is used to indicate the ratio of the number of third sample images obtained by parallel movement to the number of first sample images in the first image set, and the second sub-parameter is used to indicate the ratio of the first distance to the width or height of the first bounding box, and the first distance refers to the distance from the first bounding box to the third bounding box.
[0272] S607 , pre-training the initial first key point recognition model based on the first image set, the second sample image, and the third sample image to obtain a pre-trained first key point recognition model.
[0273] S608 : Move the boundary box in the first sample image in parallel according to the second parameter to obtain a fourth sample image.
[0274] The fourth sample image includes the target object and the fourth bounding box, the second parameter includes a third sub-parameter and a fourth sub-parameter, the third sub-parameter is used to indicate the ratio of the number of fourth sample images obtained by parallel movement to the number of first sample images in the first image set, and the fourth sub-parameter is used to indicate the ratio of the second distance to the width or height of the first bounding box, and the second distance refers to the distance from the first bounding box to the fourth bounding box.
[0275] S609: Train the pre-trained first key point recognition model based on the first image set, the second sample image, and the fourth sample image to obtain a trained first key point recognition model.
[0276] S610: Acquire a fourth image set, where the fourth image set includes a seventh sample image, and the seventh sample image is an image obtained by a preset gesture for a target object.
[0277] S611. Train the initial second key point recognition model based on the fourth image set to obtain a trained second key point recognition model, where the second key point recognition model is used to recognize key points corresponding to the preset gesture.
[0278] S612: Correct the parameters in the trained first key point recognition model based on the parameters in the trained second key point recognition model to obtain an updated first key point recognition model.
[0279] The training method of the neural network model provided in the embodiment of the present application has the same implementation principle and beneficial effects as the above Figures 7 to 16 The embodiments shown are similar and will not be described again here.
[0280] It should be understood that, although each step in the flow chart in the above-described embodiment is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is clear explanation in this article, the execution of these steps does not have strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flow chart may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0281] It is understandable that in order to implement the above functions, the electronic device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of the present application.
[0282] The embodiment of the present application can divide the functional modules of the electronic device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one module. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. It should be noted that the names of the modules in the embodiment of the present application are schematic and are not limited to the names of the modules in actual implementation.
[0283] Fig.18 A schematic diagram of the structure of a training device for a neural network model provided in an embodiment of the present application.
[0284] It should be understood that the training device 600 of the neural network model can execute Figures 7 to 17 The training method of the neural network model shown; the training device 600 of the neural network model includes: an acquisition unit 610 and a processing unit 620.
[0285] The acquisition unit 610 is used to acquire a first image set, where the first image set includes a plurality of first sample images, the first sample images include a target object, and a first bounding box obtained by detecting the target object;
[0286] The processing unit 620 is used to move the first bounding box of the first sample image in the first image set to obtain a second image set;
[0287] The processing unit 620 is used to train the initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, and the first key point recognition model is used to recognize key points in the target object.
[0288] In one embodiment, the second image set includes a second sample image and a third sample image, and the processing unit 620 is specifically used to randomly move the first bounding box in the first sample image to obtain the second sample image, where the second sample image includes the target object and the second bounding box obtained by randomly moving the first bounding box; parallel move the bounding box in the first sample image to obtain the third sample image, where the third sample image includes the target object and the third bounding box obtained by parallel moving the first bounding box, where the third bounding box is obtained by parallel moving the first bounding box along the X-axis, or the third bounding box is obtained by parallel moving the first bounding box along the Y-axis.
[0289] In one embodiment, the processing unit 620 is specifically used to parallel move the first bounding box in the first sample image according to the first parameter to obtain the third sample image, the first parameter includes a first sub-parameter and a second sub-parameter, the first sub-parameter is used to indicate the ratio of the number of third sample images obtained by parallel movement to the number of first sample images in the first image set, and the second sub-parameter is used to indicate the ratio of the first distance to the width or height of the first bounding box, and the first distance refers to the distance from the first bounding box to the third bounding box; based on the first image set, the second sample image and the third sample image, the initial first key point recognition model is pre-trained to obtain the pre-trained first key point recognition model ; The bounding box in the first sample image is moved in parallel according to the second parameter to obtain a fourth sample image, the fourth sample image includes a target object and a fourth bounding box, the second parameter includes a third sub-parameter and a fourth sub-parameter, the third sub-parameter is used to indicate the ratio of the number of fourth sample images obtained by parallel movement to the number of first sample images in the first image set, and the fourth sub-parameter is used to indicate the ratio of the second distance to the width or height of the first bounding box, and the second distance refers to the distance from the first bounding box to the fourth bounding box; based on the first image set, the second sample image and the fourth sample image, the pre-trained first key point recognition model is trained to obtain the trained first key point recognition model.
[0290] In one embodiment, the first sub-parameter is greater than the third sub-parameter, and the second sub-parameter is greater than the fourth sub-parameter.
[0291] In one embodiment, the processing unit 620 is specifically used to remove the image area outside the third bounding box in the third sample image to obtain an updated third sample image; pre-train the initial first key point recognition model based on the first image set, the second sample image and the updated third sample image to obtain a pre-trained first key point recognition model.
[0292] In one embodiment, the acquisition unit 610 is also used to acquire a third image set, which includes a fifth sample image, and the fifth sample image is an image captured by a camera device; the processing unit 620 is also used to obtain a first sample image based on the fifth sample image, depth information corresponding to the fifth sample image, and a target detection model, and the target detection model is used to mark a first bounding box of the target object.
[0293] In one embodiment, the processing unit 620 is specifically used to obtain a sixth sample image based on the fifth sample image and the target detection model, wherein the sixth sample image includes a first bounding box; and perform a three-dimensional conversion on the sixth sample image based on depth information corresponding to the fifth sample image to obtain the first sample image.
[0294] In one embodiment, the processing unit 620 is specifically configured to perform three-dimensional coordinate transformation on the sixth sample image by using the first formula and the depth information corresponding to the fifth sample image to obtain the first sample image, and the first formula includes:
[0295] x = X / W;
[0296] y=Y / H;
[0297] z=(Z-Z0) / W;
[0298] Among them, x is the coordinate of the pixel point on the x-axis in the first sample image, y is the coordinate of the pixel point on the y-axis in the first sample image, z is the coordinate of the pixel point on the z-axis in the first sample image, X is the coordinate of the pixel point corresponding to x in the sixth sample image, Y is the coordinate of the pixel point corresponding to y in the sixth sample image, Z is the coordinate of the pixel point corresponding to z in the sixth sample image, Z0 represents the coordinate of the target key point on the z-axis in the sixth sample image, W represents the width of the first bounding box, and H represents the height of the first bounding box.
[0299] In one embodiment, the target object comprises a hand in the image.
[0300] In one embodiment, the first key point recognition model is used to recognize hand key points in an image.
[0301] In one embodiment, the acquisition unit 610 is also used to acquire a fourth image set, which includes a seventh sample image, and the seventh sample image is an image obtained for a preset gesture of the target object; the processing unit 620 is also used to train the initial second key point recognition model based on the fourth image set to obtain a trained second key point recognition model, and the second key point recognition model is used to identify key points corresponding to the preset gesture; based on the parameters in the trained second key point recognition model, the parameters in the trained first key point recognition model are corrected to obtain an updated first key point recognition model.
[0302] In one embodiment, the first key point recognition model includes a first network and a first regression layer, the second key point recognition model includes a first network and a second regression layer, and the dimension of the first regression layer is higher than the dimension of the second regression layer.
[0303] The training device for the neural network model provided in this embodiment is used to execute the training method for the neural network model of the above-mentioned embodiment. The technical principle and technical effect are similar and will not be described in detail here.
[0304] It should be noted that the training device 600 of the neural network model is embodied in the form of a functional unit. The term "unit" here can be implemented in the form of software and / or hardware, and is not specifically limited to this.
[0305] For example, a "unit" may be a software program, a hardware circuit, or a combination of the two that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (e.g., a shared processor, a dedicated processor, or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combined logic circuit, and / or other suitable components that support the described functions.
[0306] Therefore, the units of each example described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.
[0307] Fig.19 A structural schematic diagram of an electronic device provided by the present application is shown. Fig.19 The dotted line in indicates that the unit or the module is optional. The electronic device 700 can be used to implement the training method of the neural network model described in the above method embodiment.
[0308] The electronic device 700 includes one or more processors 701, which can support the training method of the neural network model in the method embodiment of the electronic device 700. The processor 701 can be a general-purpose processor or a special-purpose processor. For example, the processor 701 can be a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.
[0309] The processor 701 may be used to control the electronic device 700, execute software programs, and process data of the software programs. The electronic device 700 may also include a communication unit 705 to implement input (reception) and output (transmission) of signals.
[0310] For example, the electronic device 700 may be a chip, the communication unit 705 may be an input and / or output circuit of the chip, or the communication unit 705 may be a communication interface of the chip, and the chip may be a component of a terminal device or other electronic devices.
[0311] For another example, the electronic device 700 may be a terminal device, and the communication unit 705 may be a transceiver of the terminal device, or the communication unit 705 may be a transceiver circuit of the terminal device.
[0312] The electronic device 700 may include one or more memories 702 on which a program 704 is stored. The program 704 can be executed by the processor 701 to generate instructions 703, so that the processor 701 executes the impedance matching method described in the above method embodiment according to the instructions 703.
[0313] Optionally, data may be stored in the memory 702. Optionally, the processor 701 may read data stored in the memory 702. The data may be stored at the same storage address as the program 704, or may be stored at a different storage address than the program 704.
[0314] The processor 701 and the memory 702 may be provided separately or integrated together; for example, integrated on a system on chip (SOC) of a terminal device.
[0315] Exemplarily, the memory 702 can be used to store the relevant program 704 of the training method of the neural network model provided in the embodiment of the present application, and the processor 701 can be used to call the relevant program 704 of the training method of the neural network model stored in the memory 702 when training the neural network model, and execute the training method of the neural network model of the embodiment of the present application; including: acquiring a first image set, the first image set including multiple first sample images, the first sample image including a target object, and a first bounding box obtained by detecting the target object; moving the first bounding box of the first sample image in the first image set to obtain a second image set; training the initial first key point recognition model based on the first image set and the second image set to obtain a trained first key point recognition model, and the first key point recognition model is used to identify key points in the target object.
[0316] Exemplarily, the memory 702 can be used to store the relevant program 704 of the training method of the neural network model provided in the embodiment of the present application, and the processor 701 can be used to call the relevant program 704 of the training method of the neural network model stored in the memory 702 when training the neural network model, and execute the training method of the neural network model of the embodiment of the present application; including: training the initial first key point recognition model based on the first image set and the second image set to obtain the trained first key point recognition model; acquiring a fourth image set, the fourth image set including a seventh sample image, the seventh sample image is an image obtained for a preset gesture of the target object; training the initial second key point recognition model based on the fourth image set to obtain a trained second key point recognition model, the second key point recognition model is used to identify the key points corresponding to the preset gesture; based on the parameters in the trained second key point recognition model, correct the parameters in the trained first key point recognition model to obtain an updated first key point recognition model, wherein the first image set includes multiple first sample images, the first sample image includes a target object, and a first bounding box obtained by detecting the target object; the sample image in the second image set is obtained by moving the first bounding box of the first sample image in the first image set.
[0317] The present application also provides a computer program product, which, when executed by processor 701, implements the training method of the neural network model described in any method embodiment of the present application.
[0318] The computer program product may be stored in the memory 702 , for example, a program 704 , which is converted into an executable target file that can be executed by the processor 701 after preprocessing, compiling, assembling, and linking.
[0319] The present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a computer, the training method of the neural network model described in any method embodiment of the present application is implemented. The computer program can be a high-level language program or an executable target program.
[0320] The computer-readable storage medium is, for example, a memory 702. The memory 702 may be a volatile memory or a nonvolatile memory, or the memory 702 may include both a volatile memory and a nonvolatile memory. Among them, the nonvolatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0321] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0322] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0323] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0324] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0325] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0326] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0327] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0328] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A training method for a neural network model, characterized in that: The method comprises: Acquire a first image set, wherein the first image set includes a plurality of first sample images, the first sample images include a target object, and a first bounding box obtained by detecting the target object; Moving the first bounding box of the first sample image in the first image set to obtain a second image set; An initial first key point recognition model is trained based on the first image set and the second image set to obtain a trained first key point recognition model, where the first key point recognition model is used to recognize key points in the target object.
2. The method according to claim 1, characterized in that The second image set includes a second sample image and a third sample image, and moving the bounding box of the first sample image in the first image set to obtain the second image set includes: randomly moving the first bounding box in the first sample image to obtain a second sample image, wherein the second sample image includes the target object and a second bounding box obtained by randomly moving the first bounding box; The bounding box in the first sample image is moved in parallel to obtain the third sample image, wherein the third sample image includes the target object and a third bounding box obtained by parallel moving the first bounding box, wherein the third bounding box is obtained by parallel moving the first bounding box along the X-axis, or the third bounding box is obtained by parallel moving the first bounding box along the Y-axis.
3. The method according to claim 2, characterized in that The parallel moving of the boundary box in the first sample image to obtain the third sample image includes: parallel moving the first bounding box in the first sample image according to a first parameter to obtain the third sample image, wherein the first parameter includes a first sub-parameter and a second sub-parameter, the first sub-parameter is used to indicate a ratio of the number of third sample images obtained by the parallel movement to the number of the first sample images in the first image set, and the second sub-parameter is used to indicate a ratio of a first distance to a width or a height of the first bounding box, the first distance being a distance from the first bounding box to the third bounding box; The step of training the initial first key point recognition model based on the first image set and the second image set to obtain the trained first key point recognition model includes: Pre-training the initial first key point recognition model based on the first image set, the second sample image and the third sample image to obtain a pre-trained first key point recognition model; parallel moving the bounding box in the first sample image according to a second parameter to obtain a fourth sample image, wherein the fourth sample image includes the target object and a fourth bounding box, the second parameter includes a third sub-parameter and a fourth sub-parameter, the third sub-parameter is used to indicate a ratio of the number of fourth sample images obtained by the parallel movement to the number of the first sample images in the first image set, the fourth sub-parameter is used to indicate a ratio of a second distance to a width or a height of the first bounding box, and the second distance refers to a distance from the first bounding box to the fourth bounding box; The pre-trained first key point recognition model is trained based on the first image set, the second sample image and the fourth sample image to obtain the trained first key point recognition model.
4. The method according to claim 3, characterized in that: The first sub-parameter is greater than the third sub-parameter, and the second sub-parameter is greater than the fourth sub-parameter.
5. The method according to claim 3 or 4, characterized in that: Pre-training the initial first key point recognition model based on the first image set, the second sample image, and the third sample image to obtain a pre-trained first key point recognition model includes: Removing an image area outside the third bounding box in the third sample image to obtain an updated third sample image; The initial first key point recognition model is pre-trained based on the first image set, the second sample image and the updated third sample image to obtain the pre-trained first key point recognition model.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Acquire a third image set, wherein the third image set includes a fifth sample image, and the fifth sample image is an image captured by a camera device; The first sample image is obtained based on a fifth sample image, depth information corresponding to the fifth sample image, and a target detection model, where the target detection model is used to mark a first bounding box of the target object.
7. The method according to claim 6, characterized in that Obtaining the first sample image based on the fifth sample image, the depth information corresponding to the fifth sample image, and the target detection model includes: Obtain a sixth sample image based on the fifth sample image and the object detection model, wherein the sixth sample image includes the first bounding box; The sixth sample image is three-dimensionally converted based on the depth information corresponding to the fifth sample image to obtain the first sample image.
8. The method according to claim 7, characterized in that The performing three-dimensional conversion on the sixth sample image based on the depth information corresponding to the fifth sample image to obtain the first sample image includes: The first formula and the depth information corresponding to the fifth sample image are used to perform three-dimensional coordinate transformation on the sixth sample image to obtain the first sample image, wherein the first formula includes: x = X / W; y=Y / H; z=(Z-Z0) / W; Among them, x is the coordinate of the pixel point on the x-axis in the first sample image, y is the coordinate of the pixel point on the y-axis in the first sample image, z is the coordinate of the pixel point on the z-axis in the first sample image, X is the coordinate of the pixel point corresponding to x in the sixth sample image, Y is the coordinate of the pixel point corresponding to y in the sixth sample image, Z is the coordinate of the pixel point corresponding to z in the sixth sample image, Z0 represents the coordinate of the target key point on the z-axis in the sixth sample image, W represents the width of the first bounding box, and H represents the height of the first bounding box.
9. The method according to any one of claims 1 to 8, characterized in that: The target object includes a hand in the image.
10. The method according to claim 9, characterized in that The first key point recognition model is used to recognize hand key points in an image.
11. The method according to any one of claims 1 to 10, characterized in that: After training the initial first key point recognition model based on the first image set and the second image set to obtain the trained first key point recognition model, the method further includes: Acquire a fourth image set, wherein the fourth image set includes a seventh sample image, and the seventh sample image is an image obtained by a preset gesture of the target object; Training an initial second key point recognition model based on the fourth image set to obtain a trained second key point recognition model, wherein the second key point recognition model is used to recognize key points corresponding to the preset gesture; The parameters in the trained first key point recognition model are corrected based on the parameters in the trained second key point recognition model to obtain an updated first key point recognition model.
12. The method according to claim 11, characterized in that The first key point recognition model includes a first network and a first regression layer, and the second key point recognition model includes the first network and a second regression layer, and the dimension of the first regression layer is higher than the dimension of the second regression layer.
13. An electronic device, characterized in that: The electronic device comprises means for performing the method according to any one of claims 1 to 12.
14. An electronic device, characterized in that: include: one or more processors; Memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and when the computer programs are executed by the one or more processors, the electronic device performs the method according to any one of claims 1 to 12.
15. A chip system, characterized in that: The chip system includes a processor for calling and running a computer program from a memory, so that an electronic device equipped with the chip system executes the method according to any one of claims 1 to 12.
16. A computer-readable storage medium comprising a computer program, characterized in that: When the computer program is executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Hand key point recognition model training method and device, recognition method and device
CN110163048A
Gesture recognition method based on Reders model
CN113743247A
Hand joint positioning method and device
CN114373191A
Key point detection method and key point detection model training method
CN115223192A
Gesture recognition method, wearable device and storage medium
CN115599211A