Hand detection model training method and device
By introducing timing information of multi-frame continuous sequence images into the hand detection model, feature merging and training is performed, the problem of low detection accuracy of deep learning models in continuous timing images is solved, and more accurate hand detection is achieved.
Patent Information
- Application Number
- CN202410064095.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
The hand detection model based on deep learning has labeling errors in the continuous timing screen, resulting in low detection accuracy and affecting the accuracy of hand detection.
By obtaining multi-frame hand images and their corresponding continuous sequence image sets, feature extraction and merging are performed, and the model is trained using the loss function of the hand detection box until the training end condition is met, and the object detection model is obtained.
The detection accuracy of the hand detection model is improved, allowing it to detect the hand position more accurately, and the stability and accuracy on the timing screen are improved.
Smart Images

Figure CN120340105A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more particularly, to a method and device for training a hand detection model. Background Art
[0002] Currently, gesture recognition plays a very important role in fields such as human-computer interaction and intelligent control. Gesture recognition usually includes a key step of hand detection. Generally, deep learning is a commonly used detection method for hand detection. However, based on the deep learning method, it is necessary to perform detection box annotation on the image of the object to be detected, and then use the annotated image for model training.
[0003] However, for continuous sequential images, there are certain errors in annotation, which will make the detection accuracy of the model not high and affect the accuracy of hand detection. Summary of the Invention
[0004] In order to improve the accuracy of the hand detection model and use the hand detection model to accurately detect the hand, an embodiment of the present application provides a method and device for training a hand detection model.
[0005] In a first aspect, an embodiment of the present application provides a method for training a hand detection model, including:
[0006] Obtain multiple frames of hand images and a first sequential image set respectively corresponding to each frame of hand image in the multiple frames of hand images, where the first sequential image set includes multiple consecutive sequential hand images before the hand image;
[0007] Input the hand image and the multiple sequential hand images in the first sequential image set into a hand detection model for feature extraction, and respectively obtain a first feature of the hand image and a second feature of the multiple sequential hand images;
[0008] Merge the first feature and the second feature to obtain a merged feature corresponding to the hand image, and output a hand detection box in the hand image by performing a convolution operation on the merged feature;
[0009] Train the hand detection model using the loss of the hand detection box until the training end condition is met to obtain a target detection model.
[0010] As an optional implementation manner of an embodiment of the present application, the inputting the hand image and the multiple sequential hand images in the first sequential image set into a hand detection model for feature extraction, and respectively obtaining a first feature of the hand image and a second feature of the multiple sequential hand images includes:
[0011] Input the hand image into the hand detection model, and obtain the first feature of the hand image through downsampling;
[0012] Input multiple frames of sequential hand images in the first temporal image set into the hand detection model, and obtain the second feature of the multiple frames of sequential hand images through downsampling.
[0013] As an optional implementation manner of an embodiment of the present application, the inputting the hand image and multiple frames of sequential hand images in the first temporal image set into the hand detection model for feature extraction, and respectively obtaining the first feature of the hand image and the second feature of the multiple frames of sequential hand images includes:
[0014] Determine the first position of the hand detection box in the hand image, crop the hand image according to the first position, and input the cropped hand image into the hand detection model for feature extraction to obtain the first feature of the hand image;
[0015] For the multiple frames of sequential hand images in the first temporal image set, determine the second position of the hand detection box in each frame of sequential hand image, crop each frame of sequential hand image according to the second position, and input the cropped each frame of sequential hand image into the hand detection model for feature extraction to obtain the second feature of the multiple frames of sequential hand images.
[0016] As an optional implementation manner of an embodiment of the present application, the determining the first position of the hand detection box in the hand image includes:
[0017] Obtain the second temporal image set corresponding to the hand image, where the number of frames of consecutive sequential hand images included in the second temporal image set is greater than the number of frames of consecutive sequential hand images included in the first temporal image set;
[0018] Obtain the coordinates of the hand detection box in each frame of sequential hand image in the second temporal image set;
[0019] By fitting the coordinates of the hand detection box in each frame of sequential hand image in the second temporal image set, obtain the moving speed and acceleration of the hand in the hand image;
[0020] Calculate the first position of the hand detection box in the hand image according to the moving speed and acceleration of the hand.
[0021] As an optional implementation manner of an embodiment of the present application, for the multiple frames of sequential hand images in the first temporal image set, determining the second position of the hand detection box in each frame of sequential hand image, cropping each frame of sequential hand image according to the second position, and inputting the cropped each frame of sequential hand image into the hand detection model includes:
[0022] For each frame sequence of hand images in the first time-series image set, determine the second position of the hand detection box in each frame sequence of hand images, and crop each frame sequence of hand images according to the second position;
[0023] Perform masking processing on the hand regions in each cropped frame sequence of hand images to obtain a hand mask, and input the hand mask into a hand detection model.
[0024] As an optional implementation manner of an embodiment of the present application, training the hand detection model using the loss of the hand detection box until the training end condition is satisfied to obtain a target detection model includes:
[0025] Determine the regression loss and classification loss of the hand detection box;
[0026] Perform weighted summation on the regression loss and the classification loss to obtain a target loss function;
[0027] Train the hand detection model using the target loss function until the training end condition is satisfied to obtain a target detection model.
[0028] As an optional implementation manner of an embodiment of the present application, after obtaining the target detection model, the method further includes:
[0029] Obtain a hand image stream;
[0030] Input the hand image stream into the target detection model, and use the target detection model to extract the first hand feature of the current frame in the hand image stream and the second hand features of a preset number of hand image frames before the current frame in the hand image stream;
[0031] Use the target detection model to merge the first hand feature and the second hand features, and perform a convolution operation on the merged hand features to output the hand detection box of the current frame in the hand image stream.
[0032] In a second aspect, an embodiment of the present application provides a training device for a hand detection model, including
[0033] An acquisition module, configured to acquire a training data set, where the training data set includes multiple frames of hand images and a first time-series image set corresponding to each frame of hand image in the multiple frames of hand images, and the first time-series image set includes multiple consecutive sequence hand images before the hand image;
[0034] An extraction module, configured to input the hand images and the multiple frame sequence hand images in the first time-series image set into a hand detection model for feature extraction, and respectively obtain the first feature of the hand images and the second features of the multiple frame sequence hand images;
[0035] A processing module, configured to merge the first feature and the second feature to obtain a merged feature corresponding to the hand image, and output a hand detection frame in the hand image by performing a convolution operation on the merged feature;
[0036] A training module, configured to train the hand detection model by using the loss of the hand detection frame until a training end condition is met, to obtain a target detection model.
[0037] As an optional implementation manner of an embodiment of the present application, the extraction module is specifically configured to input the hand image into the hand detection model, and obtain a first feature of the hand image through downsampling;
[0038] Input multiple frames of sequential hand images in the first sequential image set into the hand detection model, and obtain a second feature of the multiple frames of sequential hand images through downsampling.
[0039] As an optional implementation manner of an embodiment of the present application, the apparatus further includes: a cropping module; configured to determine a first position of a hand detection frame in the hand image, and crop the hand image according to the first position;
[0040] The extraction module is specifically configured to input the cropped hand image into the hand detection model for feature extraction, to obtain a first feature of the hand image;
[0041] The cropping module is further configured to, for multiple frames of sequential hand images in the first sequential image set, determine a second position of a hand detection frame in each frame of sequential hand image, and crop each frame of sequential hand image according to the second position;
[0042] The extraction module is specifically configured to input the cropped frames of sequential hand images into the hand detection model for feature extraction, to obtain a second feature of the multiple frames of sequential hand images.
[0043] As an optional implementation manner of an embodiment of the present application, the cropping module is specifically configured to obtain a second sequential image set corresponding to the hand image, where the number of frames of consecutive sequential hand images included in the second sequential image set is greater than the number of frames of consecutive sequential hand images included in the first sequential image set;
[0044] Obtain coordinates of a hand detection frame in each frame of sequential hand image in the second sequential image set;
[0045] By fitting the coordinates of the hand detection frame in each frame of sequential hand image in the second sequential image set, obtain the moving speed and acceleration of the hand in the hand image;
[0046] Calculate the first position of the hand detection box in the hand image according to the moving speed and acceleration of the hand.
[0047] As an optional implementation manner of the embodiment of the present application, the cropping module is specifically configured to determine the second position of the hand detection box in each frame sequence hand image in the first time sequence image set, and crop each frame sequence hand image according to the second position;
[0048] Perform a masking process on the hand region in each cropped frame sequence hand image to obtain a hand mask, and input the hand mask into the hand detection model.
[0049] As an optional implementation manner of the embodiment of the present application, the training module is specifically configured to determine the regression loss and classification loss of the hand detection box;
[0050] Perform a weighted sum on the regression loss and the classification loss to obtain an objective loss function;
[0051] Use the objective loss function to train the hand detection model until the training end condition is met to obtain a target detection model.
[0052] As an optional implementation manner of the embodiment of the present application, the device further includes:
[0053] A detection module, configured to obtain a hand image stream after obtaining the target detection model;
[0054] Input the hand image stream into the target detection model, and use the target detection model to extract the first hand feature of the current frame in the hand image stream and the second hand features of a preset number of hand image frames before the current frame in the hand image stream;
[0055] Use the target detection model to merge the first hand feature and the second hand features, and perform a convolution operation on the merged hand features to output the hand detection box of the current frame in the hand image stream.
[0056] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory is used to store a computer program, and the processor is used to execute the training method of the hand detection model according to the first aspect or any optional implementation manner of the first aspect when calling the computer program.
[0057] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the training method of the hand detection model according to the first aspect or any optional implementation manner of the first aspect.
[0058] The technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:
[0059] The embodiment of the present application provides a training method, device, electronic device and computer-readable storage medium for a hand detection model. The method includes: obtaining multiple frames of hand images and a first time-series image set corresponding to each frame of hand image in the multiple frames of hand images, where the first time-series image set includes multiple consecutive sequence hand images before the hand image; inputting the hand image and the multiple frames of sequence hand images in the first time-series image set into the hand detection model for feature extraction, respectively obtaining a first feature of the hand image and a second feature of the multiple frames of sequence hand images; merging the first feature and the second feature to obtain a merged feature corresponding to the hand image, and outputting a detection result of the hand image by performing a convolution operation on the merged feature; training the hand detection model using the detection result until a training end condition is satisfied to obtain a target detection model. By introducing multiple consecutive sequence hand images before the hand image to train the model, the time-series information of the hand image is introduced during the model training process, the possible positions of the hand detection frame in the hand image can be estimated, and the model structure can be made more focused on detecting in the area where the hand detection frame may appear, improving the detection accuracy of the target detection model and making the detection result of the target detection model for the hand more accurate. Description of the Drawings
[0060] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0061] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0062] Figure 1 It is a flowchart of a training method for a hand detection model provided according to one or more embodiments of the present application;
[0063] Figure 2 It is a flowchart of a training method for a hand detection model provided according to one or more embodiments of the present application;
[0064] Figure 3 It is a structural block diagram of a training device for a hand detection model provided according to one or more embodiments of the present application;
[0065] Figure 4The structural block diagram of the training device for the hand detection model provided by one or more embodiments of the present application;
[0066] Figure 5 The internal structure diagram of the electronic device provided by one or more embodiments of the present application. Detailed implementation manners
[0067] To make the objectives, implementation manners, and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part rather than all of the embodiments of the present application.
[0068] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the appended claims of the present application. In addition, although the disclosed content in the present application is introduced according to exemplary one or several examples, it should be understood that each aspect of these disclosed contents can also constitute a complete implementation manner alone. It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.
[0069] The embodiments of the present application provide a method and device for training a hand detection model. In the process of model training, in the training dataset, by introducing multiple consecutive sequence hand images before the hand image to obtain temporal information, during the process of training the model based on the training dataset, the possible positions of the hand detection boxes in the hand image can be estimated, enabling the model structure to focus more on the areas where the hand detection boxes may appear for detection, and making the detection results of the hand detection model for the hand more accurate.
[0070] The method for training the hand detection model provided by the embodiments of the present application can be executed by the electronic device provided by the embodiments of the present application, can also be implemented by the training device for the hand detection model provided by the embodiments of the present application, or can also be implemented by one or more functional entities on a vehicle. The embodiments of the present application do not make specific limitations.
[0071] The following elaborates in detail on the method for training the hand detection model provided by the embodiments of the present application through several specific embodiments.
[0072] Figure 1 The flowchart of the method for training the hand detection model provided by the embodiments of the present application. Refer to Figure 1 As shown, the method for training the hand detection model provided in this embodiment includes the following steps:
[0073] S11. Obtain multiple frames of hand images and a first sequence of images corresponding to each frame of the multiple frames of hand images respectively.
[0074] Wherein, the first sequence of images includes multiple consecutive sequence hand images before the hand image.
[0075] That is, the data in the training dataset can be divided into two parts. One part is used as labeled data, and the other part is used as the supporting data for the labeled data. The supporting data includes more image frames than the labeled data. In the embodiments of the present application, the above multiple frames of hand images are the labeled data, and the first sequence of images corresponding to each frame of hand images is the supporting data. The hand images in the first sequence of images and the hand image are consecutive sequence images. Exemplarily, an image stream sequentially includes image frames A, B, C... to K in sequence. That is, image frame B is the image frame after image frame A, image frame C is the image frame after image frame B, and image frame K is the image frame after image frame J. The number of hand images in the first sequence of images can be set according to requirements. For example, it can be 10 frames, 20 frames, etc. If the hand image is image frame K, the images in the first sequence of images corresponding to image frame K can include image frames A to J, or can include image frames B to J, or include image frames C to J, etc. The last frame of hand image in the first sequence of images is the image that is consecutive with the hand image and is the previous frame of the hand image.
[0076] S12. Input the hand image and the multiple frames of sequence hand images in the first sequence of images into a hand detection model for feature extraction, and respectively obtain a first feature of the hand image and a second feature of the multiple frames of sequence hand images.
[0077] Exemplarily, input the hand image into the hand detection model, and obtain the first feature of the hand image through downsampling; input the multiple frames of sequence hand images in the first sequence of images into the hand detection model, and obtain the second feature of the multiple frames of sequence hand images through downsampling.
[0078] Wherein, the hand detection model may include multiple convolutional layers, multiple pooling layers, and multiple fully connected layers, etc.
[0079] The input hand image and the multiple frames of sequence hand images are images of the same size. Exemplarily, input the hand image and the multiple frames of sequence hand images corresponding to the hand image into the hand detection model, and perform feature extraction through convolutional layers and pooling layers to obtain the first feature of the hand image and the second feature of each frame of sequence hand images in the multiple frames of sequence hand images.
[0080] S13. Merge the first feature and the second feature to obtain the merged feature corresponding to the hand image, and output the hand detection box in the hand image by performing a convolution operation on the merged feature.
[0081] The merging (concat) includes addition, subtraction, multiplication, and division. In this embodiment, merging the first feature and the second feature means adding the first feature and the second feature on the same channel to obtain the merged feature corresponding to the hand image, and outputting the hand detection box in the hand image by performing a convolution operation on the merged feature.
[0082] Merging the first feature and the second feature is equivalent to endowing the first feature with temporal information, and adding the position information where the hand detection box may appear in the image in the Backbone part of the network to improve the accuracy of the detected hand image.
[0083] Among them, the detection result includes a regression result and a classification result. The regression result includes the position information of the hand detection box in the hand image, including X, Y, W, and H, that is, the X coordinate, Y coordinate, length, and width of the hand detection box, and the classification result includes the target object (hand) and the image background.
[0084] S14. Train the hand detection model using the loss of the hand detection box until the training end condition is met to obtain the target detection model.
[0085] Training the hand detection model may include: determining the regression loss and classification loss of the hand detection box, performing weighted summation on the regression loss and the classification loss to obtain the target loss function; training the hand detection model using the target loss function until the training end condition is met to obtain the target detection model. Among them, the regression loss can be the IOU (Intersection-over-Union) loss, and the IOU is the ratio of the intersection area between the predicted box and the true box to the union area, and the classification loss can be the cross-entropy loss.
[0086] To avoid overfitting of the hand detection model, in the embodiments of the present application, multiple frames of sequential hand images in the first temporal image set are randomly removed and reversed, so that the trained target detection model can still run independently without temporal information.
[0087] In the embodiments of the present application, the detection of the most likely position is calculated by using the regression loss to follow up the position of the temporal information feature, so as to ensure the recall of the effective hand. By means of the classification loss, the hand detection frames closer to the expected position are encouraged, and the hand detection frames farther from the expected position are punished, so as to improve the sensitivity of the hand detection model to the temporal information. For consecutive frames, the hand detection model can be made to more attentively detect the position where the hand appears in the current frame of hand image according to the position information of the hand appearing in the temporal process, thus improving the accuracy and recall rate of hand detection.
[0088] The embodiments of the present application provide a training method for a hand detection model. Multiple frames of hand images and a first temporal image set respectively corresponding to each frame of hand image in the multiple frames of hand images are obtained. The first temporal image set includes multiple consecutive sequence hand images before the hand image. The hand image and the multiple sequence hand images in the first temporal image set are input into the hand detection model for feature extraction, and a first feature of the hand image and a second feature of the multiple sequence hand images are respectively obtained. The first feature and the second feature are combined to obtain a combined feature corresponding to the hand image. By performing a convolution operation on the combined feature, a detection result of the hand image is output. The hand detection model is trained by using the detection result until the training end condition is satisfied, and a target detection model is obtained. In the embodiments of the present application, by introducing multiple consecutive sequence hand images before the hand image to train the model, the temporal information of the hand image is introduced in the model training process, the possible position of the hand detection frame in the hand image can be estimated, the model structure can be made to more attentively detect in the area where the hand detection frame may appear, the detection accuracy of the target detection model is improved, and the detection result of the target detection model for the hand is more accurate.
[0089] Figure 2 For the flowchart of the training method for the hand detection model provided by another embodiment of the present application, in Figure 1 the embodiment shown, step S12 can be implemented through the following steps S21 to S21. Referring to Figure 2 shown, Figure 2 the same or similar steps as those in the embodiment shown in Figure 1 will not be described in detail again.
[0090] S21. Determine a first position of the hand detection frame in the hand image, crop the hand image according to the first position, and input the cropped hand image into the hand detection model for feature extraction to obtain the first feature of the hand image.
[0091] Among them, the first position of the hand detection box in the hand image can be obtained in the following manner: obtaining a second time-series image set corresponding to the hand image, where the number of frames of consecutive sequence hand images included in the second time-series image set is greater than the number of frames of consecutive sequence hand images included in the first time-series image set; obtaining the coordinates of the hand detection box in each frame of the consecutive sequence hand images in the second time-series image set; fitting the coordinates of the hand detection box in each frame of the consecutive sequence hand images to obtain the moving speed and acceleration of the hand in the hand image; calculating the first position of the hand detection box in the hand image according to the moving speed and acceleration of the hand.
[0092] Exemplarily, the first time-series image set includes ten frames of images before the hand image, and the second time-series image set may include thirty to sixty frames of images before the hand image, and the number of frames of consecutive sequence hand images included in the second time-series image set is greater than the number of frames of consecutive sequence hand images included in the first time-series image set. The hand positions in the consecutive sequence hand images of the second time-series image set represent the motion information of the hand in consecutive frames.
[0093] S22. For multiple frames of consecutive sequence hand images in the first time-series image set, determine the second position of the hand detection box in each frame of the consecutive sequence hand images, crop each frame of the consecutive sequence hand images according to the second position, and input the cropped consecutive sequence hand images into a hand detection model for feature extraction to obtain the second features of the multiple frames of consecutive sequence hand images.
[0094] Among them, the multiple frames of consecutive sequence hand images input into the hand detection model can be the hand masks corresponding to each frame of the consecutive sequence hand images. Exemplarily, for each frame of the consecutive sequence hand images in the first time-series image set, determine the second position of the hand detection box in each frame of the consecutive sequence hand images, crop each frame of the consecutive sequence hand images according to the second position; perform a masking process on the hand regions in the cropped consecutive sequence hand images to obtain hand masks, and input the hand masks into the hand detection model.
[0095] In this embodiment, the method for obtaining the second position of the hand detection box in the multiple frames of consecutive sequence hand images can refer to the method for obtaining the first position of the hand detection box in the hand image described above, and will not be elaborated here.
[0096] Through the training method of the hand detection model provided by the embodiments of the present application, for long-distance scenarios, the range of the hand in the hand image is relatively small. In order to maintain stable overhead and real-time performance, the input of the network model generally remains at a size of 320x320. Currently, the size of a single image data obtained by a TOF camera is 640x480, and the data size of a single IR image is 1600x1300. Therefore, when ensuring that the image input to the model is 320x320 in size, it is equivalent to performing downsampling, which is likely to cause missed detection of the small-sized hand information at a long distance. In this embodiment, by inputting the cropped hand image into the hand detection model, it can be ensured that the information of the hand position will not be lost due to downsampling.
[0097] Further, after obtaining the target detection model, the continuous frames can be detected based on the target detection model. Exemplarily, a hand image stream is obtained; the hand image stream is input into the target detection model, and the target detection model is used to extract the first hand feature of the current frame in the hand image stream and the second hand features of a preset number of hand image frames before the current frame in the hand image stream; the target detection model is used to merge the first hand feature and the second hand features, and perform a convolution operation on the merged hand features, and output the detection result of the current frame in the hand image stream.
[0098] The target detection model obtained based on the training method of the hand detection model provided by the embodiments of the present application can accurately detect the hand in the continuous frames, solve the problem of obvious user perception of unsmooth or inaccurate operations caused by missed detection or misdetection of a certain frame in the sequential frames, and improve the stability and accuracy in the sequential frames.
[0099] Based on the same inventive concept, as an implementation of the above method, the embodiments of the present application also provide a training device for the hand detection model provided in the above embodiments. The device embodiments correspond to the foregoing method embodiments. For the convenience of reading, the details in the foregoing method embodiments will not be described one by one in the device embodiments of the present application. However, it should be clear that the training device for the hand detection model in this embodiment can correspondingly implement all the contents in the foregoing method embodiments.
[0100] Figure 3 The structural schematic diagram of the training device for the hand detection model provided by an embodiment of the present application is as Figure 3 shown. The training device 300 for the hand detection model provided in this embodiment includes:
[0101] An acquisition module 310, configured to acquire a training data set, where the training data set includes multiple frames of hand images and a first time-series image set corresponding to each frame of hand image in the multiple frames of hand images, and the first time-series image set includes multiple consecutive sequence hand images before the hand image.
[0102] An extraction module 320 is configured to input the hand image and multiple frames of sequential hand images in the first timing image set into a hand detection model for feature extraction, and respectively obtain a first feature of the hand image and a second feature of the multiple frames of sequential hand images.
[0103] A processing module 330 is configured to merge the first feature and the second feature to obtain a merged feature corresponding to the hand image, and output a hand detection box in the hand image by performing a convolution operation on the merged feature.
[0104] A training module 340 is configured to train the hand detection model by using the loss of the hand detection box until a training end condition is satisfied, and obtain a target detection model.
[0105] As an optional implementation manner of an embodiment of the present application, the extraction module 320 is specifically configured to input the hand image into the hand detection model, and obtain the first feature of the hand image through downsampling; input multiple frames of sequential hand images in the first timing image set into the hand detection model, and obtain the second feature of the multiple frames of sequential hand images through downsampling.
[0106] Figure 4 The structural schematic diagram of a training device for a hand detection model provided by an embodiment of the present application is as Figure 4 shown. The training device for the hand detection model provided in this embodiment further includes, on the basis of the structure shown in Figure 3 the structure shown:
[0107] A cropping module 410 is configured to determine a first position of a hand detection box in the hand image, and crop the hand image according to the first position;
[0108] The extraction module 320 is specifically configured to input the cropped hand image into the hand detection model for feature extraction, and obtain the first feature of the hand image.
[0109] The cropping module 410 is further configured to, for multiple frames of sequential hand images in the first timing image set, determine a second position of a hand detection box in each frame of sequential hand image, and crop each frame of sequential hand image according to the second position.
[0110] The extraction module 320 is specifically configured to input the cropped frames of sequential hand images into the hand detection model for feature extraction, and obtain the second feature of the multiple frames of sequential hand images.
[0111] As an optional implementation manner of an embodiment of the present application, the cutting module 410 is specifically configured to obtain a second sequence of time-series images corresponding to the hand image, where the number of frames of consecutive sequence hand images included in the second sequence of time-series images is greater than the number of frames of consecutive sequence hand images included in the first sequence of time-series images; obtain the coordinates of the hand detection frames in each frame of the sequence hand images in the second sequence of time-series images; obtain the moving speed and acceleration of the hand in the hand image by fitting the coordinates of the hand detection frames in each frame of the sequence hand images in the second sequence of time-series images; calculate a first position of the hand detection frame in the hand image according to the moving speed and acceleration of the hand.
[0112] As an optional implementation manner of an embodiment of the present application, the cutting module 410 is specifically configured to determine a second position of the hand detection frame in each frame of the sequence hand images in the first sequence of time-series images, and cut each frame of the sequence hand images according to the second position; perform a masking process on the hand regions in the cut frames of the sequence hand images to obtain a hand mask, and input the hand mask into the hand detection model.
[0113] As an optional implementation manner of an embodiment of the present application, the training module 340 is specifically configured to determine a regression loss and a classification loss of the hand detection frame; perform a weighted sum on the regression loss and the classification loss to obtain a target loss function; use the target loss function to train the hand detection model until a training end condition is satisfied, to obtain a target detection model.
[0114] As an optional implementation manner of an embodiment of the present application, the apparatus further includes: a detection module 420, configured to obtain a hand image stream after obtaining the target detection model; input the hand image stream into the target detection model, and use the target detection model to extract a first hand feature of the current frame in the hand image stream and second hand features of a preset number of hand image frames before the current frame in the hand image stream; use the target detection model to merge the first hand feature and the second hand features, and perform a convolution operation on the merged hand features, and output a hand detection frame of the current frame in the hand image stream.
[0115] In one embodiment, an electronic device is provided, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the hand detection model training methods described in the above method embodiments are implemented.
[0116] Exemplarily, Figure 5 is a schematic structural diagram of the electronic device provided in the embodiment of the present application. As Figure 5As shown in the figure, the electronic device provided in this embodiment includes: a memory 51 and a processor 52. The memory 51 is used to store a computer program. The processor 52 is used to execute the steps in the training method of the hand detection model provided in the above method embodiment when calling the computer program. The implementation principle and technical effects are similar and will not be elaborated here. Those skilled in the art can understand that Figure 5 The structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0117] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of any one of the training methods of the hand detection model described in the above method embodiment.
[0118] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.
[0119] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0120] For the sake of convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for better explaining the principles and actual applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A training method for a hand detection model, characterized in that, Including: Obtaining multiple frames of hand images and a first time-series image set respectively corresponding to each frame of hand image in the multiple frames of hand images, where the first time-series image set includes multiple consecutive sequence hand images before the hand image; Inputting the hand image and the multiple frames of sequence hand images in the first time-series image set into a hand detection model for feature extraction to respectively obtain a first feature of the hand image and a second feature of the multiple frames of sequence hand images; Combining the first feature and the second feature to obtain a combined feature corresponding to the hand image, and outputting a hand detection box in the hand image by performing a convolution operation on the combined feature; Training the hand detection model using the loss of the hand detection box until a training end condition is satisfied to obtain a target detection model.
2. The method according to claim 1, wherein The step of inputting the hand image and the multiple frames of sequence hand images in the first time-series image set into a hand detection model for feature extraction to respectively obtain a first feature of the hand image and a second feature of the multiple frames of sequence hand images includes: Inputting the hand image into the hand detection model and obtaining a first feature of the hand image through downsampling; Inputting the multiple frames of sequence hand images in the first time-series image set into the hand detection model and obtaining a second feature of the multiple frames of sequence hand images through downsampling.
3. The method according to claim 1, characterized in that, The step of inputting the hand image and the multiple frames of sequence hand images in the first time-series image set into a hand detection model for feature extraction to respectively obtain a first feature of the hand image and a second feature of the multiple frames of sequence hand images includes: Determining a first position of a hand detection box in the hand image, cropping the hand image according to the first position, and inputting the cropped hand image into the hand detection model for feature extraction to obtain a first feature of the hand image; For the multiple frames of sequence hand images in the first time-series image set, determining a second position of a hand detection box in each frame of sequence hand image, cropping each frame of sequence hand image according to the second position, and inputting the cropped each frame of sequence hand image into the hand detection model for feature extraction to obtain a second feature of the multiple frames of sequence hand images.
4. The method according to claim 3, characterized in that The step of determining a first position of a hand detection box in the hand image includes: Obtaining a second time-series image set corresponding to the hand image, where the number of frames of consecutive sequence hand images included in the second time-series image set is greater than the number of frames of consecutive sequence hand images included in the first time-series image set; Obtaining the coordinates of the hand detection box in each frame of sequence hand image in the second time-series image set; Obtaining the moving speed and acceleration of the hand in the hand image by fitting the coordinates of the hand detection box in each frame of sequence hand image in the second time-series image set; Calculating a first position of the hand detection box in the hand image according to the moving speed and acceleration of the hand.
5. The method according to claim 3, characterized in that, For the multi-frame sequential hand images in the first temporal image set, determining the second positions of the hand detection frames in each frame of the sequential hand images, and cropping each frame of the sequential hand images according to the second positions, and inputting the cropped frames of the sequential hand images into a hand detection model, includes: For each frame of the sequential hand images in the first temporal image set, determining the second positions of the hand detection frames in each frame of the sequential hand images, and cropping each frame of the sequential hand images according to the second positions; Performing a masking process on the hand regions in the cropped frames of the sequential hand images to obtain a hand mask, and inputting the hand mask into the hand detection model.
6. The method according to claim 1, wherein Training the hand detection model using the loss of the hand detection frame until a training end condition is satisfied to obtain a target detection model, includes: Determining the regression loss and classification loss of the hand detection frame; Performing a weighted sum on the regression loss and the classification loss to obtain a target loss function; Training the hand detection model using the target loss function until a training end condition is satisfied to obtain a target detection model.
7. The method according to any one of claims 1-6, characterized in that, After obtaining the target detection model, the method further includes: Obtaining a hand image stream; Inputting the hand image stream into the target detection model, and using the target detection model to extract the first hand feature of the current frame in the hand image stream and the second hand features of a preset number of hand image frames before the current frame in the hand image stream; Using the target detection model to merge the first hand feature and the second hand features, and performing a convolution operation on the merged hand features to output the hand detection frame of the current frame in the hand image stream.
8. A training device for a hand detection model, characterized in that, Includes: An acquisition module, configured to acquire a training data set, where the training data set includes multiple hand images and a first temporal image set corresponding to each hand image in the multiple hand images, and the first temporal image set includes multiple consecutive sequential hand images before the hand image; An extraction module, configured to input the hand image and the multiple frame sequential hand images in the first temporal image set into a hand detection model for feature extraction, and respectively obtain a first feature of the hand image and a second feature of the multiple frame sequential hand images; A processing module, configured to merge the first feature and the second feature to obtain a merged feature corresponding to the hand image, and output the hand detection frame in the hand image by performing a convolution operation on the merged feature; A training module, configured to train the hand detection model using the loss of the hand detection frame until a training end condition is satisfied to obtain a target detection model.
9. An electronic device, comprising: A memory and a processor, where the memory stores a computer program, and is characterized in that when the processor executes the computer program, it implements the training method of the hand detection model according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the training method of the hand detection model according to any one of claims 1 to 6.
11. A vehicle, characterized in that, The vehicle is configured with the training device of the hand detection model described in claim 8, or the electronic device described in claim 9, or the storage medium described in claim 10.