Gesture recognition method and apparatus
Patent Information
- Application Number
- CN202310345218.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-03-31
AI Technical Summary
[0004]本发明提供一种姿态识别方法及装置,用以解决现有技术针对人体姿态识别的过程中,由于图像中的干扰因素,使得姿态识别的准确率较低的技术问题
[0033]本发明提供的姿态识别方法及装置,通过对目标图像进行姿态识别之前,基于高斯函数,获取目标图像的高斯分布热力图,并基于获取的高斯分布热力图,实现目标图像中的关键点数据的确定,减少了目标图像中的干扰数据,提升了对目标图像中目标对象特征的提取。基于标准化函数对关键点数据进行标准化处理,可以实现对关键点数据进行去量纲的目的,基于标准化处理后的关键点数据进行姿态识别,提升了姿态识别的准确率。
Smart Images

Figure CN116503899B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a pose recognition method and apparatus. Background Technology
[0002] Most existing methods for recognizing human postures involve directly inputting images containing human postures into neural network models to achieve posture recognition.
[0003] Existing pose recognition methods that directly identify images containing human poses using neural network models face challenges due to interference from real-life scenarios, such as different scenes, targets within those scenes, highly free human poses, and densely populated areas. These factors increase the difficulty of human pose recognition and result in lower accuracy. Summary of the Invention
[0004] This invention provides a posture recognition method and apparatus to solve the technical problem that the accuracy of posture recognition is low in the process of human posture recognition due to interference factors in the image.
[0005] This invention provides a posture recognition method, comprising:
[0006] Based on the Gaussian function, a Gaussian distribution heatmap of the target image is determined, and based on the Gaussian distribution heatmap, key point data in the target image are determined, wherein the target image contains a target object;
[0007] Based on the standardization function, the key point data is standardized to obtain standardized key point data;
[0008] The pose of the target object is determined by performing pose recognition based on the standardized key point data.
[0009] According to a pose recognition method provided by the present invention, the key point data is standardized based on a standardization function, including:
[0010] Based on the Z-score function, the keypoint data is globally standardized to obtain globally standardized keypoint data.
[0011] Based on the Gaussian distribution heatmap, local part data of each target are determined from the key point data after global standardization, and local standardization is performed on the local part data of each target based on the Z-score function to obtain the standardized key point data.
[0012] According to a posture recognition method provided by the present invention, posture recognition is performed based on the standardized key point data to determine the posture of the target object, including:
[0013] The standardized key point data is input into the pose recognition model to obtain the pose of the target object output by the pose recognition model. The pose recognition model is trained based on image samples and their corresponding human pose labels.
[0014] According to a pose recognition method provided by the present invention, a method for training a pose recognition model includes:
[0015] Construct a set of human pose images;
[0016] Based on the Gaussian function, the Gaussian distribution heatmap of each image in the human pose image set is determined, and based on the Gaussian distribution heatmap of each image, the key point data corresponding to each image is determined;
[0017] Based on the standardization function, the key point data of each image are standardized to obtain a set of standardized key point data.
[0018] The image samples and the test set of the pose recognition model are determined from the key point data set;
[0019] The initial pose recognition model is trained based on the image samples and their corresponding human pose labels to obtain the pose recognition model.
[0020] According to a pose recognition method provided by the present invention, after determining the image samples and the test set of the pose recognition model from the key point data set, the method further includes:
[0021] Based on the test set, the pose recognition model is tested to obtain the test results corresponding to the test set;
[0022] The precision of the test results is determined to be within a preset precision range, which is based on the recall, precision, and accuracy of the posture recognition model training.
[0023] According to the pose recognition method provided by the present invention, after constructing a set of human pose images, the method further includes:
[0024] The position of the human body trunk in each image in the human posture image set is determined, and the images in each image whose human body trunk position is not upright are to be adjusted.
[0025] Based on a preset rotation angle, the image to be adjusted is rotated so that the human target in each image of the human posture image set is in an upright posture.
[0026] The present invention also provides a posture recognition device, comprising:
[0027] The key point data determination module is used to determine the Gaussian distribution heatmap of the target image based on the Gaussian function, and to determine the key point data in the target image based on the Gaussian distribution heatmap, wherein the target image contains a target object;
[0028] The standardization processing module is used to standardize the key point data based on the standardization function to obtain standardized key point data.
[0029] The recognition module is used to perform posture recognition based on the standardized key point data to determine the posture of the target object.
[0030] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the above-described posture recognition methods.
[0031] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the posture recognition method as described above.
[0032] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the posture recognition method as described above.
[0033] The pose recognition method and apparatus provided by this invention obtain a Gaussian distribution heatmap of the target image based on a Gaussian function before pose recognition. Based on this heatmap, key point data in the target image are determined, reducing interference data and improving the extraction of target object features. Standardizing the key point data using a standardization function removes dimensions, and performing pose recognition based on the standardized key point data improves the accuracy of pose recognition. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly described below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a flowchart illustrating the posture recognition method provided by the present invention;
[0036] Figure 2 This is a schematic diagram of the training system module structure provided by the present invention;
[0037] Figure 3 This is a schematic diagram of the model training process provided by the present invention;
[0038] Figure 4 This is a flowchart of the deconvolution process provided by the present invention;
[0039] Figure 5 This is a flowchart of the upsampling summation provided by the present invention;
[0040] Figure 6 This is a schematic diagram of the posture recognition device provided by the present invention;
[0041] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0043] Figure 1 This is a schematic flowchart of the pose recognition method provided by the present invention. (Refer to...) Figure 1 The pose recognition method provided by this invention may include:
[0044] Step 110: Based on the Gaussian function, determine the Gaussian distribution heatmap of the target image, and based on the Gaussian distribution heatmap, determine the key point data in the target image, wherein the target image contains the target object;
[0045] Step 120: Based on the standardization function, the key point data is standardized to obtain standardized key point data;
[0046] Step 130: Perform pose recognition based on the standardized key point data to determine the pose of the target object.
[0047] The subject executing the posture recognition method provided by this invention can be an electronic device, a component within an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. For example, a mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., while a non-mobile electronic device can be a server, network attached storage (NAS), or personal computer (PC), etc. This invention does not impose specific limitations.
[0048] The technical solution of this invention will be described in detail below using the example of a computer executing the posture recognition method provided by this invention.
[0049] In step 110, a target image is acquired, containing a target object with the pose to be identified. A two-dimensional Gaussian heatmap of the target image is determined based on a Gaussian function. Key point data in the target image are then determined based on the Gaussian heatmap.
[0050] A target object can contain multiple targets, and multiple targets can contain the target object. For example, in an outdoor environment, there may be targets such as pedestrians, vehicles, and buildings. If the target object is determined to be a pedestrian, the pose of the target object in the target image can be determined based on subsequent analysis of the target image.
[0051] After acquiring the target image, a Gaussian function is used to obtain the Gaussian distribution heatmap of the target image. The Gaussian function used to obtain the Gaussian distribution heatmap is:
[0052]
[0053] Where A is the peak value of the Gaussian curve; x0 and y0 are the image parameters of the target image; and x and y are the image information of the target image.
[0054] After obtaining the Gaussian distribution heatmap of the target image, the key point data in the target image can be determined based on the image display distribution in the Gaussian distribution heatmap.
[0055] It's understandable that the target image contains multiple targets. Besides the targets, the image also contains a lot of noise or interference from the environment. By obtaining a Gaussian distribution heatmap of the target image based on a Gaussian function, the key point data of each target in the image can be determined.
[0056] Optionally, before generating the Gaussian distribution heatmap, key point information P of the human body parts in the target image can be obtained. i for:
[0057]
[0058] Among them, P i P_i represents the edge probability of key point information of human body parts in the target image; N is the normalization coefficient; i and j in P_(i|j) represent two key points respectively; P_(i|j) is the conditional probability from key point to key point, which is also the relationship model between points.
[0059] Based on formula (2), the key point information of human body parts in the target image is obtained, and the obtained key points are used as the center of the circle;
[0060] Set the Gaussian radius. The calculation method is to calculate the proportion of the normalized distance between the detected key point and its corresponding ground truth that is less than the set threshold. The Gaussian radius that can be selected is 0.1 times the industry standard size. After selecting the Gaussian radius, the Gaussian function is used to obtain the Gaussian distribution heatmap of the target image.
[0061] In step 120, the key point data determined in step 110 is standardized based on the standardization function to obtain standardized key point data.
[0062] Standardization functions are used to standardize key data points, ensuring that the processed data falls within a specific range. Standardization functions can include the Z-score, normalization functions, etc. The Z-score function is a process of dividing the difference between a number and the mean by the standard deviation. In statistics, a standard score is the sign of the standard deviation of an observation or data point's value from the mean of the observed or measured values.
[0063] When the standardization function is Z-score, the Z-score processing formula for standardizing keypoint data based on Z-score is:
[0064]
[0065] Where z is the standard score; σ is the standard deviation; μ is the mean; and x is a specific score.
[0066] Based on the z-score processing formula, the processed data falls within a specific interval, resulting in standardized keypoint data. Using the horizontal and vertical axes of the image as a reference, the human image is rotated clockwise (condition 1, indicating the human is in an upright posture) to make the human torso horizontal with the image coordinates, and counterclockwise (condition 0, indicating the human is inverted) to complete the global standardization of the human body.
[0067] It is understandable that standardizing key point data based on standardization functions can achieve the purpose of removing dimensions from key point data and improving the data quality.
[0068] In step 130, after obtaining the standardized key point data, pose recognition is performed based on the standardized key point data to determine the pose of the target object.
[0069] Optionally, the pose recognition model can be an FCN (Fully Convolutional Networks) model. Key point data is input into the trained FCN model, and the pose of the target object is determined based on the trained FCN model.
[0070] Understandably, before determining the pose of the target object, the data in the acquired target image containing the target object is filtered to identify key point data within the target image. Standardizing the key point data using a standardization function further removes dimensions from the filtered data, improving the data quality for subsequent pose recognition.
[0071] The pose recognition method provided in this invention obtains a Gaussian distribution heatmap of the target image based on a Gaussian function before pose recognition. Based on this heatmap, key point data in the target image is determined, reducing interference data and improving the extraction of target object features. Standardizing the key point data using a standardization function removes dimensions, and performing pose recognition based on this standardized data improves the accuracy of pose recognition.
[0072] In one embodiment, the keypoint data is standardized based on a standardization function, including: performing global standardization on the keypoint data based on a standard score Z-score function to obtain globally standardized keypoint data; determining local part data of each target from the globally standardized keypoint data based on the Gaussian distribution heatmap, and performing local standardization on the local part data of each target based on the Z-score function to obtain the standardized keypoint data.
[0073] After acquiring the key point data of the target image, global and local standardization processing is performed on the key point data to achieve standardization of the key point data of the target image.
[0074] The Z-score, also known as the standard score, is the difference between a number and the mean, divided by the standard deviation. In statistics, the standard score is the sign of the standard deviation of an observed or measured value from the mean.
[0075] The keypoint data is globally standardized using the Z-score function. The Z-score function is as follows:
[0076]
[0077] Where z is the standard score; σ is the standard deviation; μ is the mean; and x is a specific score.
[0078] Based on the Z-score function, the key point data is standardized by scaling it proportionally to make it fall within a specific range, thereby achieving the purpose of removing dimensions from the key point data and improving its data quality.
[0079] Furthermore, after performing global standardization on the keypoint data, local standardization can be applied to the globally standardized keypoint data. Based on the Gaussian distribution heatmap of the target image, local part data of each target can be determined from the globally standardized keypoint data.
[0080] It's understandable that the target image contains multiple targets. Besides the targets, the image also contains a lot of noise or interference from the environment. By obtaining a Gaussian distribution heatmap of the target image based on a Gaussian function, the keypoint data of each target can be determined. Therefore, based on the Gaussian distribution heatmap of the target image, local data of each target can be determined from the globally standardized keypoint data.
[0081] Based on the Z-score function, local standardization is performed on the local data of each target, and the resulting data is used as the standardized key point data.
[0082] The pose recognition method provided in this embodiment of the invention performs global and local standardization processing on the key point data of the target image after acquiring the key point data, so as to achieve standardization processing of the key point data of the target image, and to remove the dimensions of the filtered data, thereby improving the data quality of the data subsequently used for pose recognition.
[0083] In one embodiment, performing pose recognition based on the standardized keypoint data to determine the pose of the target object includes: inputting the standardized keypoint data into a pose recognition model to obtain the pose of the target object output by the pose recognition model, wherein the pose recognition model is trained based on image samples and their corresponding human pose labels.
[0084] After obtaining the standardized key point data, the standardized key point data is input into the pose recognition model to determine the pose of the target object.
[0085] Optionally, the pose recognition model can be an FCN (Fully Convolutional Networks) model. Key point data is input into the trained FCN model, and the pose of the target object is determined based on the trained FCN model.
[0086] Before inputting the standardized keypoint data into the pose recognition model, the original pose recognition model needs to be trained. During the training process, a training dataset can be obtained. This dataset consists of multiple image samples and their corresponding human pose labels. Training the original pose recognition model based on this dataset enables the model to identify human pose targets from images containing multiple targets.
[0087] The posture recognition method provided in this embodiment of the invention determines the posture of the target object by inputting the standardized key point data into the posture recognition model after obtaining the standardized key point data.
[0088] In one embodiment, a method for training a pose recognition model includes: constructing a set of human pose images; determining a Gaussian distribution heatmap for each image in the set of human pose images based on a Gaussian function, and determining key point data corresponding to each image based on the Gaussian distribution heatmap; standardizing the key point data of each image based on a standardization function to obtain a standardized key point data set; determining the image samples and the test set of the pose recognition model from the key point data set; and training an initial pose recognition model based on the image samples and their corresponding human pose labels to obtain the pose recognition model.
[0089] During the training of the pose recognition model, the images in the acquired human pose image set can be processed before being used in the training process. Based on the Gaussian function, a Gaussian distribution heatmap is determined, thereby identifying the key point data corresponding to each image. Based on the standardization function, the key point data of each image is standardized and dimensionless.
[0090] Optionally, the training process for the pose recognition model can be completed using a training system consisting of multiple modules, such as... Figure 2 The schematic diagram of the training system module structure provided by the present invention is shown. The system includes an image acquisition module 210, a heat map acquisition and processing module 220, a pre-training data processing module 230, and a model training module 240.
[0091] The set of human pose images acquired by the image acquisition module 210 for training is input into the heatmap acquisition and processing module 220 to obtain Gaussian distribution heatmaps for each image. After processing by the pre-training data processing module 230, the model training module 240 is started to build the model and save the trained model.
[0092] The image acquisition module 210 is used to acquire a set of human pose images, most of which are from publicly available datasets. For example, the LPS dataset contains 2,000 images of humans in motion, with 14 joints, including full-body images of a single person. The FLIC dataset contains 20,000 full-body images of a single person, with scenes selected from film and television works. The MPII dataset contains 25,000 images of single and multiple people in everyday life. The MSCOCO dataset contains full-body images of multiple people. The AL Challenge dataset contains full-body images of multiple people.
[0093] The heatmap processing module 220 uses software to obtain Gaussian distribution heatmaps for each image using a Gaussian function, processing the images into pixel-level information. However, it is not limited to using only the Gaussian function to obtain heatmaps; other methods can also be used to process images into pixel-level information.
[0094] The final preprocessing of the data for training is the main content of the pre-training data processing module 230. It mainly involves rotating the image to make the Gaussian heatmap of the human body object in a relatively symmetrical shape. The whole is processed first, and then the local parts are adjusted.
[0095] The model training module 240 selects a pose recognition model for training. It covers everything from building the network model to outputting the results. The entire training process is completed, and the training results are saved.
[0096] Optionally, when the pose recognition model is a fully convolutional neural network model, the specific training process can be determined by [the relevant authority / organization]. Figure 3 The schematic diagram of the model training process provided by this invention is shown below:
[0097] Step 310: Approximately 30,000 images can be selected and stored in the folder to be processed as a collection of human pose images.
[0098] The source of the image can be automatically obtained using Python, such as the BeautifulSoup method, the lxml library, the PhantomJS method, or the Selenium method.
[0099] You can also collect public datasets such as LSP, FLIC, MPLL, MSCOCO, and AI Challenge; or use images taken by a camera for subsequent model training.
[0100] After constructing a human pose image set based on various datasets, data augmentation processing is performed on the images in the human pose image set.
[0101] Rotate the images in the human pose image set, with the horizontal direction as the reference point, in both clockwise and counterclockwise directions, with the rotation range being [0°, 90°].
[0102] The images in the human pose image set are randomly scaled, and the scaling ratio can be selected as [0.8, 1.2].
[0103] The image is stretched and shrunk in both the horizontal and vertical directions, with a range of [0.9, 1, 1]. This processing addresses the problem of asymmetric human body data and expands the number of images in the human pose image set.
[0104] Step 320: Generate Gaussian distribution heatmaps for each image.
[0105] Establish a two-dimensional coordinate system. Based on formula (5), obtain the key point information of human body parts, and use the obtained key points as the center of the circle;
[0106]
[0107] Among them, P i P_i represents the edge probability of key point information of human body parts in the target image; N is the normalization coefficient; i and j in P_(i|j) represent two key points respectively; P_(i|j) is the conditional probability from key point to key point, which is also the relationship model between points.
[0108] We set the Gaussian radius because the evaluation metric for the public datasets LSP, MPII, and FLIC is PCK (Percentage of Correct Keypoints), which is calculated by determining the proportion of normalized distances between detected keypoints and their corresponding ground truth that are less than a set threshold. Therefore, we selected a Gaussian radius that is 0.1 times the PCK.
[0109] Based on the Gaussian function, a Gaussian distribution heatmap is determined for each image in the human pose image set. Then, based on the Gaussian distribution heatmap of each image, the corresponding keypoint data set for each image is determined. The formula for the Gaussian function is:
[0110]
[0111] Where A is the peak value of the Gaussian curve; x0 and y0 are the image parameters of the target image; and x and y are the image information of the target image.
[0112] Image samples and a test set for the fully convolutional neural network model were determined from the keypoint dataset.
[0113] Step 330: Perform global standardization on the key point data of each image. The heatmap obtained in Step 320 is then standardized or normalized to remove dimensions and improve data quality. The Z-Score function is selected for data standardization. Z-Score standardization scales the data proportionally to ensure it falls within a specific range. The standardization formula is:
[0114]
[0115] Where z is the standard score; σ is the standard deviation; μ is the mean; and x is a specific score.
[0116] After standardizing the data, the main trunk of the human body is located using formula (8).
[0117]
[0118] Among them, P i P_i represents the edge probability of key point information of human body parts in the target image; N is the normalization coefficient; i and j in P_(i|j) represent two key points respectively; P_(i|j) is the conditional probability from key point to key point, which is also the relationship model between points.
[0119] After locating the main body of the human figure, using the horizontal and vertical axes of the images in the human posture image set as a reference, the human figure image is rotated clockwise (condition 1, indicating the human figure is in an upright posture) to make the main body of the human figure and the image coordinates horizontal. Then, the human figure image is rotated counter-clockwise (condition 0, indicating the human figure is inverted). After rotating the images where the main body of the human figure is not in an upright posture, the pixel distribution in the subsequent determination of the Gaussian distribution heatmap is made relatively symmetrical in the coordinate system, and the processed data is saved.
[0120] Step 340 involves performing local standardization on the key point data of each image, repeating the operation of step 330 to narrow down the scope to local parts of the human body.
[0121] Step 350: Train the fully convolutional neural network model.
[0122] To build the neural network framework, since the previous operations were all at the pixel level of the image, we chose the FCN fully convolutional neural network and the TensorFlow end-to-end machine learning platform.
[0123] The FCN network uses the glob library to obtain batch data (a batch is used to define the number of samples to be processed before updating the internal model parameters). The FCN network mainly consists of two parts: a fully convolutional part and a deconvolutional part. The training of the FCN network is mainly completed in three stages: feedforward neural network, deconvolution, and upsampling improved by the skip structure.
[0124] The feedforward neural network selects five blocks of VGG16, namely block1, block2, block3, block4 and block5. Each block contains several convolutional layers and pooling layers. Taking block4 as an example, block4 contains three convolutional layers conv3-512 and one pooling layer maxpool. The input image size is 224*224.
[0125] After deconvolution, the final output image after 5 blocks of VGG16 becomes a 7*7 image, with the number of channels doubling from 64 to 128 to 256 to 512, and the width and height of the image halved from 224 to 112 to 56 to 28 to 14 to 7. Figure 4 This is a flowchart of the deconvolution process provided by the present invention, showing the transformation process from block1 to block2. The working principle is to first fill the feature points of the original feature map, and then slide the convolution kernel on the original feature map to obtain a larger feature map. Since the image size entering the deconvolution operation is 7*7 after the VGG16 full convolution process, in order to obtain the storage size corresponding to the original image, the image is restored to the original size of 224*224. The deconvolution process with an original size of 3*3 is as follows: input size is 3*3, kernel size is 3*3, stride = 2, padding = 1, then the output size is: o = 2*(3-1) + 3 - 2*1 + 1 = 6. Therefore, the parameters of the corresponding 7*7 image are: o = s(i-1) + k-2p (o = 224, i = 7);
[0126] The upsampling process involves upsampling a low-resolution feature map to restore it to a high-resolution version. This is achieved by summing the new feature layer with the previous feature layer of the same size. For example... Figure 5 The flowchart of upsampling summation provided by this invention selects a 14*14*512 image obtained by training with VGG16 and deconvolution operation, converts it to 14*14*21 through a convolutional weight layer with a 1*1 kernel, and adds it to the obtained deconvolution feature map to realize the determination of an upsampling unit.
[0127] Step 360 involves adjusting and optimizing the model's parameters. This is to ensure the heatmaps obtained in steps 330 and 340 better adapt to the model. High-quality images are used to train the model to achieve high F1 scores, and parameter optimization is performed. The learning rate is a hyperparameter of deep neural networks, determining whether and when the objective function can effectively converge to a local minimum. Learning strategies (learning_rate_strage) are divided into fixed learning rate, segmented learning rate, and adaptive learning rate. In the FCN deep learning network, the learning rate was continuously tried and adjusted, ultimately using learning_rate_strage = 0.0001 for training. batch_size is the number of samples selected in one training iteration, affecting model optimization and speed. eval_batch_size is the number of samples selected in one validation iteration. Through continuous adjustment, a batch_size of 64 was finally chosen. Experiments showed that accuracy increases with the increase of the eval_batch_size parameter. The model saving parameter indicates how often the model is saved.
[0128] Step 370 involves evaluating the trained model. If it meets the current criteria, proceed to step 380; otherwise, continue training and proceed to step 350. The evaluation is based on recall, precision, accuracy, and F1 score. The evaluation criteria are continuously reviewed during training to ensure the best model is saved.
[0129] Step 380 is to save the trained model. After training, the model is saved to a folder.
[0130] The posture recognition method provided in this embodiment of the invention trains the posture recognition model before inputting standardized key point data into the posture recognition model, thus providing a foundation for determining the posture of the target object based on the trained posture recognition.
[0131] In one embodiment, after determining the image samples and the test set of the pose recognition model from the key point data set, the method further includes: testing the pose recognition model based on the test set to obtain test results corresponding to the test set; and determining that the precision of the test results is within a preset precision range, wherein the precision is determined based on the recall, precision, and accuracy of the pose recognition model training.
[0132] After obtaining the test set, the pose recognition model is tested based on the test set to verify the training results of the pose recognition model.
[0133] During the testing of the pose recognition model, it was determined that the precision of the test results of the pose recognition model was within a preset precision range. The precision was determined based on the recall, accuracy, and precision of the pose recognition model during training.
[0134] The precision F1 is:
[0135]
[0136] Recall rate is:
[0137]
[0138] Precession is:
[0139]
[0140] The accuracy is:
[0141]
[0142] In this context, TP stands for True Positive, which is predicted to be 1 and actually is 1, indicating a correct prediction; FP stands for False Positive, which is predicted to be 1 and actually is 0, indicating an incorrect prediction; FN stands for False Negative, which is predicted to be 0 and actually is 1, indicating an incorrect prediction; and TN stands for True Negative, which is predicted to be 0 and actually is 0, indicating a correct prediction.
[0143] The posture recognition method provided in this invention, after training a posture recognition model, determines that the precision of the model's test results is within a preset precision range. If the precision of the posture recognition model's test results is not within the preset precision range, the training process is adjusted to ensure the accuracy of subsequent posture recognition by the model.
[0144] In one embodiment, after constructing the human posture image set, the method further includes: determining the position of the human body trunk in each image of the human posture image set, and determining that the human body trunk position in each image is an image to be adjusted that is not upright; and rotating the image to be adjusted based on a preset rotation angle so that the human target in each image of the human posture image set is in an upright posture.
[0145] After constructing a set of human pose images, since the human poses contained in the set have various positions and orientations in the images, there are cases where the position of the human trunk is different from the horizontal direction of the image. In order to improve the recognition accuracy of the model after training, it is necessary to adjust the human trunk position in the human pose image set for non-upright postures.
[0146] Use formula (13) to locate the main trunk of the human body.
[0147]
[0148] Among them, P i P_i represents the edge probability of key point information of human body parts in the target image; N is the normalization coefficient; i and j in P_(i|j) represent two key points respectively; P_(i|j) is the conditional probability from key point to key point, which is also the relationship model between points.
[0149] After locating the main body position, the images in the human posture image set are used as references in the horizontal and vertical directions to determine the images to be adjusted where the main body position is not upright. Based on a preset rotation angle, the images to be adjusted are rotated so that the human target in each image in the human posture image set is in an upright posture.
[0150] Optionally, and based on a preset rotation angle, the image to be adjusted can be rotated as follows: The human image is rotated clockwise (condition 1, indicating the human body is in an upright posture) to make the human torso horizontal with the image coordinates; the human image is rotated counter-clockwise (condition 0, indicating the human body is inverted in the image). Rotating the image to be adjusted when the human torso is not in an upright posture ensures that the pixel distribution in the subsequent determination of the Gaussian heatmap achieves a relatively symmetrical distribution in the coordinate system.
[0151] The posture recognition method provided in this embodiment of the invention determines the position of the human body trunk in each image of the human posture image set after constructing the human posture image set, and identifies the image to be adjusted as the human body trunk position is not upright; based on a preset rotation angle, the image to be adjusted is rotated, thereby improving the recognition accuracy of the model obtained by subsequent training.
[0152] Figure 6 This is a schematic diagram of the posture recognition device provided by the present invention, as shown below. Figure 6 As shown, the device includes:
[0153] The key point data determination module 610 is used to determine the Gaussian distribution heatmap of the target image based on the Gaussian function, and to determine the key point data in the target image based on the Gaussian distribution heatmap, wherein the target image contains a target object;
[0154] The standardization processing module 620 is used to perform standardization processing on the key point data based on the standardization function to obtain standardized key point data;
[0155] The recognition module 630 is used to perform posture recognition based on the standardized key point data to determine the posture of the target object.
[0156] The pose recognition device provided in this invention obtains a Gaussian distribution heatmap of the target image based on a Gaussian function before performing pose recognition. Based on this heatmap, it determines key point data in the target image, reducing interference data and improving the extraction of target object features. Standardizing the key point data using a standardization function removes dimensions, and performing pose recognition based on the standardized key point data improves the accuracy of pose recognition.
[0157] In one embodiment, the standardization processing module 620 is specifically used for:
[0158] Based on a standardization function, the key point data is standardized, including:
[0159] Based on the Z-score function, the keypoint data is globally standardized to obtain globally standardized keypoint data.
[0160] Based on the Gaussian distribution heatmap, local part data of each target are determined from the key point data after global standardization, and local standardization is performed on the local part data of each target based on the Z-score function to obtain the standardized key point data.
[0161] In one embodiment, the identification module 630 is specifically used for:
[0162] Based on the standardized keypoint data, pose recognition is performed to determine the pose of the target object, including:
[0163] The standardized key point data is input into the pose recognition model to obtain the pose of the target object output by the pose recognition model. The pose recognition model is trained based on image samples and their corresponding human pose labels.
[0164] In one embodiment, the identification module 630 is further configured to:
[0165] Training methods for pose recognition models include:
[0166] Construct a set of human pose images;
[0167] Based on the Gaussian function, the Gaussian distribution heatmap of each image in the human pose image set is determined, and based on the Gaussian distribution heatmap of each image, the key point data corresponding to each image is determined;
[0168] Based on the standardization function, the key point data of each image are standardized to obtain a set of standardized key point data.
[0169] The image samples and the test set of the pose recognition model are determined from the key point data set;
[0170] The initial pose recognition model is trained based on the image samples and their corresponding human pose labels to obtain the pose recognition model.
[0171] In one embodiment, the identification module 630 is further configured to:
[0172] After determining the image samples and the test set of the pose recognition model from the key point data set, the method further includes:
[0173] Based on the test set, the pose recognition model is tested to obtain the test results corresponding to the test set;
[0174] The precision of the test results is determined to be within a preset precision range, which is based on the recall, precision, and accuracy of the posture recognition model training.
[0175] In one embodiment, the identification module 630 is further configured to:
[0176] After constructing the human pose image set, the following is also included:
[0177] The position of the human body trunk in each image in the human posture image set is determined, and the images in each image whose human body trunk position is not upright are to be adjusted.
[0178] Based on a preset rotation angle, the image to be adjusted is rotated so that the human target in each image of the human posture image set is in an upright posture.
[0179] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute an attitude recognition method, which includes:
[0180] Based on the Gaussian function, a Gaussian distribution heatmap of the target image is determined, and based on the Gaussian distribution heatmap, key point data in the target image are determined, wherein the target image contains a target object;
[0181] Based on the standardization function, the key point data is standardized to obtain standardized key point data;
[0182] The pose of the target object is determined by performing pose recognition based on the standardized key point data.
[0183] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0184] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the posture recognition method provided by the above methods, the method comprising:
[0185] Based on the Gaussian function, a Gaussian distribution heatmap of the target image is determined, and based on the Gaussian distribution heatmap, key point data in the target image are determined, wherein the target image contains a target object;
[0186] Based on the standardization function, the key point data is standardized to obtain standardized key point data;
[0187] The pose of the target object is determined by performing pose recognition based on the standardized key point data.
[0188] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned attitude recognition methods, the method comprising:
[0189] Based on the Gaussian function, a Gaussian distribution heatmap of the target image is determined, and based on the Gaussian distribution heatmap, key point data in the target image are determined, wherein the target image contains a target object;
[0190] Based on the standardization function, the key point data is standardized to obtain standardized key point data;
[0191] The pose of the target object is determined by performing pose recognition based on the standardized key point data.
[0192] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0193] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pose recognition method, characterized in that, include: Obtain key point information of human body parts in the target image; For each human body key point corresponding to the key point information, a Gaussian function is applied with the key point as the center and a selected preset Gaussian radius to determine the Gaussian distribution heatmap of the target image, and based on the Gaussian distribution heatmap, the key point data in the target image is determined, wherein the target image contains the target object; Based on the Z-score function, the keypoint data is globally standardized to obtain globally standardized keypoint data. Based on the Gaussian distribution heatmap, local part data of each target are determined from the key point data after global standardization, and local standardization is performed on the local part data of each target based on the Z-score function to obtain standardized key point data. The standardized key point data is input into the pose recognition model to obtain the pose of the target object output by the pose recognition model. The pose recognition model is trained based on image samples and their corresponding human pose labels. The training method for the pose recognition model includes: Construct a set of human posture images; determine the position of the human body trunk in each image in the set of human posture images, and identify the images in which the human body trunk is not in an upright posture and need to be adjusted; rotate the images to be adjusted based on a preset rotation angle so that the human target in each image in the set of human posture images is in an upright posture. Based on the Gaussian function, the Gaussian distribution heatmap of each image in the human pose image set is determined, and based on the Gaussian distribution heatmap of each image, the key point data corresponding to each image is determined; Based on the standardization function, the key point data of each image are standardized to obtain a set of standardized key point data. The image samples and the test set of the pose recognition model are determined from the key point data set; The initial pose recognition model is trained based on the image samples and their corresponding human pose labels to obtain the pose recognition model.
2. The pose recognition method according to claim 1, characterized in that, After determining the image samples and the test set of the pose recognition model from the key point data set, the method further includes: Based on the test set, the pose recognition model is tested to obtain the test results corresponding to the test set; The precision of the test results is determined to be within a preset precision range, which is based on the recall, precision, and accuracy of the posture recognition model training.
3. A posture recognition device, characterized in that, include: The key point data determination module is used to obtain key point information of human body parts in the target image; For each human body key point corresponding to the key point information, a Gaussian function is applied with the key point as the center and a selected preset Gaussian radius to determine the Gaussian distribution heatmap of the target image, and based on the Gaussian distribution heatmap, the key point data in the target image is determined, wherein the target image contains the target object; The standardization module is used to perform global standardization on the keypoint data based on the Z-score function to obtain globally standardized keypoint data; based on the Gaussian distribution heatmap, it determines the local part data of each target from the globally standardized keypoint data, and performs local standardization on the local part data of each target based on the Z-score function to obtain standardized keypoint data. The recognition module is used to input the standardized key point data into the pose recognition model to obtain the pose of the target object output by the pose recognition model. The pose recognition model is trained based on image samples and their corresponding human pose labels. The training method for the pose recognition model includes: Construct a set of human posture images; determine the position of the human body trunk in each image in the set of human posture images, and identify the images in which the human body trunk is not in an upright posture and need to be adjusted; rotate the images to be adjusted based on a preset rotation angle so that the human target in each image in the set of human posture images is in an upright posture. Based on the Gaussian function, the Gaussian distribution heatmap of each image in the human pose image set is determined, and based on the Gaussian distribution heatmap of each image, the key point data corresponding to each image is determined; Based on the standardization function, the key point data of each image are standardized to obtain a set of standardized key point data. The image samples and the test set of the pose recognition model are determined from the key point data set; The initial pose recognition model is trained based on the image samples and their corresponding human pose labels to obtain the pose recognition model.
4. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the posture recognition method as described in claim 1 or 2.
5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the posture recognition method as described in claim 1 or 2.
6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the posture recognition method as described in claim 1 or 2.
Citation Information
Patent Citations
Attitude recognition method and system based on thermodynamic diagram and offset vector and storage medium
CN111191622A