A human key point detection method and device of a camera

CN115578754BActive Publication Date: 2026-09-25FUJIAN STAR NET WISDOM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211181092.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2026-09-25
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

[0003]现有技术中,人体关键点检测模型通常模型尺寸较大,推理时间过长,难以在摄像头等低容量低算力的嵌入式设备上部署应用

Benefits of technology

[0041]1、通过模型设计,缩减了模型大小,并且提高了模型的推理速度;通过模型转换,在尽量避免精度损失的情况下,节省了摄像头的存储空间;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115578754B_ABST
    Figure CN115578754B_ABST
Patent Text Reader

Abstract

The application discloses a kind of camera human body key point detection method and device, it is related to embedded device key point detection technical field.Camera is built-in neural network processor, the embodiment of the application uses lightweight trunk network to reduce model size, accelerates model inference;Utilize upsampling, reduce model output size, further accelerate model inference;Through web crawler, training data is obtained, and training set content is enriched;Through quantifying model data, optimize model performance in the case where model accuracy loss is minimized as far as possible;Again through model coding compression, more save device space;Using neural network processor for model inference operation acceleration.The human body key point detection method and device of a kind of camera provided by the application, through lightweight human body key point detection model, the computing power of neural network processor is mobilized, and the human body key point detection of camera is realized to realize the human body key point detection of camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of key point detection technology for embedded devices, and in particular to a method and apparatus for detecting human key points in a camera. Background Technology

[0002] With the further development of computer vision technology, deep learning neural networks have achieved significant results in image recognition and pose detection tasks. Pose detection begins with feature extraction from the image, determining the positions of key human body points to complete pose detection. On cameras, to improve performance for artificial intelligence applications, neural network processors (NPUs) are being used to accelerate the processing of specific formulas in AI algorithms.

[0003] In existing technologies, human keypoint detection models are typically large in size and have excessively long inference times, making them difficult to deploy on low-capacity, low-computing-power embedded devices such as cameras. Therefore, designing and training a lightweight human keypoint detection model, and further optimizing space utilization within a camera and leveraging the computing power of the neural network processor, has become a technical problem that needs to be solved. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and device for human body key point detection by a camera. By designing and training a lightweight human body key point detection model, and further optimizing the use of space in the camera and mobilizing the computing power of the neural network processor, the human body key point detection of the camera can be realized.

[0005] In a first aspect, the present invention provides a method for detecting key human points using a camera, comprising:

[0006] Step 10: Construct a lightweight human keypoint model, including input / output nodes, backbone network, upsampling module, and output node;

[0007] Step 20: Obtain human key point images and annotation files as training data, and then divide them into training set and test set;

[0008] Step 30: Configure the training parameters, and then train the lightweight human keypoint model using the training set;

[0009] Step 40: Validate the human keypoint detection model using the test set. If the validation result is not up to standard, return to step 20 to increase the amount of training data or return to step 30 to adjust the training parameters; if the validation result is up to standard, proceed to step 50.

[0010] Step 50: Convert the trained model file and weight data into a general neural network model format file. Then set the model conversion preprocessing and quantization method parameters, model input nodes, model input size and model output nodes. Quantize and compress the model format file into a model format that can be used by the neural network processor built into the camera. Finally, output the converted human key point detection model.

[0011] Step 70: Deploy the converted human key point detection model onto the camera device;

[0012] Step 80: Process the original image captured by the camera to obtain the input image of the model, call the neural network processor, perform human key point detection through the converted human key point detection model, convert the output heat map information into human key point coordinates and display it on the original image of the camera.

[0013] Furthermore, the method also includes:

[0014] Step 60: Evaluate the accuracy, performance, and memory usage of the converted model. If the evaluation results meet the requirements, proceed to step 70; if the evaluation results do not meet the requirements, return to step 50 to optimize the model conversion parameters.

[0015] Furthermore, in step 40, human image data is crawled using web crawlers to increase the amount of training data.

[0016] Furthermore, in step 50, the model conversion preprocessing and quantization method parameters include batch size, input mean, input normalized value, image channel order, horizontal / vertical merging, quantization type, quantization parameter optimization algorithm, number of quantization algorithm iterations, and model optimization level.

[0017] Furthermore, step 80 specifically includes:

[0018] Step 81: Convert the original image captured by the camera auxiliary stream into a common image processing format, use the human detection function to obtain the coordinate information of the human detection box on the original image; enlarge the human detection box coordinate information by a specified factor to crop out the human image; scale the cropped human image to the specified input size of the model.

[0019] Step 82: Configure the input and output attributes of the neural network processor, call the neural network processor interface to perform human key point detection model inference, and obtain a human key point heat map array;

[0020] Step 83: For each human body key point, find the number with the largest hotspot value on its corresponding heat map array, and convert its index into heat map coordinates; convert the heat map coordinates into human body key point coordinates on the cropped human body image, and then convert the human body key point coordinates of the cropped human body image into human body key point coordinates on the original image to obtain the coordinate information of all human body key points on the original image.

[0021] Step 84: Draw the human body key points on the main image of the camera based on the coordinate information of the human body key points.

[0022] Furthermore, step 84 specifically includes:

[0023] A queue is established, and the coordinate information of human key points obtained from the auxiliary stream is added to the queue. In the main processing flow, the coordinate information of human key points is retrieved from the queue. Based on the coordinates, the pixel value of the corresponding image pixel and the surrounding pixels is changed to change the color of the pixel, thereby achieving the effect of drawing human key points.

[0024] In a second aspect, the present invention provides a human key point detection device for a camera, comprising: the camera having a built-in neural network processor, and the device comprising: a model building module, a data processing module, a model training module, a model verification module, a model conversion module, a model deployment module, and a key point output module;

[0025] The model building module is used to build a lightweight human keypoint model, including input and output nodes, backbone network, upsampling module and output node;

[0026] The data processing module is used to acquire human key point images and annotation files as training data, and then divide them into training set and test set;

[0027] The model training module is used to configure the training parameters and then train the lightweight human keypoint model using the training set.

[0028] The model validation module is used to validate the human keypoint detection model using a test set. If the validation result is unsatisfactory, it returns to the data processing module to increase the amount of training data or returns to the model training module to adjust the training parameters. If the validation result is satisfactory, it proceeds to the model conversion module.

[0029] The model conversion module is used to convert the trained model files and weight data into a common neural network model format file. Then, it sets the model conversion preprocessing and quantization method parameters, model input nodes, model input size and model output nodes, quantizes and compresses the model format file into a model format that can be used by the neural network processor built into the camera, and finally outputs the converted human key point detection model.

[0030] The model deployment module is used to deploy the converted human keypoint detection model onto camera devices;

[0031] The key point output module is used to process the raw images captured by the camera to obtain the input image of the model, call the neural network processor, perform human key point detection through the converted human key point detection model, and convert the output heat map information into human key point coordinates and display them on the raw image of the camera.

[0032] Furthermore, the device also includes a model evaluation module for evaluating the accuracy, performance, and memory usage of the converted model. When the evaluation results meet the requirements, the device proceeds to the model deployment module; when the evaluation results do not meet the requirements, the device returns to the model conversion module to optimize the model conversion parameters.

[0033] Furthermore, the key point output module specifically includes:

[0034] The input image processing module is used to convert the raw images captured by the camera auxiliary stream into a common image processing format, use the human detection function to obtain the coordinate information of the human detection box on the raw image; enlarge the human detection box coordinate information by a specified factor to crop out the human image; and scale the cropped human image to the specified input size of the model.

[0035] The detection module is called to configure the input and output attributes of the neural network processor, and the neural network processor interface is called to perform human key point detection model inference to obtain a human key point heat map array.

[0036] The coordinate transformation module is used to find the number with the largest hotspot value on the corresponding heat map array for each human body key point, and convert its index into heat map coordinates; convert the heat map coordinates into human body key point coordinates on the cropped human body image, and then convert the human body key point coordinates of the cropped human body image into human body key point coordinates on the original image, so as to obtain the coordinate information of all human body key points on the original image.

[0037] The key point drawing module is used to draw human key points on the main image of the camera based on the coordinate information of human key points.

[0038] Furthermore, the key point drawing module is specifically used for:

[0039] A queue is established, and the coordinate information of human key points obtained from the auxiliary stream is added to the queue. In the main processing flow, the coordinate information of human key points is retrieved from the queue. Based on the coordinates, the pixel value of the corresponding image pixel and the surrounding pixels is changed to change the color of the pixel, thereby achieving the effect of drawing human key points.

[0040] The technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:

[0041] 1. Through model design, the model size was reduced and the model's inference speed was improved; through model conversion, the camera's storage space was saved while minimizing accuracy loss.

[0042] 2. It makes better use of the neural network processor, which accelerates the processing algorithm and improves the performance of the human key point detection model on the camera.

[0043] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0045] Figure 1 This is a schematic diagram of the overall process for detecting key human points according to an embodiment of the present invention;

[0046] Figure 2 This is a schematic diagram of the upsampling neural network structure according to an embodiment of the present invention;

[0047] Figure 3 This is a schematic diagram of the human body key point detection results according to an embodiment of the present invention;

[0048] Figure 4 This is a schematic diagram of the model conversion process according to an embodiment of the present invention;

[0049] Figure 5 This is a flowchart of the method in Embodiment 1 of the present invention;

[0050] Figure 6 This is a schematic diagram of the device in Embodiment 2 of the present invention. Detailed Implementation

[0051] This invention provides a method and apparatus for detecting human key points using a camera. By designing and training a lightweight human key point detection model, and further optimizing the use of space in the camera and mobilizing the computing power of the neural network processor, the detection of human key points in the camera is achieved.

[0052] The overall concept of this invention is as follows:

[0053] The embodiments of the present invention employ multiple methods, including a design method for a lightweight human keypoint detection model, a training method for a lightweight human keypoint detection model, a conversion method for the corresponding format of the model's neural network processor, and a method for utilizing the camera's neural network processor.

[0054] The design methodology of the lightweight human keypoint detection model features: using a lightweight backbone network to reduce model size and accelerate model inference; and utilizing upsampling to reduce the model output size, further accelerating model inference. The training methodology of the lightweight human keypoint detection model features: acquiring training data through web crawling to enrich the training set content. The format conversion method for the model's neural network processor features: quantizing model data to optimize model performance while minimizing accuracy loss; and compressing model encoding to save device space. The utilization method of the camera's neural network processor features: using computer vision image processing to modify image content attributes; and using the neural network processor to accelerate model inference computation.

[0055] Specific operation process of this invention embodiment:

[0056] Step 1: Design a lightweight human keypoint model using MXNet as the deep learning framework. Design the model's neural network structure, including input / output nodes, the backbone network, and the upsampling module. See [link to overall human keypoint detection workflow] for details. Figure 1 The specific steps are as follows:

[0057] Step 1.1: Set the model input size to 256 pixels high, 192 pixels wide, and 3 channels;

[0058] Step 1.2: Use MobileNetv2 as the backbone network and add a series of neural network layers after the backbone network, including 2D convolutional layers (Conv2D), data normalization layers (BatchNormalization), rectified linear activation function (ReLU), and 2D convolutional transpose (Conv2DTranspose).

[0059] Step 1.3: Implement the upsampling module to reduce the size of the final output heatmap by lowering the image resolution. See the network structure below. Figure 2 ;

[0060] Step 1.4: Normalize the model output into a heatmap array containing 17 human keypoints, 64 pixels high and 48 pixels wide.

[0061] Step 2: Prepare human body keypoint images and annotation files as training data. The specific steps are as follows:

[0062] Step 2.1: Collect human images from the coco2017 dataset;

[0063] Step 2.2: Prepare the corresponding human body key point annotation file. The annotation file marks the locations of the human body key points as follows: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. Each key point records its horizontal and vertical coordinates on the image.

[0064] Step 2.3: If the data in the coco17 dataset is still insufficient, crawl human image data through web scraping, preprocess the collected human images, exclude human images with indistinct features, align the human images to adapt to the model input, and write corresponding annotation files for the images.

[0065] Step 3: Train the model using the training data obtained in Step 2 and the model designed in Step 1, dividing the data into training and test sets. The specific steps are as follows:

[0066] Step 3.1: Train 128 images simultaneously in each batch, and set the size of the input image to 256 pixels high and 192 pixels wide;

[0067] Step 3.2: Based on the experience gained from multiple experiments, the training parameters are set as follows: the initial learning rate is set to 0.001, the number of training iterations is set to 140, the learning rate is set to 0.0001 at the 90th iteration, the learning rate is set to 0.00001 at the 120th iteration, and the learning rate is set to 0.000001 at the 140th iteration.

[0068] Step 3.3: The evaluation metric is set as heatmap accuracy, the loss function is minimum mean squared error loss (L2 loss), and the optimizer is trained using adaptive moment estimation (Adam).

[0069] Step 3.4: Configuration complete, start model training.

[0070] Step 4: After training, validate the human keypoint detection model using a test set. See the detection results below. Figure 3 The specific steps are as follows:

[0071] Step 4.1: If the accuracy is not high enough, it may be due to insufficient training data leading to overfitting. In step 2, use web crawling to obtain more data and add it to the training set for retraining.

[0072] Step 4.2: Alternatively, adjust the training parameters, modify the learning rate, learning rate of change, and number of training iterations, and repeatedly experiment to verify the training effect. Repeat the training and verification until the accuracy reaches the required metric.

[0073] Step 5: Convert the human keypoint detection model, see [link / reference] Figure 4 The specific steps are as follows:

[0074] Step 5.1: Using the model file and weight data obtained from training in Step 3, convert them to the general neural network model format ONNX;

[0075] Step 5.2: Convert the ONNX file to a more space-efficient model format that can be used by neural network processors. First, initialize the model conversion environment.

[0076] Step 5.3: Set the conversion model parameters, batch size 100, input mean 123.675, 116.28, 103.53, input normalized values ​​58.82, 58.82, 58.82, image channel order RGB, set horizontal merging, set quantization type to u8 type asymmetric quantization, use normal algorithm for quantization parameter optimization, set the number of quantization algorithm iterations to 3, and perform 3 levels of optimization on the model;

[0077] Step 5.4: Configure the human keypoint detection ONNX model to be converted, and set the input nodes, input size, and output nodes;

[0078] Step 5.5: After configuration, build the converted model, and then quantize and compress the model;

[0079] Step 5.6: Finally, output and save the converted human key point detection model.

[0080] Step 6: Evaluate the accuracy, performance, and memory usage of the converted model, and optimize the model conversion parameters. The specific steps are as follows:

[0081] Step 6.1: Compare the inference results of the transformed model with the inference results of the original frame, and use cosine distance to evaluate the similarity between the two results;

[0082] Step 6.2: Analyze the quantization accuracy of the converted model, save the intermediate results of each layer in both the non-quantized and quantized cases, evaluate their similarity using cosine distance, observe the changes in similarity at each layer, and identify which operators are not quantization-friendly.

[0083] Step 6.3: Perform performance evaluation on the transformed model, evaluate the time consumption information of each layer during model runtime, and calculate the overall time consumption of the model;

[0084] Step 6.4: Perform a memory evaluation on the converted model to check the memory usage during model inference on the development board;

[0085] Step 6.5: After the evaluation is completed, add the converted model to the camera's burn-in file. After burn-in, the camera can use the human key point detection model.

[0086] Step 7: In order to use the human keypoint detection model on the camera, first prepare the model input on the camera. The specific steps are as follows:

[0087] Step 7.1: Convert the NV12 raw image captured by the camera auxiliary stream into RGB24 format, and use the human detection function to obtain the coordinate information of the detection box of the human body on the raw image;

[0088] Step 7.2: Using the coordinate information of the detection box of the human body in the original image, enlarge it by 1.25 times and crop the human body image;

[0089] Step 7.3: Scale the cropped human image to the input size of the model, 256 pixels high and 192 pixels wide. The model input is now ready.

[0090] Step 8: After preparing the input image for the model, call the neural network processor for inference. The specific steps are as follows:

[0091] Step 8.1: First, configure the input and output attributes of the neural network processor. The node quantization type is UINT8, and the node format order is height-width-channel.

[0092] Step 8.2: After setting the neural network processor properties, you can call the neural network processor API to perform human keypoint detection model inference.

[0093] Step 9: After the human keypoint detection model inferences, the obtained model output is post-processed. The specific steps are as follows:

[0094] Step 9.1: The 17 human key points correspond to 17 heat map arrays. For each human key point, find the number with the largest heat value on its corresponding heat map array and convert its index into heat map coordinates.

[0095] Step 9.2: Convert the heat map coordinates to the coordinates of human key points on the cropped human image, and then convert the coordinates of human key points on the cropped human image to the coordinates of human key points on the original image.

[0096] Step 9.3: Complete the conversion of the coordinates of all 17 human body key points to obtain the coordinate information of all human body key points on the original image.

[0097] Step 10: After steps 7, 8, and 9, the image from the auxiliary camera stream is processed by the human keypoint detection model to obtain the coordinate information of human keypoints in the image. Now, based on the coordinate information of the human keypoints, the human keypoints are drawn on the main camera stream. The specific steps are as follows:

[0098] Step 10.1: Create a queue and add the human body key point coordinate information obtained from the auxiliary flow to the queue;

[0099] Step 10.2: In the main processing flow, retrieve the coordinate information of human body key points from the queue, and change the pixel value of the corresponding image pixel and the surrounding pixels according to its coordinates to change the color of the pixel, thereby achieving the effect of drawing human body key points.

[0100] Step 10.3: After drawing, the main image from the camera will have the markings of key human body points.

[0101] After applying the human keypoint detection model to a camera, a case study is conducted for verification. The steps are as follows:

[0102] 1. The camera is powered on and initialized.

[0103] 2. On a computer device, use the RTSP protocol on a media player to obtain the mainstream streaming media from the camera and play the camera image.

[0104] 3. To enable the camera lens to capture the human body;

[0105] 4. At this time, the camera view displayed on the media player will mark 17 key points of the human body.

[0106] Example 1

[0107] This embodiment provides a method for detecting key human points using a camera, wherein the camera has a built-in neural network processor, such as... Figure 5 As shown, the method includes:

[0108] Step 10: Construct a lightweight human keypoint model, including input / output nodes, backbone network, upsampling module, and output node;

[0109] Step 20: Obtain human key point images and annotation files as training data, and then divide them into training set and test set;

[0110] Step 30: Configure the training parameters, and then train the lightweight human keypoint model using the training set;

[0111] Step 40: Validate the human keypoint detection model using the test set. If the validation result is not up to standard, return to step 20 to increase the amount of training data (e.g., crawl human image data through web crawler) or return to step 30 to adjust the training parameters. If the validation result is up to standard, proceed to step 50.

[0112] Step 50: Convert the trained model file and weight data into a general neural network model format file. Then set the model conversion preprocessing and quantization method parameters, model input nodes, model input size and model output nodes. Quantize and compress the model format file into a model format that can be used by the neural network processor built into the camera. Finally, output the converted human key point detection model.

[0113] Step 70: Deploy the converted human key point detection model onto the camera device;

[0114] Step 80: Process the original image captured by the camera to obtain the input image of the model, call the neural network processor, perform human key point detection through the converted human key point detection model, convert the output heat map information into human key point coordinates and display it on the original image of the camera.

[0115] In one possible implementation, the method further includes:

[0116] Step 60: Evaluate the accuracy, performance, and memory usage of the converted model. If the evaluation results meet the requirements, proceed to step 70; if the evaluation results do not meet the requirements, return to step 50 to optimize the model conversion parameters.

[0117] Before deploying the model to the camera, evaluating its accuracy, performance, and memory usage can ensure that the converted model meets application requirements.

[0118] In one possible implementation, in step 50, the model conversion preprocessing and quantization method parameters include batch size, input mean, input normalized value, image channel order, horizontal / vertical merging, quantization type, quantization parameter optimization algorithm, number of quantization algorithm iterations, and model optimization level.

[0119] In one possible implementation, step 80 specifically includes:

[0120] Step 81: Convert the original image captured by the camera auxiliary stream into a common image processing format, use the human detection function to obtain the coordinate information of the human detection box on the original image; enlarge the human detection box coordinate information by a specified factor to crop out the human image; scale the cropped human image to the specified input size of the model.

[0121] Step 82: Configure the input and output attributes of the neural network processor, call the neural network processor interface to perform human key point detection model inference, and obtain a human key point heat map array;

[0122] Step 83: For each human body key point, find the number with the largest hotspot value on its corresponding heat map array, and convert its index into heat map coordinates; convert the heat map coordinates into human body key point coordinates on the cropped human body image, and then convert the human body key point coordinates of the cropped human body image into human body key point coordinates on the original image to obtain the coordinate information of all human body key points on the original image.

[0123] Step 84: Based on the coordinate information of the human body's key points, draw the human body's key points on the main image of the camera:

[0124] A queue is established, and the coordinate information of human key points obtained from the auxiliary stream is added to the queue. In the main processing flow, the coordinate information of human key points is retrieved from the queue. Based on the coordinates, the pixel value of the corresponding image pixel and the surrounding pixels is changed to change the color of the pixel, thereby achieving the effect of drawing human key points.

[0125] Based on the same inventive concept, this application also provides an apparatus corresponding to the method in Embodiment 1, as detailed in Embodiment 2.

[0126] Example 2

[0127] This embodiment provides a human key point detection device for a camera, wherein the camera has a built-in neural network processor, such as... Figure 6 As shown, the device includes: a model building module, a data processing module, a model training module, a model validation module, a model conversion module, a model deployment module, and a key point output module;

[0128] The model building module is used to build a lightweight human keypoint model, including input and output nodes, backbone network, upsampling module and output node;

[0129] The data processing module is used to acquire human key point images and annotation files as training data, and then divide them into training set and test set;

[0130] The model training module is used to configure the training parameters and then train the lightweight human keypoint model using the training set.

[0131] The model validation module is used to validate the human keypoint detection model using a test set. If the validation result is unsatisfactory, it returns to the data processing module to increase the amount of training data or returns to the model training module to adjust the training parameters. If the validation result is satisfactory, it proceeds to the model conversion module.

[0132] The model conversion module is used to convert the trained model files and weight data into a common neural network model format file. Then, it sets the model conversion preprocessing and quantization method parameters, model input nodes, model input size and model output nodes, quantizes and compresses the model format file into a model format that can be used by the neural network processor built into the camera, and finally outputs the converted human key point detection model.

[0133] The model deployment module is used to deploy the converted human keypoint detection model onto camera devices;

[0134] The key point output module is used to process the raw images captured by the camera to obtain the input image of the model, call the neural network processor, perform human key point detection through the converted human key point detection model, and convert the output heat map information into human key point coordinates and display them on the raw image of the camera.

[0135] Furthermore, the device also includes a model evaluation module for evaluating the accuracy, performance, and memory usage of the converted model. When the evaluation results meet the requirements, the device proceeds to the model deployment module; when the evaluation results do not meet the requirements, the device returns to the model conversion module to optimize the model conversion parameters.

[0136] Furthermore, the key point output module specifically includes:

[0137] The input image processing module is used to convert the raw images captured by the camera auxiliary stream into a common image processing format, use the human detection function to obtain the coordinate information of the human detection box on the raw image; enlarge the human detection box coordinate information by a specified factor to crop out the human image; and scale the cropped human image to the specified input size of the model.

[0138] The detection module is called to configure the input and output attributes of the neural network processor, and the neural network processor interface is called to perform human key point detection model inference to obtain a human key point heat map array.

[0139] The coordinate transformation module is used to find the number with the largest hotspot value on the corresponding heat map array for each human body key point, and convert its index into heat map coordinates; convert the heat map coordinates into human body key point coordinates on the cropped human body image, and then convert the human body key point coordinates of the cropped human body image into human body key point coordinates on the original image, so as to obtain the coordinate information of all human body key points on the original image.

[0140] The key point drawing module is used to draw human key points on the main image of the camera based on the coordinate information of human key points.

[0141] Furthermore, the key point drawing module is specifically used for:

[0142] A queue is established, and the coordinate information of human key points obtained from the auxiliary stream is added to the queue. In the main processing flow, the coordinate information of human key points is retrieved from the queue. Based on the coordinates, the pixel value of the corresponding image pixel and the surrounding pixels is changed to change the color of the pixel, thereby achieving the effect of drawing human key points.

[0143] Since the apparatus described in Embodiment 2 of the present invention is an apparatus used to implement the method of Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the apparatus based on the method described in Embodiment 1 of the present invention, and therefore will not be described again here. All apparatuses used in the method of Embodiment 1 of the present invention fall within the scope of protection of the present invention.

[0144] This invention reduces model size and improves inference speed through model design; saves camera storage space by minimizing accuracy loss through model conversion; and makes better use of neural network processors to accelerate processing algorithms, thereby improving the performance of the human key point detection model on the camera.

[0145] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0146] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0147] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0148] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0149] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for detecting key human points using a camera, characterized in that: The camera has a built-in neural network processor, and the method includes: Step 10: Construct a lightweight human keypoint model, including input / output nodes, backbone network, upsampling module, and output node; Step 20: Obtain human key point images and annotation files as training data, and then divide them into training set and test set; Step 30: Configure the training parameters, and then train the lightweight human keypoint model using the training set; Step 40: Validate the human keypoint detection model using the test set. If the validation result is not up to standard, return to step 20 to increase the amount of training data or return to step 30 to adjust the training parameters; if the validation result is up to standard, proceed to step 50. Step 50: Convert the trained model file and weight data into a general neural network model format file. Then set the model conversion preprocessing and quantization method parameters, model input nodes, model input size and model output nodes. Quantize and compress the model format file into a model format that can be used by the neural network processor built into the camera. Finally, output the converted human key point detection model. Step 70: Deploy the converted human key point detection model onto the camera device; Step 80: Process the original image captured by the camera to obtain the input image of the model, call the neural network processor, perform human key point detection through the converted human key point detection model, convert the output heat map information into human key point coordinates and display them on the original image of the camera. Specifically, step 80 includes: Step 81: Convert the original image captured by the camera auxiliary stream into a common image processing format, use the human detection function to obtain the coordinate information of the human detection box on the original image; enlarge the human detection box coordinate information by a specified factor to crop out the human image; scale the cropped human image to the specified input size of the model. Step 82: Configure the input and output attributes of the neural network processor, call the neural network processor interface to perform human key point detection model inference, and obtain a human key point heat map array; Step 83: For each human body key point, find the number with the largest hotspot value on its corresponding heat map array, and convert its index into heat map coordinates; convert the heat map coordinates into human body key point coordinates on the cropped human body image, and then convert the human body key point coordinates of the cropped human body image into human body key point coordinates on the original image to obtain the coordinate information of all human body key points on the original image. Step 84: Draw human key points on the main image of the camera based on the coordinate information of human key points. Specifically, this includes: establishing a queue and adding the coordinate information of human key points obtained from the auxiliary stream to the queue; in the main processing flow, retrieving the coordinate information of human key points from the queue, and changing the pixel value of the corresponding image pixel and the surrounding pixels according to its coordinates to change the color of the pixel, thereby achieving the effect of drawing human key points.

2. The method according to claim 1, characterized in that, The method further includes: Step 60: Evaluate the accuracy, performance, and memory usage of the converted model. If the evaluation results meet the requirements, proceed to step 70; if the evaluation results do not meet the requirements, return to step 50 to optimize the model conversion parameters.

3. The method according to claim 1, characterized in that: In step 40, human image data is crawled through a web crawler to increase the amount of training data.

4. The method according to claim 1, characterized in that: In step 50, the model conversion preprocessing and quantization method parameters include batch size, input mean, input normalized value, image channel order, horizontal / vertical merging, quantization type, quantization parameter optimization algorithm, number of quantization algorithm iterations, and model optimization level.

5. A human key point detection device for a camera, characterized in that: The camera has a built-in neural network processor, and the device includes: a model building module, a data processing module, a model training module, a model verification module, a model conversion module, a model deployment module, and a key point output module; The model building module is used to build a lightweight human keypoint model, including input and output nodes, backbone network, upsampling module and output node; The data processing module is used to acquire human key point images and annotation files as training data, and then divide them into training set and test set; The model training module is used to configure the training parameters and then train the lightweight human keypoint model using the training set. The model validation module is used to validate the human keypoint detection model using a test set. If the validation result is unsatisfactory, it returns to the data processing module to increase the amount of training data or returns to the model training module to adjust the training parameters. If the validation result is satisfactory, it proceeds to the model conversion module. The model conversion module is used to convert the trained model files and weight data into a common neural network model format file. Then, it sets the model conversion preprocessing and quantization method parameters, model input nodes, model input size and model output nodes, quantizes and compresses the model format file into a model format that can be used by the neural network processor built into the camera, and finally outputs the converted human key point detection model. The model deployment module is used to deploy the converted human keypoint detection model onto camera devices; The key point output module is used to process the raw images captured by the camera to obtain the input image of the model, call the neural network processor, perform human key point detection through the converted human key point detection model, and convert the output heat map information into human key point coordinates and display them on the raw image of the camera. Specifically, the key point output module includes: The input image processing module is used to convert the raw images captured by the camera auxiliary stream into a common image processing format, use the human detection function to obtain the coordinate information of the human detection box on the raw image; enlarge the human detection box coordinate information by a specified factor to crop out the human image; and scale the cropped human image to the specified input size of the model. The detection module is called to configure the input and output attributes of the neural network processor, and the neural network processor interface is called to perform human key point detection model inference to obtain a human key point heat map array. The coordinate transformation module is used to find the number with the largest hotspot value on the corresponding heat map array for each human body key point, and convert its index into heat map coordinates; convert the heat map coordinates into human body key point coordinates on the cropped human body image, and then convert the human body key point coordinates of the cropped human body image into human body key point coordinates on the original image, so as to obtain the coordinate information of all human body key points on the original image. The key point drawing module is used to draw human key points on the main image of the camera based on the coordinate information of human key points. Specifically, it includes: establishing a queue and adding the human key point coordinate information obtained from the auxiliary stream to the queue; in the main processing flow, retrieving the human key point coordinate information from the queue, and changing the pixel value at the corresponding image pixel position and the surrounding pixels according to the coordinates to change the color of the pixel, thereby achieving the effect of drawing human key points.

6. The apparatus according to claim 5, characterized in that, The device also includes a model evaluation module, which evaluates the accuracy, performance, and memory usage of the converted model. When the evaluation results meet the requirements, the device proceeds to the model deployment module; when the evaluation results do not meet the requirements, the device returns to the model conversion module to optimize the model conversion parameters.

Citation Information

Patent Citations

  • Method for detecting human postures

    CN111291593A

  • Lightweight improved target detection method and device based on Rockchip micro platform

    CN113065555A