A method and device for detecting human key points
By detecting the two-dimensional position information of the human body key points from the current and previous N-frame cabin images, and combining the timing information and parameter information of the image acquisition equipment, the three-dimensional position information of the human body key points is determined, which solves the problem of detecting the three-dimensional position information of the human body key points in the cabin scene, and achieves high-precision human body posture and behavior monitoring.
Patent Information
- Application Number
- CN202010892894.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-08-31
AI Technical Summary
The prior art is difficult to accurately detect three-dimensional position information of key points in human bodies in cabin scenes.
By detecting the two-dimensional position information of the human body key points from the current and previous N-frame cabin images, and combining the timing information and parameter information of the image acquisition device, the three-dimensional position information of the human body key points is determined.
The accurate determination of the three-dimensional position information of the human body key points in the target scene is achieved, and the accuracy of human body posture and behavior monitoring is improved.
Smart Images

Figure CN114201985B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular, to a method and device for detecting human key points. Background Art
[0002] Human key point detection is a classic task in computer vision, generally referring to that a computer can detect each human key point of a human body, such as each key point of the head, each key point of the shoulders, each key point of the arms, each key point of the hands, and each key point of the body and legs. Detecting human key points can provide a technical basis for scenarios such as monitoring and human-computer interaction.
[0003] In a related human key point detection process, generally, a two-dimensional image is used as an input, and the two-dimensional coordinates of the human key points of each human body included in the two-dimensional image are identified from the two-dimensional image. With the improvement of the detection accuracy requirements in various scenarios, in some scenarios of human key point detection of a human body, there is a need to determine the three-dimensional position information of the human key points of the human body for more accurate detection of human postures and behaviors. For example, in a cabin scenario, it is necessary to obtain the three-dimensional position information of the human key points of the human body in order to monitor the posture and / or behavior of the human body.
[0004] Then, how to provide a method for detecting human key points to obtain the three-dimensional position information of human key points has become an urgent problem to be solved. Summary of the Invention
[0005] The present invention provides a method and device for detecting human key points to determine the three-dimensional position information of the human key points of a human body in a target scenario. The specific technical solutions are as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for detecting human key points, the method including:
[0007] Detect the two-dimensional position information of the human key points of each human body from the obtained current cabin image, where the current cabin image is an image collected by an image acquisition device for a target cabin;
[0008] Obtain the two-dimensional position information of the human key points of each human body detected from each of the previous N cabin images, where the previous N cabin images are the previous N images of the current cabin image collected by the image acquisition device, and N is a positive integer;
[0009] Based on the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, the timing information between the current cabin image and the previous N cabin images, and the parameter information of the image acquisition device, determine the three-dimensional position information of the human key points corresponding to each human body in the cabin coordinate system corresponding to the target cabin.
[0010] Optionally, the step of detecting the two-dimensional position information of the human key points of each human body from the obtained current cabin image includes:
[0011] From the obtained current cabin image, determine the area where each human body is located and intercept it to obtain the human area image corresponding to each human body;
[0012] Based on the human area image corresponding to each human body and the target human key point detection model, determine the heat map corresponding to each human key point in the human area image corresponding to each human body, where the target human key point detection model is: a model trained based on multiple groups of sample human body image pairs and their corresponding calibration information, and each sample human body image pair is a human body image corresponding to the same human body with different translation values, and the calibration information includes the calibration translation value between the two images in the corresponding sample human body image pair;
[0013] Based on the heat map corresponding to each human key point in the human area image corresponding to each human body, determine the two-dimensional position information of each human key point in the human area image corresponding to each human body.
[0014] Optionally, before the step of determining the two-dimensional position information of the human key points of each human body in the human area image corresponding to each human body based on the human area image corresponding to each human body and the target human key point detection model, the method further includes:
[0015] The process of training the target human key point detection model, and the process includes:
[0016] Obtain an initial human key point detection model;
[0017] Obtain multiple groups of sample human body image pairs and their corresponding calibration information, where the calibration information includes the calibration position information corresponding to each human key point in each image of the corresponding sample human body image pair and the calibration translation value between the two images in the corresponding sample human body image pair;
[0018] Using the multiple groups of sample human body image pairs and the calibration position information corresponding to each human body key point in each image of the sample human body image pairs in the calibration information and the calibration translation value between the two images in the corresponding sample human body image pairs, train the initial human body key point detection model until the initial human body key point detection model reaches a preset convergence condition, and determine the target human body key point detection model.
[0019] Optionally, the step of using the multiple groups of sample human body image pairs and the calibration position information corresponding to each human body key point in each image of the sample human body image pairs in the calibration information and the calibration translation value between the two images in the corresponding sample human body image pairs to train the initial human body key point detection model until the initial human body key point detection model reaches a preset convergence condition and determine the target human body key point detection model includes:
[0020] Batch the multiple groups of sample human body image pairs to obtain different batches of sample human body image pairs;
[0021] For each group of sample human body image pairs in each batch, input each image in the group of sample human body image pairs in the batch into the initial human body key point detection model to obtain a heat map corresponding to each human body key point in each image of the group of sample human body image pairs;
[0022] Based on the heat maps corresponding to the human body key points in each image of the group of sample human body image pairs in the batch, determine the predicted position information of the human body key points in each image of the group of sample human body image pairs, and the predicted translation value between the human body key points with corresponding relationships in the two images of the group of sample human body image pairs;
[0023] Based on the predicted position information and calibration position information of the human body key points in each image of all the sample human body image pairs in the batch, and the predicted translation value and calibration translation value between the human body key points with corresponding relationships in the two images of all the sample human body image pairs in the batch, determine the current loss value corresponding to the initial human body key point detection model;
[0024] Judge whether the current loss value corresponding to the initial human body key point detection model exceeds a preset loss threshold;
[0025] If it is judged that the current loss value corresponding to the initial human body key point detection model exceeds the preset loss threshold, adjust the model parameters of the initial human body key point detection model, and return to execute the step of inputting each image in the group of sample human body image pairs in the batch into the initial human body key point detection model for each group of sample human body image pairs in different batches to obtain a heat map corresponding to each human body key point in each image of the group of sample human body image pairs;
[0026] If it is determined that the current loss value corresponding to the initial human key point detection model does not exceed the preset loss threshold, it is determined that the initial human key point detection model meets the preset convergence condition, and the target human key point detection model is determined.
[0027] Optionally, the step of determining the three-dimensional position information of the human key points corresponding to each human in the coordinate system of the target cabin based on the two-dimensional position information of the human key points of each human in the current cabin image, the two-dimensional position information of the human key points of each human in the previous N cabin images, the time sequence information between the current cabin image and the previous N cabin images, and the parameter information of the image acquisition device includes:
[0028] Based on the two-dimensional position information of the human key points of each human in the current cabin image, the two-dimensional position information of the human key points of each human in the previous N cabin images, the pixel area of the human area image corresponding to each human, the detection visibility information corresponding to each human key point, and the time sequence information between the current cabin image and the previous N cabin images, the similarity value between each human key point in each adjacent two images of the current cabin image and the previous N cabin images is determined in sequence, where the visibility information corresponding to each human key point is information indicating whether the corresponding human key point is occluded;
[0029] Based on the similarity value between each human key point in each adjacent two images of the current cabin image and the previous N cabin images, the human key points of the same human in the current cabin image and the previous N cabin images are determined;
[0030] Based on the two-dimensional position information of the human key points of each human in the current cabin image and the previous N cabin images, the target three-dimensional human key point position prediction model, the parameter information of the image acquisition device, and the preset human activity range constraint condition, the three-dimensional position information of the human key points corresponding to each human is determined, where the target three-dimensional human key point position prediction model is a model trained based on the two-dimensional position information of the sample human key points of the sample human sorted in time sequence and their corresponding calibrated three-dimensional position information.
[0031] Optionally, the step of determining the three-dimensional position information of the human key points corresponding to each human based on the two-dimensional position information of the human key points of each human in the current cabin image and the previous N cabin images, the target three-dimensional human key point position prediction model, the parameter information of the image acquisition device, and the preset human activity range constraint condition includes:
[0032] Based on the two-dimensional position information of each human body key point of each human body in the current cabin image and the previous N cabin images, and the target three-dimensional human body key point position prediction model, determine the three-dimensional position information of each human body key point of each human body in the human body coordinate system corresponding to the human body;
[0033] For each human body, based on the three-dimensional position information of each human body key point of the human body in the human body coordinate system corresponding to the human body, the parameter information of the image acquisition device, and the preset human body activity range constraint condition corresponding to the human body, determine the three-dimensional position information of each human body key point of the human body in the cabin coordinate system corresponding to the target cabin.
[0034] In a second aspect, an embodiment of the present invention provides a human body key point detection device, and the device includes:
[0035] A detection module, configured to detect the two-dimensional position information of the human body key points of each human body from the obtained current cabin image, where the current cabin image is an image collected by an image acquisition device for the target cabin;
[0036] An obtaining module, configured to obtain the two-dimensional position information of the human body key points of each human body detected from each of the previous N cabin images, where the previous N cabin images are: the previous N images of the current cabin image collected by the image acquisition device, and N is a positive integer;
[0037] A first determination module, configured to determine the three-dimensional position information of the human body key points of each human body in the cabin coordinate system corresponding to the target cabin based on the two-dimensional position information of the human body key points of each human body in the current cabin image, the two-dimensional position information of the human body key points of each human body in the previous N cabin images, the timing information between the current cabin image and the previous N cabin images, and the parameter information of the image acquisition device.
[0038] Optionally, the detection module 310 includes:
[0039] A first determination unit, configured to determine and intercept the area where each human body is located from the obtained current cabin image to obtain a human body area image corresponding to each human body;
[0040] A second determination unit, configured to determine heat maps corresponding to each human body key point in the human body region image corresponding to each human body based on the human body region image corresponding to each human body and a target human body key point detection model, where the target human body key point detection model is a model trained based on multiple groups of sample human body image pairs and their corresponding calibration information, and each sample human body image pair is a human body image corresponding to the same human body with different translation values, and the calibration information includes the calibration translation value between the two images in the corresponding sample human body image pair;
[0041] A third determination unit, configured to determine two-dimensional position information of each human body key point in the human body region image corresponding to each human body based on the heat maps corresponding to each human body key point in the human body region image corresponding to each human body.
[0042] Optionally, the apparatus further includes:
[0043] A model training module, configured to train the target human body key point detection model before determining the two-dimensional position information of each human body key point in the human body region image corresponding to each human body based on the human body region image corresponding to each human body and the target human body key point detection model. The model training module includes:
[0044] A first acquisition unit, configured to acquire an initial human body key point detection model;
[0045] A second acquisition unit, configured to acquire multiple groups of sample human body image pairs and their corresponding calibration information, where the calibration information includes the calibration position information corresponding to each human body key point in each image of the corresponding sample human body image pair and the calibration translation value between the two images in the corresponding sample human body image pair;
[0046] A training unit, configured to train the initial human body key point detection model by using the multiple groups of sample human body image pairs and the calibration position information corresponding to each human body key point in each image of the corresponding sample human body image pair and the calibration translation value between the two images in the corresponding sample human body image pair until the initial human body key point detection model reaches a preset convergence condition, and determine the target human body key point detection model.
[0047] Optionally, the training unit is specifically configured to
[0048] Batch the multiple groups of sample human body image pairs to obtain different batches of sample human body image pairs;
[0049] For each group of sample human body image pairs in each batch, input each image in the group of sample human body image pairs in the batch into the initial human body key point detection model to obtain heat maps corresponding to each human body key point in each image of the group of sample human body image pairs in the batch;
[0050] Based on the heat maps corresponding to the human key points in each image of the pair of human body images of this group within this batch, determine the predicted position information of the human key points in each image of the pair of human body images of this group, and the predicted translation value between the human key points with corresponding relationships in the two images of the pair of human body images of this group;
[0051] Based on the predicted position information and calibrated position information of the human key points in each image of all pairs of human body images within this batch, and the predicted translation value and calibrated translation value between the human key points with corresponding relationships in the two images of all pairs of human body images within this batch, determine the current loss value corresponding to the initial human key point detection model;
[0052] Determine whether the current loss value corresponding to the initial human key point detection model exceeds a preset loss threshold;
[0053] If it is determined that the current loss value corresponding to the initial human key point detection model exceeds the preset loss threshold, adjust the model parameters of the initial human key point detection model, and return to execute the step of inputting each image in the pair of human body images of this group within this batch into the initial human key point detection model for each pair of human body images of each group within different batches to obtain the heat maps corresponding to each human key point in each image of the pair of human body images of this group;
[0054] If it is determined that the current loss value corresponding to the initial human key point detection model does not exceed the preset loss threshold, determine that the initial human key point detection model reaches the preset convergence condition, and determine the target human key point detection model.
[0055] Optionally, the first determination module includes:
[0056] A fourth determination unit, configured to sequentially determine the similarity values between each human key point between every two adjacent images among the current cabin image and the previous N cabin images based on the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, the pixel area of the human body region image corresponding to each human body, the detection visibility information corresponding to each human key point, and the timing information between the current cabin image and the previous N cabin images, where the visibility information corresponding to each human key point is information indicating whether the corresponding human key point is occluded;
[0057] A fifth determination unit, configured to determine the human key points of the same human body in the current cabin image and the previous N cabin images based on the similarity values between each human key point between every two adjacent images among the current cabin image and the previous N cabin images;
[0058] A sixth determination unit, configured to determine three-dimensional position information corresponding to each human body key point of each human body based on two-dimensional position information of each human body key point of each human body in the current cabin image and the previous N cabin images, a target three-dimensional human body key point position prediction model, parameter information of the image acquisition device, and a preset human body activity range constraint condition, where the target three-dimensional human body key point position prediction model is a model trained based on two-dimensional position information of various human body key points of a sample human body sorted in time sequence and their corresponding calibrated three-dimensional position information.
[0059] Optionally, the sixth determination unit is specifically configured to
[0060] Determine three-dimensional position information of each human body key point of each human body in the human body coordinate system corresponding to the human body based on the two-dimensional position information of each human body key point of each human body in the current cabin image and the previous N cabin images and the target three-dimensional human body key point position prediction model;
[0061] For each human body, determine three-dimensional position information of each human body key point of the human body corresponding to the cabin coordinate system corresponding to the target cabin based on the three-dimensional position information of each human body key point of the human body in the human body coordinate system corresponding to the human body, the parameter information of the image acquisition device, and the preset human body activity range constraint condition corresponding to the human body.
[0062] As can be seen from the above, a method and device for detecting human body key points provided by an embodiment of the present invention detect two-dimensional position information of human body key points of each human body from the obtained current cabin image, where the current cabin image is an image collected by an image acquisition device for a target cabin; obtain two-dimensional position information of human body key points of each human body detected from each of the previous N cabin images, where the previous N cabin images are the previous N images of the current cabin image collected by the image acquisition device, and N is a positive integer; determine three-dimensional position information of human body key points of each human body corresponding to the cabin coordinate system corresponding to the target cabin based on the two-dimensional position information of human body key points of each human body in the current cabin image, the two-dimensional position information of human body key points of each human body in the previous N cabin images, the time sequence information between the current cabin image and the previous N cabin images, and the parameter information of the image acquisition device.
[0063] By applying the embodiments of the present invention, based on the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, and the timing information between the current cabin image and the previous N cabin images, the three-dimensional position information of the human key points of each human body can be determined. Furthermore, in combination with the parameter information of the image acquisition device, the three-dimensional position information of the human key points corresponding to each human body in the cabin coordinate system corresponding to the target cabin can be determined, so as to realize the determination of the three-dimensional position information of the human key points of the human body in the target scene, that is, the target cabin. Of course, any product or method for implementing the present invention does not necessarily need to achieve all the above-mentioned advantages at the same time.
[0064] The innovative points of the embodiments of the present invention include:
[0065] 1. Based on the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, and the timing information between the current cabin image and the previous N cabin images, the three-dimensional position information of the human key points of each human body can be determined. Furthermore, in combination with the parameter information of the image acquisition device, the three-dimensional position information of the human key points corresponding to each human body in the cabin coordinate system corresponding to the target cabin can be determined, so as to realize the determination of the three-dimensional position information of the human key points of the human body in the target scene, that is, the target cabin.
[0066] 2. The target human key point detection model is trained by each pair of sample human body images and the calibration information corresponding to each pair of sample human body images, which includes the calibration translation value between the two images in the corresponding pair of sample human body images, so that the target human key point detection model can overcome the problem that the position of the detected human key points deviates due to translational jitter. Furthermore, through the target human key point detection model, the two-dimensional position information of the human key points in each human body region image corresponding to each human body with higher accuracy can be obtained, providing a basis for the accurate determination of the three-dimensional position information of the human key points in the subsequent stage.
[0067] 3. In the process of training the target human key point detection model, the current loss value corresponding to the initial human key point detection model is determined through the predicted position information and the calibrated position information of the human key points in each image of the pair of sample human body images, and the predicted translation value and the calibrated translation value between the human key points with corresponding relationships in the two images of the pair of sample human body images. Furthermore, based on the current loss value, the model parameters of the initial human key point detection model are adjusted. Furthermore, until the initial human key point detection model reaches the preset convergence condition, the target human key point detection model is obtained, so that the target human key point detection model can overcome the problem that the position of the detected human key points deviates due to translational jitter, and ensure the accuracy of the detection result of the target human key point detection model.
[0068] 4. Determine the two-dimensional position information of the human key points of each human from the two-dimensional position information of the human key points of the human in the current cabin image and the previous N cabin images. Based on the two-dimensional position information of the human key points of each human, use the target three-dimensional human key point position prediction model to determine the three-dimensional position information of the human key points of each human in the human coordinate system where they are located. Combine the parameter information of the image acquisition device and the preset human activity range constraint conditions to determine the three-dimensional position information of the human key points of each human in the cabin coordinate system corresponding to the target cabin, so as to realize the determination of the real physical position information of the human key points of each human in the target cabin, and provide more meaningful position information for subsequent tasks. Brief Description of the Drawings
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0070] Figure 1 It is a schematic flowchart of a method for detecting human key points provided by an embodiment of the present invention;
[0071] Figure 2 It is a schematic flowchart of the training process of a target human key point detection model provided by an embodiment of the present invention;
[0072] Figure 3 It is a schematic structural diagram of a device for detecting human key points provided by an embodiment of the present invention. Detailed Embodiments
[0073] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0074] It should be noted that the terms "include" and "have" and any variations thereof in the embodiments of the present invention and the drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0075] The present invention provides a method and device for detecting human key points to determine the three-dimensional position information of human key points of a human body in a target scene. The embodiments of the present invention will be described in detail below.
[0076] Figure 1 FIG. 4 is a schematic flowchart of a method for detecting human key points provided by an embodiment of the present invention. The method may include the following steps:
[0077] S101: Detect the two-dimensional position information of the human key points of each human body from the obtained current cabin image.
[0078] Wherein, the current cabin image is an image collected by an image acquisition device for the target cabin.
[0079] The method for detecting human key points provided by the embodiments of the present invention can be applied to any electronic device with computing power, and the electronic device can be a terminal or a server. In one implementation, the functional software for implementing the method for detecting human key points can exist in the form of a separate client software, or can exist in the form of a plugin of the current relevant client software, which is all acceptable.
[0080] Wherein, the target cabin can be the cabin of a target vehicle, and correspondingly, the cabin coordinate system corresponding to the target cabin mentioned later can be the vehicle body coordinate system of the vehicle where the target cabin is located; or, the target cabin can be the cabin of a target ship, and correspondingly, the cabin coordinate system corresponding to the target cabin mentioned later can be the hull coordinate system of the target ship. When the target cabin is the cabin of a target vehicle, in one implementation, the electronic device can be an in-vehicle device, and the electronic device is installed in the cabin of the target vehicle. In another implementation, the electronic device can be a non-vehicle-mounted device.
[0081] The electronic device is connected to the image acquisition device and can obtain the image collected by the image acquisition device for the target cabin. The image acquisition device can be an infrared imaging camera. In one case, in order to monitor the human body in the target cabin, the image acquisition device can be arranged in front of the target cabin to photograph the entire cabin. For example: when the target cabin is the cabin of a target vehicle, the image acquisition device can be arranged at the rearview mirror in the cabin of the target vehicle.
[0082] The image acquisition device can collect images of the target cabin in real time, obtain the cabin images, and send them to the electronic device. The electronic device obtains the cabin images collected by the image acquisition device in real time for the target cabin as the current cabin images, and uses a preset human key point detection algorithm to detect the current cabin images to detect the two-dimensional position information of the human key points of each human body from the current cabin images. In one implementation, the preset human key point detection algorithm can be any type of detection algorithm in the related art. In another implementation, the preset human key point detection algorithm can be the target human key point detection model obtained by pre-training provided in the embodiments of the present invention. The training method of the target human key point detection model will be introduced later for the sake of clear layout.
[0083] In one case, considering that people are generally in a sitting position inside the target cabin, the human key points may include the key points of the upper body and head of the human body. In one case, the human key points may include, but are not limited to: nose key point, left shoulder key point, right shoulder key point, left elbow key point, right elbow key point, left wrist key point, right wrist key point, left hip key point, and right hip key point.
[0084] S102: Obtain the two-dimensional position information of the human key points of each human body detected from each of the first N cabin images.
[0085] Among them, the first N cabin images are the first N images of the current cabin images collected by the image acquisition device, and N is a positive integer.
[0086] In this step, the electronic device can obtain the first N cabin images of the current cabin images collected by the image acquisition device for the target cabin, and can obtain the two-dimensional position information of the human key points of each human body detected from each of the first N cabin images. In one implementation manner, the two-dimensional position information of the human key points of each human body detected from each of the first N cabin images may be: the two-dimensional position information of the human key points of each human body detected by the electronic device based on the preset human key point detection algorithm for each of the first N cabin images.
[0087] N is a positive integer, and its specific value can be set by the staff according to actual needs or can be the default setting of the electronic device.
[0088] The detected human key points can correspond to semantic information. When the human key points include nose key point, left shoulder key point, right shoulder key point, left elbow key point, right elbow key point, left wrist key point, right wrist key point, left hip key point and right hip key point, the semantic information corresponding to the human key points can include: nose key point, left shoulder key point, right shoulder key point, left elbow key point, right elbow key point, left and right wrist key points, right wrist key point, left hip key point and right hip key point.
[0089] S103: Based on the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, the timing information between the current cabin image and the previous N cabin images, and the parameter information of the image acquisition device, determine the three-dimensional position information of the human key points corresponding to each human body in the cabin coordinate system corresponding to the target cabin.
[0090] In this step, the electronic device can determine the two-dimensional position information of each human key point of each human body in each image based on the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, and the timing information between the current cabin image and the previous N cabin images, where the image includes the current cabin image and the previous N cabin images. Furthermore, based on the two-dimensional position information of each human key point of each human body in each image sorted according to the timing information between the images, determine the three-dimensional position information of each human key point corresponding to each human body in the human coordinate system of the human body where it is located; furthermore, use the internal and external parameters of the image acquisition device and the three-dimensional position information of each human key point corresponding to each human body in the human coordinate system of the human body where it is located to determine the three-dimensional position information of each human key point of the human body in the cabin coordinate system corresponding to the target cabin, so as to determine the true physical position information of each human key point of each human body in the target cabin.
[0091] By applying the embodiments of the present invention, the three-dimensional position information of the human key points of each human body can be determined based on the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, and the timing information between the current cabin image and the previous N cabin images. Furthermore, in combination with the parameter information of the image acquisition device, the three-dimensional position information of the human key points corresponding to each human body in the cabin coordinate system corresponding to the target cabin can be determined, so as to determine the three-dimensional position information of the human key points of the human body in the target scenario, that is, the target cabin.
[0092] In another embodiment of the present invention, S101 may include the following steps 011-013:
[0093] 011: From the obtained current cabin image, determine the regions where each human body is located and crop them to obtain the human body region images corresponding to each human body.
[0094] 012: Based on the human body region images corresponding to each human body and the target human body key point detection model, determine the heat maps corresponding to each human body key point in the human body region images corresponding to each human body.
[0095] Among them, the target human body key point detection model is: a model trained based on multiple groups of sample human body image pairs and their corresponding calibration information. Each sample human body image pair is a human body image corresponding to the same human body with different translation values, and the calibration information includes the calibration translation value between the two images in the corresponding sample human body image pair.
[0096] 013: Based on the heat maps corresponding to each human body key point in the human body region images corresponding to each human body, determine the two-dimensional position information of each human body key point in the human body region images corresponding to each human body.
[0097] In this implementation, the electronic device can adopt a Top-down detection scheme to detect the two-dimensional position information of each human body key point in the current cabin image. The electronic device can first detect the regions where each human body is located from the current cabin image based on a preset human body detection algorithm, and crop the detected regions of each human body from the current cabin image to obtain the human body region images corresponding to each human body. Among them, the preset human body detection algorithm can be any human body detection algorithm in the related art that can detect human bodies from images, which will not be elaborated here.
[0098] The electronic device inputs the cropped human body region images corresponding to each human body into the target human body key point detection model trained based on multiple groups of sample human body image pairs and their corresponding calibration information. The target human body key point detection model detects the human body region images corresponding to each human body to determine the heat maps corresponding to each human body key point in the human body region images corresponding to each human body, that is, the target human body key point detection model outputs the heat maps corresponding to each human body key point in the human body region images corresponding to each human body.
[0099] Each human body key point corresponds to a heat map. The response value of the pixel points in the heat map can represent the possibility that the pixel point is the corresponding human body key point. The larger the response value of the pixel points in the heat map, the greater the possibility that the pixel point is the corresponding human body key point. Each heat map can correspond to the semantic information of a human body key point to distinguish the heat maps corresponding to different human body key points.
[0100] Subsequently, for the heat maps corresponding to each human key point in the human body area images corresponding to each human body, the electronic device can determine the two-dimensional position information of the human key point corresponding to the heat map based on the response values of the pixels in the heat map. One implementation can determine the two-dimensional position information of the human key point corresponding to the heat map based on the position of the pixel with the highest response value in the heat map. Another implementation can determine the two-dimensional position information of the human key point corresponding to the heat map based on the intermediate position of the positions of the top M pixels with the highest response values in the heat map. All of these are possible.
[0101] In another embodiment of the present invention, the present invention provides a training process for a target human key point detection model. Before step 012, the method may further include the following steps:
[0102] The process of training the target human key point detection model is as Figure 2 shown, and the process may include:
[0103] S201: Obtain an initial human key point detection model.
[0104] S202: Obtain multiple groups of sample human body image pairs and their corresponding calibration information.
[0105] Among them, the calibration information includes the calibration position information corresponding to each human key point in each image of the corresponding sample human body image pair and the calibration translation value between the two images in the corresponding sample human body image pair.
[0106] S203: Use multiple groups of sample human body image pairs, the calibration position information corresponding to each human key point in each image of the corresponding sample human body image pair in the calibration information, and the calibration translation value between the two images in the corresponding sample human body image pair to train the initial human key point detection model until the initial human key point detection model reaches a preset convergence condition, and determine the target human key point detection model.
[0107] Among them, the initial human key point detection model can be a neural network model based on deep learning.
[0108] In one implementation, the process of obtaining multiple groups of sample human body image pairs and their corresponding calibration information may be as follows: Obtain sample images collected for the target cabin or other cabins. The annotator manually or through a specific annotation tool marks the area where the human body is located in each sample image. Among them, the area where the human body is located in the sample image is marked by a rectangular box. The annotator manually or through a specific annotation tool marks a new rectangular box corresponding to each area where the human body is located in each sample image again. The new rectangular box is a box obtained by randomly translating the rectangular box of the marked area where the human body is located. For the convenience of description, the rectangular box of the marked area where the human body is located is hereinafter referred to as the first rectangular box, and the new rectangular box obtained by randomly translating the rectangular box of the marked area where the human body is located is hereinafter referred to as the second rectangular box.
[0109] For each sample image, based on the marked first rectangular box and second rectangular box, extract the corresponding regional images from the sample image. The regional images corresponding to the first rectangular box and the second rectangular box with a corresponding relationship are used as a group of sample human body image pairs. At the same time, record the translation value between the first rectangular box and the second rectangular box with a corresponding relationship as the calibration translation value. For each sample human body image in each group of sample human body image pairs, mark each human key point of the human body contained therein to obtain the calibration position information corresponding to each human key point, so as to obtain the calibration information corresponding to each group of sample human body image pairs. Among them, the calibration position information corresponding to each human key point in the calibration information is the position information in its corresponding sample human body image.
[0110] Based on the first rectangular box and the second rectangular box marked in the sample image, when extracting the corresponding regional images from the sample image, that is, the process of the sample human body image, record the position information of the regional images corresponding to the first rectangular box and the second rectangular box in the sample image. Furthermore, through the position information of each human key point in its corresponding sample human body image in the calibration information, and the position information of the sample human body images corresponding to the first rectangular box and the second rectangular box in the sample image, the position information of each human key point in the sample image can be determined.
[0111] Subsequently, the electronic device can input multiple groups of sample human body image pairs into the initial human key point detection model in batches, so as to train the initial human key point detection model through multiple groups of sample human body image pairs, and the calibration position information corresponding to each human key point in each image of each group of sample human body image pairs and the calibration translation value between the two images in each group of sample human body image pairs in the calibration information corresponding to each group of sample human body image pairs until the initial human key point detection model reaches the preset convergence condition, and determine the target human key point detection model.
[0112] In one implementation, the calibration information corresponding to each sample human body image pair further includes the calibration visibility information corresponding to each human body key point in each image of the sample human body image pair, where the calibration visibility information is information indicating whether the corresponding human body key point is occluded. In one case, the value range of the calibration visibility information can be [0, 1]. The larger the value of the calibration visibility information, the more visible the corresponding human body key point is, that is, the less severe the occluded situation is.
[0113] Correspondingly, in the case where the calibration information corresponding to each sample human body image pair further includes the calibration visibility information corresponding to each human body key point in each image of the sample human body image pair, when the trained target human body key point detection model outputs the heat maps corresponding to each human body key point in the output image, it can also output the visibility information corresponding to each human body key point.
[0114] By using the calibration translation values between the two images in the sample human body image pair included in each sample human body image pair and its corresponding calibration information, a target human body key point detection model that can overcome the problem of deviation in the predicted positions of human body key points caused by translational jitter can be trained, which can improve the accuracy of the detection results of the target human body key point detection model to a certain extent.
[0115] In another embodiment of the present invention, S203 may include the following steps 021-027:
[0116] 021: Batch the multiple groups of sample human body image pairs to obtain different batches of sample human body image pairs.
[0117] 022: For each group of sample human body image pairs in each batch, input each image in the group of sample human body image pairs in the batch into the initial human body key point detection model to obtain the heat maps corresponding to each human body key point in each image of the group of sample human body image pairs.
[0118] 023: Based on the heat maps corresponding to the human body key points in each image of the group of sample human body image pairs in the batch, determine the predicted position information of the human body key points in each image of the group of sample human body image pairs, and the predicted translation value between the human body key points with corresponding relationships in the two images of the group of sample human body image pairs. Among them, the human body key points with corresponding relationships are the human body key points corresponding to the same semantic information.
[0119] 024: Based on the predicted position information and calibration position information of the human body key points in each image of all the sample human body image pairs in the batch, and the predicted translation value and calibration translation value between the human body key points with corresponding relationships in the two images of all the sample human body image pairs in the batch, determine the current loss value corresponding to the initial human body key point detection model.
[0120] 025: Determine whether the current loss value corresponding to the initial human key point detection model exceeds a preset loss threshold.
[0121] 026: If it is determined that the current loss value corresponding to the initial human key point detection model exceeds the preset loss threshold, adjust the model parameters of the initial human key point detection model, and return to execute 022;
[0122] 027: If it is determined that the current loss value corresponding to the initial human key point detection model does not exceed the preset loss threshold, determine that the initial human key point detection model reaches the preset convergence condition, and determine the target human key point detection model.
[0123] In this implementation manner, multiple groups of sample human image pairs can be first divided into batches to obtain multiple groups of sample human image pairs in each batch. Furthermore, for each group of sample human image pairs in each batch, input the group of sample human image pairs into the initial human key point detection model to obtain the heat maps corresponding to each human key point in each image of the group of sample human image pairs. Furthermore, for each group of sample human image pairs in this batch, process the heat maps corresponding to the human key points in each image of the group of sample human image pairs to obtain the predicted position information of the human key points in each image of the group of sample human image pairs. And use the predicted position information of the corresponding same human key points in the two images of the group of sample human image pairs to calculate the difference between the predicted position information of the corresponding same human key points in the two images, and obtain the deviation between the predicted position information of the corresponding same human key points in the two images of the group of sample human image pairs, which is the predicted translation value between the human key points with corresponding relationships. Among them, the predicted position information of the corresponding same human key points is the predicted position information of the human key points with the same corresponding semantic information.
[0124] Among them, in one case, the process of processing the heat maps corresponding to the human key points in each image of the group of sample human image pairs to obtain the predicted position information of the human key points in each image of the group of sample human image pairs can be: use the soft-argmax operation to process the heat maps corresponding to the human key points in each image of the group of sample human image pairs to obtain the predicted position information of the human key points in each image of the group of sample human image pairs. Among them, the soft-argmax operation can be represented by the following formula:
[0125]
[0126] Among them, i is the index value of the i-th pixel point of the heat map corresponding to an image in the sample human image pair, that is, the position information in the heat map corresponding to an image in the sample human image pair, x iis the response value of the i-th pixel of the heat map corresponding to the image in the sample human image pair, that is, the specific value at the position information i in the heat map corresponding to one image in the sample human image pair. p is the index value of the p-th pixel of the heat map corresponding to one image in the sample human image pair, x p is the response value of the p-th pixel of the heat map corresponding to the image in the sample human image pair, represents the sum of the exponents corresponding to all pixels of the heat map. For example, the exponent corresponding to the p-th pixel is The value ranges of i and p are [1, Q], and Q is determined based on the size of the heat map.
[0127] After the heat map corresponding to each image in the sample human image pair undergoes the soft-argmax operation, it can be converted into the predicted position information of the human key points corresponding to the heat map represented in the form of two-dimensional coordinates. Furthermore, based on the predicted position information of the human key points corresponding to the heat map represented in the form of the converted two-dimensional coordinates, and the size ratio between the heat map and the image in the corresponding sample human image pair, the predicted position information of the human key points in the image corresponding to the heat map can be determined. This entire conversion process is differentiable and can be used during model training.
[0128] Furthermore, after obtaining the predicted position information of the human key points in the heat map corresponding to each image in all the sample human image pairs within a batch, that is, obtaining the predicted position information of the human key points in each image in all the sample human image pairs within a batch, and the predicted translation values between the human key points with corresponding relationships in the two images of the set of sample human image pairs, based on the predicted position information, calibration position information, and the first loss function of the human key points in each image in all the sample human image pairs within this batch, the first position loss value corresponding to this batch is determined. Among them, it can be: for each image within this batch, based on the predicted position information, calibration position information, and the first loss function of the human key points in this image, determine the first loss value corresponding to this image. Then, take the sum of the first loss values corresponding to all the images within this batch, or the average value of the sum of the first loss values corresponding to all the images within this batch, as the first position loss value corresponding to this batch.
[0129] Among them, the first position loss value can be called the heat map loss value, and the first loss function can be the MSE Loss function. The expression of the MSE Loss function can be:
[0130]
[0131] Among them, MSE represents the first position loss value, y jhwIt represents the predicted position information of the j-th human key point in the heat map corresponding to one of the sample human body image pairs. It represents the calibrated position information of the j-th human key point in the heat map corresponding to one of the sample human body image pairs. H and W are the size information of the heat map corresponding to this image in the sample human body image pair, and S is the number of human key points in the heat map corresponding to this image in the sample human body image pair.
[0132] Furthermore, the electronic device determines the predicted position information of the human key points in each image of each sample human body image pair in this batch based on the predicted position information of the human key points in the heat map corresponding to each image of each sample human body image pair in this batch; furthermore, based on the predicted position information of the human key points in each image of each sample human body image pair in this batch, and the position information of each image of each sample human body image pair in its corresponding sample image recorded in advance, it determines the predicted position information of the human key points in each image of each sample human body image pair in its corresponding sample image; and based on the calibrated position information of the human key points in each image of each sample human body image pair in this batch, and the position information of each image of each sample human body image pair in its corresponding sample image recorded in advance, it determines the calibrated position information of the human key points in each image of each sample human body image pair in its corresponding sample image; based on the predicted position information and calibrated position information of the human key points in each image of all sample human body image pairs in this batch in their corresponding sample images, and the second loss function, it determines the second position loss value corresponding to this batch.
[0133] Among them, the process of determining the second position loss value corresponding to this batch based on the predicted position information and calibrated position information of the human key points in each image of all sample human body image pairs in this batch in their corresponding sample images, and the second loss function, can be:
[0134] For each image in this batch, based on the predicted position information, calibrated position information and the second loss function of the human key points in this image in its corresponding sample image, it determines the second loss value corresponding to this image. Furthermore, the sum of the second loss values corresponding to all images in this batch, or the average value of the sum of the second loss values corresponding to all images in this batch, is used as the second position loss value corresponding to this batch.
[0135] The second position loss value can be key point loss. In one case, the second loss function can be the L1Loss function.
[0136] Moreover, based on the predicted translation values and the calibrated translation values between the human key points with corresponding relationships in the two images of all the sample human body image pairs within the batch, and the third loss function, the electronic device determines the stability loss value stable loss corresponding to the batch. Among them, the third loss function can be the L1 Loss function.
[0137] The process of determining the stability loss value corresponding to the batch can be as follows: for each image within the batch, based on the predicted translation values and the calibrated translation values between the human key points with corresponding relationships in the image and the third loss function, determine the third loss value corresponding to the image, and then use the sum of the third loss values corresponding to all the images within the batch, or the average value of the sum of the third loss values corresponding to all the images within the batch, as the stability loss value corresponding to the batch.
[0138] Subsequently, based on the first position loss value, the second position loss value, and the stability loss value corresponding to the batch described above, the electronic device determines the current loss value corresponding to the initial human key point detection model. For example: it can be the sum of the above three as the current loss value corresponding to the initial human key point detection model.
[0139] The electronic device compares the current loss value corresponding to the initial human key point detection model with a preset loss threshold to determine whether the current loss value corresponding to the initial human key point detection model exceeds the preset loss threshold; if it is determined that the current loss value corresponding to the initial human key point detection model exceeds the preset loss threshold, it is considered that the current initial human key point detection model has reached the preset convergence condition, and then based on a preset optimization function, such as the gradient descent method, adjust the model parameters of the initial human key point detection model, and return to execute 022; if it is determined that the current loss value corresponding to the initial human key point detection model does not exceed the preset loss threshold, it is determined that the initial human key point detection model has reached the preset convergence condition, and the target human key point detection model is determined.
[0140] In another embodiment of the present invention, S103 may include the following steps 031-033:
[0141] 031: Based on the two-dimensional position information of the human key points of each human in the current cabin image, the two-dimensional position information of the human key points of each human in the previous N frame cabin images, the pixel area of the human area image corresponding to each human, the detection visibility information corresponding to each human key point, and the temporal information between the current cabin image and the previous N frame cabin images, sequentially determine the similarity values between the human key points of each adjacent two images between the current cabin image and the previous N frame cabin images.
[0142] Among them, the visibility information corresponding to each human key point is information characterizing whether the corresponding human key point is occluded.
[0143] 032: Determine the human key points of the same human body in the current cabin image and the previous N cabin images based on the similarity values between the human key points of each adjacent pair of images among the current cabin image and the previous N cabin images.
[0144] 033: Determine the three-dimensional position information corresponding to each human key point of each human body based on the two-dimensional position information of each human key point of each human body in the current cabin image and the previous N cabin images, the target three-dimensional human key point position prediction model, the parameter information of the image acquisition device, and the preset human activity range constraint conditions.
[0145] Among them, the target three-dimensional human key point position prediction model is a model trained based on the two-dimensional position information of the sample human key points of the sample human body sorted in time sequence and their corresponding calibrated three-dimensional position information.
[0146] In this implementation manner, the electronic device needs to determine the human key points of the same human body from the human key points of each human body in different frame images, that is, the current cabin image and the previous N images. Furthermore, the three-dimensional position information corresponding to each human key point of each human body is determined by using the two-dimensional position information of the human key points of the same human body in different frame images and the time sequence information between different frame images.
[0147] In one case, the OKS measurement method can be used to calculate the similarity values between the human key points of each human body in different frame images, that is, the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, the pixel area of the human region image corresponding to each human body, the detection visibility information corresponding to each human key point, and the time sequence information between the current cabin image and the previous N cabin images are used to sequentially determine the similarity values between the human key points of each adjacent pair of images among the current cabin image and the previous N cabin images.
[0148] Specifically, it can be: For each human key point of the target human body in the Mth frame, perform the following operations to determine the human key points of the same human body as those of the target human body in the Mth frame in the (M + 1)th frame; where the target human body is each human body in the Mth frame, and both the Mth frame and the (M + 1)th frame belong to the current cabin image and the previous N cabin images, the Mth frame is any frame image in the current cabin image and the previous N cabin images, and the value range of M is [1, N + 1].
[0149] Using the two-dimensional position information of the human key points of the target human body in the M-th frame, the two-dimensional position information of the human key points corresponding to the same semantic information as the human key points of the target human body in each human body in the (M + 1)-th frame, the pixel area of the human body region image corresponding to the target human body, and the detection visibility information corresponding to the human key points of the target human body, determine the similarity value between the human key points of the target human body in the M-th frame and the human key points corresponding to the same semantic information as the human key points of the target human body in each human body in the (M + 1)-th frame; and so on, determine the similarity value between each human key point of the target human body in the M-th frame and the human key points corresponding to the same semantic information as each human key point of the target human body in each human body in the (M + 1)-th frame; furthermore, based on the similarity values between each human key point of the target human body in the M-th frame and the human key points corresponding to the same semantic information as each human key point of the target human body in each human body in the (M + 1)-th frame, determine whether there are human key points that are the same human body as each human key point of the target human body in the M-th frame among the human key points of all human bodies in the (M + 1)-th frame. When it is determined that there are, determine the human key points that are the same human body as each human key point of the target human body in the M-th frame, that is, determine the human body that is the same human body as the target human body in the M-th frame from all human bodies in the (M + 1)-th frame.
[0150] The formula (1) used in this process can be:
[0151]
[0152] Among them, OKS represents the similarity value between the j-th human key point of the target human body in the M-th frame of the current cabin image and the j-th human key point of the currently calculated human body in the (M + 1)-th frame among the first N frame cabin images. j is the serial number of the j-th human key point of the target human body in the M-th frame, and d j represents the Euclidean distance between the two-dimensional position information of the j-th human key point of the target human body in the M-th frame and the two-dimensional position information of the j-th human key point of the currently calculated human body in the (M + 1)-th frame; s represents the pixel area of the human body region image corresponding to the target human body in the M-th frame; K j is a hyperparameter, and v j represents the detection visibility information corresponding to the j-th human key point of the target human body in the M-th frame, which is the output of the above-mentioned target human key point detection model; a is a preset value, which is determined according to the calibration visibility information and detection visibility information corresponding to each human key point in the image of the sample human body image pair for training the target human key point detection model; among them, v j>a indicates that the j-th human key point of the target human body is visible, i.e., not occluded; the target human body is any human body in the M-th frame among the current cabin image and the first N cabin images, where when M takes values in [1, N], the M-th frame can refer to the n-th cabin image among the first N cabin images, and the value of n is in [1, N]; when M takes N + 1, the M-th frame can refer to the current cabin image.
[0153] For example, the current cabin image and the first N cabin images include images 1, 2, and 3 sorted in order of acquisition time. Among them, image 1 includes the human key points corresponding to human body A and the human key points corresponding to human body B; image 2 includes the human key points corresponding to human body C and the human key points corresponding to human body D; image 3 includes the human key points corresponding to human body E and the human key points corresponding to human body F. First, for each human key point of human body A in image 1, search for the human key points of the same human body from the human key points corresponding to C and the human key points corresponding to D in image 2. Taking the human key point 1 of human body A in image 1 as an example:
[0154] Using the above formula (1), the two-dimensional position information of the human key point 1 of human body A in image 1, and the two-dimensional position information of the human key point 1 of human body C in image 2, the pixel area of the human body region image corresponding to human body A, and the detection visibility information corresponding to the human key point 1 of human body A, determine the similarity value between the human key point 1 of human body A in image 1 and the human key point 1 of human body C in image 2.
[0155] And using the above formula (1), the two-dimensional position information of the human key point 1 of human body A in image 1, and the two-dimensional position information of the human key point 1 of human body D in image 2, the pixel area of the human body region image corresponding to human body A, and the detection visibility information corresponding to the human key point 1 of human body A, determine the similarity value between the human key point 1 of human body A in image 1 and the human key point 1 of human body D in image 2.
[0156] If the similarity value between the human key point 1 of human body A in image 1 and the human key point 1 of human body C in image 2 is greater than the similarity value between the human key point 1 of human body A in image 1 and the human key point 1 of human body D in image 2, it can be considered that the human key point 1 of human body C in image 2 is the most similar to the human key point 1 of human body A in image 1.
[0157] If each human key point of the human body C in Image 2 is the most similar to each human key point of the human body A in Image 1, or the number of human key points of the human body C in Image 2 that are the most similar to each human key point of the human body A in Image 1 exceeds a certain threshold, then it can be considered that the human body C in Image 2 and the human body A in Image 1 are the same human body, and the corresponding human key points of the human body C in Image 2 and the human key points of the human body A in Image 1 are the human key points of the same human body.
[0158] Furthermore, it is possible to determine the human body in Image 2 that is the same human body as the human body B in Image 1 from the human body C and the human body D in Image 2.
[0159] Subsequently, from the human body E and the human body F in Image 3, the human bodies that are the same human body as the human body C and the human body D in Image 2 are determined respectively. The process is the same as the process of determining the human body that is the same human body as the human body A in Image 1 from the human body in Image 2, which will not be elaborated here.
[0160] After the electronic device determines the human key points of the same human body in the current cabin image and the previous N cabin images, for each human key point of each human body, based on the two-dimensional position information of each human key point of the human body in the current cabin image and the previous N cabin images, the target three-dimensional human key point position prediction model, the parameter information of the image acquisition device, and the preset human activity range constraint condition, the three-dimensional position information corresponding to each human key point of the human body is determined.
[0161] Among them, the target three-dimensional human key point position prediction model is a model trained based on the two-dimensional position information of each sample human key point of the sample human body sorted in time series and its corresponding calibrated three-dimensional position information. The training process of the target three-dimensional human key point position prediction model can refer to the training process of the three-dimensional human key point position prediction model in the related technology, which will not be elaborated here.
[0162] In another embodiment of the present invention, the 033 may include the following steps 0331-0332:
[0163] 0331: Based on the two-dimensional position information of each human key point of each human body in the current cabin image and the previous N cabin images and the target three-dimensional human key point position prediction model, determine the three-dimensional position information corresponding to each human key point of each human body in the human coordinate system corresponding to the human body.
[0164] 0332: For each human body, based on the three-dimensional position information corresponding to each human key point of the human body in the human coordinate system corresponding to the human body, the parameter information of the image acquisition device, and the preset human activity range constraint condition corresponding to the human body, determine the three-dimensional position information corresponding to each human key point of the human body in the cabin coordinate system corresponding to the target cabin.
[0165] In this implementation manner, for each human body key point of each human body, the electronic device can determine the two-dimensional position information of each human body key point of the human body in the current cabin image and the previous N cabin images from the current cabin image and the previous N cabin images; based on the timing information between the current cabin image and the previous N cabin images, sort the two-dimensional position information of each human body key point of the human body in the current cabin image and the previous N cabin images in the order of the acquisition time of the current cabin image and the previous N cabin images, and obtain a two-dimensional position information sequence corresponding to each human body key point of the human body sorted in the order of the acquisition time of the current cabin image and the previous N cabin images; perform normalization processing on the two-dimensional position information sequence corresponding to each human body key point of the human body, and obtain a normalized two-dimensional position information sequence corresponding to each human body key point of the human body; input the normalized two-dimensional position information sequence corresponding to each human body key point of the human body into the target three-dimensional human body key point position prediction model, and obtain the three-dimensional position information of each human body key point of the human body in the human body coordinate system corresponding to the human body.
[0166] Among them, the human body coordinate system is a coordinate system with the center point of the connection line between the left and right hip joints of the corresponding human body, that is, the left hip joint key point and the right hip joint key point, as the coordinate origin.
[0167] Furthermore, in order to obtain the three-dimensional position information corresponding to each human body key point of each human body in the real physical world, it is necessary to convert the three-dimensional position information of each human body key point of the human body in the human body coordinate system corresponding to the human body to the cabin coordinate system corresponding to the target cabin, that is, to obtain the three-dimensional position information of each human body key point of the human body in the cabin coordinate system corresponding to the target cabin.
[0168] In this conversion process, the preset human body activity range constraint condition corresponding to the human body can be introduced, that is, the movement range of the coordinate origin of the human body coordinate system where the human body is located is constrained to be within a specific area of the seat plane where the human body is located.
[0169] Then, for each detected human body in the current cabin image, calculate the reprojection error. That is, for each detected human body in the current cabin image, based on the three-dimensional position information of each human body key point of the human body in the human body coordinate system where the human body is located, and the parameter information of the image acquisition device, project the space points corresponding to each human body key point of the human body in the human body coordinate system where the human body is located onto the current cabin image, and constrain the movement range of the coordinate origin of the human body coordinate system where the human body is located to the characteristic area of the seat plane where the human body is located. Calculate the distance error between the projection position information of the projection points of the space points corresponding to each human body key point of the human body in the corresponding human body coordinate system of the human body in the current cabin image and the two-dimensional position information of each human body key point of the human body in the current cabin image, that is, the reprojection error.
[0170] Among them, in the process of projecting the space points corresponding to each human body key point of the human body in the human body coordinate system where the human body is located onto the current cabin image based on the three-dimensional position information of each human body key point of the human body in the human body coordinate system where the human body is located and the parameter information of the image acquisition device, it is necessary to first convert the three-dimensional position information of each human body key point of the human body in the human body coordinate system where the human body is located to the device coordinate system of the image acquisition device based on the three-dimensional position information of each human body key point of the human body in the human body coordinate system where the human body is located, so as to obtain the three-dimensional position information of each human body key point of the human body in the device coordinate system of the image acquisition device; furthermore, based on the parameter information of the image acquisition device and the three-dimensional position information of each human body key point of the human body in the device coordinate system of the image acquisition device, project the space points corresponding to each human body key point of the human body in the human body coordinate system where the human body is located onto the current cabin image. Correspondingly, in the process of converting the three-dimensional position information of each human body key point of the human body in the human body coordinate system where the human body is located to the device coordinate system of the image acquisition device to obtain the three-dimensional position information of each human body key point of the human body in the device coordinate system of the image acquisition device, it is necessary to determine the position conversion relationship between the human body coordinate system where the human body is located and the device coordinate system of the image acquisition device, and the scaling scale between the human body coordinate system where the human body is located and the device coordinate system of the image acquisition device.
[0171] In view of this, through the above process of calculating the reprojection error, the position conversion relationship between the body coordinate system where the body is located and the device coordinate system of the image acquisition device can be determined, as well as the scaling factor between the body coordinate system where the body is located and the device coordinate system of the image acquisition device; if the reprojection error between the projection position information of the spatial points corresponding to the respective body key points of the body in the body coordinate system corresponding to the body in the current cabin image and the two-dimensional position information of the respective body key points of the body in the current cabin image is minimized, it can be considered that the optimal position conversion relationship between the body coordinate system where the body is located and the device coordinate system of the image acquisition device and the scaling factor between the body coordinate system where the body is located and the device coordinate system of the image acquisition device are obtained.
[0172] That is, in the above process of calculating the reprojection error for each detected body in the current cabin image, by adjusting the position conversion relationship between the body coordinate system where the body is located and the device coordinate system of the image acquisition device and the scaling factor between the body coordinate system where the body is located and the device coordinate system of the image acquisition device, the adjustment of the three-dimensional position information of the respective body key points of the body in the device coordinate system of the image acquisition device is realized. When the reprojection error between the projection position information of the spatial points corresponding to the respective body key points of the body in the body coordinate system corresponding to the body in the current cabin image and the two-dimensional position information of the respective body key points of the body in the current cabin image is minimized, it can be considered that the three-dimensional position information of the respective body key points of the body corresponding to the body coordinate system corresponding to the body, based on the position conversion relationship between the body coordinate system where the body is located and the device coordinate system of the image acquisition device and the scaling factor between the body coordinate system where the body is located and the device coordinate system of the image acquisition device, the converted three-dimensional position information in the device coordinate system of the image acquisition device is the optimal three-dimensional position information.
[0173] Furthermore, based on the optimal position conversion relationship between the body coordinate system where the body is located and the device coordinate system of the image acquisition device, the scaling factor between the body coordinate system where the body is located and the device coordinate system of the image acquisition device, and the three-dimensional position information of the respective body key points of the body corresponding to the body coordinate system where the body is located, the three-dimensional position information of the respective body key points of the body corresponding to the device coordinate system of the image acquisition device is determined.
[0174] One implementation is that it can be considered that the device coordinate system of the image acquisition device coincides with the cabin coordinate system corresponding to the target cabin. That is, the three-dimensional position information of the respective body key points of the body corresponding to the device coordinate system of the image acquisition device is directly used as the three-dimensional position information of the respective body key points of the body corresponding to the cabin coordinate system corresponding to the target cabin.
[0175] In another implementation, if it is considered that the device coordinate system of the image acquisition device does not coincide with the corresponding cabin coordinate system of the target cabin, the coordinate transformation relationship between the device coordinate system of the image acquisition device and the corresponding cabin coordinate system of the target cabin can be pre-calibrated. After determining the three-dimensional position information of each human body key point of the human body in the device coordinate system of the image acquisition device, based on the three-dimensional position information of each human body key point of the human body in the device coordinate system of the image acquisition device and the pre-calibrated coordinate transformation relationship between the device coordinate system of the image acquisition device and the corresponding cabin coordinate system of the target cabin, the three-dimensional position information of each human body key point of the human body in the device coordinate system of the image acquisition device is transformed to the corresponding cabin coordinate system of the target cabin, and the three-dimensional position information of each human body key point of the human body in the corresponding cabin coordinate system of the target cabin is obtained.
[0176] To facilitate the determination of the three-dimensional position information of each human body key point corresponding to the image acquisition device's device coordinate system, the process of calculating the reprojection error for each detected human body in the current cabin image can be transformed into a constrained optimization process. Among them, the optimization objective equation of this optimization process can be expressed as:
[0177]
[0178] Among them, P is the target value of the optimization objective equation, M1 represents the projection matrix of the image acquisition device, that is, the parameter information of the image acquisition device, represents projecting the spatial points corresponding to each human body key point of the human body that has been transformed to the device coordinate system of the image acquisition device onto the current cabin image; scale represents the scaling factor to be solved, which is used to unify the scale of the device coordinate system of the image acquisition device and the scale of the human body coordinate system where the human body is located; o represents the moving vector to be solved, that is, the position transformation relationship between the above-mentioned human body coordinate system where the human body is located and the device coordinate system of the image acquisition device; v j represents the three-dimensional position information of the j-th human body key point of the human body in its human body coordinate system where the human body is located; represents the two-dimensional position information of the j-th human body key point of the human body in the current cabin image, and K is the number of human body key points of the human body.
[0179] The above optimization process is constrained by the preset human body activity range constraint condition f(o) = a1x + b1y + c1z + d1 = 0 corresponding to the human body. Among them, a1, b1, c1, and d1 are the coefficients of the seat plane where the human body is located, and their values are preset values; the (x, y, z) substituted into the preset human body activity range constraint condition corresponding to the human body is: the three-dimensional position information of the coordinate origin of the human body coordinate system where the human body is located in the device coordinate system of the image acquisition device;
[0180] In another implementation, in order to further improve the accuracy of the true physical position information of each human body key point of each human body in the determined target cabin, the preset human body activity range constraint condition may further include a limitation on the activity range of the coordinate origin (x, y, z) of the human body coordinate system where the human body is located on a plane, that is, the limited value range of the coordinate origin (x, y, z) of the human body coordinate system where the human body is located can be set, and combined with the limited value range of the coordinate origin (x, y, z) of the human body coordinate system where the human body is located, and f(o) = a1x + b1y + c1z + d1 = 0, the above optimization process is jointly constrained.
[0181] Through the above solution process of the optimization process with constraints, the values of scale and o that minimize P are obtained. Furthermore, based on the obtained values of scale and o, and the three-dimensional position information of each human body key point of the human body in the human body coordinate system where the human body is located, the three-dimensional position information of each human body key point of the human body in the device coordinate system of the image acquisition device is determined, and based on the three-dimensional position information of each human body key point of the human body in the device coordinate system of the image acquisition device, the three-dimensional position information of each human body key point of the human body in the cabin coordinate system corresponding to the target cabin is determined.
[0182] Corresponding to the above method embodiment, an embodiment of the present invention provides a device, as Figure 3 shown, the device may include:
[0183] A detection module 310, configured to detect the two-dimensional position information of the human body key points of each human body from the obtained current cabin image, where the current cabin image is an image collected by an image acquisition device for the target cabin;
[0184] An obtaining module 320, configured to obtain the two-dimensional position information of the human body key points of each human body detected from each of the previous N cabin images, where the previous N cabin images are: the previous N images of the current cabin image collected by the image acquisition device, and N is a positive integer;
[0185] A first determination module 330, configured to determine the three-dimensional position information corresponding to the human body key points of each human body in the cabin coordinate system corresponding to the target cabin based on the two-dimensional position information of the human body key points of each human body in the current cabin image, the two-dimensional position information of the human body key points of each human body in the previous N cabin images, the timing information between the current cabin image and the previous N cabin images, and the parameter information of the image acquisition device.
[0186] By applying the embodiments of the present invention, based on the two-dimensional position information of the human key points of each human body in the current cabin image, the two-dimensional position information of the human key points of each human body in the previous N cabin images, and the timing information between the current cabin image and the previous N cabin images, the three-dimensional position information of the human key points of each human body can be determined. Furthermore, by combining the parameter information of the image acquisition device, the three-dimensional position information of the human key points corresponding to each human body in the cabin coordinate system corresponding to the target cabin can be determined, so as to realize the determination of the three-dimensional position information of the human key points of the human body in the target scene, that is, the target cabin.
[0187] In another embodiment of the present invention, the detection module 310 includes:
[0188] A first determination unit (not shown in the figure), configured to determine and intercept the area where each human body is located from the obtained current cabin image to obtain a human body area image corresponding to each human body;
[0189] A second determination unit (not shown in the figure), configured to determine a heat map corresponding to each human key point in the human body area image corresponding to each human body based on the human body area image corresponding to each human body and the target human key point detection model, where the target human key point detection model is a model trained based on multiple groups of sample human body image pairs and their corresponding calibration information. Each sample human body image pair is a human body image corresponding to the same human body with different translation values, and the calibration information includes the calibration translation value between the two images in the corresponding sample human body image pair;
[0190] A third determination unit (not shown in the figure), configured to determine the two-dimensional position information of each human key point in the human body area image corresponding to each human body based on the heat map corresponding to each human key point in the human body area image corresponding to each human body.
[0191] In another embodiment of the present invention, the device further includes:
[0192] A model training module (not shown in the figure), configured to train the target human key point detection model before determining the two-dimensional position information of each human key point in the human body area image corresponding to each human body based on the human body area image corresponding to each human body and the target human key point detection model. The model training module includes:
[0193] A first acquisition unit (not shown in the figure), configured to acquire an initial human key point detection model;
[0194] A second acquisition unit (not shown in the figure) is configured to acquire multiple groups of sample human body image pairs and their corresponding calibration information, where the calibration information includes the calibration position information corresponding to each human key point in each image of the corresponding sample human body image pair and the calibration translation value between the two images in the corresponding sample human body image pair;
[0195] A training unit (not shown in the figure) is configured to use the multiple groups of sample human body image pairs, the calibration position information corresponding to each human key point in each image of the corresponding sample human body image pair in the calibration information, and the calibration translation value between the two images in the corresponding sample human body image pair to train the initial human key point detection model until the initial human key point detection model reaches a preset convergence condition, and determine the target human key point detection model.
[0196] In another embodiment of the present invention, the training unit is specifically configured to
[0197] Batch the multiple groups of sample human body image pairs to obtain different batches of sample human body image pairs;
[0198] For each group of sample human body image pairs in each batch, input each image in the group of sample human body image pairs in the batch into the initial human key point detection model to obtain the heat maps corresponding to each human key point in each image of the group of sample human body image pairs;
[0199] Based on the heat maps corresponding to the human key points in each image of the group of sample human body image pairs in the batch, determine the predicted position information of the human key points in each image of the group of sample human body image pairs, and the predicted translation value between the human key points with corresponding relationships in the two images of the group of sample human body image pairs;
[0200] Based on the predicted position information and calibration position information of the human key points in each image of all the sample human body image pairs in the batch, and the predicted translation value and calibration translation value between the human key points with corresponding relationships in the two images of all the sample human body image pairs in the batch, determine the current loss value corresponding to the initial human key point detection model;
[0201] Judge whether the current loss value corresponding to the initial human key point detection model exceeds a preset loss threshold;
[0202] If it is judged that the current loss value corresponding to the initial human key point detection model exceeds the preset loss threshold, adjust the model parameters of the initial human key point detection model, and return to execute the step of inputting each image in the group of sample human body image pairs in the batch into the initial human key point detection model to obtain the heat maps corresponding to each human key point in each image of the group of sample human body image pairs for each group of sample human body image pairs in different batches;
[0203] If it is determined that the current loss value corresponding to the initial human key point detection model does not exceed the preset loss threshold, it is determined that the initial human key point detection model reaches the preset convergence condition, and the target human key point detection model is determined.
[0204] In another embodiment of the present invention, the first determination module 330 includes:
[0205] A fourth determination unit (not shown in the figure), configured to sequentially determine the similarity values between the human key points of each human in the current cabin image and the human key points of each human in the previous N cabin images, based on the two-dimensional position information of the human key points of each human in the current cabin image, the two-dimensional position information of the human key points of each human in the previous N cabin images, the pixel area of the human region image corresponding to each human, the detection visibility information corresponding to each human key point, and the timing information between the current cabin image and the previous N cabin images, where the visibility information corresponding to each human key point is information indicating whether the corresponding human key point is occluded;
[0206] A fifth determination unit (not shown in the figure), configured to determine the human key points of the same human in the current cabin image and the previous N cabin images based on the similarity values between the human key points of each human in the current cabin image and the previous N cabin images;
[0207] A sixth determination unit (not shown in the figure), configured to determine the three-dimensional position information corresponding to the human key points of each human based on the two-dimensional position information of the human key points of each human in the current cabin image and the previous N cabin images, the target three-dimensional human key point position prediction model, the parameter information of the image acquisition device, and the preset human activity range constraint condition, where the target three-dimensional human key point position prediction model is a model trained based on the two-dimensional position information of the sample human key points of the sample human sorted in time sequence and their corresponding calibrated three-dimensional position information.
[0208] In another embodiment of the present invention, the sixth determination unit is specifically configured to
[0209] Based on the two-dimensional position information of the human key points of each human in the current cabin image and the previous N cabin images and the target three-dimensional human key point position prediction model, determine the three-dimensional position information of the human key points of each human in the human coordinate system corresponding to the human;
[0210] For each human body, based on the three-dimensional position information of each human body key point of the human body in the human body coordinate system corresponding to the human body, the parameter information of the image acquisition device, and the preset human body activity range constraint condition corresponding to the human body, determine the three-dimensional position information of each human body key point of the human body in the cabin coordinate system corresponding to the target cabin.
[0211] The above system and device embodiments correspond to the system embodiment and have the same technical effects as the method embodiment. For specific descriptions, please refer to the method embodiment. The device embodiment is obtained based on the method embodiment. For specific descriptions, please refer to the method embodiment section and will not be elaborated here. Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily required for implementing the present invention.
[0212] Those of ordinary skill in the art can understand that the modules in the device in the embodiment can be distributed in the device in the embodiment according to the description in the embodiment, or can be correspondingly changed and located in one or more devices different from this embodiment. The modules in the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0213] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting human key points, characterized in that, the method includes: Detecting the two-dimensional position information of the human key points of each human from the obtained current cabin image, where the current cabin image is an image collected by an image acquisition device for the target cabin; Obtaining the two-dimensional position information of the human key points of each human detected from each of the previous N cabin images, where the previous N cabin images are: the previous N images of the current cabin image collected by the image acquisition device, and N is a positive integer; Based on the two-dimensional position information of the human key points of each human in the current cabin image, the two-dimensional position information of the human key points of each human in the previous N cabin images, the pixel area of the human area image corresponding to each human, the detection visibility information corresponding to each human key point, and the timing information between the current cabin image and the previous N cabin images, successively determining the similarity value between each human key point between every two adjacent images in the current cabin image and the previous N cabin images, where the visibility information corresponding to each human key point is information indicating whether the corresponding human key point is occluded; Based on the similarity value between each human key point between every two adjacent images in the current cabin image and the previous N cabin images, determining the human key points of the same human in the current cabin image and the previous N cabin images; Based on the two-dimensional position information of each human key point of each human in the current cabin image and the previous N cabin images, the target three-dimensional human key point position prediction model, the parameter information of the image acquisition device, and the preset human activity range constraint conditions, determining the three-dimensional position information corresponding to each human key point of each human, where the target three-dimensional human key point position prediction model is: a model trained based on the two-dimensional position information of the sample human key points of the sample human sorted in time sequence and their corresponding calibrated three-dimensional position information.
2. The method according to claim 1, characterized in that, the step of detecting the two-dimensional position information of the human key points of each human from the obtained current cabin image includes: Determining and intercepting the area where each human is located from the obtained current cabin image to obtain the human area image corresponding to each human; Based on the human area image corresponding to each human and the target human key point detection model, determining the heat map corresponding to each human key point in the human area image corresponding to each human, where the target human key point detection model is: a model trained based on multiple groups of sample human image pairs and their corresponding calibration information, and each sample human image pair is a human image corresponding to the same human with different translation values, and the calibration information includes the calibration translation value between the two images in the corresponding sample human image pair; Based on the heat map corresponding to each human key point in the human area image corresponding to each human, determining the two-dimensional position information of each human key point in the human area image corresponding to each human.
3. The method according to claim 2, characterized in that, Before the step of determining the two-dimensional position information of each human key point in the human region image corresponding to each human body based on the human region image corresponding to each human body and the target human key point detection model, the method further includes: The process of training the target human key point detection model, the process includes: Obtain an initial human key point detection model; Obtain multiple groups of sample human body image pairs and their corresponding calibration information, where the calibration information includes the calibration position information corresponding to each human key point in each image of the corresponding sample human body image pair and the calibration translation value between the two images in the corresponding sample human body image pair; Use the multiple groups of sample human body image pairs and the calibration position information corresponding to each human key point in each image of the corresponding sample human body image pair in the calibration information and the calibration translation value between the two images in the corresponding sample human body image pair to train the initial human key point detection model until the initial human key point detection model reaches a preset convergence condition, and determine the target human key point detection model.
4. The method according to claim 3, wherein, The step of using the multiple groups of sample human body image pairs and the calibration position information corresponding to each human key point in each image of the corresponding sample human body image pair in the calibration information and the calibration translation value between the two images in the corresponding sample human body image pair to train the initial human key point detection model until the initial human key point detection model reaches a preset convergence condition and determine the target human key point detection model includes: Batch the multiple groups of sample human body image pairs to obtain different batches of sample human body image pairs; For each group of sample human body image pairs in each batch, input each image in the group of sample human body image pairs in the batch into the initial human key point detection model to obtain the heat map corresponding to each human key point in each image of the group of sample human body image pairs; Based on the heat map corresponding to the human key point in each image of the group of sample human body image pairs in the batch, determine the predicted position information of the human key point in each image of the group of sample human body image pairs, and the predicted translation value between the human key points with corresponding relationships in the two images of the group of sample human body image pairs; Based on the predicted position information and calibration position information of the human key point in each image of all sample human body image pairs in the batch, and the predicted translation value and calibration translation value between the human key points with corresponding relationships in the two images of all sample human body image pairs in the batch, determine the current loss value corresponding to the initial human key point detection model; Judge whether the current loss value corresponding to the initial human key point detection model exceeds a preset loss threshold; If it is determined that the current loss value corresponding to the initial human key point detection model exceeds the preset loss threshold, adjust the model parameters of the initial human key point detection model, and return to execute the step of inputting each image in each group of sample human body image pairs in different batches into the initial human key point detection model to obtain the heat maps corresponding to each human key point in each image of each group of sample human body image pairs; If it is determined that the current loss value corresponding to the initial human key point detection model does not exceed the preset loss threshold, determine that the initial human key point detection model reaches the preset convergence condition, and determine the target human key point detection model.
5. The method according to claim 1, wherein, the step of determining the three-dimensional position information corresponding to each human key point of each human based on the two-dimensional position information of each human key point of each human in the current cabin image and the first N frame cabin images, the target three-dimensional human key point position prediction model, the parameter information of the image acquisition device, and the preset human activity range constraint condition includes: Based on the two-dimensional position information of each human key point of each human in the current cabin image and the first N frame cabin images and the target three-dimensional human key point position prediction model, determine the three-dimensional position information of each human key point of each human in the human coordinate system corresponding to the human; For each human, based on the three-dimensional position information of each human key point of the human in the human coordinate system corresponding to the human, the parameter information of the image acquisition device, and the preset human activity range constraint condition corresponding to the human, determine the three-dimensional position information of each human key point of the human in the cabin coordinate system corresponding to the target cabin.
6. A human key point detection device, wherein, the device includes: a detection module configured to detect the two-dimensional position information of the human key points of each human from the obtained current cabin image, where the current cabin image is an image collected by an image acquisition device for a target cabin; an obtaining module configured to obtain the two-dimensional position information of the human key points of each human detected from each of the first N frame cabin images, where the first N frame cabin images are the first N frame images of the current cabin image collected by the image acquisition device, and N is a positive integer; a first determination module including a fourth determination unit, a fifth determination unit, and a sixth determination unit; the fourth determination unit is configured to sequentially determine the similarity values between the human key points between every two adjacent images in the current cabin image and the first N frame cabin images based on the two-dimensional position information of the human key points of each human in the current cabin image, the two-dimensional position information of the human key points of each human in the first N frame cabin images, the pixel area of the human region image corresponding to each human, the detection visibility information corresponding to each human key point, and the time sequence information between the current cabin image and the first N frame cabin images, where the visibility information corresponding to each human key point is information indicating whether the corresponding human key point is blocked; The fifth determination unit is configured to determine the human key points of the same human body in the current cabin image and the previous N cabin images based on the similarity values between the human key points of each adjacent two images in the current cabin image and the previous N cabin images; The sixth determination unit is configured to determine the three-dimensional position information corresponding to the human key points of each human body based on the two-dimensional position information of the human key points of each human body in the current cabin image and the previous N cabin images, the target three-dimensional human key point position prediction model, the parameter information of the image acquisition device, and the preset human activity range constraint condition, where the target three-dimensional human key point position prediction model is: a model trained based on the two-dimensional position information of the sample human key points of the sample human body sorted in time series and their corresponding calibrated three-dimensional position information.
7. The device according to claim 6, wherein, the detection module includes: The first determination unit is configured to determine and intercept the regions where each human body is located from the obtained current cabin image to obtain the human body region images corresponding to each human body; The second determination unit is configured to determine the heat maps corresponding to the human key points in the human body region images corresponding to each human body based on the human body region images corresponding to each human body and the target human key point detection model, where the target human key point detection model is: a model trained based on multiple groups of sample human body image pairs and their corresponding calibration information, and each sample human body image pair is a human body image corresponding to the same human body with different translation values, and the calibration information includes the calibration translation value between the two images in the corresponding sample human body image pair; The third determination unit is configured to determine the two-dimensional position information of the human key points in the human body region images corresponding to each human body based on the heat maps corresponding to the human key points in the human body region images corresponding to each human body.
8. The device according to claim 7, wherein, the device further includes: The model training module is configured to train the target human key point detection model before determining the two-dimensional position information of the human key points in the human body region images corresponding to each human body based on the human body region images corresponding to each human body and the target human key point detection model. The model training module includes: The first acquisition unit is configured to acquire an initial human key point detection model; The second acquisition unit is configured to acquire multiple groups of sample human body image pairs and their corresponding calibration information, where the calibration information includes the calibration position information corresponding to the human key points in each image of the corresponding sample human body image pair and the calibration translation value between the two images in the corresponding sample human body image pair; A training unit, configured to use the multiple groups of sample human body image pairs and the calibration position information corresponding to each human body key point in each image of the sample human body image pairs in the calibration information and the calibration translation value between the two images in the corresponding sample human body image pair to train the initial human body key point detection model until the initial human body key point detection model reaches a preset convergence condition, and determine the target human body key point detection model.
9. The device according to claim 8, wherein, the training unit is configured to, for each group of sample human body image pairs, input each image in the group of sample human body image pairs into the initial human body key point detection model to obtain a heat map corresponding to each human body key point in each image of the group of sample human body image pairs; Based on the heat maps corresponding to the human body key points in each image of the group of sample human body image pairs, determine the predicted position information of the human body key points in each image of the group of sample human body image pairs, and the predicted translation value between the human body key points with corresponding relationships in the two images of the group of sample human body image pairs; Based on the predicted position information and the calibration position information of the human body key points in each image of the group of sample human body image pairs, and the predicted translation value and the calibration translation value between the human body key points with corresponding relationships in the two images of the group of sample human body image pairs, determine the current loss value corresponding to the initial human body key point detection model; Judge whether the current loss value corresponding to the initial human body key point detection model exceeds a preset loss threshold; If it is judged that the current loss value corresponding to the initial human body key point detection model exceeds the preset loss threshold, adjust the model parameters of the initial human body key point detection model, and return to execute the step of inputting each image in the group of sample human body image pairs into the initial human body key point detection model to obtain a heat map corresponding to each human body key point in each image of the group of sample human body image pairs; If it is judged that the current loss value corresponding to the initial human body key point detection model does not exceed the preset loss threshold, determine that the initial human body key point detection model reaches the preset convergence condition, and determine the target human body key point detection model.
Citation Information
Patent Citations
Indoor monitoring device
US20200098128A1