Human body posture estimation and positioning method, controller, system and storage medium

The camera acquires environmental images and inertial data, performs three-dimensional construction and attitude estimation, which solves the problems of low accuracy and inaccurate positioning of human poses in the prior art, and realizes high-precision human pose estimation and positioning in complex environments.

CN120071435APending Publication Date: 2025-05-30HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510073141.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing human posture estimation method has self-occlusion problem, making it difficult to capture the overall posture, the accuracy of posture estimation is low, and it is difficult to accurately position the characters in 3D scenes.

Method used

The camera deployed on the target object acquires the environment image, performs three-dimensional construction to obtain camera pose and depth data, generates a dense three-dimensional point cloud map, and combines inertial data for pose estimation and calibration, and ultimately achieves the positioning of the target object in any scene.

Benefits of technology

It improves the accuracy of human posture estimation, avoids occlusion problems, and achieves accurate positioning of the target object in any scene, and is suitable for complex environments with large terrain fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071435A_ABST
    Figure CN120071435A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a human body posture estimation and positioning method, a controller, a system and a storage medium, and belongs to the technical field of computer vision. The method comprises the following steps: acquiring a target number of environment images through a camera deployed on a target object; performing three-dimensional construction according to each environment image to obtain target camera pose and target depth data; performing dense point cloud mapping according to the target number of depth data and the target number of target camera poses to generate a dense three-dimensional point cloud map; performing dimension conversion on the dense three-dimensional point cloud map to obtain an elevation map; acquiring inertial data, performing posture estimation according to the inertial data, and determining an initial human body posture; calibrating the initial human body posture according to the elevation map to obtain a target human body posture; and positioning according to the target human body posture and the elevation map to obtain a target data set. According to the embodiment of the invention, the target object can be positioned in any scene, and the precision of human body posture estimation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular, to a human pose estimation and positioning method, a controller, a system, and a storage medium. Background Art

[0002] Camera-based human pose estimation (HPE) is a technology that simulates human visual perception. This technology is often closely combined with 3D scene reconstruction to show the interaction between the human body and the scene and to achieve the positioning of virtual characters in a 3D scene.

[0003] Current human pose estimation methods use downward-facing cameras to capture human poses. However, this method has the problem of self-occlusion, making it difficult to capture the overall pose and resulting in low accuracy of pose estimation. Moreover, current human pose estimation and 3D scene reconstruction methods are difficult to achieve the positioning of a person in a 3D scene. For example, situations may occur where a person penetrates the ground of the 3D scene (such as the ground in a non-flat terrain scene) or a person hangs in the air above the ground of the 3D scene. Summary of the Invention

[0004] The main objective of the embodiments of the present application is to propose a human pose estimation and positioning method, a controller, a system, and a storage medium, aiming to improve the accuracy of human pose estimation and to be able to achieve the positioning of a target object in any scene.

[0005] To achieve the above objective, a first aspect of the embodiments of the present application proposes a human pose estimation and positioning method, the method comprising:

[0006] Obtaining a target number of environmental images through a camera deployed on a target object; the environmental images are used to represent the environment where the target object is located;

[0007] Performing three-dimensional construction based on each of the environmental images to obtain a target camera pose and target depth data;

[0008] Performing dense point cloud mapping based on the target number of the depth data and the target number of the target camera poses to generate a dense three-dimensional point cloud map;

[0009] Performing dimension conversion on the dense three-dimensional point cloud map to obtain an elevation map; wherein, the dimension of the elevation map is less than three dimensions;

[0010] Obtaining inertial data of the target object;

[0011] Performing pose estimation based on the inertial data to determine an initial human pose;

[0012] Calibrate the initial human pose according to the elevation map to obtain the target human pose;

[0013] Locate according to the target human pose and the elevation map to obtain a target data set.

[0014] In some embodiments, the three-dimensional construction according to each environmental image to obtain the target camera pose and the target depth data includes:

[0015] Perform pose recognition on the environmental image to obtain the initial camera pose;

[0016] Perform feature recognition on the environmental image to obtain image feature points;

[0017] Perform depth recognition on the environmental image to obtain the initial depth data;

[0018] Construct a function according to a preset positioning and mapping algorithm, the initial camera pose, the image feature points, and the initial depth data to obtain a loss function;

[0019] Update the pose according to the loss function and the initial camera pose to obtain the target camera pose;

[0020] Update the depth according to the loss function and the initial depth data to obtain the target depth data.

[0021] In some embodiments, the construction of a function according to a preset positioning and mapping algorithm, the initial camera pose, and the environmental image to obtain a loss function includes:

[0022] Construct a relationship between any two of the initial camera poses in the target quantity to obtain relationship data between the poses;

[0023] Construct a function according to the relationship data between the poses, the initial camera pose, the image feature points, and the initial depth data to obtain a sub-loss function;

[0024] Sum the sub-loss functions to obtain the loss function.

[0025] In some embodiments, the loss function is as shown in the following formula:

[0026] E total = E repr + λ·E inert

[0027] The E total represents the loss function, the E repr represents the reprojection error, the E inert represents the inertial error, and λ represents the weight, and λ ∈ R;

[0028] The definition of E repr is as follows:

[0029]

[0030] where i represents the i-th frame of environmental image, j represents the j-th frame of environmental image, and ε represents the set of environmental images; represents the optical flow matching error between the i-th frame of environmental image and the j-th frame of environmental image; ∏ c (·) represents the pinhole projection function, represents the inverse projection function of the pinhole projection function; G ij represents the relationship data between the poses between the i-th frame of environmental image and the j-th frame of environmental image; u i represents the image feature points of the i-th frame of environmental image; d i represents the initial depth data of the i-th frame of environmental image; ∑ij represents the confidence weight;

[0031] The definition of G ij is as follows:

[0032]

[0033] where G j represents the matrix corresponding to the initial camera pose of the j-th frame of environmental image, represents the inverse matrix of the matrix corresponding to the initial camera pose of the i-th frame of environmental image.

[0034] In some embodiments, after performing pose recognition on the environmental image to obtain the initial camera pose, the method further includes:

[0035] Aligning the coordinate systems according to the human body coordinate system in which the inertial data is located and the camera coordinate system in which the environmental image is located to obtain a target coordinate system;

[0036] Performing coordinate transformation on the initial camera pose according to the target coordinate system to obtain a first camera pose;

[0037] Replacing the initial camera pose with the first camera pose.

[0038] In some embodiments, after performing attitude estimation according to the inertial data to determine the initial human body attitude, the method further includes:

[0039] Calibrating the initial human body attitude according to a preset attitude joint constraint relationship and the target camera pose to obtain a calibrated human body attitude; the attitude joint constraint relationship characterizes the position constraint relationship between the target object and the camera;

[0040] Replace the initial human pose with the calibrated human pose.

[0041] In some embodiments, after updating the pose according to the loss function and the initial camera pose to obtain the target camera pose, the method further includes:

[0042] Updating the pose of the target camera through a preset loop detection algorithm to obtain a second camera pose;

[0043] Replace the target camera pose with the second camera pose.

[0044] To achieve the above object, a second aspect of the embodiments of the present application provides a controller, the controller includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the human pose estimation and positioning method described in the first aspect above.

[0045] To achieve the above object, a third aspect of the embodiments of the present application provides a human pose estimation and positioning system, the system includes: a camera, an inertial measurement unit, and the controller described in the second aspect above.

[0046] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the human pose estimation and positioning method described in the first aspect above.

[0047] The human body pose estimation and positioning method, controller, system and storage medium proposed in this application obtain environmental images of a target number through a camera deployed on a target object, and perform three-dimensional construction based on each environmental image to obtain target camera poses and target depth data; perform dense point cloud mapping based on the target number of depth data and the target number of target camera poses to generate a dense three-dimensional point cloud map. It can be seen that the dense three-dimensional point cloud map generated according to the depth data and the target camera pose can more accurately and comprehensively reflect the environment around the target object, and can be better applied to human body pose estimation in complex environments with large terrain undulations; moreover, after obtaining inertial data, perform pose estimation according to the inertial data of the target object to determine the initial human body pose, which can avoid occlusion problems (for example, if a downward-facing head-mounted camera is used to capture the pose of the target object, there will be a perspective occlusion problem and the overall pose of the target object cannot be collected), improving the reliability and accuracy of the human body pose estimation method; obtain an elevation map according to the dense three-dimensional point cloud map, and then calibrate the initial human body pose with the elevation map to obtain the target human body pose, and perform positioning according to the target human body pose and the elevation map to obtain a target data set. This method can achieve the positioning of the target object in any scenario and improve the accuracy of human body pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of the human body pose estimation and positioning method provided by an embodiment of the present application;

[0049] Figure 2 is provided by an embodiment of the present application Figure 1 flowchart of step 102 in;

[0050] Figure 3 is provided by an embodiment of the present application Figure 2 flowchart of step 204 in;

[0051] Figure 4 is a schematic hardware structure diagram of a controller provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0055] First, several nouns involved in this application are analyzed as follows:

[0056] 3D scene reconstruction: It is a technology that scans or photographs an object or scene through a camera and uses computer algorithms to convert two-dimensional photo data into an interactive three-dimensional model. For example, real building or terrain data is obtained, and then 3D scene reconstruction is carried out through technologies such as 3D modeling to generate an interactive three-dimensional model. This technology is widely used in fields such as computer vision and virtual reality.

[0057] Monocular camera: A monocular camera is a camera system that uses a single lens to capture a scene. This camera system has only one lens, so it is called a monocular camera.

[0058] Monocular image: The image obtained by a monocular camera is called a monocular image, which is relative to binocular images (two images for the same scene) and multiocular images (multiple images for the same scene).

[0059] Pose: It refers to the position and orientation of an object in space, such as the position and orientation of an object in a Cartesian coordinate system.

[0060] 3D point cloud map: It is used to represent objects or scenes in three-dimensional space and can describe the three-dimensional structure and features of the real world through a discrete three-dimensional point cloud data set.

[0061] SLAM (Simultaneous Localization and Mapping, synchronous positioning and mapping): It is a technology used to construct a map of the surrounding environment in real time and perform positioning on the map. SLAM technology can include SLAM algorithms based on graph optimization and SLAM algorithms based on deep learning (such as optical flow method).

[0062] Optical flow method: It is a method used to detect moving targets in the field of computer vision and can estimate the motion information of an object by analyzing the changes in pixel values in an image sequence.

[0063] Loop detection: Also known as closed-loop detection, it is a part of the SLAM technology. During the SLAM mapping process, if only key frames at adjacent times are considered, cumulative errors are likely to occur over time. The loop detection algorithm reduces the cumulative errors in the map construction process by identifying whether the target object has returned to an environment that has been previously captured, thereby improving the accuracy of the target object's positioning.

[0064] The human pose estimation and positioning method, controller, system, and storage medium provided by the embodiments of the present application will be specifically described through the following embodiments. First, the human pose estimation and positioning method in the embodiments of the present application will be described.

[0065] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0066] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0067] The human pose estimation and positioning method provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the human pose estimation and positioning method, etc., but is not limited to the above forms.

[0068] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0069] It should be noted that in each specific embodiment of this application, when it comes to performing relevant processing based on data related to the user's body pose data, historical pose data, and user location information, etc., which are related to the user's identity or characteristics, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when this application embodiment needs to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of this application embodiment will be obtained.

[0070] Figure 1 It is an optional flowchart of the human body pose estimation and positioning method provided by an embodiment of this application. Figure 1 The method in [it] can include but is not limited to steps 101 to 108.

[0071] Step 101, obtain a target number of environmental images through a camera deployed on the target object; the environmental images are used to characterize the environment where the target object is located.

[0072] Step 102, perform three-dimensional construction based on each environmental image to obtain the target camera pose and target depth data.

[0073] Step 103, perform dense point cloud mapping based on the target number of depth data and the target number of target camera poses to generate a dense three-dimensional point cloud map.

[0074] Step 104, perform dimension conversion on the dense three-dimensional point cloud map to obtain an elevation map; wherein, the dimension of the elevation map is less than three dimensions.

[0075] Step 105, obtain the inertial data of the target object.

[0076] Step 106: Perform attitude estimation based on inertial data to determine the initial human body attitude.

[0077] Step 107: Calibrate the initial human body attitude according to the elevation map to obtain the target human body attitude.

[0078] Step 108: Perform positioning based on the target human body attitude and the elevation map to obtain the target data set.

[0079] The beneficial effects of the embodiments of the present application include but are not limited to: obtaining environmental images of a target quantity through a camera deployed on a target object, and performing three-dimensional construction based on each environmental image to obtain the target camera pose and target depth data; performing dense point cloud mapping based on the target quantity of depth data and the target quantity of target camera poses to generate a dense three-dimensional point cloud map. It can be seen that the dense three-dimensional point cloud map generated according to the depth data and the target camera pose can more accurately and comprehensively reflect the environment around the target object, and can be better applied to human body attitude estimation in complex environments with large terrain undulations; moreover, after obtaining inertial data, performing attitude estimation based on the inertial data of the target object to determine the initial human body attitude can avoid occlusion problems (for example, if a downward-facing head-mounted camera is used to capture the attitude of the target object, there will be a perspective occlusion problem and the overall attitude of the target object cannot be collected), improving the reliability and accuracy of the human body attitude estimation method; obtaining an elevation map based on the dense three-dimensional point cloud map, then calibrating the initial human body attitude with the elevation map to obtain the target human body attitude, and performing positioning based on the target human body attitude and the elevation map to obtain the target data set. This method can achieve the positioning of the target object in any scenario and improve the accuracy of human body attitude estimation.

[0080] In step 101 of some embodiments, the camera can be a monocular camera for collecting videos or images from a first-person perspective. For example, a head-mounted monocular RGB camera can be used to collect a color video from a first-person perspective, and key frames can be extracted from the color video to obtain multiple RGB (red, green, blue) images, which are the above-mentioned environmental images. Other types of cameras can also be selected, and are not limited thereto.

[0081] It should be noted that the target object can be a target person or a humanoid robot. The target quantity can be set according to actual needs. The camera poses corresponding to any two of the environmental images of the target quantity should be different.

[0082] In step 103 of some embodiments, map construction can be performed according to the RGB-SLAM algorithm, the depth data of the target quantity, and the target camera poses of the target quantity to generate a dense three-dimensional point cloud map. It should be noted that the RGB-SLAM algorithm is an algorithm that uses RGB (red, green, blue) images and depth information for simultaneous localization and mapping (SLAM). It can use color information to enhance the understanding and recognition of the environment, thereby obtaining an accurate dense three-dimensional point cloud map.

[0083] It should be noted that since the dense three-dimensional point cloud map is generated based on the depth data of the target quantity and the target camera poses of the target quantity, compared with the sparse point cloud map, the embodiments of the present application can be better applied to different scenarios. For example, human pose localization can be achieved in non-flat terrain scenarios.

[0084] In step 104 of some embodiments, the dense three-dimensional point cloud map can be converted into a height map through a point cloud data processing algorithm. For example, through a deep learning algorithm, the dense three-dimensional point cloud map is divided into multiple voxels (such as divided into cubes), and the average height of each voxel is calculated. Then, a height map is generated based on the average height of the multiple voxels. Then (see step 107), the initial human pose is calibrated using the height map. Compared with directly calibrating the pose using a dense three-dimensional point cloud map composed of discrete points, using a continuous height map can more accurately calibrate the initial human pose, thereby more accurately determining the relationship between the target object and the environment represented by the height map.

[0085] Specifically, the height map can be a 2.5D height map.

[0086] In step 105 of some embodiments, the pose of the target object can be estimated through an inertial measurement unit (IMU) to obtain inertial data. For example, six inertial measurement units can be used to synchronously collect the acceleration and orientation of the target object at a frequency of 60 Hz, and the acceleration and orientation are used as inertial data.

[0087] In step 106 of some embodiments, the inertial data can be used for pose estimation through a deep learning model to determine the initial human pose.

[0088] In step 107 of some embodiments, the initial human pose and the height map can be matched through a feature matching algorithm to determine the contact point between the initial human pose and the height map. Then, based on the contact point, the pose of the target object is located in the height map to obtain the target human pose, avoiding situations such as the person corresponding to the initial human pose penetrating the ground in the height map or the person hanging in the air above the ground, thereby improving the authenticity and reliability of human pose estimation.

[0089] In step 108 of some embodiments, a target data set is obtained based on the target human pose and the elevation map. Since the storage space occupied by the elevation map is smaller than that required by the dense three-dimensional point cloud map, compared with directly using the dense three-dimensional point cloud map and the target human pose for positioning, the size of the target data set can be reduced, the data storage cost can be lowered, and the calculation speed of the positioning method can be improved. It should be noted that the target data set is used to represent the position and human pose of the target object in the environment.

[0090] In some embodiments, after step 101, the environmental image can be preprocessed, such as performing a distortion removal process, so as to improve the image quality.

[0091] In some embodiments, after step 105, the inertial data can be filtered to obtain the filtered inertial data, thereby reducing noise. Then, the initial human pose can be determined according to the filtered inertial data for pose estimation.

[0092] Please refer to Figure 2 , in some embodiments, step 102 may include but is not limited to steps 201 to 206:

[0093] Step 201, perform pose recognition on the environmental image to obtain the initial camera pose;

[0094] Step 202, perform feature recognition on the environmental image to obtain image feature points;

[0095] Step 203, perform depth recognition on the environmental image to obtain the initial depth data;

[0096] Step 204, construct a function according to the preset positioning and mapping algorithm, the initial camera pose, the image feature points, and the initial depth data to obtain a loss function;

[0097] Step 205, update the pose according to the loss function and the initial camera pose to obtain the target camera pose;

[0098] Step 206, update the depth according to the loss function and the initial depth data to obtain the target depth data.

[0099] The advantage of this embodiment is that by constructing a loss function and using the loss function to update the camera pose and depth data, the target camera pose and target depth data are obtained, thereby improving the accuracy of human pose estimation and positioning.

[0100] In step 201 of some embodiments, the initial camera pose can characterize the position and orientation of the camera when collecting environmental images. For example, if the target object is moving and the camera deployed on the target object is also moving, during the movement, different environmental images collected by the camera correspond to different initial camera poses.

[0101] In step 202 of some embodiments, all pixel points of the environmental image can be used as image feature points. Since the image feature points are used to construct the loss function, the embodiments of the present application can improve the accuracy of the depth data determined according to the loss function. Image feature points can also be determined by other methods, and are not limited thereto.

[0102] In step 203 of some embodiments, the initial depth data can characterize the distances from points on the surfaces of objects in the environment to the camera;

[0103] In step 204 of some embodiments, the positioning and mapping algorithm includes the Multi-View Bundle Adjustment (MDBA) algorithm and the deep learning-based SLAM algorithm. Specifically, a function can also be constructed through the optical flow method, the initial camera pose, the image feature points, and the initial depth data to obtain the loss function.

[0104] In step 205 of some embodiments, the loss function is minimized, and the initial camera pose corresponding to the minimum loss function is used as the target camera pose.

[0105] In step 206 of some embodiments, the loss function is minimized, and the initial depth data corresponding to the minimum loss function is used as the target depth data.

[0106] Please refer to Figure 3 , in some embodiments, step 204 may include but is not limited to:

[0107] Step 301, constructing a relationship based on any two initial camera poses among the target quantities to obtain pose relationship data;

[0108] Step 302, constructing a function based on the pose relationship data, the initial camera pose, the image feature points, and the initial depth data to obtain a sub-loss function;

[0109] Step 303, summing the sub-loss functions to obtain the loss function.

[0110] The advantage of this embodiment is that, based on the pose relationship data, the initial camera pose, the image feature points, and the initial depth data, the construction of the loss function is realized, so that the camera pose and the depth data are updated through the loss function to obtain the target camera pose and the target depth data, improving the accuracy of human pose estimation and positioning.

[0111] In some embodiments, the loss function is as shown in the following formula:

[0112] E total = E repr + λ·E inert ,

[0113] E total represents the loss function, E repr represents the reprojection error, E inert represents the inertial error, λ represents the weight, and λ ∈ R;

[0114] E repr is defined as:

[0115]

[0116] where i represents the i-th frame of the environmental image, j represents the j-th frame of the environmental image, and ε represents the set of environmental images; represents the optical flow matching error between the i-th frame of the environmental image and the j-th frame of the environmental image; ∏ c (·) represents the pinhole projection function, represents the inverse projection function of the pinhole projection function; G ij represents the relationship data between the poses between the i-th frame of the environmental image and the j-th frame of the environmental image; u i represents the image feature points of the i-th frame of the environmental image; d i represents the initial depth data of the i-th frame of the environmental image; ∑ij represents the confidence weight;

[0117] G ij is defined as:

[0118]

[0119] where G j represents the matrix corresponding to the initial camera pose of the j-th frame of the environmental image, represents the inverse matrix of the matrix corresponding to the initial camera pose of the i-th frame of the environmental image.

[0120] The advantage of this embodiment is that by updating the camera pose and depth data through the loss function, the target camera pose and target depth data are obtained, improving the accuracy of human pose estimation and positioning.

[0121] It should be noted that the set of environmental images includes all environmental images; represents the inverse matrix of the matrix G i corresponding to the initial camera pose of the i-th frame of the environmental image; where G i ∈ SE(3), and G j∈ SE(3); SE(3) represents the three-dimensional rotation group, which can specifically be a set of multiple matrices; the confidence weight can be represented as a diagonal matrix.

[0122] In some embodiments, the reprojection error in the loss function can also be minimized by the Dense Bundle Adjustment (DBA) algorithm to obtain the minimum reprojection error, thereby obtaining the minimum loss function. It should be noted that the inertial error mainly stems from the accuracy limitation of the inertial measurement unit.

[0123] In some embodiments, after step 201, the human pose estimation and positioning method further includes:

[0124] Align the coordinate systems based on the human body coordinate system where the inertial data is located and the camera coordinate system where the environmental image is located to obtain the target coordinate system;

[0125] Perform coordinate system transformation on the initial camera pose according to the target coordinate system to obtain the first camera pose;

[0126] Replace the initial camera pose with the first camera pose.

[0127] The advantage of this embodiment is that through coordinate system alignment and coordinate system transformation, the first camera pose corresponding to the target coordinate system is obtained, and the initial camera pose is replaced with the first camera pose. Then, the loss function is constructed using the initial camera pose, and further, the camera pose and depth data are updated through the loss function to obtain the target camera pose and target depth data, improving the accuracy of human pose estimation and positioning.

[0128] In some embodiments, after step 106, the method further includes:

[0129] Calibrate the initial human pose according to the preset pose joint constraint relationship and the target camera pose to obtain the calibrated human pose; the pose joint constraint relationship characterizes the position constraint relationship between the target object and the camera;

[0130] Replace the initial human pose with the calibrated human pose.

[0131] The advantage of this embodiment is that according to the preset pose joint constraint relationship and the target camera pose, the initial human pose is calibrated to obtain the calibrated human pose, and the initial human pose is replaced with the calibrated human pose, improving the accuracy of human pose estimation.

[0132] It should be noted that since the camera is deployed on the target object (for example, the target object wears a head-mounted camera), there is a pose joint constraint relationship between the camera pose and the human pose, and the pose joint constraint relationship can be used to represent the position relationship between the camera and the target object.

[0133] In some embodiments, after step 205, the human pose estimation and positioning method further includes:

[0134] Performing pose update on the target camera pose through a preset loop detection algorithm to obtain a second camera pose;

[0135] Replacing the target camera pose with the second camera pose.

[0136] The advantage of this embodiment is that the pose of the target camera is updated through the loop detection algorithm, improving the accuracy of human pose estimation. The target camera pose is used for dense point cloud mapping to obtain a target data set, enabling the positioning of the target object in any scenario.

[0137] An embodiment of the present application also provides a controller, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned human pose estimation and positioning method. The controller can include any intelligent terminal such as a tablet computer or an in-vehicle computer.

[0138] Please refer to Figure 4 , Figure 4 which illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0139] A processor 401, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0140] A memory 402, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402 and are called by the processor 401 to execute the human pose estimation and positioning method of the embodiments of the present application;

[0141] An input / output interface 403, which is used to implement information input and output;

[0142] A communication interface 404, which is used to implement the communication interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WI-FI, Bluetooth, etc.);

[0143] A bus 405, which transmits information between various components of the device (such as a processor 401, a memory 402, an input / output interface 403, and a communication interface 404);

[0144] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 achieve communication connections with each other inside the device through the bus 405.

[0145] The embodiment of this application also provides a human body pose estimation and positioning system, which can implement the above-mentioned human body pose estimation and positioning method. This system includes: a camera, an inertial measurement unit, and the above-mentioned controller.

[0146] The specific implementation manner of this human body pose estimation and positioning system is basically the same as the specific embodiment of the above-mentioned human body pose estimation and positioning method, and will not be elaborated here.

[0147] The embodiment of this application also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program, and when this computer program is executed by a processor, it implements the above-mentioned human body pose estimation and positioning method.

[0148] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0149] The embodiments described in the embodiments of this application are for more clearly explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0150] Those skilled in the art can understand that the technical solutions shown in the figure do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than those shown in the figure, or combine certain steps, or different steps.

[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0152] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0153] In the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0154] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0155] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0156] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0158] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0159] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A method for estimating and positioning a human body posture, characterized in that: The method comprises: Acquire a target number of environmental images through a camera deployed on the target object; the environmental images are used to characterize the environment in which the target object is located; Perform three-dimensional construction based on each of the environment images to obtain target camera position and target depth data; Perform dense point cloud mapping according to the depth data of the target number and the target camera poses of the target number to generate a dense three-dimensional point cloud map; Performing dimension conversion on the dense three-dimensional point cloud map to obtain an elevation map; wherein the dimension of the elevation map is less than three dimensions; Acquiring inertial data of the target object; Performing posture estimation based on the inertial data to determine an initial human body posture; Calibrate the initial human body posture according to the elevation map to obtain a target human body posture; Positioning is performed according to the target human body posture and the elevation map to obtain a target data set.

2. The method for estimating and positioning a human body posture according to claim 1, characterized in that: The three-dimensional construction is performed according to each of the environment images to obtain the target camera pose and target depth data, including: Performing posture recognition on the environment image to obtain an initial camera posture; Recognize features of the environment image to obtain image feature points; Performing depth recognition on the environment image to obtain initial depth data; Constructing a function according to a preset positioning and mapping algorithm, the initial camera pose, the image feature points and the initial depth data to obtain a loss function; Performing pose update according to the loss function and the initial camera pose to obtain the target camera pose; The depth is updated according to the loss function and the initial depth data to obtain the target depth data.

3. The method for estimating and positioning a human body posture according to claim 2, characterized in that: The function is constructed according to the preset positioning and mapping algorithm, the initial camera pose and the environment image to obtain a loss function, including: Establishing a relationship between any two of the initial camera poses in the target number to obtain relationship data between the poses; Constructing a function according to the relationship data between the postures, the initial camera posture, the image feature points and the initial depth data to obtain a sub-loss function; The sub-loss functions are summed to obtain the loss function.

4. The method for estimating and positioning a human body posture according to claim 2, characterized in that: The loss function is shown in the following formula: AND total =And repr +λ·E inert , The E total Represents the loss function, the E repr represents the reprojection error, the E inert represents the inertia error, λ represents the weight, and λ∈R; The E repr is defined as: Among them, i represents the i-th frame environment image, j represents the j-th frame environment image, and ε represents the environment image set; represents the optical flow matching error between the i-th frame environment image and the j-th frame environment image; c (·) represents the pinhole projection function, represents the inverse projection function of the pinhole projection function; G ij Represents the relationship data between the positions and postures of the i-th frame environment image and the j-th frame environment image; u i Represents the image feature points of the i-th frame environment image; d i represents the initial depth data of the i-th frame environment image; ∑ij represents the confidence weight; The G ij is defined as: Among them, G j Represents the matrix corresponding to the initial camera pose of the j-th frame environment image, The inverse matrix of the matrix corresponding to the initial camera pose of the i-th frame environment image.

5. The method for estimating and positioning a human body posture according to claim 2, characterized in that: After performing posture recognition on the environment image to obtain an initial camera posture, the method further includes: Aligning coordinate systems according to the human body coordinate system in which the inertial data is located and the camera coordinate system in which the environmental image is located, to obtain a target coordinate system; Performing a coordinate system transformation on the initial camera pose according to the target coordinate system to obtain a first camera pose; The initial camera pose is replaced with the first camera pose.

6. The method for estimating and positioning a human body posture according to any one of claims 1 to 5, characterized in that: After performing posture estimation according to the inertial data to determine the initial human body posture, the method further includes: According to a preset posture joint constraint relationship and the target camera posture, the initial human body posture is calibrated to obtain a calibrated human body posture; the posture joint constraint relationship represents the position constraint relationship between the target object and the camera; The initial human body posture is replaced by the calibrated human body posture.

7. The method for estimating and positioning a human body posture according to claim 2, characterized in that: After performing pose updating according to the loss function and the initial camera pose to obtain the target camera pose, the method further includes: The target camera pose is updated by using a preset loop closure detection algorithm to obtain a second camera pose; The target camera pose is replaced with the second camera pose.

8. A controller, characterized in that: The controller includes a memory and a processor, the memory stores a computer program, and the processor implements the human body posture estimation and positioning method according to any one of claims 1 to 7 when executing the computer program.

9. A human body posture estimation and positioning system, characterized in that: The system comprises: a camera, an inertial measurement unit and the controller of claim 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the human body posture estimation and positioning method according to any one of claims 1 to 7 is implemented.