Man-machine co-fusion system
By combining augmented reality technology, inertial sensors and multimodal data fusion in the human-computer inclusion system, the problems of high cost, insufficient accuracy and complex deployment of human posture and interactive behavior capture in dynamic environments are solved, and high-precision spatial positioning and dynamic interactive tracking are achieved.
Patent Information
- Application Number
- CN202510120855.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
When the prior art captures human posture, global positioning and interactive behavior in a dynamic environment, there are problems such as high cost, insufficient accuracy and complex deployment, and it is difficult to form a unified system solution.
It provides a human-computer inclusion system, including augmented reality positioning module, human posture capture module and data fusion module. Through augmented reality technology, inertial sensors and multimodal data fusion, it realizes high-precision positioning and dynamic interactive tracking of the human body.
It realizes high-precision spatial positioning and human posture capture, reduces costs, simplifies deployment, and establishes a two-way real-time system that supports virtual and realistic interaction with high-precision data transmission, realizing the seamless integration and interaction between the physical world and the virtual world.
Smart Images

Figure CN119987557A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to but is not limited to the fields of artificial intelligence and situational awareness technology, and in particular to a human-machine collaborative system based on egocentric global tracking, positioning and perception. Background Art
[0002] Augmented reality (AR) refers to the process of superimposing information or images provided by a computer system with real-world information and presenting it to users, thereby enhancing users' perception of the real world. The key point is that the information or images are superimposed on the real world, which is a "real within virtual" performance effect. For users, it is equivalent to "enhancing" their understanding and perception of the real world.
[0003] Capturing and modeling the interaction between humans and scenes in a dynamic environment is a complex and important research topic, covering the fields of robotics, augmented reality, virtual reality (VR), and digital twins. The relevant technologies are mainly concentrated in optical motion capture, vision-based methods, inertial sensor (IMU)-based technologies, and traditional positioning methods. These technologies have their own advantages and disadvantages in terms of accuracy, scope of application, and cost. However, they generally have problems such as complex deployment, high cost, or poor environmental adaptability, which limits their widespread application.
[0004] In general, related technologies have problems such as high cost, insufficient accuracy and complex deployment when capturing human posture, global positioning and interactive behavior in dynamic environments. At the same time, these technologies are often studied independently, making it difficult to form a unified system solution. Summary of the invention
[0005] The present application provides a human-machine collaborative system that can form a unified system, improve positioning accuracy, and reduce costs.
[0006] The embodiment of the present invention provides a human-machine integration system, including: an augmented reality positioning module, a human posture capture module, and a data fusion module; wherein: The augmented reality positioning module is used to obtain the real-time positioning information of the augmented reality device in the physical space through augmented reality technology; The human posture capture module is used to capture the human posture data based on the inertial sensor IMU, and constrain it through pre-set physical constraints to ensure the accuracy and naturalness of posture reconstruction; The data fusion module is used to combine the acquired real-time positioning information with the obtained posture data to realize the absolute position tracking and real-time correction of the human body, so as to ensure the consistency of the absolute position and motion posture of the human body in the virtual and actual scenes.
[0007] In an exemplary embodiment, it also includes a human-scene dynamic interaction tracking module, which is used to use augmented reality technology to realize the interaction between the virtual human body and the real scene, and dynamically record the scene changes; including: an interactive behavior modeling submodule, a data recording submodule, and a virtual-real matching submodule; wherein, The interactive behavior modeling submodule is used to add physical constraints and logical rules to simulate the physical behaviors in real interactions to capture the interactive behaviors between the human body and the scene; The data recording submodule is used to record the state changes of objects during scene interaction in real time through custom scripts; The virtual-reality matching submodule is used to synchronize the interaction effects with the physical model in the virtual scene.
[0008] In an exemplary embodiment, the augmented reality positioning module is used to: By combining the simultaneous positioning and mapping SLAM provided by the augmented reality device with the multi-marker registration technology, the position and posture in the three-dimensional space are identified to obtain the real-time positioning information of the augmented reality device in the physical space.
[0009] In an exemplary embodiment, the augmented reality positioning module includes: a positioning data acquisition submodule, a three-dimensional registration and alignment submodule, and a coordinate system alignment and correction submodule, wherein: A positioning data acquisition submodule, used to collect environmental data in real time through the augmented reality device, and generate the real-time positioning information through the positioning function in the SLAM; A three-dimensional registration and alignment submodule, for spatially registering the multiple markers based on a perspective n-point problem (PnP) algorithm, and establishing alignment between a device coordinate system and a world coordinate system; The coordinate system alignment and correction submodule is used to combine the multi-marker correction and spatial coordinate adjustment technology to dynamically update the mapping relationship between the coordinate system of the augmented reality device and the world coordinate system to ensure that the augmented reality device maintains consistency in spatial position and direction during dynamic interaction.
[0010] In an exemplary embodiment, the multiple identifiers include: a QR code or a marking point.
[0011] In an exemplary embodiment, the human body posture capture module is used to: The motion data collected by the IMU is transmitted to a real-time 3D engine and a development platform through a high-level programming language to dynamically present the human body posture, so that the posture of the virtual human body in the augmented reality device is synchronized with the real human body in real time.
[0012] In an exemplary embodiment, the human body posture capture module includes: an inertial sensor acquisition submodule, a posture estimation submodule, and a physical constraint optimization submodule, wherein: The inertial sensor acquisition submodule is used to collect acceleration and angular velocity data in real time through an IMU worn on key parts of the human body to capture the movement information of the human body; The posture estimation submodule is used to calculate the posture data of the human body based on the parameterized three-dimensional human body model SMPL and the deep learning algorithm, and in combination with the acceleration and angular velocity data collected in real time; the posture data of the human body includes the joint position and the motion state; The physical constraint optimization submodule is used to introduce one or any combination of the following physical constraints: joint angle limit, motion coherence, and dynamic consistency, to constrain the captured posture data to ensure the accuracy and naturalness of the captured posture.
[0013] In an exemplary embodiment, the physical constraint optimization submodule may also be used to set ground contact and sliding detection constraints to optimize the stability of the human body's posture in the environment.
[0014] In an exemplary embodiment, the data fusion module includes a mapping submodule and a dynamic correction submodule; wherein, A mapping submodule, used for establishing an initial mapping relationship between the augmented reality device and the human skeleton through static calibration; The dynamic correction submodule is used to dynamically update the position information of the skeleton root node by using the data collected by the IMU and the nonlinear state estimation method.
[0015] In an exemplary embodiment, the nonlinear state estimation method is an extended Kalman filter (EKF) algorithm.
[0016] The human-machine collaborative system based on egocentric global tracking, positioning and perception provided in the embodiment of the present application realizes the estimation of the position and motion posture of the human body in a large 3D scene through a small number of wearable sensors, realizes high-precision spatial positioning, and ensures the precise position and posture of the head display in three-dimensional space, thereby realizing stable tracking of the positioning system.
[0017] Furthermore, the human-machine integration system provided in the embodiment of the present application uses augmented reality technology to assist in the dynamic interactive tracking of people and scenes, establishes a two-way real-time system for virtual and real interaction that supports high-precision data transmission, and truly realizes the seamless integration and interaction of the physical world and the virtual world.
[0018] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to provide further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0020] Figure 1 This is a schematic diagram of the composition of the human-machine integration system in the embodiment of the present application; Figure 2 This is a schematic diagram of the implementation process of the human-machine integration system in the embodiment of the present application; Figure 3 A schematic diagram of the implementation process of the three-dimensional registration method based on multiple markers in the augmented reality positioning module in the embodiment of the present application; Figure 4 This is a schematic diagram of the SMPL human body model and the sensor wearing position in the system in the embodiment of the present application; Figure 5 Schematic diagram of the implementation process of data fusion in the embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present application more clear, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other arbitrarily without conflict.
[0022] In a typical configuration of the present application, a computing device includes one or more processors (CPU), an input / output interface, a network interface, and a memory.
[0023] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0024] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0025] The steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. Also, although a logical sequence is shown in the flowchart, in some cases, the steps shown or described can be performed in a sequence different from that shown here.
[0026] Optical motion capture technology uses a multi-camera system and external markers to achieve high-precision capture of human motion through a 3D reconstruction algorithm, and is widely used in filmmaking and professional animation production. However, this method is sensitive to ambient light conditions and performs poorly in strong light, shadows, or low light environments. At the same time, the deployment and calibration of multiple cameras are complex and costly, making it unsuitable for non-professional users or dynamic wide-area scenes.
[0027] The vision-based method does not rely on external markers, and uses a single camera or multiple cameras combined with a deep learning algorithm to estimate human posture. Although this method has improved accuracy and flexibility, it is easily affected by occlusion problems in dynamic scenes, and the computational complexity is high, making it difficult to meet real-time requirements. In addition, the visual method is highly dependent on feature points, and ambient light and dynamic changes will significantly affect its effectiveness.
[0028] Motion capture technology based on inertial sensors (IMU) has gradually attracted attention due to its portability, low cost and simple deployment. IMU can capture human motion data without relying on external devices and is suitable for dynamic and diverse scenarios. However, due to the drift error of IMU itself, its accuracy is limited when used alone, and it is difficult to provide global position information of the human body. Therefore, inertial sensors usually need to be combined with other technologies, such as visual positioning or static calibration, to improve the overall performance.
[0029] In the field of positioning technology, traditional methods usually rely on external positioning systems such as GPS, LiDAR, or multiple cameras. Although these methods can provide high-precision location information, they are less effective in indoor or occluded environments, and have high costs and deployment requirements. In recent years, visual inertial odometry (VIO) and simultaneous localization and mapping (SLAM) technologies have gradually become mainstream. They achieve autonomous positioning and scene modeling of devices by fusing data from cameras and IMUs. Although VIO and SLAM have advantages in flexibility and convenience, they still face the problems of cumulative errors and insufficient adaptability to dynamic environments.
[0030] In order to solve at least one of the above problems, the embodiment of the present application provides a human posture capture and human-scene interaction modeling system based on egocentric sensors (such as cameras and IMUs), which combines visual and inertial data for multimodal fusion. It can not only make up for the limitations of related technologies, but also has the characteristics of low cost, flexible deployment and strong environmental adaptability, so as to provide more practical solutions for fields such as robots, AR / VR and digital twins.
[0031] Figure 1 The structure of the human-machine integration system in the embodiment of the present application is shown in FIG. Figure 1 As shown, it may include: an augmented reality positioning module, a human posture capture module, and a data fusion module; wherein, The augmented reality positioning module is used to obtain real-time positioning information of the augmented reality device in the physical space through augmented reality technology.
[0032] In one embodiment, the augmented reality positioning module can be used to: identify high-precision position and posture in three-dimensional space through simultaneous localization and mapping (SLAM) provided by the augmented reality device in combination with multi-marker registration technology to obtain real-time positioning information of the augmented reality device in the physical space, and provide a reference benchmark for human posture capture and scene interaction.
[0033] The human posture capture module is used to capture the posture data of the human body based on inertial sensors, and constrain it through pre-set physical constraints to ensure the accuracy and naturalness of posture reconstruction.
[0034] In one embodiment, the human posture capture module can be used to: the motion data collected based on the IMU can be transmitted to a real-time 3D engine and development platform such as Unity3D or other visualization environments through a high-level programming language such as Python program to dynamically present the human posture; in an augmented reality device (such as HoloLens2), the posture of the virtual human body is synchronized with the real human body in real time, which means that the movements of the virtual human body are consistent with the movements of the real human body in time (almost no delay), and the posture reflects the posture of the real human body as accurately as possible. For example, when the user raises his right hand, the right hand of the virtual human model will immediately make the same hand-raising movement, which enhances the interactive experience. Here, Python is a high-level programming language that is widely used in scenarios such as data processing, algorithm development, and system integration; Unity is a popular real-time 3D engine and development platform that is widely used in game development, AR, VR, and 3D model visualization.
[0035] The data fusion module is used to combine the acquired real-time positioning information with the obtained posture data to realize the absolute position tracking and real-time correction of the human body, so as to ensure the consistency of the absolute position and motion posture of the human body in the virtual and actual scenes.
[0036] In one embodiment, the augmented display device may include, but is not limited to, a head-mounted augmented reality device (HMD for short).
[0037] In an exemplary embodiment, the augmented reality positioning module may include a positioning data acquisition submodule, a three-dimensional registration and alignment submodule, and a coordinate system alignment and correction submodule, wherein: The positioning data acquisition submodule is used to collect environmental data in real time through the visual sensor of the augmented reality device such as the head display, and generate the real-time positioning information of the augmented reality device in the physical space through the positioning function in the SLAM; A 3D registration and alignment submodule is used to spatially register multiple markers based on a perspective n-point problem (PnP) algorithm to establish alignment between the device coordinate system and the world coordinate system. In one embodiment, the multiple markers may include, but are not limited to, QR codes or marker points, and the multiple markers are used to enhance the accuracy of the alignment, thereby ensuring that the superposition effect of the virtual content and the real scene is accurate; The coordinate system alignment and correction submodule is used to combine multi-marker correction and spatial coordinate adjustment technology to dynamically update the mapping relationship between the coordinate system of augmented reality devices such as head-mounted displays and the world coordinate system to ensure that augmented reality devices such as head-mounted displays maintain consistency in spatial position and direction during dynamic interaction, thereby achieving stability and accuracy in device positioning.
[0038] In an embodiment of the present application, the augmented reality positioning module collects environmental data in real time through the visual sensor of an augmented reality device such as a head-mounted augmented reality device, uses its SLAM technology to achieve spatial positioning, and obtains real-time positioning data of the head display; combined with the multi-marker registration method based on the PnP algorithm, the device coordinate system and the world coordinate system are aligned and registered, thereby achieving high-precision spatial positioning and ensuring the precise position and posture of the head display in three-dimensional space.
[0039] In an exemplary embodiment, the human body posture capture module may include an inertial sensor acquisition submodule, a posture estimation submodule, and a physical constraint optimization submodule, wherein: The inertial sensor acquisition submodule is used to collect acceleration and angular velocity data in real time and capture human motion information through IMUs worn on key parts of the human body, such as 6 IMUs worn on major joints of the human body; The posture estimation submodule is used to calculate the posture data of the human body, including joint positions and motion states, based on the parameterized 3D human body model (SMPL, Skinned Multi-Person Linear) and deep learning algorithm, combined with the acceleration and angular velocity data collected in real time; The physical constraint optimization submodule is used to introduce one or any combination of the following physical constraints: joint angle limit, motion coherence, dynamic consistency, etc., to constrain the captured posture data to ensure the accuracy and naturalness of the captured posture. Furthermore, in one embodiment, the physical constraint optimization submodule can also be used to set ground contact and sliding detection constraints to optimize the stability of the human body's posture in the environment.
[0040] The range of motion of human joints is limited. For example, the knee joint cannot bend significantly in the opposite direction, and the rotation angle of the arm is also limited. In the embodiment of the present application, by setting the limit of the joint angle (such as the upper and lower limit range), it is possible to avoid the generated posture model from having erroneous postures that do not conform to the anatomical characteristics of the human body, prevent unnatural postures from appearing, and ensure that the virtual human body looks real and credible.
[0041] Human body movement is continuous and smooth, rather than sudden jumps or discontinuous changes. In the embodiment of the present application, the joint movement trajectory of the model is constrained by optimizing the algorithm to make it conform to the time continuity characteristics of human body movement, making the movement of the virtual human body smoother and more fluent, avoiding jitter or sudden displacement.
[0042] Human body movements need to comply with the laws of mechanics, such as gravity, inertia, friction, etc. In the embodiments of the present application, by introducing these dynamic constraints, it is possible to ensure that the movements of the virtual human body comply with physical logic (e.g., the natural expression of the force on the feet when running), enhance the realism of the virtual human body, especially in fast-moving or complex action scenes, to avoid looking distorted or contrary to common sense.
[0043] Accurately matching the posture of the virtual human body (digital model) with the user's posture in the real world, for example, when the user raises his arm, the virtual human body's arm is displayed synchronously at the same position and angle, which can enhance the immersive and interactive experience of augmented reality or virtual reality.
[0044] When the human body interacts with the ground, such as standing or walking, in the embodiment of the present application, it will also detect whether the feet are in correct contact with the ground, and avoid the model from "floating" or "sliding" (such as the soles of the feet leaving the ground or abnormal movement), so as to enhance the realism of the interaction between the virtual human body and the scene and make the movements look natural.
[0045] In the embodiment of the present application, the human posture capture module collects joint motion data in real time by wearing a small number of IMUs, such as 6, on key parts of the human body; and based on the SMPL human body model and deep learning algorithm, combined with the acceleration and angular velocity data provided by the IMU, the three-dimensional posture of the human body is solved in real time, including motion state estimation and joint position reconstruction. At the same time, physical constraints such as joint angle restrictions, motion coherence and dynamic consistency are introduced into the generated posture model to ensure the accuracy and naturalness of posture reconstruction, thereby achieving precise synchronization between the virtual human body and the actual scene. Furthermore, by visually displaying on the PC side, constraints such as ground contact and sliding detection are set at the same time.
[0046] In an exemplary embodiment, the data fusion module may include a mapping submodule and a dynamic correction submodule, wherein: The mapping submodule is used to establish the initial mapping relationship between the augmented reality device and the human skeleton through static calibration, such as the position transformation between the head and the pelvis. In one embodiment, the relationship between the head and the skeleton root node (such as the pelvis) can be statically calibrated, and the position change between the head and the pelvis can be dynamically updated through the human kinematics algorithm.
[0047] The dynamic correction submodule is used to dynamically update the position information of the skeleton root node (such as the pelvis) using the data collected by the IMU and nonlinear state estimation methods such as the extended Kalman filter (EKF) algorithm. In one embodiment, the IMU data, i.e., the motion data of the IMU sensor, is used as the observed quantity, and the real-time positioning data, i.e., the absolute positioning data provided by the augmented reality headset, is used as the predicted quantity (i.e., the reference quantity) to dynamically optimize the position information of the skeleton root node to ensure the continuity and accuracy of the posture and position. The real-time synchronization of the human body posture and spatial position is ensured by the fusion of the IMU data and the real-time positioning information in the embodiment of the present application.
[0048] The human-machine collaborative system provided in the embodiment of the present application is a human body posture capture and human-scene interaction modeling system. Through a small number of wearable sensors (IMU sensors and augmented reality head-mounted devices), it realizes the estimation of the position and motion posture of the human body in large 3D scenes, realizes high-precision spatial positioning, ensures the precise position and posture of the head display in three-dimensional space, and thus realizes stable tracking of the positioning system.
[0049] In an exemplary embodiment, the human-machine integration system provided in the embodiment of the present application may also include a human-scene dynamic interaction tracking module, which is used to: Augmented reality technology is used to realize the interaction between the virtual human body and the real scene, and the scene changes are dynamically recorded, thereby providing a dynamic update function for the interaction between the virtual human body and the environment, and achieving a high degree of synchronization between the virtual scene and the real scene.
[0050] In one embodiment, the human-scene dynamic interaction tracking module may include: The interactive behavior modeling submodule is used to add physical constraints and logic rules to simulate the physical behaviors in real interactions to capture the interactive behaviors between the human body and the scene, such as pushing objects, changing the state of objects, etc. In one embodiment, gravity, friction, elasticity, etc. can be added as physical constraints through the physics engine to simulate real interactive effects. In one embodiment, custom logic rules can be used to record information such as the force, path, and speed of an object.
[0051] The data recording submodule is used to record the state changes of objects (such as path, speed, force, etc.) during scene interaction in real time through custom scripts.
[0052] The virtual-reality matching submodule is used to synchronously update the interaction effects with the physical model in the virtual scene, providing data support for subsequent virtual scene analysis and modeling.
[0053] In the embodiment of the present application, augmented reality technology is used to provide dynamic update capabilities for the interaction between the captured virtual human body and the environment, and to record the changes in the scene (such as the movement or state change of objects) that occur with user interaction in real time. The human scene dynamic interaction tracking module simulates the physical behavior in the interaction process, such as gravity, elasticity, etc., by adding physical constraints and logical rules, and records information such as force, path, speed, etc. through custom scripts. In this way, the precise match of the interaction effect between the virtual scene and the real scene is ensured, and support is provided for subsequent analysis and virtual environment modeling.
[0054] Furthermore, the human-machine integration system provided in the embodiment of the present application uses augmented reality technology to assist in the dynamic interactive tracking of people and scenes, establishes a two-way real-time system for virtual and real interaction that supports high-precision data transmission, and truly realizes the seamless integration and interaction of the physical world and the virtual world.
[0055] The human-machine integration system provided in the embodiment of the present application is a deep combination of augmented reality, inertial sensors, data fusion and dynamic interaction technology: using augmented reality equipment and multi-marker registration to achieve high-precision alignment of virtual and real scenes; combining IMU sensors with deep learning algorithms to achieve accurate reconstruction of human posture; multi-sensor fusion ensures high consistency between the virtual human body and the real environment; and through physical behavior simulation and logical rules, enhances the interaction effect of virtual and real scenes. It achieves high-precision synchronization and dynamic interaction between the virtual human body and the actual scene, and is suitable for multiple fields such as virtual reality, motion capture, and human-computer interaction.
[0056] In one exemplary embodiment, in combination Figure 2 , the augmented reality device HoloLens2 can be used to realize real-time positioning of the head display and precise alignment with the virtual space, providing a basis for the superposition of virtual content and real scenes. In one embodiment, the Microsoft HoloLens2 augmented reality device, PC (64-bit Windows 10), Unity3D, VisualStudio, Mixed Reality Toolkit (MRTK) and Vuforia engine can be used to develop and implement the construction of the augmented reality positioning module. The goal of the augmented reality positioning module is to obtain the posture data of the HoloLens2 device in real time, and receive, process and send data through the PC to achieve alignment and registration of the virtual space and the real scene. The construction process of the augmented reality positioning module generally includes: Use Unity3D and Visual Studio to build a development environment, create an augmented reality application that supports HoloLens2, and combine the Mixed Reality Toolkit (MRTK) to simplify the development of HoloLens2. At the same time, use the Vuforia engine to identify and register markers in the real world. The application running on HoloLens2 collects the device's posture data (including position and orientation) in real time through the MRTK interface and sends it to the PC through a network connection. In one embodiment, in order to ensure the stability and efficiency of data transmission, the TCP / IP protocol is used to establish real-time communication between HoloLens2 and PC.
[0057] During the actual operation, you first need to initialize the system coordinates and set the current user's position as the origin. That is to say, set the user position when HoloLens is started as the coordinate origin to ensure the consistency between the virtual space and the real scene, which is convenient for subsequent three-dimensional registration and alignment. After initialization, start the application and enter the normal operation state. After successfully connecting to the server, enter the three-dimensional registration mode, and provide interactive operations through the user interface (UI). In the UI, the user can select and confirm the position of the markers one by one, and complete the virtual-real alignment through the multi-marker registration algorithm. In one embodiment, the UI interaction can be operated using classic virtual buttons and control panels, and the user can start the registration process by clicking a virtual button. Each marker is selected and confirmed through a virtual button to ensure that the position of the marker is accurately calibrated. During the three-dimensional registration process, the system uses a three-dimensional registration algorithm for multiple markers, such as Figure 3 As shown, the Vuforia engine identifies multiple markers in the real world and aligns their coordinates with those in the virtual space to achieve high-precision alignment between the virtual space and the real scene. The selection and confirmation of markers are performed one by one. The user can manually confirm the alignment position of each marker. After all markers are confirmed, the registration end mode is entered. After all markers are confirmed, the system automatically calculates and completes the alignment between the virtual space and the real scene.
[0058] Figure 3 Demonstrated alignment of virtual scenes with real scenes through augmented reality devices (such as HoloLens2) and recognition and registration of markers, including: First, confirm the core object, that is, obtain the first sub-object as the core object, and record its position and rotation information as a reference and basis for subsequent calculations. In one embodiment, the camera of an augmented reality device (such as HoloLens2) can be used to scan a marker (such as a QR code), select the first recognized marker as the core object, and record its position and rotation information before transformation as a reference and basis for subsequent transformations. The information (position and posture) of the core object will be used to guide the alignment of other sub-objects.
[0059] Next, select the sub-objects, that is, traverse and select other markers except the core object, and process them one by one. In one embodiment, for each sub-object, determine its position and rotation information relative to the core object. The markers can be identified by the Vuforia engine or similar tools and aligned with the core object.
[0060] Then, the rotation matrix and rotation angle are calculated, that is, the new position and posture of each sub-object are calculated, and the corresponding 4×4 matrix, namely Matrix4×4, is constructed. In one embodiment, the absolute position of the sub-object in the virtual space can be derived based on the position and posture information (position and rotation) of the core object; the rotation angle and transformation matrix are calculated to describe the spatial transformation relationship from the core object to the sub-object.
[0061] Finally, the main object is adjusted, that is, the overall coordinate system is adjusted to the position and posture of the core object to ensure the alignment of the virtual scene with the actual scene. In one embodiment, a rotation matrix and a translation matrix can be applied to adjust the positions of all sub-objects relative to the core object to ensure that the final virtual space layout is consistent with the actual scene.
[0062] Figure 3 The process shown achieves efficient three-dimensional registration of virtual space and real scene through the confirmation, posture calculation and position adjustment of markers one by one, providing a reliable spatial positioning basis for augmented reality applications.
[0063] In one exemplary embodiment, in combination Figure 2 , we can use the SMPL skeleton model and the motion capture algorithm based on sparse sensors to achieve real-time reconstruction of human posture in Unity, and combine physical constraints and inverse kinematics (IK) technology to ensure that the generated posture conforms to natural movement and physical rules, thereby realizing the construction of the human posture capture module. The goal of the human posture capture module is to achieve real-time 3D posture reconstruction of the human body based on the motion capture algorithm of the IMU sensor and the SMPL model, and ensure the naturalness and physical consistency of the posture through physical constraints. The construction process of the human posture capture module generally includes: First, the SMPL model is used to extract the human skeleton and perform kinematic analysis. The SMPL model is a parametric human model based on the bone structure. The SMPL model is used for skin binding to generate high-quality skeletons and postures, such as Figure 4 As shown. Migrate the SMPL skeleton model to Unity, and combine it with model development tools and plug-ins to ensure that it can be rendered and animated in the Unity environment. Use commercial sensors, such as the NoitomPN3 sensor, to collect motion data of key nodes of the human body (such as the joints where the IMU is located). The sensor data is transmitted through the receiver, and the intermediate data broadcast function is used to transmit the data to the PC. In one embodiment, in order to read the sensor data, a communication protocol interface such as C++ can be used. The interface is responsible for converting the raw data obtained from the sensor into usable joint position data to decode the raw sensor data into joint position data. For subsequent posture estimation and reconstruction, the algorithm code is written in Python and integrated through Unity's communication interface. See Figure 2 The data processing and visual debugging process shown in .
[0064] Then, a fast human posture estimation algorithm is applied based on the SMPL model to infer the positions of all joints of the body (such as Figure 4 As shown), a continuously differentiable and trainable model is used to inversely analyze the action, so that the human body posture can be reasonably estimated based on the joint motion data, and the obtained whole-body joint posture data does not include movement data. In other words, in the human body posture estimation algorithm based on the SMPL model and IMU, the derived joint posture data only represents the relative motion and posture of the human body joints (such as bending angle, rotation angle, direction, etc.), and does not include the global position change of the human body as a whole (that is, the overall displacement or movement path in space). In the embodiment of the present application, the posture estimation algorithm does not need to process complex overall motion data (movement trajectory), but efficiently focuses on the local motion between joints.
[0065] When acquiring joint data and driving the SMPL model to move, physical constraints are applied to ensure that the generated posture conforms to real physical rules, avoid joints from penetrating the ground, and maintain reasonable postures and angles; the inverse kinematics (IK) algorithm is used to ensure that the feet maintain contact with the ground during movement, that is, the feet always remain on the ground during movement to avoid penetrating the ground; further, constraints such as joint angle restrictions, motion coherence, and dynamic consistency can be introduced to ensure natural postures and smooth movements. After adjusting all joint data, in Unity, the SMPL model updates its posture in real time according to the incoming joint data to ensure that the movement of the virtual character conforms to actual physical constraints.
[0066] In an exemplary embodiment, after the construction of the augmented reality positioning module and the human body posture capture module is completed, each of them is initialized and calibrated, and combined with Figure 5 , which may include: In the augmented reality positioning module, the position of HoloLens when it is started can be used as the origin of the camera coordinate system , directly obtain the preliminary positioning data of the headset in the world coordinate system In the process of multi-marker 3D registration, the camera coordinate system is aligned with the external world coordinate system using the marker position as the alignment reference. When HoloLens is started, the human posture capture module is started to ensure that the user stands at the coordinate origin of the augmented reality system and maintains a static standard posture (such as T-shaped posture) for a period of time. The IMU defaults to its startup position as the initial position. The pelvic position in the SMPL model is usually used as the reference origin of the human skeleton, that is, the root node. The captured skeleton data is based on the relative position of the origin of the human coordinate system. During the system initialization phase, the user maintains a static posture, and the relative posture transformation matrix between the camera coordinate system and the human coordinate system is calibrated. Assume that the initial translation from the head to the pelvis is , ,in and is the rotation matrix of the head and pelvis at the time of initialization, then the initial pose matrix is: Then, after the calibration is completed, the global pose matrix of the human body root node in the world coordinate system is obtained as follows: .
[0067] In an exemplary instance, when the data fusion module executes the transfer of the global posture of the head to the root node (pelvic position), based on the real-time motion capture data of the human body and the acquired global positioning data, a low-pass filter is used to remove the local motion interference of the head, and the constraint relationship of the skeleton model is used to transfer the global posture of the head to the root node to ensure the stability and rationality of the root node position.
[0068] Since the human body posture estimation algorithm does not consider the rotation data reconstruction of the human body's end joints, that is, the augmented reality device is worn on the head, its local rotation cannot be synchronized with the head movement estimated by the posture. Therefore, in the embodiment of the present application, Figure 5As shown, the acquired global positioning data will be low-pass filtered to eliminate the influence of high-frequency local rotation gestures, such as quick nodding and shaking of the head, on the root node position calculation, and only retain the lower-frequency overall head movement information. In one embodiment, using a low-pass filter to smooth the head posture data may include: filtering the rotation part of the posture data using quaternion interpolation to avoid interpolation distortion caused by direct operation of Euler angles; using a standard one-dimensional low-pass filter for the translation part of the posture data.
[0069] For head position and the rotation quaternion Filter the head positioning data to remove local high-frequency movements (such as nodding and shaking the head). In three-dimensional space, the rotation quaternion It is a mathematical tool for representing rotation, often used to describe the orientation of an object. The filtering is implemented as follows: The translation part uses the position filtering formula: ,in, is the smoothing factor, is the filtering time constant, is the sampling time interval.
[0070] The rotation part is smoothed using quaternion spherical linear interpolation (Slerp), and the rotation filter formula is: ,in, is the quaternion spherical linear interpolation function, is the interpolation coefficient. By adjusting The filter cutoff frequency controls the degree to which high-frequency motion is suppressed.
[0071] The global pose of the camera after filtering is .
[0072] like Figure 5 As shown in the figure, when the human posture capture module is running, the real-time relative position of the head to the root node is derived based on the estimated joint data. That is, the human posture capture module will provide the rotation matrix of the head relative to the root node in real time. And the relative translation of the head to the root node that changes in real time , so the real-time relative posture relationship is obtained as follows: ; use The relative pose relationship and the filtered global pose of the camera , calculate the global pose of the root node for: .
[0073] In one exemplary embodiment, in combination Figure 5As shown in the figure, the data fusion module is dominated by visual positioning data and IMU is used as a loosely coupled auxiliary correction for short-term motion changes. The visual positioning data is the global pose obtained by data fusion after being transmitted from the head to the root node. The IMU correction module generates state estimates independently of the visual positioning data. It selects the raw outputs of the accelerometer and gyroscope of the sensor worn at the root node to calculate the short-term posture change, obtains the rotation posture update by integrating the angular velocity, and obtains the displacement update by integrating the acceleration twice (the influence of gravity needs to be removed). The output is the posture change in a relatively short time window. That is, the IMU is regarded as a short-term motion change. The correction only occurs at the final fusion level, so there is no need to deeply integrate the raw data (such as feature points or acceleration / angular velocity).
[0074] Taking the visual positioning data as the prediction input, the state prediction equation is constructed as: ,in, The global position data provided by vision is the current global state (position, posture) used as input to the prediction model, is the system noise. At each update, the current state is estimated visually .
[0075] The IMU-corrected observations, i.e., the acceleration and angular velocity data provided by the IMU, are used as the input of the observation model to correct the deviation of the visual prediction: ,in, is the measurement value of the IMU, is the observation model of the system, is the measurement noise. By observing the error , calculate the Kalman gain .
[0076] Figure 5 The Kalman filter data fusion shown in the figure includes: using the extended Kalman filter (EKF) framework to fuse visual predictions and IMU measurements. The global pose provided by the vision in the prediction stage is used as , the state formula after updating the fusion in the correction stage is: .
[0077] In an exemplary embodiment, the process of tracking dynamic interaction between people and scenes may include: Create a digital 3D model based on the real scene and the interactive objects in the scene, supplement the construction and development of the interactive functions in the scene and deploy it on hololens2. During operation, the camera of HoloLens is used to capture environmental information in real time and identify the interactive objects in the environment, calibrate and track the position, posture and state of these objects. When people interact with objects in the scene, the state of the object is dynamically updated, and the interactive data is transmitted to the Unity engine on the PC in real time, such as the movement, rotation, and state change of the object (such as the opening and closing of the door, the pushing and pulling of the chair). When the user operates the object (such as moving, rotating, pushing and pulling), in the virtual scene, based on the interactive data, the physical properties of the object (such as position, rotation angle, force, etc.) are updated, and the state changes of the virtual object are synchronized using Unity's physical engine to ensure that physical behaviors (such as gravity, friction, etc.) are reasonably simulated. In one embodiment, for objects with more complex interactive operations (such as sliding doors), predefined physical constraints can be used to ensure the naturalness and stability of the interaction process. During the dynamic interaction tracking of people and scenes, the action data of each interaction is recorded, including information such as object position, motion trajectory, force conditions, etc. At the same time, detailed interaction records are generated through scripts and algorithms to support later data analysis and playback. As well as using these data, user behavior analysis can be performed, or the interaction data can be used to optimize subsequent virtual environment design.
[0078] Although the embodiments disclosed in this application are as above, the contents described are only embodiments adopted to facilitate understanding of this application and are not intended to limit this application. Any technician in the field to which this application belongs can make any modifications and changes in the form and details of implementation without departing from the spirit and scope disclosed in this application, but the scope of patent protection of this application shall still be based on the scope defined in the attached claims.
Claims
1. A human-machine collaborative system, characterized in that: include: Augmented reality positioning module, human posture capture module, data fusion module; among them, The augmented reality positioning module is used to obtain the real-time positioning information of the augmented reality device in the physical space through augmented reality technology; The human posture capture module is used to capture the human posture data based on the inertial sensor IMU, and constrain it through pre-set physical constraints to ensure the accuracy and naturalness of posture reconstruction; The data fusion module is used to combine the acquired real-time positioning information with the obtained posture data to realize the absolute position tracking and real-time correction of the human body, so as to ensure the consistency of the absolute position and motion posture of the human body in the virtual and actual scenes.
2. The human-machine symbiosis system according to claim 1 further comprises a human-scene dynamic interaction tracking module, which is used to realize the interaction between the virtual human body and the real scene by using augmented reality technology and dynamically record the scene changes; including: Interactive behavior modeling submodule, data recording submodule, virtual-real matching submodule; among them, The interactive behavior modeling submodule is used to add physical constraints and logical rules to simulate the physical behaviors in real interactions to capture the interactive behaviors between the human body and the scene; The data recording submodule is used to record the state changes of objects during scene interaction in real time through custom scripts; The virtual-reality matching submodule is used to synchronize the interaction effects with the physical model in the virtual scene.
3. The human-machine collaborative system according to claim 1 or 2, wherein: The augmented reality positioning module is used for: By combining the simultaneous positioning and mapping SLAM provided by the augmented reality device with the multi-marker registration technology, the position and posture in the three-dimensional space are identified to obtain the real-time positioning information of the augmented reality device in the physical space.
4. The human-machine collaborative system according to claim 3, wherein: The augmented reality positioning module includes: a positioning data acquisition submodule, a three-dimensional registration and alignment submodule, and a coordinate system alignment and correction submodule, wherein: A positioning data acquisition submodule, used to collect environmental data in real time through the augmented reality device, and generate the real-time positioning information through the positioning function in the SLAM; A three-dimensional registration and alignment submodule, for spatially registering the multiple markers based on a perspective n-point problem (PnP) algorithm, and establishing alignment between a device coordinate system and a world coordinate system; The coordinate system alignment and correction submodule is used to combine the multi-marker correction and spatial coordinate adjustment technology to dynamically update the mapping relationship between the coordinate system of the augmented reality device and the world coordinate system to ensure that the augmented reality device maintains consistency in spatial position and direction during dynamic interaction.
5. The human-machine collaborative system according to claim 4, wherein: The multiple identifiers include: a QR code or a marking point.
6. The human-machine collaborative system according to claim 1 or 2, wherein: The human body posture capture module is used for: The motion data collected by the IMU is transmitted to a real-time 3D engine and a development platform through a high-level programming language to dynamically present the human body posture, so that the posture of the virtual human body in the augmented reality device is synchronized with the real human body in real time.
7. The human-machine collaborative system according to claim 6, wherein: The human body posture capture module includes: an inertial sensor acquisition submodule, a posture estimation submodule, and a physical constraint optimization submodule, wherein: The inertial sensor acquisition submodule is used to collect acceleration and angular velocity data in real time through an IMU worn on key parts of the human body to capture the movement information of the human body; The posture estimation submodule is used to calculate the posture data of the human body based on the parameterized three-dimensional human body model SMPL and the deep learning algorithm, and in combination with the acceleration and angular velocity data collected in real time; the posture data of the human body includes the joint position and the motion state; The physical constraint optimization submodule is used to introduce one or any combination of the following physical constraints: joint angle limit, motion coherence, and dynamic consistency, to constrain the captured posture data to ensure the accuracy and naturalness of the captured posture.
8. The human-machine collaborative system according to claim 7, wherein: The physical constraint optimization submodule may also be used to set ground contact and sliding detection constraints to optimize the stability of the human body's posture in the environment.
9. The human-machine collaborative system according to claim 1 or 2, wherein: The data fusion module includes a mapping submodule and a dynamic correction submodule; wherein, A mapping submodule, used for establishing an initial mapping relationship between the augmented reality device and the human skeleton through static calibration; The dynamic correction submodule is used to dynamically update the position information of the skeleton root node by using the data collected by the IMU and the nonlinear state estimation method.
10. The human-machine collaborative system according to claim 9, wherein: The nonlinear state estimation method is an extended Kalman filter (EKF) algorithm.
Citation Information
Patent Citations
Virtual reality system space positioning method based on depth perception
CN111007939A
Robot pose positioning method and system
CN111136660A
Fusion positioning system and method based on optical tracking and inertial tracking
CN111947650A
Scene fusion positioning system and method based on virtual reality technology
CN117274124A
Motion capture method and device based on AR glasses and electronic equipment
CN119068151A