Method for generating inertial measurement data from video

The method generates synthetic IMU data from video using virtual sensors on 3D body meshes, addressing data scarcity and enhancing model generalization and privacy in multimodal AI models.

DE102024208707A1Pending Publication Date: 2026-03-12ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024208707
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing multimodal AI models face data scarcity issues due to the lack of comprehensive multimodal datasets, particularly in IMU sensor recordings, which hinders their development and effectiveness.

Method used

A method to generate synthetic IMU data from video footage using virtual sensor placements on 3D body meshes, allowing for customizable and controlled data generation, including acceleration, angular velocity, and orientation measurements.

Benefits of technology

Enables cost-effective and scalable data creation, enhances model generalization, and improves data privacy by reducing reliance on real-world data collection, while providing flexible and tailored datasets for various applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method for generating synthetic inertial measurement unit (IMU) data comprises: accessing a video dataset containing a multitude of video frames depicting at least one human performing a physical activity; using a pose estimation model, extracting three-dimensional (3D) body mesh data from the multitude of video frames, where the 3D body mesh data represents a human posture in each of the multiple video frames; generating an animation of the 3D body mesh data, where the animation comprises a sequence of 3D body meshes representing the human movement; defining a multitude of virtual sensor positions on the 3D body mesh; and generating synthetic inertial data for each virtual sensor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for generating synthetic inertial measurement unit (IMU) data by analyzing human movement in video material, the virtual placement of sensors on a 3D model of the body and the calculation of sensor measurements based on the simulated movements. State of the art

[0002] Recent advances in multimodal AI foundation models (FM) have sparked a wave of innovation and excitement in the field. One such groundbreaking model is ImageBind [1], which ingeniously integrates data from five different modalities into a single embedding space. This integration enables the model to perform new functions, such as generating images from audio clips. Other notable multimodal foundation models, such as MetaTransformer [2] and mPLUG-2 [3], are also making significant strides by proposing different model architectures designed for modularized multimodal foundations. These models are particularly exciting because they go beyond traditional image and text processing to include other modalities such as IMUs (inertial measurement units), audio, and more.This expansion has the potential to revolutionize our interactions with machines and improve our engagement with the world around us in ways that were previously unimaginable.

[0003] A major challenge for multimodal startup models, however, is the lack of high-quality, multimodal data. While action videos often contain extensive video and audio recordings, they rarely include detailed text descriptions or captions, and even more rarely IMU sensor recordings. Conversely, datasets rich in IMU sensor signals typically offer only brief text descriptions and lack comprehensive data from other modalities. This discrepancy in data availability hinders the development and effectiveness of multimodal startup models.

[0004] Innovative solutions to this data scarcity are crucial for the continuous development and success of these models. By overcoming the challenges of data integration and expanding the availability of multimodal datasets, we can unlock the full potential of these innovative models and pave the way for a new era of machine interaction and perception.

[0005] Detecting and inferring human activity using sensors is a challenging task with broad applications in security, sports, and virtual reality. Collecting multimodal datasets of human activity is a complex and expensive undertaking that is generally highly application-specific. The data collected is typically insufficient to cover a wide range of scenarios for a generalizable business model.

[0006] Several existing studies address the generation of virtual IMU data, as in [4] and [5]. However, both studies have demonstrated significant gaps between synthetic and real IMU data for human activity detection applications. Both existing solutions rely on video / mocap solutions to generate 2D skeletal poses / 3D body meshes and lack flexibility in generating additional poses.

[0007] Another existing line of work involves synthetic body poses by prompting, as described in [6]. The entire pipeline is based solely on synthetic data. It is difficult to compare with reality.

[0008] It is an object of the present invention to establish a synthetic IMU data pipeline (also referred to as a virtual IMU) which has the flexibility to use both real and synthetic pose data and to generate IMU data at multiple body sites. [1] [2305.05665] ImageBind: An embedding space to bind them all (arxiv.org) [2] [2307.10802] Meta-Transformer: A unified framework for multimodal learning (arxiv.org) [3] [2302.00402] mPLUG-2: A modularized multimodal startup model using text, image and video (arxiv.org) [4] [2006.05675] IMUTube: Automatic extraction of virtual body accelerometers from video for the detection of human activities (arxiv.org) [5] [2202.10562] CROMOSim: A deep learning-based, cross-modality inertial measurement simulator (arxiv.org) [6] [PDF] IMUGPT 2.0: Language-based cross-modality transfer for sensor-based detection of human activity I Semantic Scholar Advantages of the invention

[0009] This invention offers several key advantages. First, it directly addresses the challenge of data scarcity by enabling the creation of synthetic IMU data from readily available video footage. This eliminates the need for expensive and time-consuming real-world data acquisition, making the process significantly more cost-effective and scalable. Furthermore, the method allows for customization and control over the generated data. Users can define precise sensor placements, simulate specific movement types, and even adjust body shapes to generate tailored datasets for their respective applications.

[0010] Beyond the advantages of synthetic data, its use can lead to improved generalization in trained models. By augmenting training datasets with both real and synthetic IMU data, models can learn more robust representations of human movement and potentially improve their performance on unseen real-world data. This approach also enhances data privacy by reducing reliance on collecting sensitive real-world movement data from individuals. Disclosure of the invention

[0011] In a first aspect of the invention, a method for generating synthetic inertial measurement unit (IMU) data is proposed. The method begins with a step of accessing a video dataset comprising a plurality of video frames depicting at least one person performing a physical activity. The next step involves using a pose estimation model that extracts three-dimensional (3D) body mesh data from the plurality of video frames. Preferably, a method based on the Skinned Multi-Person Linear Model (SMPL) is used for pose estimation; see, for example, https: / / smpl.is.tue.mpg. / . The 3D body mesh data represents a human posture in each of the plurality of video frames. The next step involves generating an animation of the 3D body mesh data, wherein the animation comprises a sequence of 3D body meshes representing the human movement.

[0012] The next step involves defining multiple virtual sensor positions on the 3D body mesh, with each virtual sensor location corresponding to a position on the body. A virtual sensor can be understood as a software-based system that estimates or predicts physical quantities using mathematical models, algorithms, and data from other sensors, rather than directly measuring the quantities with physical sensors. The virtual sensor can be an IMU sensor. Preferably, the virtual sensor measures the following quantities: acceleration and / or angular velocity. It should be noted that IMUs are generally often combined with other sensors, such as GPS or magnetometers, in a process called sensor fusion to obtain accurate position and orientation data. Therefore, the virtual sensor could also be configured to output sensor measurements from GPS or magnetometers.It should also be noted that the virtual sensor can output measurements that characterize the orientation, position, and movement of an object.

[0013] The next step is optional and involves tracking the movement of the multiple virtual sensor locations throughout the animation.

[0014] The next step involves generating synthetic inertial measurement data for each virtual sensor location based on the tracked motion. The synthetic inertial measurement data includes at least one of the following: linear acceleration, angular velocity, or orientation.

[0015] Exemplary embodiments of the invention are explained in more detail with reference to the following figures. The figures show: Fig. a schematic representation of a method, a device, a storage medium according to embodiments of the invention. Description of the embodiments

[0016] Fig. shows a flowchart of a method (10) for generating synthetic inertial measurement data.

[0017] The procedure begins with a step of accessing (S21) a video dataset comprising a multitude of video frames depicting at least one person performing a physical activity. In other words, the dataset can include human subjects performing a wide range of activities.

[0018] The following step involves (S22) the use of a pose estimation model that extracts three-dimensional (3D) body mesh data from the multitude of video frames, where the 3D body mesh data represents a human posture in each of the multitude of video frames. Existing pose recognition models exist that fit 3D body meshes of the person for each frame of each video and create 3D mesh animations.

[0019] The next step (S23) involves generating an animation of the 3D body mesh data, where the animation comprises a sequence of 3D body meshes representing human movement. The generated body mesh animations can either be used directly for IMU data generation or used to generate additional synthetic animations with available standard solutions. Furthermore, the extracted body meshes can be further modified to incorporate additional body shapes, either for the original activities shown in the input videos or for the additional synthetic animations.

[0020] The generation of additional synthetic animations by prompting, such as the generation of pose animations via text, is known, as described in Jiang, Biao, et al., “Motiongpt: Human Movement as a Foreign Language.” Advances in Neural Information Processing Systems 36 (2024). Descriptive texts, such as “show me a person doing kickboxing,” are used as prompts for the model to generate additional poses with extracted / modified 3D body meshes.

[0021] The next step involves defining (S24) a plurality of virtual sensor locations on the 3D body mesh, where each virtual sensor location corresponds to a position on the body. More specifically, subsets of vertices on the body mesh are selected as locations for the virtual sensors, such as head, wrist, ankle, chest, etc. Additionally, information about the bone orientation and joint configuration of the 3D model can also be used to anchor the virtual sensors for greater consistency of sensor position relative to the body.

[0022] This is followed by a step of optional tracking (S25) movement of the multitude of virtual sensor locations during the animation.

[0023] The next step involves generating (S26) synthetic inertial measurement data for each virtual sensor location based on the tracked motion. Generate IMU data, including acceleration, angular velocity, orientation, linear velocity, etc., for each virtual sensor, based on the sensor movement during each animation, generated / extracted from the previous steps. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited non-patent literature

[0000] Jiang, Biao, et al. described. “Motiongpt: Human Movement as a Foreign Language.” Advances in Neural Information Processing Systems 36 (2024

[0020]

Claims

[1] A computer-implemented method (10) for generating synthetic inertial measurement unit data, the method comprising: Access to (S21) a video dataset comprising a multitude of video images depicting at least one human being performing a physical activity; Using (S22) a pose estimation model, three-dimensional (3D) body mesh data are extracted from the multitude of video images, with the 3D body mesh data representing a human body posture in each of the multitude of video images; Generating (S23) an animation of the 3D body mesh data, wherein the animation comprises a sequence of 3D body meshes representing the movement of the human; Define (S24) a plurality of virtual sensor positions on the 3D body mesh, where each virtual sensor position corresponds to a position on the body; and Generating (S26) synthetic inertial data for each virtual sensor position. [2] Method according to claim 1, further comprising: generating additional animations of the 3D body mesh data based on at least one text prompt describing a desired physical activity. [3] Method according to claim 1 or 2, wherein the definition of the plurality of virtual sensor locations comprises: using at least one of the bone orientation or joint configuration data from the 3D body mesh data to determine the placement of the virtual sensor locations. [4] Method according to one of claims 1-3, further comprising: adapting a body shape of the 3D body mesh data to generate the animation. [5] Method according to any one of claims 1-4, further comprising: adding the generated synthetic inertial unit data to a training dataset for training artificial intelligence models. [6] Computer program configured to cause a computer to execute the method according to any one of claims 1 to 5, including all its steps, when the computer program is executed by a processor. [7] Machine-readable storage medium on which the computer program according to claim 6 is stored [8] System for carrying out a method according to any one of claims 1-5.