Human body action data set construction method and electronic equipment

By automatically filtering and processing multi-source video data from the Internet, a high-quality human motion dataset is generated, solving the problems of insufficient data and inaccurate estimation, and realizing the construction of an efficient and accurate human motion dataset.

CN120932039APending Publication Date: 2025-11-11BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511052989.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies have limited data collection for human motion datasets, unmet demand for large multimodal model data, lack of standardized automated processing procedures, and inaccurate human pose estimation under noise and occlusion conditions.

Method used

Human motion datasets are generated by screening, processing, and generating data from multi-source video data on the Internet through an automated process. Tools such as YOLO model, CLIFF, PIXIE, MeTRAbs, and RoHM are used for human detection and motion feature extraction to generate high-quality motion description labels and optimize the dataset.

Benefits of technology

It improves the efficiency and quality of human motion dataset generation, ensures the accuracy and robustness of the dataset, and is suitable for a variety of computer vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932039A_ABST
    Figure CN120932039A_ABST
Patent Text Reader

Abstract

The invention discloses a human motion data set construction method and electronic equipment, and belongs to the technical field of video processing, and the method comprises the steps: automatically obtaining multi-source original video data, and screening out the original video data meeting a first condition as target video data; performing human body detection and key point extraction based on the target video data to obtain human body action feature data; generating a human body action description label based on the human body action feature data; generating a human body action data set based on the human body action feature data and the human body action description label; and optimizing the human body action data set to obtain a target human body action data set. According to the method, the high-quality human body action data set is extracted, screened, processed and generated from various video data of the Internet through an automatic process, and the generation efficiency of the human body action data set and the quality of the data set are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of video processing technology, and specifically relates to a method for constructing a human motion dataset and an electronic device. Background Technology

[0002] Human motion datasets are the cornerstone of research and development of human behavior recognition technologies based on human motion data. However, the following challenges remain in the creation of human motion datasets and automated video data processing: 1) Current methods typically extract motion from standard, deliberately collected datasets. This type of data is limited in quantity and cannot support the current data requirements for multimodal large-scale models in human motion generation. 2) There is no standard automated processing workflow that can handle datasets from all sources and transform them into usable text2motion datasets. 3) Accurately estimating human pose remains a challenge, especially in the presence of noise and occlusion, particularly in monocular RGB(-D) videos.

[0003] To address the aforementioned issues, this application proposes a method for constructing a human motion dataset and an electronic device. Summary of the Invention

[0004] To address the problems existing in the prior art, this application provides a method for constructing a human motion dataset and an electronic device to solve the technical problems of insufficient data collection, lack of variety in the types of data, lack of automation in the processing flow, and inaccurate estimation of human motion in the prior art.

[0005] The technical effect to be achieved in this application is accomplished through the following solution:

[0006] Firstly, this application provides a method for constructing a human motion dataset, including:

[0007] Automatically acquire raw video data from multiple sources, and filter out the raw video data that meets the first condition as the target video data:

[0008] Human body detection and key point extraction are performed based on the target video data to obtain human motion feature data;

[0009] Generate human motion description tags based on the aforementioned human motion feature data;

[0010] A human motion dataset is generated based on the human motion feature data and human motion description labels.

[0011] Optimize the human motion dataset to obtain the target human motion dataset.

[0012] In some embodiments, the raw video data that meets the first condition includes at least one of the following:

[0013] Each moment includes raw video data of at least one human body;

[0014] The original video data where the occluded portion of the human body is less than the occlusion threshold;

[0015] Raw video data in which the human body occupies a larger area than a certain threshold;

[0016] Raw video data whose length exceeds a length threshold;

[0017] Raw video data with a frame rate greater than the frame rate threshold;

[0018] Raw video data with the camera stationary;

[0019] Raw video data excluding shot transitions.

[0020] In some embodiments, the YOLO model is used to filter the raw video data.

[0021] In some embodiments, when the raw video data for the first condition includes raw video data where the camera position has not moved, the raw video data where the camera position has not moved is filtered out as follows:

[0022] N frames are uniformly extracted from the original video data, where N is an integer greater than 1;

[0023] And calculate the optical flow based on the N frames;

[0024] Set an optical flow threshold, and determine the original video data whose mean optical flow is less than the optical flow threshold as the original video data where the camera position has not moved.

[0025] In some embodiments, human detection and key point extraction are performed based on the target video data to obtain human motion feature data, including:

[0026] Use CLIFF or PIXIE to detect human bodies from target video data and extract 3D human body model data;

[0027] MeTRAbs was used to determine 3D human pose and extract skeleton data from target video data;

[0028] 3D human body model data and skeleton data are used as human motion feature data.

[0029] In some embodiments, generating human motion description tags based on the human motion feature data includes:

[0030] Analyze human motion feature data using deep learning models to generate human motion description labels;

[0031] The human motion description tags include motion type, body part, and occlusion status.

[0032] In some embodiments, RoHM is used to optimize the human motion dataset.

[0033] In some embodiments, after optimizing the human motion dataset to obtain the target human motion dataset, the method further includes:

[0034] Publish the target human motion dataset.

[0035] Secondly, this application provides an electronic device, the electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the aforementioned methods.

[0036] Thirdly, this application provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement any of the foregoing methods.

[0037] This application provides a method and electronic device for constructing human motion datasets. The method extracts, filters, processes and generates high-quality human motion datasets from various types of video data on the Internet through an automated process, thereby improving the generation efficiency and quality of human motion datasets. Attached Figure Description

[0038] To more clearly illustrate the embodiments of this application or the existing technical solutions, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a flowchart of a method for constructing a human motion dataset according to an embodiment of this application;

[0040] Figure 2 This is a schematic block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in one or more embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0043] Significant progress has been made in several technologies related to human motion datasets and automated video data processing. Here are some key current technological advancements:

[0044] Multimodal learning: Multimodal learning improves task performance by combining data from different sources (such as vision and speech). Innovative methods such as Orthogonal Order Fusion (OSF) and the BalanceMLA framework address the problems of modality imbalance and inconsistent learning rates by dynamically balancing the contributions of different modalities and optimizing the learning process.

[0045] Video analytics and real-time processing: Video analytics technology has entered the era of big data and artificial intelligence. Utilizing deep learning technology for video analysis can achieve higher accuracy and real-time performance. However, computational cost and algorithm complexity remain challenges.

[0046] Human pose estimation: HuMoR et al. demonstrated strong robustness and the ability to handle occlusion problems through video pose estimation.

[0047] Large-scale human motion datasets, such as the Motion-X dataset, provide large-scale 3D full-body motion data, covering diverse scenes and motion sequences, and offer rich data resources for human motion generation research.

[0048] Deep learning-based action recognition: Deep learning methods such as Temporal Convolutional Networks (TCN) and Transformer-based models perform well in action recognition, but the complexity of actions, background and occlusion, and temporal dependencies remain challenges.

[0049] RoHM (Robust Human Motion Reconstruction) is a method for robust 3D human motion reconstruction in the presence of noise and occlusion. It utilizes the iterative and denoising characteristics of a diffusion model to reconstruct complete and reasonable motion. Various non-limiting embodiments of this application are described in detail below with reference to the accompanying drawings.

[0050] Despite some progress in existing technologies, many problems remain in the creation of human motion datasets and automated video data processing. Therefore, the human motion dataset construction method described in this application is needed to create human motion datasets.

[0051] First, refer to Figure 1 The method for constructing the human motion dataset in this application is described in detail below:

[0052] S1: Automatically acquire raw video data from multiple sources and filter out the raw video data that meets the first condition as the target video data;

[0053] S2: Perform human body detection and key point extraction based on the target video data to obtain human motion feature data;

[0054] S3: Generate human motion description tags based on the human motion feature data;

[0055] S4: Generate a human motion dataset based on the human motion feature data and human motion description labels;

[0056] S5: Optimize the human motion dataset to obtain the target human motion dataset.

[0057] The above method extracts, filters, processes, and generates high-quality human motion datasets from various types of video data on the Internet through an automated process, thereby improving the efficiency of human motion dataset generation and the quality of the datasets.

[0058] For example, automatically acquiring raw video data from multiple sources includes using automated tools to extract video data from the Internet, which includes videos containing human motion from various samples from multiple sources; for example, selecting Internet video data sources such as Kinetics-TPS Dataset and InternVid, and downloading and transmitting the selected data sources by developing automated scripts.

[0059] In some embodiments, the raw video data that meets the first condition includes at least one of the following:

[0060] Each moment includes raw video data of at least one human body;

[0061] The original video data where the occluded portion of the human body is less than the occlusion threshold;

[0062] Raw video data in which the human body occupies a larger area than a certain threshold;

[0063] Raw video data whose length exceeds a length threshold;

[0064] Raw video data with a frame rate greater than the frame rate threshold;

[0065] Raw video data with the camera stationary;

[0066] Raw video data excluding shot transitions.

[0067] The first condition for video length can also be replaced with raw video data with a video length between 20s and 30s.

[0068] For example, the above-mentioned multiple original video data that meet the first condition can be arbitrarily combined. For instance, the original video data that meet the first condition include:

[0069] Raw video data with the camera position unchanged, and raw video data excluding lens transitions.

[0070] The above is an example only, and all combinations of the first conditions are within the scope of protection of this application, and will not be listed here one by one.

[0071] For example, the occlusion threshold, area threshold, length threshold, and frame rate threshold can be set according to the actual situation. For example, the length threshold can be set to 20 seconds, the occlusion threshold can be set to 30%, the area threshold can be set to 60%, and the frame rate threshold can be set to 10 frames per second. These are just examples and are not limited to this.

[0072] In some embodiments, the YOLO model is used to filter the raw video data.

[0073] In some embodiments, when the raw video data for the first condition includes raw video data where the camera position has not moved, the raw video data where the camera position has not moved is filtered out as follows:

[0074] N frames are uniformly extracted from the original video data, where N is an integer greater than 1, such as N being 150, 300, etc.

[0075] And calculate the optical flow based on the N frames;

[0076] Set an optical flow threshold to identify raw video data with an average optical flow value less than the threshold as raw video data where the camera position has not moved. For example, the optical flow threshold could be 1, but this is just an example and can be adjusted according to the actual situation.

[0077] In some embodiments, human detection and key point extraction are performed based on the target video data to obtain human motion feature data, including:

[0078] Use CLIFF or PIXIE to detect human bodies from target video data and extract 3D human body model data;

[0079] MeTRAbs was used to determine 3D human pose and extract skeleton data from target video data;

[0080] 3D human body model data and skeleton data are used as human motion feature data.

[0081] In some embodiments, generating human motion description tags based on the human motion feature data includes:

[0082] Analyze human motion feature data using deep learning models to generate human motion description labels;

[0083] The human motion description tags include motion type, body part, and occlusion status.

[0084] For example, the action type can be an overview of the human body's current behavior (e.g., walking, sitting, interacting, etc.); body parts can include the upper body, lower body, head, etc.; occlusion situations can include the human body being occluded by 10%, or the lower body of a specific human body being occluded by 5%, etc.

[0085] For example, the description of each body part can be further specified. For the upper body: their arms hang loosely at their sides, shoulders are slightly back, and chest is open. The torso is upright with minimal movement, indicating a calm and neutral posture. Lower body: feet are firmly planted on the ground, shoulder-width apart. Knees are slightly bent, and weight is evenly distributed between the legs.

[0086] In some embodiments, RoHM is used to optimize the human motion dataset; specifically including:

[0087] Use tools such as RoHM to correct for noisy or occluded human motion data.

[0088] In some embodiments, after optimizing the human motion dataset to obtain the target human motion dataset, the method further includes:

[0089] The target human motion dataset will be published to make it a resource for research and application.

[0090] The method for constructing a human motion dataset and the electronic device described in this application have the following advantages:

[0091] 1) High efficiency: Automated data processing greatly improves the efficiency of creating large-scale human motion datasets.

[0092] 2) Accuracy: Deep learning models and advanced human motion analysis tools ensure high accuracy of the dataset.

[0093] 3) Scalability: This method supports data extraction from different data sources and is easy to extend and adapt to different application scenarios.

[0094] 4) High value: The generated dataset can be used for various computer vision tasks such as action recognition and behavior analysis, and has high research value.

[0095] 5) Robustness: It can process videos from any source.

[0096] It should be noted that the methods of one or more embodiments of this application can be executed by a single device, such as a computer or server. The methods of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to complete the process. In such a distributed scenario, one of these devices may execute only one or more steps of the methods of one or more embodiments of this application, and the multiple devices will interact with each other to complete the method described.

[0097] It should be noted that the above description describes specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0098] Based on the same inventive concept, and corresponding to any of the above embodiments, this application also discloses an electronic device.

[0099] Specifically, Figure 2 The diagram illustrates the hardware structure of an electronic device for constructing a human motion dataset according to this embodiment. The device may include a processor 410, a memory 420, an input / output interface 430, a communication interface 440, and a bus 450. The processor 410, memory 420, input / output interface 430, and communication interface 440 are interconnected internally via the bus 450.

[0100] The processor 410 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0101] The memory 420 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 420 can store the operating system and other applications. When the technical solutions provided in the embodiments of this application are implemented by software or firmware, the relevant program code is stored in the memory 420 and is called and executed by the processor 410.

[0102] Input / output interface 430 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0103] The communication interface 440 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (e.g., USB, Ethernet cable, etc.) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth, etc.).

[0104] Bus 450 includes a pathway for transmitting information between various components of the device (e.g., processor 410, memory 420, input / output interface 430, and communication interface 440).

[0105] It should be noted that although the above-described device only shows the processor 410, memory 420, input / output interface 430, communication interface 440, and bus 450, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this application, and not necessarily all the components shown in the figures.

[0106] The electronic devices described above are used to implement the corresponding human motion dataset construction method in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0107] Based on the same inventive concept, corresponding to any of the above embodiments, one or more embodiments of this application also provide a computer-readable storage medium storing computer instructions for causing the computer to execute the human motion dataset construction method as described in any of the above embodiments.

[0108] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0109] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the human motion dataset construction method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0110] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0111] Additionally, to simplify the description and discussion, and to avoid obscuring one or more embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring one or more embodiments of this application, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which one or more embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) are set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that one or more embodiments of this application may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0112] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0113] One or more embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this application should be included within the protection scope of this application.

Claims

1. A method for constructing a human motion dataset, characterized in that, include: Automatically acquire raw video data from multiple sources, and filter out the raw video data that meets the first condition as the target video data. Human body detection and key point extraction are performed based on the target video data to obtain human motion feature data. Generate human motion description tags based on the aforementioned human motion feature data; A human motion dataset is generated based on the human motion feature data and human motion description labels. Optimize the human motion dataset to obtain the target human motion dataset.

2. The method for constructing a human motion dataset as described in claim 1, characterized in that, The raw video data that meets the first condition includes at least one of the following: Each moment includes raw video data of at least one human body. The original video data where the occluded portion of the human body is less than the occlusion threshold; Raw video data in which the human body occupies a larger area than a certain threshold; Raw video data whose length exceeds a length threshold; Raw video data with a frame rate greater than the frame rate threshold; Raw video data with the camera stationary; Raw video data excluding shot transitions.

3. The method for constructing a human motion dataset as described in claim 1 or 2, characterized in that, Use the YOLO model to filter the raw video data.

4. The method for constructing a human motion dataset as described in claim 3, characterized in that, If the raw video data for the first condition includes raw video data where the camera position has not moved, the raw video data where the camera position has not moved is filtered out using the following method: N frames are uniformly extracted from the original video data, where N is an integer greater than 1; And calculate the optical flow based on the N frames; Set an optical flow threshold, and determine the original video data whose mean optical flow is less than the optical flow threshold as the original video data where the camera position has not moved.

5. The method for constructing a human motion dataset as described in claim 4, characterized in that, Human body detection and key point extraction are performed based on the target video data to obtain human motion feature data, including: Use CLIFF or PIXIE to detect human bodies from target video data and extract 3D human body model data; MeTRAbs was used to determine 3D human pose and extract skeleton data from target video data; 3D human body model data and skeleton data are used as human motion feature data.

6. The method for constructing a human motion dataset as described in claim 5, characterized in that, Generate human motion description tags based on the aforementioned human motion feature data, including: Analyze human motion feature data using deep learning models to generate human motion description labels; The human motion description tags include motion type, body part, and occlusion status.

7. The method for constructing a human motion dataset as described in claim 6, characterized in that, The human motion dataset was optimized using RoHM.

8. The method for constructing a human motion dataset as described in claim 7, characterized in that, After optimizing the human motion dataset to obtain the target human motion dataset, the process further includes: Publish the target human motion dataset.

9. An electronic device, the electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs that can be executed by one or more processors to implement the method as described in any one of claims 1 to 8.