Virtual IMU self-supervision human body activity identification method, system and device and medium

By using a virtual IMU self-supervised human activity recognition method, diverse training samples are generated and a pre-set task classifier is used to overcome the limitations of virtual action sequences in terms of expressive style and individual differences, thereby improving the accuracy and robustness of action recognition and reducing data collection costs.

CN121167518APending Publication Date: 2025-12-19GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511083022.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing technologies that synthesize virtual action sequences based on a single data generation mechanism have limitations in terms of expressive style, action details, and individual differences. They are difficult to cover the diverse manifestations of the same action under different performers and contextual conditions, thus limiting the model's generalization ability in complex real-world environments.

Method used

The virtual IMU self-supervised human activity recognition method generates three-dimensional action sequences using a preset language model and action generation mechanism. Combined with virtual IMU calculation, a training sample set is obtained to extract encoder parameters, and a preset task classifier is used to determine the action category, avoiding excessive reliance on actual sensor data and generating diverse training samples.

Benefits of technology

It improves the semantic coverage of virtual data, enhances the accuracy and robustness of the model in action category discrimination during comparative learning, reduces data collection costs, and generates virtual data that better matches the characteristics of actual actions, thereby improving the accuracy of action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167518A_ABST
    Figure CN121167518A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a virtual IMU self-supervised human body activity identification method, system and device and a medium. The method comprises the steps of obtaining a to-be-identified action category; processing according to the to-be-recognized action category and a preset language model to obtain a natural language action description; generating a three-dimensional action sequence according to the natural language action description and a preset action generation mechanism; inputting the three-dimensional action sequence into a target virtual IMU (Inertial Measurement Unit) for calculation to obtain virtual IMU data; obtaining a training sample set according to the virtual IMU data to extract a preset encoder to obtain encoder parameters; and determining the type of the action to be identified according to the encoder parameters and a preset task classifier. According to the embodiment of the invention, a large number of diversified training samples can be generated, and the generated virtual data better fits the actual action characteristics, so that the quality of the training samples is improved, the to-be-recognized action category can be accurately recognized, and the accuracy and robustness of action recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, system, device and medium for identifying human activities using a virtual IMU (Integrated Virtual Machine). Background Technology

[0002] In related technologies, virtual action sequences synthesized based on a single data generation mechanism (such as 3D skeletal animation or video pose estimation) are limited in terms of expressive style, action details and individual differences. They are difficult to cover the diverse manifestations of the same action under different performers and different contexts, thus limiting the model's generalization ability in complex real-world environments. Summary of the Invention

[0003] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method, system, device, and medium for the recognition of human activities using a virtual IMU (In-Musical Unit), aiming to improve the semantic coverage of virtual data and enhance the accuracy of the model in action category discrimination during contrastive learning.

[0004] In a first aspect, embodiments of this application provide a method for identifying self-supervised human activity using a virtual IMU, including: Obtain the action category to be identified; The natural language action description is obtained by processing the action category to be identified and the preset language model. Based on the natural language action description and the preset action generation mechanism, a three-dimensional action sequence is generated; The three-dimensional action sequence is input into the target virtual IMU for calculation to obtain virtual IMU data; The training sample set is obtained based on the virtual IMU data to extract the preset encoder parameters. The action category to be identified is determined based on the encoder parameters and the preset task classifier.

[0005] According to some embodiments of this application, the target virtual IMU is obtained through the following steps: Get the target action category name; The target action category name and the preset language model are processed to obtain a natural language action description; A first action sequence is obtained based on the natural language action description and the preset first action generation mechanism, wherein the preset first action generation mechanism is obtained through a preset VAE model; A second action sequence is obtained based on the natural language action description and the preset second action generation mechanism, wherein the preset second action generation mechanism is obtained through a preset diffusion model; Virtual IMU data is extracted using the first action sequence and the second action sequence; Training samples are generated based on the virtual IMU data, the first action sequence, and the second action sequence; The initial virtual IMU is trained based on the training samples to obtain the target virtual IMU.

[0006] According to some embodiments of this application, the step of obtaining a training sample set based on the virtual IMU data to extract encoder parameters from a preset encoder includes: A training sample set is obtained based on the virtual IMU data, wherein the training samples include positive sample pairs and negative sample pairs; The training samples are input into the convolutional layer of a preset encoder for training to obtain a global feature representation, wherein the preset encoder includes multiple convolutional layers; The global feature representation is input into a preset encoder to obtain the contrast embedding space; By comparing the embedding space, the positive sample pairs, and the negative sample pairs, the similarity of the positive sample pairs and the similarity of the negative sample pairs are obtained. The encoder parameters are obtained based on the similarity between the positive sample pairs and the negative sample pairs.

[0007] According to some embodiments of this application, determining the action category to be identified based on the encoder parameters and a preset task classifier includes: The initial number of action categories is obtained based on the encoder parameters and the preset task classifier; The initial number of action categories is updated based on the training sample set to obtain the target number of action categories; The action category to be identified is determined based on the number of target action categories.

[0008] According to some embodiments of this application, when the three-dimensional action sequence is a multi-segment initial action sequence of first duration, the step of inputting the three-dimensional action sequence into a target virtual IMU for calculation to obtain virtual IMU data includes: The three-dimensional action sequence is input into the target virtual IMU for calculation to obtain multiple three-dimensional position changes; Time-continuous acceleration data are obtained based on multiple changes in the three-dimensional position; The time-continuous acceleration data is denoised to obtain virtual IMU data.

[0009] According to some embodiments of this application, the preset action generation mechanism includes: An algorithm for generating action semantics based on the natural language action description; And / or, an action template generation model based on the natural language action description.

[0010] According to some embodiments of this application, obtaining the training sample set based on the virtual IMU data includes: The virtual IMU data is input into the first action sequence and the second action sequence to obtain positive sample pairs and negative sample pairs. The positive sample pairs represent sample pairs of data instances that are semantically consistent but have different generation styles, and the negative sample pairs represent sample pairs that differ in action category or semantics. Based on the positive sample pairs and the negative sample pairs, a training sample set is obtained.

[0011] Secondly, embodiments of this application provide a virtual IMU self-monitored human activity recognition system, characterized in that it includes: The acquisition module is used to acquire the action category to be identified; The processing module is used to process the action category to be identified and the preset language model to obtain a natural language action description; The generation module is used to generate a three-dimensional action sequence based on the natural language action description and the preset action generation mechanism; The calculation module is used to input the three-dimensional action sequence into the target virtual IMU for calculation to obtain virtual IMU data; The extraction module is used to obtain training samples based on the virtual IMU data to extract the preset encoder parameters. The determination module is used to determine the action category to be identified based on the encoder parameters and a preset task classifier.

[0012] Thirdly, embodiments of this application provide a computer device, including: Memory, used to store programs; A processor for executing a program stored in the memory, wherein when the processor executes the program stored in the memory, the processor is configured to perform the method described in the first aspect above. Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions for performing the method described in the first aspect above.

[0013] The technical solution according to the embodiments of this application has at least the following beneficial effects: acquiring the action category to be identified; processing the action category to be identified and a preset language model to obtain a natural language action description; generating a three-dimensional action sequence based on the natural language action description and a preset action generation mechanism; inputting the three-dimensional action sequence into a target virtual IMU for calculation to obtain virtual IMU data; acquiring a training sample set based on the virtual IMU data to extract encoder parameters from a preset encoder; and determining the action category to be identified based on the encoder parameters and a preset task classifier. This application embodiment avoids excessive reliance on actual sensor data by generating virtual IMU data as training samples, reducing data acquisition costs, and simultaneously generating a large number of diverse training samples, solving the problem of insufficient data. Utilizing a preset language model and action generation mechanism, action categories can be converted into detailed three-dimensional action sequences, and then corresponding data can be generated through a virtual IMU, making the generated virtual data more closely match actual action characteristics and improving the quality of training samples. Through training the preset encoder and applying the preset task classifier, the action category to be identified can be accurately identified, improving the accuracy and robustness of action recognition.

[0014] The solutions provided in the second to fourth aspects above are used to implement or cooperate with the methods provided in the first aspect above, and therefore can achieve the same or corresponding beneficial effects as the first aspect, which will not be elaborated here.

[0015] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0016] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0017] Figure 1 A flowchart illustrating a virtual IMU self-supervised human activity recognition method provided in one embodiment of this application; Figure 2 A schematic diagram illustrating the process of obtaining a target virtual IMU according to one embodiment of this application; Figure 3 This is a schematic diagram illustrating the process of obtaining encoder parameters according to one embodiment of this application; Figure 4 A flowchart illustrating the process of determining the action category to be identified, provided as an embodiment of this application; Figure 5This is a schematic diagram of the process for obtaining virtual IMU data provided in one embodiment of this application; Figure 6 A schematic diagram illustrating the process of obtaining a training sample set according to one embodiment of this application; Figure 7 A schematic diagram of a virtual IMU self-supervised human activity recognition system provided in one embodiment of this application; Figure 8 This is a schematic diagram of a computer device provided in one embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical methods, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. It should be noted that the meaning of "multiple" (or "more than") in the description of the embodiments of this application refers to two or more, and "greater than," "less than," "exceeding," etc. are understood to exclude the number itself, while "above," "below," "within," etc. are understood to include the number itself. If "first," "second," etc. are used in the description, they are only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated. In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: the existence of a alone, the existence of b alone, the existence of c alone, the simultaneous existence of a and b, the simultaneous existence of a and c, the simultaneous existence of b and c, or the simultaneous existence of a, b, and c, where a, b, and c can be single or multiple.

[0019] In the description of this application, unless otherwise expressly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.

[0020] In some cases, the synthesized virtual action sequences based on a single data generation mechanism (such as 3D skeletal animation or video pose estimation) have limitations in terms of expressive style, action details and individual differences. They are difficult to cover the diverse manifestations of the same action under different performers and different contexts, which limits the model's generalization ability in complex real-world environments.

[0021] Based on the above, this application proposes a method, system, device, and medium for identifying human activities using a pseudo-IMU self-supervised method, aiming to improve the semantic coverage of virtual data and enhance the accuracy of the model in distinguishing action categories during comparative learning.

[0022] To facilitate understanding of the solutions in the embodiments of this application, the relevant concepts involved in the embodiments of this application will be introduced below.

[0023] VAE (Variational Autoencoder) is a generative model that combines probabilistic graphical models with deep neural networks. It learns the latent probability distribution of data through variational inference and reparameterization techniques, achieving efficient sampling and data generation.

[0024] The diffusion model is a deep learning framework based on probabilistic generative modeling. It learns the ability to reconstruct complex data distributions from simple noise distributions by simulating a bidirectional random process of gradually adding and removing noise from data, thereby generating high-fidelity and diverse samples (such as images, audio, time-series signals, etc.).

[0025] A virtual IMU (Virtual Inertial Measurement Unit) is a technology or data format that uses algorithms or simulation to simulate the acceleration and angular velocity signals output by a real IMU (Inertial Measurement Unit) without the need for physical sensors. Its core purpose is to provide low-cost, controllable, and scalable sensor data input for positioning, navigation, motion recognition, or simulation systems.

[0026] The virtual IMU self-supervised human activity recognition provided in this application embodiment can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms; the software can be an application that implements virtual IMU self-supervised human activity recognition, etc., but is not limited to the above forms. This application can be applied to numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via communication networks. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices. It should be noted that in various specific embodiments of this invention, when processing is required based on data related to the characteristics of an object (e.g., user attributes or sets of attribute information), permission or consent from the corresponding object is obtained first, and the collection, use, and processing of this data comply with relevant laws and standards. Furthermore, when the embodiments of the present invention need to obtain the attribute information of an object, they will obtain the separate permission or separate consent of the corresponding object through pop-up windows or redirection to a confirmation page. After obtaining the separate permission or separate consent of the corresponding object, they will then obtain the relevant data of the object necessary for the embodiments of the present invention to operate normally.

[0027] See Figure 1 , Figure 1This is a flowchart illustrating a method for identifying self-supervised human activity using a virtual IMU, as provided in one embodiment of this application. The method includes, but is not limited to, steps S110 to S160, which will be described in detail below. Step S110: Obtain the action category to be identified; Specifically, the category of the action to be identified is obtained through user input, preliminary sensor detection, or transmission from other external devices. The user selects "running" as the action category to be identified via the interface; or the sensor initially detects signs of rapid human movement and selects "running" as a candidate action category. Step S120: Process the action according to the action category to be identified and the preset language model to obtain a natural language action description; Specifically, the preset language model is a natural language processing model trained on a large amount of action-related text, capable of converting action categories into detailed natural language descriptions. When the action category to be identified is "running," the natural language action description obtained after processing by the preset language model could be: "During running, the human body alternates and swings its legs forward rapidly, with large strides, a slight forward lean, and the arms swinging back and forth in coordination with the leg movements to maintain balance; the overall movement rhythm is relatively fast." Step S130: Generate a three-dimensional action sequence based on the natural language action description and the preset action generation mechanism; Furthermore, the preset motion generation mechanism, based on natural language motion descriptions, generates 3D motion sequences through 3D modeling and animation generation techniques. It analyzes the motion details in the natural language descriptions, including the type of motion, initial posture, direction, motion path, rhythm, repetition count, and emotional content, and then transforms these details into motion data in 3D space. For example, based on the natural language description of "running," the 3D motion sequence generated by the preset motion generation mechanism includes complete trajectory data of the legs from push-off to forward swing, angle and speed data of the arm swings, and changes in the angle of the body's forward lean, etc., and these data are arranged in chronological order to form a continuous 3D motion sequence. Step S140: Input the three-dimensional action sequence into the target virtual IMU for calculation to obtain virtual IMU data; Furthermore, the target virtual IMU is a virtual sensor model simulating an actual inertial measurement unit, capable of calculating corresponding virtual IMU data, including acceleration information, based on the input three-dimensional motion sequence. The virtual IMU's parameter settings are consistent with the actual IMU, such as measurement range and sampling frequency, to ensure the high realism of the generated virtual IMU data. For example, when a human leg makes a rapid forward swinging motion in a three-dimensional motion sequence, the target virtual IMU will calculate the corresponding acceleration based on the changes in velocity and acceleration of this motion.

[0028] Step S150: Obtain a training sample set based on virtual IMU data to extract the preset encoder parameters; Furthermore, the virtual IMU data is divided according to certain rules to form multiple training samples, which together form a training sample set. The preset encoder is a neural network model used to extract features from the IMU data. The training samples are input into the preset encoder, and the encoder parameters are adjusted until the encoder can accurately extract motion features from the virtual IMU data. The virtual IMU data is divided into time segments, with each time segment serving as a training sample. The preset encoder uses a convolutional neural network structure. Through training, the encoder is able to extract unique features related to the "running" motion from the training samples, such as specific acceleration change patterns and frequency.

[0029] Step S160: Determine the action category to be identified based on the encoder parameters and the preset task classifier.

[0030] Furthermore, the preset task classifier is a model that classifies actions based on features extracted by the encoder. Encoder parameters are input into the preset task classifier, and the classifier calculates the recognition result of the action category to be identified. The preset task classifier can employ classification algorithms such as support vector machines and decision trees. After the encoder extracts the features of the "running" action, the preset task classifier compares these features with the features of known action categories, and finally outputs "running" as the recognition result of the action category to be identified. It is worth noting that this embodiment of the application avoids over-reliance on actual sensor data by generating virtual IMU data as training samples, reducing the cost of data acquisition. Simultaneously, it can generate a large number of diverse training samples, solving the problem of insufficient data volume. Utilizing a preset language model and action generation mechanism, action categories can be converted into detailed three-dimensional action sequences, and then corresponding data can be generated through the virtual IMU. This makes the generated virtual data more closely resemble actual action characteristics, improving the quality of training samples. Through training a preset encoder and applying a preset task classifier, the action category to be identified can be accurately recognized, improving the accuracy and robustness of action recognition, and making it suitable for various application scenarios. In one embodiment, the user inputs "walking" into the interactive interface to determine that the action category to be identified is walking. "Walking" is input into a preset language model, which has been trained on a large amount of walking-related text. The output natural language action description is: "When a person walks, their legs move forward alternately, with relatively small steps, the body remains basically upright, the arms swing naturally back and forth, and the movement rhythm is stable." A preset action generation mechanism parses the above natural language action description, determines parameters such as the trajectory of the legs, the angle of the arm swing, and the body posture, and generates a three-dimensional action sequence. This sequence records the position and angle information of each joint of the human body at each moment, with time as the axis. The three-dimensional action sequence is input into a target virtual IMU. The virtual IMU simulates the working principle of an actual IMU, calculating the acceleration and angular velocity data at each moment based on the human body's motion state in the action sequence, forming virtual IMU data. The virtual IMU data is divided into time segments of 0.5 seconds each, resulting in multiple training samples, forming a training sample set. A preset encoder uses a recurrent neural network structure, and the training samples are input into the encoder for training. Through multiple iterations of training, the encoder's weights and biases are adjusted until the encoder can accurately extract the IMU data features corresponding to walking movements, thus obtaining the encoder parameters. These encoder parameters are then input into a preset task classifier. The classifier compares the features extracted by the encoder with known walking movement features and outputs the recognition result "walking," which matches the movement category to be identified, indicating successful recognition. See Figure 2 , Figure 2 This is a schematic diagram of the process of obtaining a target virtual IMU according to an embodiment of this application; including but not limited to steps S210 to S270, which are described in turn below. Step S210: Obtain the target action category name; Specifically, the category name of the target action to be processed is obtained through user input, database access, or other information retrieval methods. For example, "squat" can be retrieved from the motion database as the target action category name; or the user can input "bend over" through the interface as the target action category name. Step S220: Process the target action category name and the preset language model to obtain a natural language action description; In one embodiment, when the target action category name is "squat", the natural language action description obtained after processing by the preset language model can be "When a person performs a squat, he stands with his feet shoulder-width apart, then slowly bends his hip, knee and ankle joints to squat down until his thighs are parallel to or slightly lower than the ground, and then slowly stands up to return to the initial standing position, keeping his back straight throughout the process." Step S230: Obtain the first action sequence based on the natural language action description and the preset first action generation mechanism, wherein the preset first action generation mechanism is obtained through a preset VAE model; It's important to note that VAE (Variational Autoencoder) models are generative models. The pre-defined first action generation mechanism is built upon a pre-trained VAE model, enabling it to parse key information from natural language action descriptions and generate corresponding first action sequences. These first action sequences contain data on the position and angles of various body parts in three-dimensional space over time. For the natural language description of "squatting," the VAE model learns the movement patterns of the body's joints during the squat, generating a first action sequence that includes data on changes in hip and knee angles from standing to squatting and then standing up.

[0031] Step S240: Based on the natural language action description and the preset second action generation mechanism, a second action sequence is obtained, wherein the preset second action generation mechanism is obtained through a preset diffusion model; Step S240: The diffusion model is another generative model. The pre-defined second action generation mechanism is built upon the trained diffusion model. Compared to the VAE model, the diffusion model has advantages in the detail and diversity of the generated data. It also parses natural language action descriptions to generate a second action sequence. For the "squat" action, the second action sequence generated by the diffusion model may have richer descriptions of body center of gravity changes and muscle exertion details, complementing the first action sequence. Step S250: Extract virtual IMU data through the first action sequence and the second action sequence; Step S260: Generate training samples based on virtual IMU data, the first action sequence, and the second action sequence; Step S270: Train the initial virtual IMU based on the training samples to obtain the target virtual IMU.

[0032] In one embodiment, for different virtual IMU sequences generated under the same semantic description, it is assumed that the action sequence generated by the diffusion model is... The action sequence generated by the VAE model is The action sequence generated by the Transformer structure is Then, positive sample pairs are constructed by pairing them up (e.g., ... , ), ( , )and( , The current sample is used as a positive sample pair to represent data instances that are semantically consistent but generated in different styles. Meanwhile, virtual IMU data from other semantic descriptions are paired with the current sample to construct negative sample pairs, reflecting differences in action categories or semantics. This construction method fully utilizes the action diversity and modal variations introduced by the multi-generation mechanism, enhancing the discriminative and generalization capabilities of contrastive learning while maintaining semantic consistency, and providing a high-quality contrastive sample structure for the pre-training stage.

[0033] In one embodiment, the initial virtual IMU is an untrained virtual sensor model. Training samples are input into the initial virtual IMU, and the model parameters are adjusted to make the output of the virtual IMU as consistent as possible with the virtual IMU data in the training samples. After multiple iterations of training, training stops when the model error reaches a preset threshold, and the target virtual IMU is obtained. During the training process, the IMU data generated by the initial virtual IMU based on the action sequence is continuously compared with the virtual IMU data extracted from the training samples. The model parameters are adjusted using the backpropagation algorithm until the difference between the two is within an acceptable range. In one embodiment, the user inputs "push-up" into the training system interface to determine the target action category as push-up. The input "push-up" is then fed into a preset language model, which has been trained on a large amount of push-up-related text. The output natural language action description is: "When doing a push-up, the hands are shoulder-width apart or slightly wider, the feet are together, the body is in a straight line, then the elbows are bent to lower the body until the chest is close to the ground, then the elbows are extended to return the body to the initial position, and this is repeated." A preset first action generation mechanism, based on a trained VAE model, parses the above natural language action description to generate a first action sequence. This sequence includes data on the changes in hand position, elbow angle, and body center of gravity trajectory over time during the push-up process; for example, during the descent phase, the elbow angle gradually decreases from 180 degrees to approximately 90 degrees. A preset second action generation mechanism, based on a trained diffusion model, parses the natural language action description to generate a second action sequence. This sequence adds more details to the first action sequence, such as subtle tremors in the shoulder muscles and details of speed changes during the descent and ascent, for example, the descent speed gradually increases while the ascent speed is relatively uniform. Virtual IMU data is extracted from both the first and second action sequences. The first action sequence and its corresponding virtual IMU data are combined into one training sample, and the second action sequence and its corresponding virtual IMU data are combined into another training sample, forming a training sample set. The training samples are input into an initial virtual IMU, which generates IMU data based on the input action sequence. This initial virtual IMU data is compared with the virtual IMU data in the training samples to calculate the error. The parameters of the initial virtual IMU are adjusted using a backpropagation algorithm. After multiple iterations of training, training stops when the error is less than a preset threshold, yielding the target virtual IMU. This target virtual IMU can accurately generate corresponding IMU data based on the input push-up action sequence. See Figure 3 , Figure 3 This is a schematic diagram of the process for obtaining encoder parameters according to an embodiment of this application; including but not limited to steps S310 to S350, which will be described in turn below. Step S310: Obtain a training sample set based on the virtual IMU data, wherein the training samples include positive sample pairs and negative sample pairs; Step S320: Input the training samples into the convolutional layer of the preset encoder for training to obtain the global feature representation, wherein the preset encoder includes multiple convolutional layers; Step S330: Input the global feature representation into the preset encoder to obtain the contrast embedding space; Step S340: By comparing the embedding space, positive sample pairs, and negative sample pairs, the similarity of positive sample pairs and the similarity of negative sample pairs are obtained. Step S350: Obtain encoder parameters based on the similarity between positive sample pairs and the similarity between negative sample pairs.

[0034] In one embodiment, action sequences obtained through the aforementioned multi-generation mechanism are imported into a virtual environment to acquire virtual IMU data. Positive sample pairs are constructed by pairwise matching of virtual data obtained with "same semantics, different generation," while other samples serve as negative samples. This achieves the construction of training data required for contrastive learning. The present invention constructs a three-layer one-dimensional convolutional neural network as an IMU encoder to extract temporal features from the signal. Then, two fully connected layers are mapped to the contrastive embedding space, and NT-Xent contrastive loss is used to optimize network parameters, maximizing the similarity between positive samples and minimizing the similarity between negative samples to obtain discriminative semantic embedding expressions. In the downstream recognition task, the encoder parameters are frozen, and only a small amount of real IMU labeled data is used to fine-tune the newly connected classification layer, achieving efficient knowledge transfer between the virtual and real data domains. This significantly reduces the dependence on large-scale real data and improves adaptability to new environments or new users.

[0035] In one embodiment, a semantically driven generation mechanism ensures that positive samples share the same text description category, and action sequences generated by multiple models under this description are considered positive sample pairs; while virtual action data generated under all different semantic categories are considered to have semantic isolation and can be used as negative samples for contrastive learning training. High-quality semantic clustering is indirectly achieved by relying on the structured control of the prompt design and the semantic fidelity of the generation model itself. All segments are processed by sliding window segmentation and normalization before entering the subsequent training module.

[0036] In one embodiment, characters are treated as actors, and specific but explainable English instructions are provided. Each animation is 10 seconds long, combining various elements such as action type, direction, and emotion to achieve detailed and accurate animation effects. For example, descriptive elements include: imagining actions in a context. Where do they begin their actions? How do they feel? Which arm or leg do they need to use? Are they interacting with any objects such as a chair or stairs? Subsequently, by inputting specific action categories (e.g., "output 10 descriptions about high knees"), multiple action semantic expressions with fine-grained differences are generated. These descriptions all follow the aforementioned structured control paradigm, containing a unified subject, action type, and multiple detailed attributes, ensuring diversity of expression methods while maintaining high-level semantic consistency. This invention treats each semantic description as a semantic instance, using it as the basis for constructing contrastive learning samples. This not only enriches the semantic representation space of each target action but also effectively improves the expression accuracy and semantic coverage of virtual samples, providing high-quality semantic control input for subsequent self-supervised training based on a virtual IMU.

[0037] In one embodiment, a constructed virtual IMU sample pair is input into the encoder network. The encoder consists of three one-dimensional convolutional layers. The first convolutional layer uses 32 channels, a kernel size of 24, a stride of 1 (default), ReLU activation, followed by L2 regularization, and 10% dropout to prevent overfitting. The second convolutional layer has 64 channels and a kernel size of 16, with the same structure as the first layer. The third layer has 96 channels and a kernel size of 8, with the same structure as the first layer. All convolutional layers use padding='valid' by default. At the end of the convolutional module, the temporal dimension is compressed through Global Max Pooling to generate a global feature representation. This feature vector is then connected to a three-layer fully connected projection head module to construct a contrastive embedding space. This module consists of: a 256-dimensional fully connected layer with ReLU activation, a 128-dimensional fully connected layer with ReLU activation, and finally outputs a 50-dimensional contrastive embedding space. The training objective of the entire network is to maximize the cosine similarity between positive sample pairs and minimize the similarity between negative samples using the NT-Xent contrastive loss function, thereby enabling the encoder to learn discriminative and semantically consistent feature representations. This structure maintains training stability while ensuring the full extraction of fine-grained action features from the IMU time series.

[0038] See Figure 4 , Figure 4 This is a flowchart illustrating the process of determining the action category to be identified according to an embodiment of this application; including but not limited to steps S410 to S430, which will be described in turn below. Step S410: Obtain the initial number of action categories based on the encoder parameters and the preset task classifier; Step S420: Update the initial number of action categories based on the training sample set to obtain the target number of action categories; Step S430: Determine the action category to be identified based on the number of target action categories.

[0039] In one embodiment, the parameters of the trained encoder convolutional layer are fixed to prevent overfitting on small samples of real data. At the same time, a new fully connected classification layer is connected after it, with the output dimension equal to the number of target action categories. A small number of real IMU labeled samples are used as training data, and only the parameters of the classification head are updated. The optimization objective is the cross-entropy classification loss. After training, it can predict the action category of new samples. By "freezing the encoder + fine-tuning the classification layer", the transfer benefits brought by virtual pre-training are maximized, and the adaptability of the model to new users and complex scenarios is improved.

[0040] See Figure 5 , Figure 5 This is a schematic diagram of the process for obtaining virtual IMU data according to an embodiment of this application; including but not limited to steps S510 to S530, which will be described in turn below. Step S510: Input the three-dimensional motion sequence into the target virtual IMU for calculation to obtain multiple three-dimensional position changes; Step S520: Obtain time-continuous acceleration data based on multiple three-dimensional position changes; Step S530: Denoise the time-continuous acceleration data to obtain virtual IMU data.

[0041] In one embodiment, a text-to-motion motion generation mechanism driven by natural language is used to receive motion descriptions and output a 3D skeletal motion sequence of approximately 10 seconds in duration, achieving multi-style motion expression under the same semantics. Each 3D skeletal motion segment is imported into a virtual simulation environment (such as Unity or Blender) to drive a standard digital human to complete a full motion demonstration. Virtual IMU nodes are placed at specified sensor locations on the virtual human body (such as the thigh, waist, and head), and the three-dimensional position changes of each node are extracted through forward kinematics calculation of joint trajectories. The second derivative is then calculated, as shown in formula (1), to obtain time-continuous acceleration data.

[0042] In one embodiment, after acquiring the acceleration data from the virtual IMU, a filtering module is introduced into the original acceleration signal to enhance the smoothness and realism of the data. A fourth-order Butterworth low-pass filter is used to denoise the triaxial acceleration sequence to effectively filter out high-frequency components introduced by animation interpolation, attitude jitter, or simulation errors. The cutoff frequency of the filter is set to 10Hz, and the sampling frequency is 100Hz. To avoid signal offset problems caused by phase delay, a forward-backward bidirectional filtering method (filtfilt) is used in the filtering process to ensure that the timing structure of the filtered signal remains unchanged on the time axis.

[0043] In one embodiment, the preset action generation mechanism includes an action semantic parsing generation algorithm based on natural language action description; and / or, an action template generation model based on natural language action description.

[0044] In one embodiment, a Named Entity Recognition (NER) model combining Bi-LSTM (Bi-LSTM) and Conditional Random Field (CRF) is used to extract data from natural language action descriptions, establish an element-parameter mapping table, convert natural language parameters into quantized values ​​required for IMU data generation, and generate IMU data sequences by calling a preset dynamics model based on the extracted action type and parameters.

[0045] See Figure 6 , Figure 6 This is a schematic diagram of the process for obtaining a training sample set according to an embodiment of the present application; including but not limited to steps S610 to S620, which will be described in turn below. Step S610: Input the virtual IMU data into the first action sequence and the second action sequence to obtain positive sample pairs and negative sample pairs. Positive sample pairs represent sample pairs of data instances that are semantically consistent but have different generation styles, while negative sample pairs represent sample pairs that differ in action category or semantics. Step S620: Based on positive sample pairs and negative sample pairs, obtain the training sample set.

[0046] See Figure 7 , Figure 7 This is a schematic diagram of a virtual IMU self-monitored human activity recognition system provided in one embodiment of this application. The virtual IMU self-monitored human activity recognition system 700 includes: The acquisition module 710 is used to acquire the action category to be identified; The processing module 720 is used to process the action category to be identified and the preset language model to obtain a natural language action description; The generation module 730 is used to generate a three-dimensional action sequence based on natural language action descriptions and a preset action generation mechanism; The calculation module 740 is used to input the three-dimensional action sequence into the target virtual IMU for calculation to obtain virtual IMU data; The extraction module 750 is used to obtain training samples from virtual IMU data to extract encoder parameters from a preset encoder. The determination module 760 is used to determine the action category to be identified based on encoder parameters and a preset task classifier.

[0047] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0048] like Figure 8 As shown, Figure 8 This is a schematic diagram of a computer device provided according to an embodiment of this application. The computer device 800 may be a server or a terminal, and its internal structure includes, but is not limited to: Memory 810 is used to store programs; The processor 820 is used to execute the program stored in the memory 810. When the processor 820 executes the program stored in the memory 810, the processor 820 is used to execute the above-mentioned virtual IMU self-supervised human activity recognition method. The processor 820 and memory 810 can be connected via a bus or other means. The memory 810, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the virtual IMU self-monitored human activity recognition method described in any embodiment of this application. The processor 820 implements the above-described virtual IMU self-monitored human activity recognition method by running the non-transitory software program and instructions stored in the memory 810. The memory 810 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function. The data storage area may store the virtual IMU self-monitored human activity recognition method described above. Furthermore, the memory 810 may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 810 may optionally include memory remotely located relative to the processor 820, which can be connected to the processor 820 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The non-transitory software program and instructions required to implement the above-described virtual IMU self-supervised human activity recognition method are stored in memory 810. When executed by one or more processors 820, the virtual IMU self-supervised human activity recognition method provided in any embodiment of this application is executed. This application also provides a computer-readable storage medium storing computer-executable instructions for executing the above-described virtual IMU self-monitored human activity recognition method. In one embodiment, the storage medium stores computer-executable instructions that are executed by one or more control processors, such as one or more processors 820 in the computer device 800, which can cause the one or more processors 820 to perform the virtual IMU self-supervised human activity recognition method provided in any embodiment of this application. The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium. The foregoing has provided a detailed description of the preferred embodiments of this application. However, this application is not limited to the above-described embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A method for recognizing human activity self-supervised by virtual IMU, characterized in that, The method comprises the following steps: acquiring an action category to be identified; processing the action category to be identified and a preset language model to obtain a natural language action description; generating a three-dimensional action sequence according to the natural language action description and a preset action generation mechanism; inputting the three-dimensional action sequence into a target virtual IMU to obtain virtual IMU data; acquiring a training sample set from the virtual IMU data to extract a preset encoder to obtain an encoder parameter; determining the action category to be identified according to the encoder parameter and a preset task classifier.

2. The method of claim 1, wherein, The target virtual IMU is obtained by the following steps: acquiring a target action category name; processing the target action category name and a preset language model to obtain a natural language action description; obtaining a first action sequence according to the natural language action description and a preset first action generation mechanism, wherein the preset first action generation mechanism is obtained through a preset VAE model; obtaining a second action sequence according to the natural language action description and a preset second action generation mechanism, wherein the preset second action generation mechanism is obtained through a preset diffusion model; extracting virtual IMU data through the first action sequence and the second action sequence; generating a training sample according to the virtual IMU data, the first action sequence and the second action sequence; training an initial virtual IMU based on the training sample to obtain the target virtual IMU.

3. The method of claim 1, wherein, The method of acquiring a training sample set from the virtual IMU data to extract a preset encoder to obtain an encoder parameter comprises the following steps: acquiring a training sample set from the virtual IMU data, wherein the training sample comprises a positive sample pair and a negative sample pair; inputting the training sample into a convolution layer of a preset encoder for training to obtain a global feature representation, wherein the preset encoder comprises a plurality of convolution layers; inputting the global feature representation into the preset encoder to obtain a contrast embedding space; obtaining a similarity of the positive sample pair and a similarity of the negative sample pair through the contrast embedding space, the positive sample pair and the negative sample pair; obtaining an encoder parameter according to the similarity of the positive sample pair and the similarity of the negative sample pair.

4. The method of claim 1, wherein, The method of determining the action category to be identified according to the encoder parameter and a preset task classifier comprises the following steps: obtaining an initial action category number according to the encoder parameter and a preset task classifier; updating the initial action category number according to the training sample set to obtain a target action category number; determining the action category to be identified according to the target action category number.

5. The method of claim 1, wherein, In the case that the three-dimensional action sequence is a plurality of initial action sequences of a first time length, the method of inputting the three-dimensional action sequence into a target virtual IMU to obtain virtual IMU data comprises the following steps: inputting the three-dimensional action sequence into a target virtual IMU to obtain a plurality of three-dimensional position changes; obtaining time-continuous acceleration data according to the plurality of three-dimensional position changes; performing denoising processing on the time-continuous acceleration data to obtain virtual IMU data.

6. The method of claim 1, wherein, The preset action generation mechanism comprises: An action semantic parsing generation algorithm based on the natural language action description; And / or, an action template generation model based on the natural language action description.

7. The method of claim 2, wherein, The virtual IMU data is used to obtain a training sample set, including: The virtual IMU data is input into the first action sequence and the second action sequence to obtain a positive sample pair and a negative sample pair, wherein the positive sample pair represents a sample pair of data instances with consistent semantics but different generation styles, and the negative sample pair represents a sample pair with differences in action categories or semantics; Based on the positive sample pair and the negative sample pair, a training sample set is obtained.

8. A system for recognizing human activity self-supervised by virtual IMU, characterized in that, It includes: An acquisition module is configured to acquire an action category to be identified; A processing module is configured to process the action category to be identified and a preset language model to obtain a natural language action description; A generation module is configured to generate a three-dimensional action sequence according to the natural language action description and a preset action generation mechanism; A calculation module is configured to input the three-dimensional action sequence into a target virtual IMU for calculation to obtain virtual IMU data; An extraction module is configured to extract a training sample from the virtual IMU data to extract a preset encoder to obtain an encoder parameter; A determination module is configured to determine the action category to be identified according to the encoder parameter and a preset task classifier.

9. A computer device, comprising: It includes: A memory is configured to store a program; A processor is configured to execute the program stored in the memory, and when the processor executes the program stored in the memory, the processor is configured to execute the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Computer executable instructions are stored, and the computer executable instructions are used to execute the method of any one of claims 1 to 7.