Data augmentation method of action recognition model training data set

Through the combination of action generation large models and visual large models, augmented data sets in the field of behavior recognition are generated, which solves the problems of small scale and poor generalization of the data set, and achieves efficient and accurate data augmentation, reducing costs.

CN120339741APending Publication Date: 2025-07-18SHENZHEN SUPERNODE NETWORK TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510212299.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the data set in the field of behavioral recognition has a small scale and poor generalization, high cost of simulation data generation, and relying on manual annotation to cause problems such as errors and low efficiency.

Method used

The target action example is generated through the action generation big model, combined with the visual big model to augment the data, generate the initial data set and generalize it, avoid manual annotation, and use the image generation control model to extract depth, pose, normal vector and line draft control parameters to generate the augmented data set.

Benefits of technology

It improves the scale and generalization of the data set, reduces the generation cost, avoids errors caused by manual annotation, and improves the accuracy and construction efficiency of the data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339741A_ABST
    Figure CN120339741A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data set augmentation, and particularly discloses a data augmentation method for a training data set of an action recognition model. According to the method, the target action instance can be generated according to actual needs by utilizing the action generation large model, the initial data set is obtained according to the target action instance, and the initial data set is generalized by virtue of the visual large model to generate the augmented data set as the target training data, so that the use value of the data in the A I field is improved; the defects that traditional simulation data is poor in generalization, small in scale and high in generation cost are overcome; and the generation method does not depend on manual data labeling, so that the problems of wrong labeling, inconsistent labeling and the like caused by manual experience errors are avoided, the defects that manual data labeling is difficult to scale, the precision is poor, the cost is high and the efficiency is low are overcome, and the data scale, the accuracy, the data generalization and the construction efficiency of the data set are improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application with the application number 202411595744.2, the application date of November 11, 2024, and the invention title of "Method for Constructing a Training Dataset for Training an Action Recognition Model". Technical Field

[0002] This application relates to the technical field of dataset augmentation, and particularly relates to a method for augmenting an action recognition model training dataset. Background Art

[0003] With the rapid development of science and technology, AI (Artificial Intelligence) has become the core of innovation. The development of machine learning and deep learning methods is the key link for current AI technology to enter the industrial field and drive major technological changes. Among them, supervised machine learning methods represented by deep learning rely on a large number of labeled high-quality datasets, which are the cornerstones for training and evaluating AI models.

[0004] Currently, in each professional subdivision field, large-scale datasets are being constructed for deep learning, including large-scale mechanical data collection, manual annotation in large-scale data factories, and generation of large-scale simulation datasets, such as in the autonomous driving industry. Among them, action recognition, as a visual problem and a basic technical field in current human-computer interaction, digital twins, and metaverse applications, is much more difficult to construct datasets than most other AI fields. Because it is difficult to collect action data in the real world, resulting in a small data scale; on the other hand, although the simulation field can generate accurately labeled action data, due to the high data dimension, it is also difficult for the simulation field to generate reasonable data with good generalization performance, often limited to the simulation rendering output of a small number of fixed actions, with high data duplication and low data generalization ability. Therefore, how to augment the dataset and thereby improve the data scale and data generalization of the dataset in the field of action recognition has become an urgent problem to be solved and has extremely high practical value. Summary of the Invention

[0005] This application provides a method for augmenting an action recognition model training dataset, aiming to improve the data scale and data generalization of the dataset in the field of action recognition.

[0006] In a first aspect, this application provides a method for augmenting an action recognition model training dataset, the method comprising:

[0007] Obtain an initial dataset;

[0008] Based on a first sub-model of the image generation control model, process the scene depth image data in the initial dataset to obtain depth control parameters;

[0009] Based on the second sub-model of the image generation control model, process the skeletal joint point information in the initial dataset to obtain pose control parameters;

[0010] Based on a preset line drawing extraction processor, process the video frames in the initial dataset to obtain a line drawing image, and based on the third sub-model of the image generation control model, extract line drawing control parameters from the line drawing image;

[0011] Based on the fourth sub-model of the image generation control model, extract normal vector control parameters from the scene normal vector image data in the initial dataset;

[0012] Use the depth control parameters, the pose control parameters, the line drawing control parameters, and the normal vector control parameters as the first model parameters;

[0013] Based on a preset sample requirement, obtain an image generation prompt, process the image generation prompt to obtain target prompt information, and use the target prompt information as the second model parameter of the image generation large model;

[0014] Based on the image generation large model, the first model parameters, and the second model parameters, generate an augmented dataset.

[0015] In a second aspect, the present application also provides a computer device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the data augmentation method of the action recognition model training dataset as described above when executing the computer program.

[0016] In a third aspect, the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is enabled to implement the data augmentation method of the action recognition model training dataset as described above.

[0017] The present application discloses a data augmentation method for an action recognition model training dataset. Based on a preset action generation large model, preset action description information is processed to generate target action instances; based on the target action instances, a scene animation is constructed, and the scene animation is processed to obtain target parameter data. Based on the target parameter data, an initial dataset is generated; the initial dataset is processed to obtain model parameters, and based on a preset image generation large model and the model parameters, an augmented dataset is generated as the target training dataset. The present application can generate target action instances according to actual needs, obtain an initial dataset based on the target action instances, and generalize the initial dataset with the help of a vision large model to generate an augmented dataset as the target training data, improving the value of data used in the field of AI, and solving the defects of poor generalization, small scale, and high generation cost of traditional simulation data; moreover, this generation method does not rely on manually labeled data, avoiding problems such as incorrect labeling and inconsistent labeling caused by manual experience errors, and solving the defects of difficulty in scaling up, poor accuracy, high cost, and low efficiency of manually labeled data, improving the data scale, accuracy, data generalization, and construction efficiency of the dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of the first embodiment of a data augmentation method for an action recognition model training dataset provided by an embodiment of the present application;

[0020] Figure 2 It is a schematic diagram of the action instance generation process of a data augmentation method for an action recognition model training dataset provided by an embodiment of the present application;

[0021] Figure 3 It is a schematic diagram of the initial dataset generation process of a data augmentation method for an action recognition model training dataset provided by an embodiment of the present application;

[0022] Figure 4 It is a schematic flowchart of the second embodiment of a data augmentation method for an action recognition model training dataset provided by an embodiment of the present application;

[0023] Figure 5 It is a schematic flowchart of the third embodiment of a data augmentation method for an action recognition model training dataset provided by an embodiment of the present application;

[0024] Figure 6 FIG. 0 is a schematic diagram of a data augmentation video generation process for a data augmentation training dataset of an action recognition model provided by an embodiment of the present application;

[0025] Figure 7 FIG. 1 is a schematic block diagram of a data augmentation device for a data augmentation training dataset of an action recognition model provided by an embodiment of the present application;

[0026] Figure 8 FIG. 2 is a schematic block diagram of a structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0028] The flowcharts shown in the accompanying drawings are only illustrative examples and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.

[0029] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless otherwise clearly specified in the context, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0030] It should also be understood that the term " / and" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0031] An embodiment of the present application provides a data augmentation method for a data augmentation training dataset of an action recognition model. Among them, the data augmentation method for the data augmentation training dataset of the action recognition model can be applied to a server, and a target action event can be generated according to actual needs. An initial dataset is obtained according to the target action event, and the initial dataset is augmented, which improves the comprehensiveness of the dataset, and there is no need for manual data annotation, simplifies the dataset construction process, and avoids problems such as incorrect annotation and inconsistent annotation caused by manual experience errors, improving the accuracy and construction efficiency of the dataset. Among them, the server can be an independent server or a server cluster.

[0032] The following will describe in detail some embodiments of the present application in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0033] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a data augmentation method for an action recognition model training dataset provided by an embodiment of the present application. The data augmentation method for the action recognition model training dataset can be applied to a server, used to generate a target action instance through an action generation large model according to actual needs, obtain an initial dataset based on the target action instance, and generalize the initial dataset with the help of a vision large model to generate an augmented dataset as the target training data, improving the utilization value of data in the AI field and solving the defects of poor generalization, small scale, and high generation cost of traditional simulation data; and this generation method does not rely on manually labeled data, avoiding problems such as incorrect labeling and inconsistent labeling caused by manual experience errors, and solving the defects of difficult large-scale, poor accuracy, high cost, and low efficiency of manually labeled data, improving the data scale, accuracy, data generalization, and construction efficiency of the dataset.

[0034] As Figure 1 shown, the data augmentation method for the action recognition model training dataset specifically includes steps S101 to S103.

[0035] S101. Based on a preset action generation large model, process the preset action description information to generate a target action instance;

[0036] In one embodiment, the action description information is a specific description of a human action event, which can be freely set by the user according to actual needs.

[0037] In one embodiment, as Figure 2 shown, Figure 2 is a schematic flowchart of an action instance generation process of a data augmentation method for an action recognition model training dataset provided by an embodiment of the present application. The action generation large model can generate a corresponding action event according to the action description information, and then perform bone redirection on the corresponding action bones in the action event to obtain a target action instance.

[0038] Specifically, first, use the action generation large model to process the action description information to generate an initial action event. Among them, the human skeleton of the initial action event is in the skeleton format of the SMPL (Skinned Multi-Person Linear Model, parameterized human model) model. Then, redirect the SMPL skeleton in the initial action event to the skeleton format of the Mixamo model (or other skeleton models and skeleton formats set by the user according to actual needs) to obtain the target action instance. Among them, the Mixamo model is the skeleton format used by the online 3D character animation design and production platform Mixamo.

[0039] In one embodiment, the action generation large model can be an MDM (Motion Diffusion Model) model, or other models set by the user according to needs. Among them, the MDM large model is a generative machine learning model based on the diffusion model, specifically used to generate realistic human action sequences according to text.

[0040] S102. Based on the target action instance, construct a scene animation, process the scene animation to obtain target parameter data, and generate an initial data set based on the target parameter data;

[0041] Further, the target parameter data includes video frames, synchronized annotation frame data, and camera data. Based on the target action instance, construct a scene animation, process the scene animation to obtain target parameter data, and generate an initial data set, including: transmit the target action instance to a preset scene, and generate the scene animation based on the motion state of the target action instance and the preset scene; render the preset scene and the scene animation based on a preset rendering engine to obtain the video frames corresponding to the scene animation, and correspond the target action instance and the scene data of the preset scene with the video frames to obtain the synchronized annotation frame data and camera data with spatio-temporal parameter alignment; obtain the initial data set based on the video frames, the synchronized annotation frame data, and the camera data.

[0042] Further, the synchronized annotation frame data includes action labels, target model bounding boxes, skeleton joint point information representing actions, scene depth image data, and scene normal vector image data.

[0043] In one embodiment, an action instance generally refers to an instance in which a specific human model or virtual character performs a specific action in 3D animation, game development, motion capture, or other related fields.

[0044] In one embodiment, the preset scene is a scene generated by the UE5 program according to a preset scene model file. Among them, the scene model refers to 3D objects that make up the game world or scene, such as buildings, trees, props, etc. Among them, UE5, that is, Unreal Engine 5, the Unreal Engine 5, is a game engine used to create high-end 3D content and interactive experiences.

[0045] In a specific embodiment, such as Figure 3 shown, Figure 3 FIG. is a schematic diagram of the initial dataset generation process of a data augmentation method for an action recognition model training dataset provided by an embodiment of the present application. First, import the fbx format action model file corresponding to the action instance obtained by redirection and the user-preset scene model file into the UE5 program. Among them, the import step can be executed through instructions. Specifically, after the skeleton redirection is completed and the fbx format file is obtained, call the import instruction to import the action model file and the scene model file into the UE5 program. Then, construct an animated character class and a character action instance, and set the public parameters of the character action instance to the character identifier and the task action list. Then, place the character action instance into the scene generated according to the scene model file, where the identifier of each character action instance in each frame is unique, and the identifier of the same action instance remains consistent throughout the action sequence. Finally, use the level sequence tool to assign the skeleton actions of the target action instance based on the mixamo skeleton to all the character action instances in the scene with clear logic. When the character action changes, convert the action parameters in its action list by inserting key frames to complete the establishment of the scene animation and obtain the scene animation.

[0046] In one embodiment, the initial dataset includes video frames, synchronized annotation frame data, and camera data. The synchronized annotation frame data includes action labels, target model bounding boxes (i.e., Figure 3 the position labels in), skeleton joint point information representing the action (i.e., Figure 3 the pose information in), scene depth image data (i.e., Figure 3 the depth information in), and scene normal vector image data (i.e., Figure 3 the normal vector information in). Specifically, in the process of processing the scene animation through UE5, obtain the video frames corresponding to the scene animation, and correspond the target action instance and the scene data of the preset scene to the video frames to obtain synchronized annotation frame data and camera data with spatio-temporal parameter alignment. Combine the video frames, synchronized annotation frame data, and camera data to obtain the initial dataset.

[0047] In one embodiment, the video frames are composed of a series of static images. In this embodiment, the static images are Figure 3The simulation data image in it. The camera data includes information such as the camera position, rotation, focal length, etc. for each frame, which can be used to understand the structure of the scene and the perspective changes.

[0048] In a specific embodiment, the steps for obtaining the scene normal vector image data include: obtaining the scene object normal vector values, creating a normal vector material based on the scene object normal vector values, and obtaining the scene normal vector image data based on the normal vector material and the scene animation.

[0049] In one embodiment, the scene object normal vector data is obtained and mapped to the RGB space.

[0050] In a specific embodiment, the scene object normal vector data, that is, the normal vector values in the x-axis, y-axis, and z-axis directions, is mapped to the G, R, and B channels of the RGB space and used as the self-illuminating color of the material to complete the creation of the normal vector material. Then, a post-processing volume block is added to the scene, its scope is set to infinite, and its post-processing material is set to an instance of the normal vector material. After that, the scene animation constructed based on the target action instance is re-recorded to obtain the final scene normal vector image data.

[0051] In a specific embodiment, the steps for obtaining the scene depth image data include: obtaining the first depth value of the scene animation, obtaining the second depth value based on the second preset formula and the first depth value, creating a depth material based on the second depth value, and obtaining the scene depth image data based on the depth material and the scene animation.

[0052] In one embodiment, the first depth value is the scene depth value X recorded in the current scene, and the second depth value is the depth value Y obtained by processing the scene depth value X with the second preset formula.

[0053] In a specific embodiment, the depth image data is obtained by creating a new depth material and processing the depth material. Specifically, using the scene depth value X recorded in the current scene as the input, the depth value Y is obtained through the second preset formula, where the second preset formula is:

[0054]

[0055] After obtaining the depth value Y, the depth change rate is adjusted using the visualization curve function in UE5, and the finally obtained depth value is mapped to the RGB space as a gray value, that is, this depth value is assigned to the R, G, and B channels at the same time to generate a uniform color, which is used as the self-illuminating color of the material to complete the creation of the depth material. Then, a post-processing volume block is added to the scene, its scope is set to infinite, and its post-processing material is set to an instance of the depth material. After that, the scene animation constructed based on the target action instance is re-recorded to obtain the final scene depth image data.

[0056] In a specific embodiment, the steps for obtaining skeletal joint point information include: performing skeletal marking on the action skeletal key points corresponding to the target action instance to determine multiple skeletal marking points of the target action instance; obtaining the world coordinate values of the skeletal marking points and the pose and frustum information of the rendering camera; obtaining a projection transformation matrix based on the pose and frustum information of the rendering camera; and converting the world coordinate values into viewport coordinate values based on the projection transformation matrix to obtain the final skeletal joint point information.

[0057] In one embodiment, skeletal marking refers to identifying the key points (such as joint points) of each bone. In this way, when the bone moves, the actions of the character can be accurately tracked and represented.

[0058] In one embodiment, after the skeletal marking is completed, the positions of these skeletal marking points can be determined for each action instance (such as walking, running, etc.). Among them, these positions define the pose of the character at a specific time point.

[0059] In one embodiment, each skeletal marking point has a world coordinate, which represents the position of the point in the entire scene. Obtain the coordinate values of each skeletal marking point in the world coordinate system.

[0060] In one embodiment, the rendering camera determines the viewing angle and position of the scene. The pose (position and orientation) describes the position and orientation of the camera in the scene. The frustum information defines the spatial range that the camera can see, that is, which objects are visible and which are occluded.

[0061] In one embodiment, the projection transformation matrix is the key to converting 3D world coordinates into 2D screen coordinates. Combining the pose and frustum information of the camera ensures that only the objects within the camera's field of view are correctly rendered.

[0062] In one embodiment, the viewport coordinate values are usually pixel positions on the screen, which determine the final position of the object in the rendered image.

[0063] In one embodiment, by converting the world coordinate values of the skeletal marking points into viewport coordinate values through the projection transformation matrix, the position data of the human skeletal marking points on the screen is obtained, that is, the human skeletal marking point data, which can be used for further rendering, animation processing, or analysis.

[0064] In a specific embodiment, through skeleton editing, 17 slots are established in the Mixamo general skeleton, corresponding to the 17-point skeleton markers of the OpenPose skeleton, which are used to locate the key points of the skeleton. During the rendering of the scene animation, the world coordinate values of the skeleton markers of each action instance, the pose of the rendering camera, and the information of the viewing frustum can be obtained at the same moment. According to the pose of the rendering camera and the information of the viewing frustum, the projection transformation matrix is calculated. According to the projection transformation matrix, the coordinates of the 17 OpenPose skeleton nodes are converted into viewport coordinates and output, and finally the human skeleton joint point information of each action instance is obtained and output.

[0065] Among them, Mixamo is a platform that provides 3D character animation, and the Mixamo general skeleton is a standard skeleton structure used to create and edit 3D character animations.

[0066] In one embodiment, a slot usually refers to a placeholder or slot for placing specific data or components. In the embodiments of the present application, establishing 17 slots means creating 17 specific positions or interfaces on the Mixamo general skeleton for matching and docking with the 17-point skeleton markers of OpenPose.

[0067] In one embodiment, OpenPose is an open-source human pose estimation library that can extract human skeleton point information from images or videos. The 17-point skeleton markers of the OpenPose skeleton refer to 17 key skeleton points that OpenPose can identify and locate, usually including the main joints of the human body, such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles.

[0068] In one embodiment, by corresponding the slots of the Mixamo general skeleton with the 17-point skeleton markers of OpenPose, accurate positioning of the key points of the 3D character skeleton can be achieved. This helps to accurately apply the human pose information recognized by OpenPose from images or videos to the Mixamo 3D character, thereby achieving a more realistic and natural animation effect.

[0069] In a specific embodiment, the steps for obtaining the target model bounding box include: rendering the scene animation based on preset rendering parameters. When rendering the basic image data, it is determined whether there is a human action instance for each pixel. When there is a human action instance, the human action instance and its action type are recorded at this pixel position, and based on the pixel position, the bounding box of the target model of each human action instance is obtained.

[0070] In one embodiment, rendering is the process of converting a three-dimensional scene into a two-dimensional image. In this process, it is necessary to render the scene animation according to preset rendering parameters (such as lighting, materials, camera perspective, etc.) to obtain basic image data, usually one or more two-dimensional images that contain all elements in the scene, including human action instances.

[0071] In one embodiment, when rendering the basic image data, it is determined whether there is a human action instance for each pixel. When there is a human action instance, the human action instance at this pixel position is recorded. At the same time, the action type of the human action instance is also recorded, such as walking, running, jumping, etc.

[0072] In one embodiment, the bounding box of a human action instance can be obtained by calculating the pixel position of the human action instance in the basic image data.

[0073] In a specific embodiment, first, a file with a string of length 11 as the file name is randomly generated, and the name of each frame is set as a string filled with 6-digit frame numbers. Parameters such as the viewport size (Xv, Yv) of the rendered frame and the rendering frame rate of 30 FPS (Frames Per Second) are also set. After configuring the rendering settings, enter the rendering process to obtain the basic image data and the discriminator information and action list information of the action instances recorded in each frame of pixels. After the recording is completed, according to the discriminator information of each action instance, record the position (x, y) of the pixels belonging to each action instance, and calculate the minimum xmin value of the pixel position in the X-axis direction, the maximum xmax value of the pixel position in the X-axis direction, the minimum ymin value of the pixel position in the Y-axis direction, and the maximum ymax value of the pixel position in the Y-axis direction. Record it as the bounding box [xmin, xmax, ymin, ymax] of the current frame's current action instance. At the same time, the action type list of the current frame's current action instance is also recorded. Finally, summarize the bounding box and action type marking data of all action instances in all frames and output them as a txt format file.

[0074] In the above embodiment, the scene animation is rendered through preset rendering parameters to obtain basic image data, and it is determined whether there is a human action instance for each frame of pixels, and the human action instance and its action type information are recorded; finally, the action instance bounding box is obtained according to the pixel position. There is no need for manual data annotation, which avoids problems such as incorrect annotation and inconsistent annotation caused by manual experience errors, and improves the accuracy of the dataset.

[0075] S103. Process the initial dataset to obtain model parameters, and generate an augmented dataset based on a preset image generation large model and the model parameters as the target training dataset.

[0076] In one embodiment, the model parameters include a first model parameter and a second model parameter. First, the target parameter data in the initial dataset is processed based on an image generation control model to obtain the first model parameter of the large image generation model. Then, the preset sample requirements are processed to obtain an image generation prompt, and according to the image generation prompt, a target prompt information is comprehensively formed, and the target prompt information is used as the second model parameter. Among them, the image generation control model can be a ControlNet model or other models set by the user according to actual needs.

[0077] In one embodiment, first, an augmented video of the target action instance is generated through a preset large image generation model, a first model parameter, and a second model parameter, and then the augmented video is processed to obtain an augmented dataset.

[0078] In a specific embodiment, according to the large image generation model, in combination with the first model parameter and the second model parameter, various augmentation operations can be performed on the original video frame sequence (i.e., the video frames in the initial dataset mentioned in the above embodiment), and the action instance video can be regenerated by using AIGC (Artificial Intelligence Generated Content) technology, so as to generate a richer and more representative augmented dataset. The augmented dataset can be used as a target training dataset for training a deep learning model to improve the performance and accuracy of the model.

[0079] In the above embodiment, the target action instance can be generated according to actual needs, the initial dataset can be obtained according to the target action instance, and the initial dataset can be generalized by means of a large vision model to generate an augmented dataset as the target training data, which improves the use value of the data in the field of AI, solves the defects of poor generalization, small scale, and high generation cost of traditional simulation data; and this generation method does not rely on manually labeled data, avoiding problems such as incorrect labeling and inconsistent labeling caused by manual experience errors, and solves the defects of difficulty in scaling up, poor accuracy, high cost, and low efficiency of manually labeled data, and improves the data scale, accuracy, data generalization, and construction efficiency of the dataset.

[0080] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a method for augmenting a training dataset of an action recognition model provided by an embodiment of the present application.

[0081] As Figure 4 shown, the method for augmenting the training dataset of the action recognition model specifically includes steps S201 to S203.

[0082] S201. Process the action description information based on the action generation large model to generate an initial action event;

[0083] S202. Perform bone redirection based on the action bones corresponding to the initial action event to obtain a target action event;

[0084] S203. Control the movement of the character model based on the target action event to generate the target action instance.

[0085] In one embodiment, performing bone redirection based on the action bones corresponding to the initial action event to obtain a target action event includes: converting the rotation weight of the original bone into a rotation matrix based on a first preset formula, and converting the rotation matrix into a quaternion, where the original bone is the bone corresponding to the initial action event; redirecting the quaternion corresponding to the original bone into a preset target bone to obtain a target action event.

[0086] In one embodiment, the action description information is a specific description of the initial action event, which can be freely set by the user according to actual needs.

[0087] In a specific embodiment, use the action generation large model to process the action description information to generate an initial action event. Among them, the human bones of the initial action event can be in the bone format of the SMPL (Skinned Multi-Person Linear Model, parametric human model) model. Then, redirect the SMPL bones in the initial action event to the bone format of the Mixamo model (or other bone models and bone formats set by the user according to actual needs) to obtain a target action event.

[0088] In a specific embodiment, clean the action data of the generated initial action event, and obtain bone rotation data with a data format of [24, 6, frames] and root node displacement data with a data format of [3, frames]. Then, convert the bone rotation data into axis / angle representation and match it one by one with 24 nodes of SMPL. Finally, obtain SMPL format action rotation data with a data format of [24, 3, frames] and root node displacement data with a data format of [3, frames]. Convert the rotation weight d6 of all the original bones into a rotation matrix [b1, b2, b3] according to the first preset formula, where the first preset formula is:

[0089]

[0090] After calculating each key rotation matrix, convert it into the quaternion mode of [w, x, y, z], where a fine adjustment is made to the redirection of the quaternion rotation values of each bone to obtain the target action event. Control the movement of the human body model according to the target action event. After the bones are bound, generate a target action instance and save the target action instance in the fbx format.

[0091] Specifically, clear the original action stack in the Mixamo model and reset the action stack and action layers. Convert each bone data of each frame of the action data recorded in SMPL format into quaternions, and redirect the quaternion rotation [w1, x1, y1, z1] of the bone Pelvis in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone Hips in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone Spine1 in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone Spine in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone Spine2 in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone Spine1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone Spine3 in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone Spine2 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone Neck in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone Neck in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone Head in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone Head in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Collar in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftShoulder in Mixamo, where w2 = w1, x2 = x1 + 0.1, y2 = y1 - 0.2, z2 = z1 + 0.4. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Collar in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightShoulder in Mixamo, where w2 = w1, x2 = x1 + 0.1, y2 = y1 + 0.2, z2 = z1 - 0.4.Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Shoulder in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftArm in Mixamo, where w2 = w1, x2 = x1 + 0.5, y2 = y1 + 0.2, z2 = z1 + 0.2. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Shoulder in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightArm in Mixamo, where w2 = w1, x2 = x1 + 0.5, y2 = y1 - 0.3, z2 = z1 - 0.2. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Elbow in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftForeArm in Mixamo, where w2 = w1, x2 = x1, y2 = y1 + 0.6, z2 = z1 + 0.1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Elbow in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightForeArm in Mixamo, where w2 = w1, x2 = x1, y2 = y1 - 0.6, z2 = z1 - 0.1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Wrist in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftHand in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Wrist in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightHand in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftHandThumb1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftHandIndex1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1.Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftHandMiddle1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftHandRing1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftHandPinky1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightHandThumb1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightHandIndex1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightHandMiddle1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightHandRing1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Hand in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightHandPinky1 in Mixamo, where w2 = w1, x2 = x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Hip in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftUpLeg in Mixamo, where w2 = w1, x2 = -x1, y2 = y1, z2 = z1.Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Knee in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftLeg in Mixamo, where w2 = w1, x2 = -x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Ankle in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftFoot in Mixamo, where w2 = w1, x2 = -x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone L_Foot in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone LeftToe_End in Mixamo, where w2 = w1, x2 = -x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Hip in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightUpLeg in Mixamo, where w2 = w1, x2 = -x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Knee in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightLeg in Mixamo, where w2 = w1, x2 = -x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Ankle in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightFoot in Mixamo, where w2 = w1, x2 = -x1, y2 = y1, z2 = z1. Redirect the quaternion rotation [w1, x1, y1, z1] of the bone R_Foot in SMPL to the quaternion rotation [w2, x2, y2, z2] of the bone RightToe_End in Mixamo, where w2 = w1, x2 = -x1, y2 = y1, z2 = z1. Complete the redirection of the action, and finally export the result in fbx format.

[0092] In the above embodiment, an initial action event is generated according to the action description information, and then, bone redirection is performed on the action bones corresponding to the initial action event to generate a target action event, and the target action event can be generated according to user needs, improving the usability and comprehensiveness of the data set.

[0093] Please refer to Figure 5 , Figure 5 is a schematic flowchart of a data augmentation method for an action recognition model training data set provided by an embodiment of the present application.

[0094] AsFigure 5 As shown in Figure 5 , the data augmentation method for the training dataset of the action recognition model specifically includes steps S301 to S303.

[0095] S301. Process the target parameter data in the initial dataset based on the image generation control model to obtain the first model parameters of the large image generation model.

[0096] Further, the process of processing the target parameter data in the initial dataset based on the image generation control model to obtain the first model parameters of the large image generation model includes: processing the scene depth image data in the initial dataset based on the first sub - model of the image generation control model to obtain depth control parameters; processing the skeletal joint point information in the initial dataset based on the second sub - model of the image generation control model to obtain pose control parameters; processing the video frames in the initial dataset based on a preset line drawing extraction processor to obtain line drawing images, and extracting line drawing control parameters from the line drawing images based on the third sub - model of the image generation control model; extracting normal vector control parameters from the scene normal vector image data in the initial dataset based on the fourth sub - model of the image generation control model; taking the depth control parameters, the pose control parameters, the line drawing control parameters, and the normal vector control parameters as the first model parameters.

[0097] In this embodiment, the image generation control model can be a ControlNet model. The ControlNet model includes a first sub - model: controlnet_Depth model, a second sub - model: controlnet_0penpose model, a third sub - model: controlnet_Lineart model, and a fourth sub - model: controlnet_Normal model; or it can be other image generation control models set by users according to actual needs, which are used to extract depth control parameters from scene depth image data, extract pose control parameters from skeletal joint point information, extract line drawing control parameters from line drawing images, and extract normal vector control parameters from scene normal vector image data.

[0098] In one embodiment, the skeletal joint point information can be used to track and identify human action instances in the scene. By combining these marked points with the basic image data, a pose image of the human skeleton can be generated, and then pose control parameters can be obtained.

[0099] In a specific embodiment, the ControlNet_Depth model is used to process the scene depth image data to extract depth control parameters. The human body bone joint point information is used to track and identify the human action instances in the scene to generate an openpose pose image. Then, the ControlNet_OpenPose model is used to extract pose control parameters from the openpose pose image. The realistic Lineart processor is used to process the images in the original video to generate line drawing images. Then, the ControlNet_Lineart model is used to extract line drawing control parameters from the line drawing images. The ControlNet_Normal model is used to extract normal vector control parameters from the scene normal vector image data.

[0100] S302. Based on the preset sample requirements, obtain an image generation prompt, process the image generation prompt to obtain target prompt information, and use the target prompt information as the second model parameter of the image generation large model.

[0101] Further, the obtaining of the image generation prompt based on the preset sample requirements includes: performing semantic parsing on the sample requirements to obtain at least one key information corresponding to the sample requirements; constructing at least one of the image generation prompts based on each of the key information; wherein the image generation prompt includes a reverse inference prompt, a style prompt, and a custom prompt.

[0102] In one embodiment, the sample requirements can be custom text obtained according to user input content, and can include descriptions or annotation information of the target training data set, such as annotations of scenes, actions, or targets.

[0103] In one embodiment, the reverse inference prompt is used to guide the model to reverse-infer the input conditions from the output. The style prompt is used to specify the style or tone of the image. The custom prompt is a prompt customized according to specific requirements. The various image generation prompts are combined to obtain the target prompt information.

[0104] S303. Based on the image generation large model, the first model parameter, and the second model parameter, generate the augmented data set, and use the augmented data set as the target training data set.

[0105] Further, the generating of the augmented data set by processing the first model parameter and the second model parameter based on the image generation large model includes: generating an augmented video of the target action instance based on the image generation large model, the first model parameter, and the second model parameter; processing the augmented video to obtain the augmented data set.

[0106] In one embodiment, by comprehensively using information such as depth control parameters, attitude control parameters, normal vector control parameters, line drawing control parameters, and augmentation control parameters, various augmentation operations can be performed on the original video frame sequence, and the action instance video can be regenerated using AIGC (Artificial Intelligence Generated Content) technology, thereby generating a richer and more representative augmented dataset. The augmented dataset can be used as the target training dataset to train the deep learning model and improve the performance and accuracy of the model.

[0107] In a specific embodiment, as Figure 6 described, Figure 6 FIG. is a schematic diagram of the data augmentation video generation process for the action recognition model training dataset provided by the embodiment of the present application. Through various ControlNet models, the scene depth image data (i.e., the depth data in Figure 6 ), the skeleton key point information representing the action (i.e., the pose estimation data in Figure 6 ), the scene normal vector image data (i.e., the normal vector data in Figure 6 ), and the line drawing data of the original video (i.e., the lineart data in Figure 6 ) are processed to obtain the first model parameters. And the reverse prompt word, style prompt word, and custom prompt word are processed to obtain the final target prompt information as the second model parameters. Load the IPAdapter, text-to-image diffusion model (such as the SD1.5 base model, DreamShaper model, etc.), and the Animatediff model to perform video generation. During the video generation process, the generated video is precisely controlled according to the first model parameters (depth control parameters, attitude control parameters, normal vector control parameters, line drawing control parameters) and the second model parameters (augmentation control parameters). The original video frame sequence containing action instances is converted into a low-dimensional abstract representation using VAE (Variational AutoEncoder) and input into the Ksampler module to generate the augmented video. Finally, the video is optimized (such as fixing hands and faces) through the SEGSDetailer method to complete the generation of the augmented video of the action event instance. Finally, the generated video is sliced into images and saved in the jpg format, and finally the dataset is expanded to obtain the expanded dataset and used as the target training dataset.

[0108] Among them, the ControlNet_Depth model in the Controlnet model is a deep learning model specifically designed to utilize depth image data to influence the image generation model. It can fully utilize the information of depth images to control the video and the results of image generation. The ControlNet_OpenPose model in the Controlnet model is a deep learning model that specifically uses human pose estimation data to influence the image generation model. It can use the key points and pose data of the human body in the image to control the video and the results of image generation. The ControlNet_Lineart model in the Controlnet model is a deep learning model that specifically uses image line drawing data to influence the image generation model. It can use the line drawing or sketch data of the image to control the video and the results of image generation. The ControlNet_Normal model in the Controlnet model is a deep learning model that specifically uses image normal vector data to influence the image generation model. It can use the normal vector data of the objects in the image to control the video and the results of image generation. IPAdapter is a deep learning model used to add additional information to influence the graphics output model. It enables the pre-trained text-to-image diffusion model to generate images through image prompts, thereby generating more controllable and accurate outputs. The SD1.5 model, namely Stable Diffusion 1.5, is a text-to-image diffusion model. The SD1.5 base model and the Animatediff model are usually used for video generation tasks. Among them, the SD1.5 base model is a versatile text-to-image diffusion model used to generate the basic structure and content of the video, while the SEGS Detailer method is used to refine and optimize the details. The Ksampler module is a tool for advanced sampling operations, usually used to control the sampling strategies and methods in the image generation process. It allows users to customize the sampling process through various parameters, generate new data samples by operating on the representation of the latent space, adjusting the conditions, and adjusting the noise level.

[0109] In the above embodiments, the initial dataset is generalized by means of the large vision model to achieve diverse augmentation, and the augmented dataset is generated as the target training data, solving the defect of poor generalization of traditional simulation data and improving the generalization ability and effectiveness of the dataset. Moreover, the method does not rely on manually labeled data, avoiding problems such as incorrect labeling and inconsistent labeling caused by manual experience errors, and solving the defects of small scale, poor accuracy, high cost, and low efficiency of manually labeled data, improving the data scale, accuracy, data generalization, and construction efficiency of the dataset.

[0110] Please refer to Figure 7 , Figure 7FIG. 0 is a schematic block diagram of a data augmentation device for an action recognition model training data set provided by an embodiment of the present application. The data augmentation device for the action recognition model training data set is used to execute the data augmentation method for the action recognition model training data set described above. Among them, the data augmentation device for the action recognition model training data set can be configured in a server.

[0111] As Figure 7 shown, the data augmentation device 400 for the action recognition model training data set includes:

[0112] A target action instance generation module 401, configured to process preset action description information based on a preset action generation large model to generate a target action instance;

[0113] An initial data set generation module 402, configured to construct a scene animation based on the target action instance, process the scene animation to obtain target parameter data, and generate an initial data set based on the target parameter data;

[0114] A target training data set generation module 403, configured to process the initial data set to obtain model parameters, and generate an augmented data set as a target training data set based on a preset image generation large model and the model parameters.

[0115] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0116] The above device can be implemented in the form of a computer program, and the computer program can run on a computer device as Figure 8 shown.

[0117] Please refer to Figure 8 , Figure 8 FIG. is a schematic structural block diagram of a computer device provided by an embodiment of the present application. The computer device can be a server.

[0118] Referring to Figure 8 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.

[0119] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any data augmentation method for the action recognition model training data set.

[0120] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0121] The internal memory provides an environment for the operation of a computer program in a non-volatile storage medium. When the computer program is executed by a processor, the processor can execute a data augmentation method for any action recognition model training dataset.

[0122] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 8 The structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0123] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0124] Among them, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0125] Based on a preset action generation large model, process the preset action description information to generate a target action instance;

[0126] Based on the target action instance, construct a scene animation, process the scene animation to obtain target parameter data, and generate an initial dataset based on the target parameter data;

[0127] Process the initial dataset to obtain model parameters, and generate an augmented dataset as the target training dataset based on a preset image generation large model and the model parameters.

[0128] In one embodiment, when the processor realizes generating a large model based on a preset action and processing the preset action description information to generate a target action instance, it is used to realize:

[0129] Generate a large model based on the action, process the action description information, and generate an initial action event;

[0130] Perform bone redirection based on the action skeleton corresponding to the initial action event to obtain a target action event;

[0131] Based on the target action event, control the movement of the character model to generate the target action instance.

[0132] In one embodiment, the target parameter data includes video frames, synchronized annotation frame data, and camera data. When the processor realizes building a scene animation based on the target action instance, processing the scene animation, obtaining the target parameter data, and generating an initial data set based on the target parameter data, it is used to realize:

[0133] Transmit the target action instance to a preset scene, and generate the scene animation based on the motion state of the target action instance and the preset scene;

[0134] Render the preset scene and the scene animation based on a preset rendering engine to obtain the video frames corresponding to the scene animation, and correspond the scene data of the target action instance and the preset scene with the video frames to obtain synchronized annotation frame data and camera data with spatio-temporal parameter alignment;

[0135] Obtain the initial data set based on the video frames, the synchronized annotation frame data, and the camera data.

[0136] In one embodiment, the synchronized annotation frame data includes action labels, target model bounding boxes, skeletal joint point information representing actions, scene depth image data, and scene normal vector image data.

[0137] In one embodiment, when the processor realizes processing the initial data set to obtain model parameters, and generating an augmented data set as the target training data set based on a preset image generation large model and the model parameters, it is used to realize:

[0138] Process the target parameter data in the initial data set based on an image generation control model to obtain the first model parameters of the image generation large model;

[0139] Obtain an image generation prompt word based on a preset sample requirement, process the image generation prompt word to obtain target prompt information, and use the target prompt information as the second model parameters of the image generation large model;

[0140] Generate the augmented data set based on the image generation large model, the first model parameters, and the second model parameters, and use the augmented data set as the target training data set.

[0141] In one embodiment, when the processor implements processing the target parameter data in the initial data set by the image generation control model to obtain the first model parameters of the image generation large model, it is used to implement:

[0142] Process the scene depth image data in the initial data set based on the first sub-model of the image generation control model to obtain depth control parameters;

[0143] Process the skeletal joint point information in the initial data set based on the second sub-model of the image generation control model to obtain pose control parameters;

[0144] Process the video frames in the initial data set based on a preset line drawing extraction processor to obtain line drawing images, and extract line drawing control parameters from the line drawing images based on the third sub-model of the image generation control model;

[0145] Extract normal vector control parameters from the scene normal vector image data in the initial data set based on the fourth sub-model of the image generation control model;

[0146] Use the depth control parameters, the pose control parameters, the line drawing control parameters, and the normal vector control parameters as the first model parameters.

[0147] In one embodiment, when the processor implements processing the first model parameters and the second model parameters based on the image generation large model to generate the augmented data set, it is used to implement:

[0148] Generate an augmented video of the target action instance based on the image generation large model, the first model parameters, and the second model parameters;

[0149] Process the augmented video to obtain the augmented data set.

[0150] In one embodiment, when the processor implements obtaining an image generation prompt based on a preset sample requirement, it is used to implement:

[0151] Perform semantic parsing on the sample requirement to obtain at least one key information corresponding to the sample requirement;

[0152] Construct at least one of the image generation prompts based on each of the key information;

[0153] Perform natural language processing on each image generation prompt to obtain the target prompt information;

[0154] Among them, the image generation prompts include reverse prompts, style prompts, and custom prompts.

[0155] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement the data augmentation method of any action recognition model training dataset provided by the embodiments of the present application.

[0156] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0157] As mentioned above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data augmentation method for an action recognition model training dataset, characterized in that Including: Obtain an initial data set; Based on the first sub-model of the image generation control model, process the scene depth image data in the initial data set to obtain depth control parameters; Based on the second sub-model of the image generation control model, process the skeletal joint point information in the initial data set to obtain pose control parameters; Based on the third sub-model of the image generation control model, extract line drawing control parameters from the line drawing image corresponding to the initial data set; Based on the fourth sub-model of the image generation control model, extract normal vector control parameters from the scene normal vector image data in the initial data set; Use the depth control parameters, the pose control parameters, the line drawing control parameters, and the normal vector control parameters as the first model parameters; Based on a preset sample requirement, obtain an image generation prompt, process the image generation prompt to obtain target prompt information, and use the target prompt information as the second model parameter of the image generation large model; Generate an augmented data set based on the image generation large model, the first model parameter, and the second model parameter.

2. The data augmentation method for the action recognition model training dataset according to claim 1, wherein The generating the augmented data set by processing the first model parameter and the second model parameter based on the image generation large model includes: Generate an augmented video of the target action instance based on the image generation large model, the first model parameter, and the second model parameter; Process the augmented video to obtain the augmented data set.

3. The data augmentation method for the action recognition model training dataset according to claim 2, wherein The processing the augmented video to obtain the augmented data set includes: Optimize the augmented video to obtain a target video; Segment and save the images of the target video to obtain the augmented data set.

4. The data augmentation method for the action recognition model training dataset according to claim 1, wherein Before extracting the line drawing control parameters from the line drawing image corresponding to the initial data set based on the third sub-model of the image generation control model, further include: Process the video frames in the initial data set based on a preset line drawing extraction processor to obtain the line drawing image.

5. The data augmentation method for the action recognition model training dataset according to claim 1, wherein The obtaining the image generation prompt based on a preset sample requirement includes: Perform semantic parsing on the sample requirement to obtain at least one key information corresponding to the sample requirement; Construct at least one image generation prompt based on each key information; Wherein, the image generation prompt includes a reverse inference prompt, a style prompt, and a custom prompt.

6. The data augmentation method for the action recognition model training dataset according to claim 1, wherein The obtaining the initial data set includes: Process a preset action description information based on a preset action generation large model to generate a target action instance; Construct a scene animation based on the target action instance, process the scene animation to obtain target parameter data, and generate an initial data set based on the target parameter data.

7. The data augmentation method for the action recognition model training dataset according to claim 6, wherein The target parameter data includes video frames, synchronized annotation frame data, and camera data, and the synchronized annotation frame data includes action labels, target model bounding boxes, skeletal joint point information representing actions, scene depth image data, and scene normal vector image data.

8. The data augmentation method for the action recognition model training data set according to any one of claims 1-7, characterized in that, The second sub-model based on the image generation control model processes the bone joint point information in the initial data set to obtain pose control parameters, including: Based on the second sub-model, the bone joint point information is identified to generate a pose image, and the pose control parameters are extracted from the pose image.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the data augmentation method for the action recognition model training data set as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the processor is caused to implement the data augmentation method for the action recognition model training data set as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image data augmentation method and device

    CN113222114A

  • Method and device for generating training data of human skeleton joint point extraction model

    CN114359445A

  • Monocular video dynamic human body three-dimensional reconstruction method based on attitude optimization

    CN116681838A