Image recognition generation method based on action recognition capture

Through the synchronous acquisition and temporal calibration of color cameras and depth cameras, combined with the space-time grouping self-attention model and knowledge graph, the consistency of action feature extraction and complex action classification are solved, and the scene adaptability and intention accuracy of generated images are achieved, which is suitable for motion training and virtual character generation.

CN120388423AInactive Publication Date: 2025-07-29SHIJIAZHUANG TIEDAO UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510583544.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently integrate color vision and depth information to improve the space-time consistency of action feature extraction, and it is difficult to optimize the self-attention model through space-time grouping strategy to improve the accuracy of complex action classification. It lacks accurate reasoning and situational correlation of action intentions, resulting in the generated image not meeting the action scenes and intentions.

Method used

The color camera and the depth camera are used to synchronize data collection, combine time synchronization calibration and efficient data transmission, and use the space-time grouping self-attention model and knowledge graph for action classification and intention reasoning, and combine the situational knowledge base and template matching to generate images.

Benefits of technology

It realizes the accurate extraction of action space and time features, improves the accuracy of complex action classification and scene adaptability of generated images. The generated images are in line with the action itself and intentions, and are suitable for applications such as sports training and virtual character generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388423A_ABST
    Figure CN120388423A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition generation method based on action recognition capture, and belongs to the technical field of image recognition and generation. Through time synchronization calibration and efficient data transmission, color and depth information synchronization is ensured, the accuracy of action space and time feature extraction is improved, meanwhile, the feature capture capacity of complex actions is enhanced through a space-time grouping self-attention model, and the feature extraction efficiency is improved by combining a knowledge graph and machine learning. According to the method, deep semantic understanding from action classification to intention reasoning is realized, and a situation knowledge base and a template matching mechanism are applied, so that the generated image not only conforms to the action, but also can accurately reflect the scene and intention of action occurrence, and the adaptability of application scenes, such as exercise training visualization and virtual role generation, is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition and generation, and in particular to an image recognition and generation method based on motion recognition capture. Background Art

[0002] With the vigorous development of computer vision technology, motion recognition and image generation technologies are showing increasingly strong application demands in many cutting-edge fields such as intelligent interaction, virtual reality, and film and television production. In the field of intelligent interaction, precise motion recognition and image generation can achieve more natural and efficient human-computer interaction, such as intelligent somatosensory games and intelligent security monitoring; in virtual reality, it can provide users with a more realistic immersive experience and enhance the realism and interactivity of virtual scenes; in film and television production, it can help to achieve more efficient special effects production and virtual character creation, reduce production costs and improve visual effects.

[0003] 1. How to efficiently integrate color vision and depth information to improve the spatiotemporal consistency of motion feature extraction?

[0004] 2. How to optimize the self-attention model through spatiotemporal grouping strategy to improve the accuracy of complex action classification?

[0005] 3. How to combine the action intention knowledge graph with machine learning algorithms to achieve accurate reasoning of action intention?

[0006] 4. How to build a contextual knowledge base and match contextual templates to improve the contextual association of action semantics to generate semantically consistent images?

[0007] In summary, in order to meet the higher requirements of action recognition and image generation technology in various fields and solve the problems existing in traditional methods, an image recognition generation method based on action recognition capture is proposed. Summary of the invention

[0008] In view of this, the embodiments of the present invention hope to provide an image recognition generation method based on motion recognition capture to solve or alleviate the technical problems existing in the prior art and at least provide a beneficial option.

[0009] The technical solution of the embodiment of the present invention is implemented as follows: a method for generating image recognition based on motion recognition capture, the method comprising:

[0010] Multimodal data acquisition and fusion: Using a camera system consisting of a color camera and a depth camera, we simultaneously capture color illumination images and depth information of human motion. Based on the visually captured images, we extract the spatial and temporal features of the human motion.

[0011] Action Classification and Motion Feature Analysis: Classify the spatial and temporal features of human actions using an intra-group self-attention model based on spatio-temporal grouping. The intra-group self-attention model based on spatio-temporal grouping adopts a self-attention mechanism combined with a spatio-temporal two-dimensional grouping strategy;

[0012] Action Intention Inference: Construct an action intention knowledge graph, and at the same time input the preliminarily classified action categories into the action intention inference model;

[0013] Situation Association Enhancement: Establish a situation knowledge base, and according to the scene information where the action occurs and the enhanced action semantics, match the corresponding situation template from the situation knowledge base to enhance the situation association of the action semantics;

[0014] Image Generation: Convert the deeply understood action semantics into control parameters for image generation, and input the control parameters into the image generation model to generate an image that conforms to the semantics.

[0015] As a further preference of this technical solution: Use a camera group composed of a color camera and a depth camera to synchronously collect the color illumination image and depth information of human actions, specifically including:

[0016] Perform time synchronization calibration on the color camera and the depth camera to ensure that the data collected by both strictly corresponds on the time axis;

[0017] Adopt a specific data interface and transmission protocol to ensure the synchronization and data integrity of the color illumination image and depth information during the collection process;

[0018] As a further preference of this technical solution: The extraction methods of the spatial and temporal features of the human action include:

[0019] For spatial feature extraction, use a convolutional neural network to process the color illumination image to identify spatial information such as the positions of human key points and limb postures in the image;

[0020] For temporal feature extraction, use the optical flow method or a recurrent neural network to analyze the consecutive frames of the color illumination image to capture the change trend of human actions in the time series;

[0021] As a further preference of this technical solution: The spatio-temporal two-dimensional grouping strategy is specifically as follows:

[0022] In the spatial dimension, divide the human action image into multiple local regions, and the pixel points in each region form a spatial group;

[0023] In the temporal dimension, divide the consecutive action frames into several time segments, and the frames in each time segment form a time group;

[0024] Apply the self-attention mechanism within the group to perform weighted fusion on the features within the spatial group and the temporal group respectively, so as to highlight the key spatial and temporal features;

[0025] As a further preference of this technical solution: The specific steps of constructing the action intention knowledge graph include:

[0026] Collect a large number of action sample data, including information such as action categories, action descriptions, action execution scenarios, etc.;

[0027] Adopt knowledge graph construction technology to model the entities in the action sample data, such as actions, scenarios, objects, etc., as well as the associations between actions and scenarios and the interaction relationships between actions and objects, to form an action intention knowledge graph;

[0028] As a further preference of this technical solution: The specific application of the action intention inference model includes:

[0029] Use machine learning algorithms to learn and train the action intention knowledge graph to obtain a basic model for action intention inference;

[0030] During the inference process, take the action category, the time and location information where the action occurs as input features, input them into the trained basic model, and calculate the specific intention of the action through the model;

[0031] As a further preference of this technical solution: The establishment of the situation knowledge base and according to the action occurrence scenario information and enhanced action semantics specifically include:

[0032] Collect situation information in various scenarios, including scenario descriptions, objects that may appear in the scenario, common associations between scenarios and actions, etc.;

[0033] Perform structured processing on the collected situation information and store it in the database to form a situation knowledge base;

[0034] As a further preference of this technical solution: The matching of the corresponding situation template from the situation knowledge base according to the action occurrence scenario information and enhanced action semantics to improve the situation association of action semantics specifically includes:

[0035] Extract features from the action occurrence scenario information to obtain a scenario feature vector;

[0036] Encode the enhanced action semantics to obtain an action semantics vector;

[0037] As a further preference of this technical solution: The conversion of the deeply understood action semantics into control parameters for image generation specifically includes:

[0038] Semantically parse the action semantics after in-depth understanding, and extract key semantic information, such as the action subject, action object, action state, situational information, etc.;

[0039] According to the preset mapping rules between semantics and control parameters, convert the key semantic information into control parameters required by the image generation model, such as parameters for color, shape, texture, layout, etc.;

[0040] As a further preference of this technical solution: The image generation model is any one of a generative adversarial network, a variational autoencoder, or a diffusion model. During the process of generating an image that conforms to the semantics, the image generation model adjusts the features of the generated image according to the input control parameters to make the generated image match the action semantics.

[0041] Due to the adoption of the above technical solutions in the embodiments of the present invention, it has the following advantages:

[0042] First, through time synchronization calibration and efficient data transmission, the present invention ensures the synchronization of color and depth information, improves the accuracy of action space and time feature extraction. At the same time, the spatio-temporal grouped self-attention model enhances the ability to capture features of complex actions. Combining knowledge graphs and machine learning, it realizes in-depth semantic understanding from action classification to intention reasoning, and uses a situational knowledge base and a template matching mechanism to make the generated image not only conform to the action itself, but also accurately reflect the scene and intention where the action occurs, improving the adaptability of the application scenario, such as sports training visualization, virtual character generation, etc.

[0043] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present invention will be readily apparent by reference to the drawings and the following detailed description. Brief Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a flowchart of an image recognition and generation method based on action recognition and capture of the present invention. Detailed Embodiments

[0046] In the following text, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present invention. Therefore, the drawings and the description are considered to be exemplary in nature rather than restrictive.

[0047] It should be noted that terms such as "first", "second", "symmetric", "array", etc. are only used for the purpose of distinguishing descriptions and position descriptions, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "symmetric", etc. can explicitly or implicitly include one or more of such features; similarly, when there is no numerical limitation on certain features in the form of words such as "two", "three", etc., it should be noted that such features also explicitly or implicitly include one or more feature quantities;

[0048] In the present invention, unless otherwise clearly specified and defined, terms such as "installation", "connection", "fixation", etc. should be understood in a broad sense; for example, it can be a fixed connection, a detachable connection, or an integral molding; it can be a mechanical connection, a direct connection, a welding connection, or an indirect connection through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the accompanying drawings of the specification in combination with specific situations.

[0049] Partial technical problems, main solutions, secondary technical problems and corresponding methods

[0050] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0051] As Figure 1 shown, the embodiments of the present invention provide an image recognition generation method based on action recognition capture, and the method includes:

[0052] Multi-modal data collection and fusion: Using a camera group composed of a color camera and a depth camera, synchronously collect the color illumination image and depth information of the human body movement; and based on the visual capture screen, extract the spatial features and temporal features of the human movement.

[0053] Action classification and motion feature analysis: Use the intra-group self-attention model based on spatio-temporal grouping to classify the spatial features and temporal features of the human movement, and the intra-group self-attention model based on spatio-temporal grouping adopts the self-attention mechanism combined with the spatio-temporal two-dimensional grouping strategy.

[0054] Action intention reasoning: Construct an action intention knowledge graph, and at the same time input the action categories obtained by preliminary classification into the action intention reasoning model.

[0055] Context - related refinement: Establish a context knowledge base, and based on the scene information where the action occurs and the enhanced action semantics, match the corresponding context templates from the context knowledge base to refine the context association of the action semantics;

[0056] Image generation: Convert the deeply understood action semantics into control parameters for image generation, and input the control parameters into the image generation model to generate images that conform to the semantics.

[0057] Specifically, use a camera group composed of a color camera and a depth camera to synchronously collect the color illumination image and depth information of human actions, specifically including:

[0058] Perform time - synchronization calibration on the color camera and the depth camera to ensure that the data collected by both is strictly corresponding on the time axis;

[0059] Adopt specific data interfaces and transmission protocols to ensure the synchronization and data integrity of the color illumination image and depth information during the acquisition process.

[0060] Specifically, the methods for extracting the spatial and temporal features of human actions include:

[0061] For spatial feature extraction, use a convolutional neural network to process the color illumination image and identify spatial information such as the positions of key points of the person and limb postures in the image;

[0062] For temporal feature extraction, use the optical flow method or a recurrent neural network to analyze consecutive frames of the color illumination image to capture the changing trend of human actions in the time series.

[0063] Specifically, the spatio - temporal two - dimensional grouping strategy is as follows:

[0064] In the spatial dimension, divide the human action image into multiple local regions, and the pixel points within each region form a spatial group;

[0065] In the temporal dimension, divide consecutive action frames into several time segments, and the frames within each time segment form a temporal group;

[0066] Apply the self - attention mechanism within the group to perform weighted fusion on the features within the spatial group and the temporal group respectively to highlight the key spatial and temporal features.

[0067] Specifically, the specific steps for constructing an action intention knowledge graph include:

[0068] Collect a large amount of action sample data, including information such as action categories, action descriptions, and action execution scenarios;

[0069] Adopt knowledge graph construction technology to model entities in action sample data, such as actions, scenes, objects, etc., as well as the association between actions and scenes and the interaction relationship between actions and objects, to form an action intention knowledge graph.

[0070] Specifically, the specific applications of the action intention inference model include:

[0071] Use machine learning algorithms to learn and train the action intention knowledge graph to obtain a basic model for action intention inference;

[0072] During the inference process, use the action category, the time and location information where the action occurs as input features, input them into the trained basic model, and calculate the specific intention of the action through the model.

[0073] Specifically, establish a situation knowledge base, and according to the scene information where the action occurs and the enhanced action semantics, specifically including:

[0074] Collect situation information in various scenarios, including scene descriptions, objects that may appear in the scene, common associations between the scene and actions, etc.;

[0075] Structurally process the collected situation information and store it in the database to form a situation knowledge base.

[0076] Specifically, according to the scene information where the action occurs and the enhanced action semantics, match the corresponding situation template from the situation knowledge base to improve the situation association of the action semantics, specifically including:

[0077] Extract feature vectors of the scene information where the action occurs to obtain scene feature vectors;

[0078] Encode the enhanced action semantics to obtain action semantic vectors.

[0079] Specifically, transform the deeply understood action semantics into control parameters for image generation, specifically including:

[0080] Semantically analyze the deeply understood action semantics, and extract key semantic information, such as action subjects, action objects, action states, situation information, etc.;

[0081] According to the preset mapping rules between semantics and control parameters, transform the key semantic information into control parameters required by the image generation model, such as color, shape, texture, layout and other parameters.

[0082] Specifically, the image generation model is any one of a generative adversarial network, a variational autoencoder or a diffusion model. During the process of generating an image that conforms to the semantics, the image generation model adjusts the features of the generated image according to the input control parameters to make the generated image match the action semantics.

[0083] In this embodiment, the specific working mode of the present invention is as follows:

[0084] S1. Multimodal data collection and fusion:

[0085] a. Hardware calibration and data synchronization: Use a hardware trigger synchronization module to synchronize the clocks of the color camera and the depth camera, with the error controlled within ±1 ms.

[0086] b. Transmit data through Gigabit Ethernet or USB3.1 interface, and use the TCP / IP protocol stack to ensure data integrity, with the transmission rate not less than 500 MB / s.

[0087] S2. Feature extraction:

[0088] a. Spatial feature extraction: Use the ResNet+OpenPose model to extract the coordinates of 18 key points, such as shoulders, elbows, wrists, etc., and construct a limb posture matrix.

[0089] b. Temporal feature extraction: Calculate the pixel displacement between adjacent frames through the optical flow method to generate an optical flow field matrix; or use the LSTM network to analyze the posture sequence of 16 consecutive frames to capture the action speed and acceleration features.

[0090] S3. Grouping strategy for training and application of the spatio-temporal grouped self-attention model:

[0091] a. Spatial grouping: Divide the human body into 6 regions according to the anatomical structure, namely the head, torso, left arm, right arm, left leg, and right leg, and generate local feature maps for each region.

[0092] b. Temporal grouping: Divide 32 consecutive frames into 4 time segments, with 8 frames in each segment, and extract cross-frame motion features.

[0093] Among them, calculate the Query, Key, and Value matrices for each spatial group and temporal group respectively, and fuse the features through the multi-head self-attention mechanism to output an action feature vector with a dimension of 1024, and construct an action intention inference model.

[0094] S4. Knowledge graph construction:

[0095] a. By collecting action samples, define entity types such as actions, scenes, objects, and relationships such as "occurs in", "involves", "belongs to", and store them using the Neo4j graph database at the same time.

[0096] b. During the inference process: Input the action category, traverse the associated nodes in the knowledge graph through the GCN model, and output the intention probability distribution.

[0097] S5. Situation association and image generation:

[0098] a. Situation knowledge base: Structurally store scene templates, each template contains scene tags such as "indoor" and "outdoor", object lists such as "treadmill" and "basketball", and action association rules such as "running" → "sports scene";

[0099] b. Image generation control: Generate control parameters after semantic parsing. For example, for "outdoor running", the dominant hue is green, and the composition can have the person in the center and the background as a runway. Input the StyleGAN3 model to generate an image with a resolution of 1024×1024.

[0100] This method realizes the full - process intelligence from action capture to semantic image generation through multi - modal fusion, spatio - temporal feature analysis, intention reasoning, and situation association. It is applicable to fields such as motion analysis, virtual reality, and film production, and has significant advantages of high precision and high robustness.

[0101] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.

Claims

1. An image recognition generation method based on action recognition capture, characterized in that, The method includes: Multi-modal data acquisition and fusion: Using a camera group composed of a color camera and a depth camera, synchronously acquire the color illumination image and depth information of human actions; and based on the visually captured images, extract the spatial and temporal features of human actions. Action classification and motion feature analysis: Use the intra-group self-attention model based on spatio-temporal grouping to classify the spatial and temporal features of human actions. The intra-group self-attention model based on spatio-temporal grouping adopts the self-attention mechanism combined with the spatio-temporal two-dimensional grouping strategy. Action intention reasoning: Construct an action intention knowledge graph, and at the same time input the preliminarily classified action categories into the action intention reasoning model. Situation association improvement: Establish a situation knowledge base, and according to the scene information where the action occurs and the enhanced action semantics, match the corresponding situation templates from the situation knowledge base to improve the situation association of action semantics. Image generation: Convert the deeply understood action semantics into control parameters for image generation, and input the control parameters into the image generation model to generate images that conform to the semantics.

2. The image recognition generation method based on action recognition capture according to claim 1, characterized in that, The step of using the camera group composed of a color camera and a depth camera to synchronously acquire the color illumination image and depth information of human actions specifically includes: Perform time synchronization calibration on the color camera and the depth camera to ensure that the data collected by both strictly corresponds on the time axis. Adopt a specific data interface and transmission protocol to ensure the synchronization and data integrity of the color illumination image and depth information during the acquisition process.

3. The image recognition generation method based on action recognition capture according to claim 1, wherein, The methods for extracting the spatial and temporal features of human actions include: For spatial feature extraction, use a convolutional neural network to process the color illumination image to identify spatial information such as the positions of key points of the person in the image and limb postures. For temporal feature extraction, use the optical flow method or a recurrent neural network to analyze the consecutive frames of the color illumination image to capture the change trend of human actions in the time series.

4. A method for generating image recognition based on action recognition capture according to claim 1, characterized in that, The spatio-temporal two-dimensional grouping strategy is specifically: In the spatial dimension, divide the human action image into multiple local regions, and the pixel points in each region form a spatial group. In the temporal dimension, divide the consecutive action frames into several time segments, and the frames in each time segment form a time group. Apply the self-attention mechanism within the group to perform weighted fusion of the features within the spatial group and the time group respectively to highlight the key spatial and temporal features.

5. The image recognition generation method based on action recognition capture according to claim 4, wherein, The specific steps for constructing the action intention knowledge graph include: Collect a large amount of action sample data, including information such as action categories, action descriptions, and action execution scenarios. Adopt knowledge graph construction technology to model the entities in the action sample data, such as actions, scenarios, objects, etc., as well as the associations between actions and scenarios and the interaction relationships between actions and objects, to form an action intention knowledge graph.

6. A method for generating image recognition based on action recognition capture according to claim 1, characterized in that, The specific application of the action intention reasoning model includes: Use machine learning algorithms to learn and train the action intention knowledge graph to obtain a basic model for action intention reasoning. During the reasoning process, use the action category, the time and location information where the action occurs as input features, input them into the trained basic model, and calculate the specific intention of the action through the model.

7. A method for generating image recognition based on action recognition capture according to claim 1, characterized in that, The establishment of the situation knowledge base specifically includes: Collect situation information in various scenarios, including scenario descriptions, objects that may appear in the scenario, common associations between scenarios and actions, etc.; Structurally process the collected situation information and store it in a database to form a situation knowledge base.

8. The image recognition generation method based on action recognition capture according to claim 1, characterized in that, According to the scenario information where the action occurs and the enhanced action semantics, match the corresponding situation template from the situation knowledge base to improve the situation association of the action semantics, which specifically includes: Extract features from the scenario information where the action occurs to obtain a scenario feature vector; Encode the enhanced action semantics to obtain an action semantics vector; Calculate the similarity between the scenario feature vector and each situation template in the situation knowledge base, and select the situation template with the highest similarity to associate with the action semantics, thereby improving the situation association of the action semantics.

9. A method for generating image recognition based on action recognition capture according to claim 1, characterized in that Converting the deeply understood action semantics into control parameters for image generation specifically includes: Semantically analyze the deeply understood action semantics, and extract key semantic information, such as action subject, action object, action state, situation information, etc.; According to the preset mapping rules between semantics and control parameters, convert the key semantic information into control parameters required by the image generation model, such as parameters for color, shape, texture, layout, etc.

10. A method for generating image recognition based on action recognition capture according to claim 1, characterized in that, The image generation model is any one of a generative adversarial network, a variational autoencoder, or a diffusion model. During the process of generating an image that conforms to the semantics, the image generation model adjusts the features of the generated image according to the input control parameters to make the generated image match the action semantics.