Timing sequence vision-based conference training scene behavior analysis method

By using a time-series visual method in the behavior analysis of Huicai scenes, combined with Swin-Transformer, CBAM and Kuffman filters, the deep learning model Mamba based on the state space model is used for timing information modeling and behavior recognition, which solves the problem of difficult real-time behavior analysis and process intervention in the existing technology, and improves analysis efficiency and accuracy.

CN120182887APending Publication Date: 2025-06-20SHAANXI ZHIBO CHUANGKE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252338.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing behavioral analysis methods for training scenarios are difficult to achieve real-time behavioral analysis and process intervention, and are divided into multiple stages of processing, which is relatively inefficient.

Method used

The behavior analysis method of scheduling scenes based on time-series vision is adopted, and the feature extraction mode of Backone+Neck+Head is combined with Swin-Transformer, CBAM and multiple detection heads are used to perform multi-scale feature extraction and adaptive adjustment; the Kuffman filter is used to combine target object features and track real-time tracking; the deep learning model Mamba based on state space model performs timing information modeling and behavior recognition.

Benefits of technology

Real-time analysis and prediction of the behavior of the training scenario is realized, the model's abstraction and expression ability of the image content is improved, the ability to detect targets at different scales is enhanced, and multi-objective tracking and processing noise and uncertainty are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182887A_ABST
    Figure CN120182887A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of conference training scene behavior analysis, and discloses a conference training scene behavior analysis method based on time sequence vision, and the method comprises the steps: employing a Backone + Neck + Head feature extraction mode for a single video frame in a time sequence through an analysis system; according to the method, in the aspect of feature extraction on a visual frame, a Swindow-transformer module is adopted to play a role similar to convolution operation in a CNN in image processing, but the Swindow-transformer module has advantages in the aspects of global relation processing, long-distance information transmission and the like based on the characteristics of a transformer structure, and the abstract ability and expression ability of the model on image content can be effectively improved; a CBAM module is adopted, and through comprehensive utilization of channel attention and space attention mechanisms, the image feature processing and learning ability of a CNN model can be improved, the network characterization ability is enhanced, and the performance and accuracy of an image processing task are improved; the four detection heads are adopted to improve the detection capability of targets with different scales, so that the problem of missing detection or poor detection effect is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of behavior analysis in meeting and training scenarios, and specifically provides a method for behavior analysis in meeting and training scenarios based on temporal vision. Background Art

[0002] Generally, the existing methods for behavior analysis in meeting and training scenarios mainly include three parts: behavior detection, candidate video generation, and behavior recognition. However, in the existing technology, the behavior analysis in meeting and training scenarios mainly detects abnormal behaviors frame by frame for the behaviors of temporal vision, records the abnormal behaviors of the target object into json, parses the recorded json frame by frame after the event, intercepts the video frame by frame for the corresponding video, and then performs temporal behavior recognition on the candidate abnormal behavior videos of the single target object after editing to determine the abnormal behavior videos in the final meeting and training. Usually, the abnormal behavior analysis is carried out after the event. It is difficult to perform real-time behavior analysis and real-time process intervention according to the behavior analysis results. In addition, the behavior analysis needs to be processed in multiple stages. Therefore, a method for behavior analysis in meeting and training scenarios based on temporal vision is proposed. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for behavior analysis in meeting and training scenarios based on temporal vision to solve the problems raised in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A method for behavior analysis in meeting and training scenarios based on temporal vision, including the following steps:

[0005] S1. Adopt a feature extraction mode of Backone + Neck + Head for a single video frame in the time series by an analysis system;

[0006] S2. Extract features of the target object in a single video frame in the time series, fuse the single target feature and the position information in the spatial domain into a two-dimensional table according to the id automatically assigned by the system, and create a corresponding feature index;

[0007] S3. Perform temporal information modeling on the above features at each moment of the temporal vision through a deep learning model Mamba based on the state space model, and then output the inference prediction of the behavior of each target object at each moment in an end2end manner to determine whether it is the start moment and whether it is the end moment, and associate the spatial position and behavior category of the target in the video in the temporal vision frame.

[0008] Preferably, in the above S1, Swin-Transformer is introduced into the backbone to perform downsampling after each feature extraction through feature fusion, increasing the receptive field of the next window attention operation on the original image, thereby performing multi-scale feature extraction on the input image.

[0009] Preferably, in the above S1, a CBAM structure that processes the channels and spatial dimensions of the input feature map is introduced into the Neck. By inferring the corresponding attention weights and multiplying these weights with the original feature map, adaptive adjustment of the features is achieved.

[0010] Preferably, in the above S1, multiple detection heads are introduced into the Head to improve the detection ability for targets of different scales.

[0011] Preferably, in the above S2, a Kaufman filter is used to obtain the same object in the front and back time series, and model a single behavior in temporal vision.

[0012] Preferably, in the above S1, the analysis system includes a visual frame feature extraction module, a target object feature combination module, and a temporal behavior recognition module;

[0013] The visual frame feature extraction module is used to adopt the feature extraction mode of Backbone + Neck + Head for a single video frame in the time series;

[0014] The target object feature combination module is used to extract features from the target object in a single video frame in the time series, adopt a Kaufman filter to obtain the same object in the front and back time series, model a single behavior in temporal vision, fuse the single target feature and the position information in the spatial domain into a two-dimensional table according to the id automatically assigned by the system, and create a corresponding feature index;

[0015] The temporal behavior recognition module is used to perform temporal information modeling on the above features at each moment of temporal vision through the deep learning model Mamba based on the state space model, and then output the end-to-end way of the action behavior to infer and predict whether the behavior of each target object at each moment is the start moment and whether it is the end moment, as well as the spatial position and behavior category of the target associated with the video in the temporal vision frame.

[0016] Preferably, the above visual frame feature extraction module is connected to the target object feature combination module, and the target object feature combination module is connected to the temporal behavior recognition module.

[0017] Preferably, the above visual frame feature extraction module includes a Backbone module, a Neck module, and a Head module. The Backbone module is connected to the Neck module, and the Neck module is connected to the Head module;

[0018] The Backbone module uses Swin-Transformer to perform downsampling after each feature extraction through feature fusion, increasing the receptive field of the next window attention operation on the original image, thereby performing multi-scale feature extraction on the input image;

[0019] The Neck module is used to infer the corresponding attention weights and multiply these weights with the original feature map to achieve adaptive adjustment of the features;

[0020] The Head module is used to introduce multiple detection heads to improve the detection ability for targets of different scales.

[0021] Preferably, the above Neck module includes a channel attention module and a spatial attention module. The channel attention module processes the channels of the input feature map, and the spatial attention module is a CBAM structure that processes the space of the input feature map.

[0022] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0023] 1. In terms of feature extraction on visual frames, the present invention uses the Swin-Transformer module, which acts as a role similar to the convolution operation in CNN in image processing. However, its characteristics based on the Transformer structure give it advantages in processing global relationships and long-distance information transmission, and can effectively improve the model's ability to abstract and express image content; it uses the CBAM module, which can improve the CNN model's ability to process and learn image features, enhance the network's representational ability, and improve the performance and accuracy of image processing tasks through the comprehensive utilization of channel attention and spatial attention mechanisms; it uses multiple detection heads to improve the detection ability for targets of different scales, with four detection heads, to solve the problems of missed detection or poor detection effects.

[0024] 2. In the target object feature combination module of the present invention, a Kalman filter is adopted. First, it can process the target tracking algorithm in real time, is applicable to application scenarios that require quick response, and has real-time performance and high efficiency. Second, it can effectively predict the state of a dynamic system, which is crucial for tracking tasks because it can accurately predict the position of the target in the next frame, and has the ability to predict dynamic systems. Third, it combines the current observation data and the previous prediction results to update the target state and perform feature fusion, and has the ability of information fusion. Fourth, through its modeling and update mechanism, it can effectively reduce these errors, thereby improving the tracking accuracy, and has the ability to handle noise and uncertainty. Fifth, it can use multiple Kalman filters to track each target object respectively, which makes the Kalman filter very suitable for dealing with complex tracking scenarios and supports multi-target tracking.

[0025] 3. In the temporal behavior recognition module of the present invention, a deep learning model Mamba based on the state space model is adopted. First, it shows significant performance advantages when processing long sequence data, can effectively process and analyze long-term dependencies, and has the ability to process long sequence tasks. Second, compared with the attention mechanism, it can not only reduce the computational complexity, but also the memory usage does not depend on the context length, and has the effect of reducing computational complexity and memory usage. Third, by introducing a selective state space, the model can selectively transmit or forget information according to the current data, thereby solving the deficiencies of previous models when dealing with discrete and information-intensive data, and has the function of introducing a selective state space. Fourth, relying on its proficient context modeling ability, the Mamba model shows strong performance and ability in context learning applications, and has the ability to improve context modeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a schematic diagram of the recognition model process of the present invention;

[0028] Figure 2 It is a flowchart of the temporal frame feature extraction of the present invention;

[0029] Figure 3 It is a schematic diagram of the single-frame behavior target feature of the present invention;

[0030] Figure 4 It is a schematic diagram of the temporal behavior target feature fusion of the present invention. Specific Embodiments

[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0032] It should be noted that the structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions that can be implemented in this application. Therefore, they do not have technical substance significance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that this application can produce and the purposes that can be achieved, should still fall within the scope that the technical content disclosed in this application can cover.

[0033] Embodiment

[0034] Please refer to Figures 1-4 , the present invention provides a technical solution: a method for analyzing the behavior of a meeting training scene based on temporal vision, including the following steps:

[0035] S1. The analysis system adopts a feature extraction mode of Backone+Neck+Head for a single video frame in the time series; Swin-Transformer is introduced in Backone to perform downsampling once after each feature extraction through feature fusion, increasing the receptive field of the next window attention operation on the original image, thereby performing multi-scale feature extraction on the input image; a CBAM structure including processing the channels of the input feature map and the space of the input feature map is introduced in Neck. By inferring the corresponding attention weights and multiplying these weights with the original feature map, the adaptive adjustment of the features is realized; multiple detection heads are introduced in Head to improve the detection ability for targets of different scales, and four detection heads are used to solve the problems of missed detection or poor detection effect.

[0036] S2. Extract features of the target object in a single video frame in the time series, adopt a Kaufman filter to obtain the same object in the front and back time series, model a single behavior in the temporal vision, fuse the single target feature and the position information in the spatial domain into a two-dimensional table according to the id automatically assigned by the system, and create a corresponding feature index;

[0037] S3. Use the deep learning model Mamba based on the state space model to model the temporal information of the above features at each moment of the temporal vision, and then output the end-to-end inference and prediction of the action behavior to determine whether the behavior of each target object at each moment is the start moment and the end moment, as well as the spatial position and behavior category of the target associated with the video in the temporal vision frame.

[0038] In S1, the analysis system includes a visual frame feature extraction module, a target object feature combination module, and a temporal behavior recognition module;

[0039] The visual frame feature extraction module is used to adopt the feature extraction mode of Backone+Neck+Head for a single video frame in the time series;

[0040] The target object feature combination module is used to extract features through the target object in a single video frame in the time series, adopt the Kaufman filter to obtain the same object in the front and back time series, and model a single behavior in the temporal vision. The modeling is as follows:

[0041]

[0042] Fuse the single target feature and the position information in the spatial domain into a two-dimensional table according to the id automatically assigned by the system, and create the corresponding feature index;

[0043] The temporal behavior recognition module is used to use the deep learning model Mamba based on the state space model to model the temporal information of the above features at each moment of the temporal vision, and then output the end-to-end inference and prediction of the action behavior to determine whether the behavior of each target object at each moment is the start moment and the end moment, as well as the spatial position and behavior category of the target associated with the video in the temporal vision frame.

[0044] The visual frame feature extraction module is connected to the target object feature combination module, and the target object feature combination module is connected to the temporal behavior recognition module.

[0045] The visual frame feature extraction module includes a backone module, a Neck module, and a Head module. The backone module is connected to the Neck module, and the Neck module is connected to the Head module;

[0046] The backone module is used for Swin-Transformer to perform downsampling after each feature extraction through feature fusion, increasing the receptive field of the next window attention operation on the original image, thereby performing multi-scale feature extraction on the input image;

[0047] The Neck module is used to infer the corresponding attention weights and multiply these weights with the original feature map to achieve adaptive adjustment of features;

[0048] The Head module is used to introduce multiple detection heads to improve the detection ability for targets of different scales.

[0049] The Neck module includes a channel attention module and a spatial attention module. The channel attention module processes the channels of the input feature map, and the spatial attention module is a CBAM structure that processes the space of the input feature map.

[0050] In summary, in terms of feature extraction on visual frames, the Swin-transformer module is adopted. It acts as a role similar to the convolution operation in CNN in image processing. However, its characteristics based on the Transformer structure give it advantages in processing global relationships, long-distance information transmission, etc., and can effectively improve the model's abstract ability and expression ability for image content; the CBAM module is adopted. Through the comprehensive utilization of channel attention and spatial attention mechanisms, it can improve the CNN model's processing and learning ability for image features, enhance the network's representation ability, and improve the performance and accuracy of image processing tasks; multiple detection heads are used to improve the detection ability for targets of different scales, with four detection heads to solve the problems of missed detection or poor detection effects.

[0051] The Kalman filter is used in the target object feature combination module. First, it can process the target tracking algorithm in real time, is suitable for application scenarios that require fast response, and has real-time performance and high efficiency; second, it can effectively predict the state of a dynamic system, which is crucial for tracking tasks because it can accurately predict the position of the target in the next frame and has the ability to predict dynamic systems; third, it combines the current observation data and the previous prediction results to update the target state and perform feature fusion, and has the ability of information fusion; fourth, through its modeling and update mechanism, it can effectively reduce these errors, thereby improving the tracking accuracy and has the ability to handle noise and uncertainty; fifth, it can use multiple Kalman filters to track each target object separately, which makes the Kalman filter very suitable for handling complex tracking scenarios and supports multi-target tracking.

[0052] In the temporal behavior recognition module, a deep learning model Mamba based on the state space model is adopted. First, it shows significant performance advantages in processing long sequence data, can effectively process and analyze long-term dependencies, and has the ability to handle long sequence tasks. Second, compared with the attention mechanism, it can not only reduce the computational complexity, but also the memory usage does not depend on the context length, which has the effect of reducing computational complexity and memory usage. Third, by introducing a selective state space, the model can selectively transmit or forget information according to the current data, thus solving the deficiencies of previous models in processing discrete and information-intensive data, and has the function of introducing a selective state space. Fourth, the Mamba model shows strong performance and capabilities in context learning applications by virtue of its proficiency in context modeling, and has the ability to improve context modeling.

[0053] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present invention can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

Claims

1. A method for analyzing behavior in a meeting scene based on temporal vision, characterized in that: The following steps are involved: S1, using the Backone+Neck+Head feature extraction mode for a single video frame in the time sequence through the analysis system; S2, extracting features from the target object in a single video frame in the time sequence, fusing the single target features and the position information in the spatial domain into a two-dimensional table according to the ID automatically assigned by the system, and creating a corresponding feature index; S3. The above features of the temporal vision at each moment are modeled through the deep learning model Mamba based on the state-space model. Then, the end2end method of outputting the action behavior is used to infer and predict whether the behavior of each target object at each moment is the start moment and the end moment, as well as the spatial position and behavior category of the target in the temporal vision frame in the video.

2. A method for analyzing behavior in a meeting scene based on temporal vision according to claim 1, characterized in that: In S1, Swin-Transformer is introduced in the backbone to perform downsampling after each feature extraction by feature fusion, which increases the receptive field of the next window attention operation on the original image, thereby performing multi-scale feature extraction on the input image.

3. A method for analyzing behavior in a meeting scene based on temporal vision according to claim 1, characterized in that: In S1, a CBAM structure containing channels for processing input feature maps and a space for processing input feature maps is introduced in Neck. The corresponding attention weights are inferred and multiplied with the original feature maps to achieve adaptive adjustment of the features.

4. A method for analyzing behavior in a meeting scene based on temporal vision according to claim 1, characterized in that: In S1, multiple detection heads are introduced in the Head to improve the detection capability of objects of different scales.

5. A method for analyzing behavior in a meeting scene based on temporal vision according to claim 1, characterized in that: In S2, the Kaffman filter is used to obtain the same object in the previous and next time series and model a single behavior in time series vision.

6. A method for analyzing behavior in a meeting scene based on temporal vision according to claim 1, characterized in that: In S1, the analysis system includes a visual frame feature extraction module, a target object feature combination module, and a temporal behavior recognition module; The visual frame feature extraction module is used to extract the feature of a single video frame in the time sequence using the Backone+Neck+Head mode; The target object feature combination module is used to extract features from the target object in a single video frame in the time sequence, use the Kaffman filter to obtain the same object in the previous and next time sequences, model a single behavior in the time sequence vision, fuse the single target feature and the position information in the airspace into a two-dimensional table according to the ID automatically assigned by the system, and create the corresponding feature index; The temporal behavior recognition module is used to model the temporal information of the above features of the temporal vision at each moment through the deep learning model Mamba based on the state-space model, and then output the action behavior in an end-to-end manner to infer and predict whether the behavior of each target object at each moment is the start moment and the end moment, as well as the spatial position and behavior category of the target in the temporal visual frame associated with the video.

7. A method for analyzing behavior in a meeting scene based on temporal vision according to claim 6, characterized in that: The visual frame feature extraction module is connected to the target object feature combination module, and the target object feature combination module is connected to the temporal behavior recognition module.

8. A method for analyzing behavior in a meeting scene based on temporal vision according to claim 6, characterized in that: The visual frame feature extraction module includes a backone module, a neck module and a head module, the backone module is connected to the neck module, and the neck module is connected to the head module; The backone module is used in Swin-Transformer to perform downsampling after each feature extraction by feature fusion, which increases the receptive field of the next window attention operation on the original image, thereby performing multi-scale feature extraction on the input image; The Neck module is used to infer the corresponding attention weights and multiply these weights with the original feature map to achieve adaptive adjustment of the features; The Head module is used to introduce multiple detection heads to improve the detection capability of objects of different scales.

9. A method for analyzing behavior in a meeting scene based on temporal vision according to claim 8, characterized in that: The Neck module includes a channel attention module and a spatial attention module. The channel attention module processes the channel of the input feature map, and the spatial attention module processes the CBAM structure of the space of the input feature map.