Motion detection method, device, equipment and storage medium

By extracting and interacting video frames and using the self-attention mechanism to process interactive features, the problem of ignoring interaction information between different subjects in the prior art is solved, and the detection accuracy of video action detection is significantly improved.

CN119649472BActive Publication Date: 2025-06-06深圳智眸未来科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510189241.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

When handling multimodal multi-input, existing video action detection technology ignores the interaction information between different subjects, resulting in a decrease in detection accuracy.

Method used

By extracting the video frames to be detected, human body features, object features and memory features are obtained, and feature interaction is performed to generate the first interactive feature, the second interactive feature and the third interactive feature. Then, the self-attention mechanism is used to process these interactive features, generate interactive attention features, and finally perform action detection based on these features.

Benefits of technology

Through feature interaction and self-attention mechanisms, the interactive features between characters and objects in video frames are fully captured, which improves the ability to understand complex action scenes and significantly improves the detection accuracy of video action detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649472B_ABST
    Figure CN119649472B_ABST
Patent Text Reader

Abstract

The present application relates to the field of motion recognition technology, and discloses a motion detection method, device, equipment and storage medium. The method comprises: extracting features from a video frame of a video to be detected to obtain human features, object features and memory features; performing feature interaction on the human features, object features and memory features to obtain first interaction features, second interaction features and third interaction features; performing self-attention mechanism processing on the first interaction features, second interaction features and third interaction features to obtain interaction attention features; performing motion detection based on the interaction attention features to obtain motion detection results. The embodiments of the present application can improve the detection accuracy of video motion detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of motion recognition technology, and in particular to a motion detection method, device, equipment and storage medium. Background Art

[0002] In artificial intelligence applications, video action detection plays a key role in many fields. Understanding the interaction between subjects is an important way to understand videos and optimize action detection.

[0003] In related technologies, video action detection uses spatial representation and temporal features in videos to infer the interaction relationship between people and the contextual environment, but ignores the interaction between different subjects. Under multi-modal and multi-input conditions, whether it is the interaction information between objects, the interaction information between people, the temporal interaction information of object changes, or the temporal interaction information of people changes, the accuracy of video action detection is reduced. Summary of the invention

[0004] The purpose of this application is to provide a motion detection method, device, equipment and storage medium, aiming to improve the detection accuracy of video motion detection.

[0005] The present application provides an action detection method, including:

[0006] Perform feature extraction on the video frames of the video to be detected to obtain human features, object features and memory features; the memory features include human features in a number of consecutive video frames;

[0007] Performing feature interaction on the human features, the object features and the memory features to obtain a first interaction feature, a second interaction feature and a third interaction feature; the first interaction feature is a feature obtained by performing feature interaction on the human features of the same video frame, the second interaction feature describes a feature obtained by performing feature interaction on the first interaction feature and the object features of the same video frame, and the third interaction feature describes a feature obtained by performing feature interaction on the first interaction feature, the second interaction feature and the memory feature;

[0008] Performing self-attention mechanism processing on the first interaction feature, the second interaction feature, and the third interaction feature to obtain an interaction attention feature;

[0009] Action detection is performed based on the interactive attention feature to obtain an action detection result.

[0010] In some embodiments, the feature extraction of the video frame of the video to be detected to obtain human features, object features and memory features includes:

[0011] Disassembling the video to be detected to obtain a plurality of video segments to be detected;

[0012] Performing edge detection on the human body contour and the object contour in the video frame of the video segment to be detected to obtain a human body contour edge detection result and an object contour edge detection result;

[0013] Perform feature extraction based on the human body contour edge detection result and the object contour edge detection result to obtain the human body features and the object features;

[0014] The human body features in a number of consecutive video frames are aggregated to obtain the memory features.

[0015] In some embodiments, the performing feature interaction on the human body feature, the object feature, and the memory feature to obtain a first interaction feature, a second interaction feature, and a third interaction feature includes:

[0016] Performing feature interaction on the human body features of the same video frame to obtain the first interaction features;

[0017] Performing feature interaction on the first interaction feature of the same video frame and the object feature of the same video frame to obtain the second interaction feature;

[0018] Perform feature interaction on the first interaction feature, the second interaction feature, and the memory feature to obtain the third interaction feature.

[0019] In some embodiments, the interaction method among the first interaction feature, the second interaction feature and the memory feature comprises:

[0020] Performing feature division on the target feature to obtain a first division feature, a second division feature and a third division feature; the target feature is the human body feature, the object feature, the memory feature, the first interaction feature and / or the second interaction feature;

[0021] Performing linear mapping on the first segmentation feature, the second segmentation feature, and the third segmentation feature to obtain a first mapping feature, a second mapping feature, and a third mapping feature;

[0022] Performing an attention weight operation on the first mapping feature and the second mapping feature to obtain a mapping attention feature;

[0023] Linearly map the mapping attention feature and the scaled third mapping feature to obtain a target interaction feature; the target interaction feature is the first interaction feature, the second interaction feature or the third interaction feature.

[0024] In some embodiments, the performing of a self-attention mechanism on the first interaction feature, the second interaction feature, and the third interaction feature to obtain an interaction attention feature includes:

[0025] Performing matrix transformation processing on the first interaction feature, the second interaction feature, and the third interaction feature to obtain corresponding attention query vectors, attention key vectors, and attention value vectors;

[0026] A self-attention mechanism is performed based on the attention query vector, the attention key vector and the attention value vector to obtain the interactive attention feature.

[0027] In some embodiments, performing action detection based on the interactive attention feature to obtain an action detection result includes:

[0028] Linearly mapping the interactive attention features to obtain an action feature sequence;

[0029] Performing feature retrieval in a feature database based on the action feature sequence to obtain a feature retrieval result;

[0030] The action detection result is determined based on the feature retrieval result.

[0031] In some embodiments, the action detection method is implemented by an action detection model, and the action detection model is obtained by training a preset deep neural network model using an asynchronous memory update algorithm.

[0032] The present application also provides an action detection device, including:

[0033] The first module is used to extract features from video frames of the video to be detected to obtain human features, object features and memory features; the memory features include human features in a number of consecutive video frames;

[0034] The second module is used to perform feature interaction on the human features, the object features and the memory features to obtain a first interaction feature, a second interaction feature and a third interaction feature; the first interaction feature is a feature obtained by performing feature interaction on the human features of the same video frame, the second interaction feature describes a feature obtained by performing feature interaction on the first interaction feature and the object features of the same video frame, and the third interaction feature describes a feature obtained by performing feature interaction on the first interaction feature, the second interaction feature and the memory feature;

[0035] A third module is used to perform self-attention mechanism processing on the first interaction feature, the second interaction feature and the third interaction feature to obtain an interaction attention feature;

[0036] The fourth module is used to perform action detection based on the interactive attention feature to obtain an action detection result.

[0037] An embodiment of the present application further provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned action detection method when executing the computer program.

[0038] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned action detection method is implemented.

[0039] The beneficial effects of the present application are as follows: by extracting human features, object features and memory features in the video to be detected, the human features of the same video frame are subjected to feature interaction to generate a first interaction feature, the first interaction feature of the same video frame and the object feature of the same video frame are subjected to feature interaction to generate a second interaction feature, the first interaction feature, the second interaction feature and the memory feature are subjected to feature interaction to obtain a third interaction feature, and action detection is performed based on the first interaction feature, the second interaction feature and the third interaction feature to obtain an action detection result. Since the human features, object features and memory features are subjected to corresponding feature interaction, the interaction features between people in the video frame, the interaction features between people and objects, and the interaction features between people and objects that change over time are fully captured, providing rich semantic clues for action detection, enhancing the ability to understand complex action scenes, and improving the detection accuracy of video action detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a diagram of the application environment of the action detection method provided in an embodiment of the present application.

[0041] Figure 2 It is a flow chart of the action detection method provided in an embodiment of the present application.

[0042] Figure 3 It is a flowchart of the specific method of step S201 provided in an embodiment of the present application.

[0043] Figure 4 It is a flowchart of the specific method of step S202 provided in an embodiment of the present application.

[0044] Figure 5 It is a flowchart of the specific method of step S204 provided in an embodiment of the present application.

[0045] Figure 6 It is a structural schematic diagram of the action detection device provided in an embodiment of the present application.

[0046] Figure 7 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown can be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second" and the like in the specification, claims and drawings are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0050] The motion detection method provided in the embodiment of the present application can be applied to Figure 1 The application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other servers. When the user of the terminal 102 wants to perform motion detection, the video to be detected can be submitted to the server 104 through the terminal 102, and then the server 104 implements the motion detection processing of the video to be detected. When the server 104 performs motion detection, it extracts features of the video frame of the video to be detected, obtains human features, object features and memory features, performs feature interaction on the human features, object features and memory features, obtains the first interactive feature, the second interactive feature and the third interactive feature, performs attention mechanism processing on the first interactive feature, the second interactive feature and the third interactive feature, obtains the interactive attention feature, performs motion detection based on the interactive attention feature, and obtains the motion detection result. The memory feature includes human features in several consecutive video frames, the first interaction feature describes the interaction information between human features in the same video frame, the second interaction feature describes the interaction information between the first interaction feature and the object feature in the same video frame, and the third interaction feature describes the interaction information between the first interaction feature, the second interaction feature and the memory feature. The terminal 102 may be, but is not limited to, various desktop computers, laptops, smart phones, tablet computers, IoT devices and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0051] Figure 2 is a flow chart of the action detection method provided by the embodiment of the present application. Figure 2 In some embodiments, the method includes but is not limited to steps S201 to S204.

[0052] Step S201 , extracting features from video frames of the video to be detected to obtain human features, object features and memory features.

[0053] The memory features include human features in several consecutive video frames. Human features refer to action features corresponding to the actions being performed by people in the video frames of the video to be detected. Object features refer to category features, state features and / or position features of objects in the video frames of the video to be detected.

[0054] The video to be detected refers to a video that needs to be action detected. It is necessary to detect the actions being performed by the characters in the video. Different actions correspond to different action categories. There may be multiple characters in the video, and different characters may have different actions corresponding to different human body characteristics.

[0055] In some embodiments, the server can obtain the video to be detected from the database, or can obtain the video to be detected uploaded by the terminal. After obtaining the video to be detected, the server can use a pre-trained feature extraction model to perform corresponding feature extraction on the video to be detected to obtain human features, object features, and memory features. The server can also use pre-trained feature extraction parameters to perform feature extraction on the video to be detected to obtain human features, object features, and memory features.

[0056] Step S202, performing feature interaction on human body features, object features and memory features to obtain a first interaction feature, a second interaction feature and a third interaction feature.

[0057] In some embodiments, the server may use a pre-trained feature interaction model to perform feature interaction on human features, object features, and memory features to obtain a first interaction feature, a second interaction feature, and a third interaction feature. When performing feature interaction, human features, object features, and memory features may be input into a pre-trained feature interaction model. In the pre-trained feature interaction model, the features to be interacted are divided into several groups and corresponding linear mapping processing is performed and attention weight features of the features to be interacted are calculated. The attention weight features are used to perform corresponding weighted operations on the features to be interacted, and finally the first interaction feature, the second interaction feature, and the third interaction feature are obtained.

[0058] The first interactive feature is a feature obtained by performing feature interaction on human features of the same video frame. In some embodiments, the human features of the same video frame may be divided into several groups, and corresponding linear mapping processing and attention weight features of the features to be interacted are performed, and the attention weight features are used to perform corresponding weighted operations on the features to be interacted, and finally the first interactive feature is obtained.

[0059] The second interaction feature describes the feature obtained by performing feature interaction between the first interaction feature of the same video frame and the object feature of the same video frame. In some embodiments, the first interaction feature of the same video frame and the object feature of the same video frame may be divided into several groups, and corresponding linear mapping processing and attention weight features of the features to be interacted are performed, and the attention weight features are used to perform corresponding weighted operations on the features to be interacted, and finally the second interaction feature is obtained.

[0060] The third interaction feature describes the feature obtained by the feature interaction of the first interaction feature, the second interaction feature and the memory feature. In some embodiments, the first interaction feature, the second interaction feature and the memory feature may be divided into several groups, and corresponding linear mapping processing and attention weight features of the features to be interacted are performed, and the attention weight features are used to perform corresponding weighted operations on the features to be interacted, and finally the third interaction feature is obtained.

[0061] Step S203, performing self-attention mechanism processing on the first interaction feature, the second interaction feature and the third interaction feature to obtain interaction attention features.

[0062] Among them, self-attention is generally in the form of QKV, which adds restrictions on the attention mechanism, maps the input to three different spaces, and generates Q, K, and V by itself. In the solution of this application, the three attention vectors Q, K, and V are extracted from the first interaction feature, the second interaction feature, and the third interaction feature through the self-attention mechanism to calculate the attention.

[0063] In some embodiments, the server may use a pre-trained self-attention mechanism model to perform self-attention mechanism processing on the first interaction feature, the second interaction feature, and the third interaction feature to obtain an interaction attention feature. In a specific implementation, the self-attention mechanism model semantically encodes the first interaction feature, the second interaction feature, and the third interaction feature, and then implements the calculation of the self-attention mechanism in the form of semantic interaction feature encoding, so that the feature representation has rich diversity and obtains the interaction attention feature.

[0064] Step S204, performing action detection based on the interactive attention feature to obtain an action detection result.

[0065] In some embodiments, the server may use a pre-trained classification module to perform action detection based on the interactive attention feature to obtain an action detection result. In a specific implementation, the interactive attention feature contains the semantic features of the fusion of the first interactive feature, the second interactive feature, and the third interactive feature. The interactive attention feature is decoded by the classification model to identify and classify the character actions in the video to be detected through the decoding process, thereby generating an action detection result corresponding to the video to be detected.

[0066] Figure 3 is a flowchart of a specific method of step S201 provided in an embodiment of the present application. Figure 3 The method includes but is not limited to steps S301 to S304.

[0067] Step S301, disassemble the video to be detected to obtain a number of video segments to be detected.

[0068] Step S302 , performing edge detection on the human body contour and the object contour in the video frame of the video segment to be detected, and obtaining a human body contour edge detection result and an object contour edge detection result.

[0069] Step S303, performing feature extraction based on the human body contour edge detection result and the object contour edge detection result to obtain human body features and object features.

[0070] Step S304, summarizing the human features in a number of consecutive video frames to obtain memory features.

[0071] In the specific implementation, the video to be detected is first disassembled in a preset manner (for example, at a certain time interval) to obtain multiple video segments to be detected. For each video segment to be detected, the fast detection module (for example, Yolov10 module) in the feature extraction model is used to perform edge detection on the human body contour and object contour in the video frame of the video segment to be detected, so as to detect the boundaries of both the person and the object through the bounding box, and obtain the human body contour edge detection result and the object contour edge detection result. Then, based on the human body contour edge detection result and the object contour edge detection result, the corresponding human body contour and object contour are cut out through the RoIAlign operation to obtain human body features and object features. After obtaining the human body features in multiple video frames, the human body features in the continuous video frames are summarized in the order of the video frames to obtain the memory features.

[0072] Figure 4 is a flowchart of a specific method of step S202 provided in an embodiment of the present application. Figure 4 The method includes but is not limited to steps S401 to S403.

[0073] Step S401, performing feature interaction on human features of the same video frame to obtain a first interaction feature.

[0074] Step S402: Perform feature interaction on the first interaction feature of the same video frame and the object feature of the same video frame to obtain a second interaction feature.

[0075] Step S403: Perform feature interaction on the first interaction feature, the second interaction feature, and the memory feature to obtain a third interaction feature.

[0076] In some embodiments, the feature interaction model may include a first interaction module, a second interaction module, and a third interaction module. In a specific implementation, the human body features of the same video frame are input into the first interaction module to perform feature interaction on the human body features of the same video frame and obtain the first interaction feature, the first interaction feature of the same video frame and the object feature of the same video frame are input into the second interaction module to perform feature interaction on the first interaction feature of the same video frame and the object feature of the same video frame and obtain the second interaction feature, and the first interaction feature, the second interaction feature, and the memory feature are input into the third interaction module to perform feature interaction on the first interaction feature, the second interaction feature, and the memory feature and obtain the third interaction feature.

[0077] In a specific embodiment, the first interaction module, the second interaction module and the third interaction module are all composed of a first linear mapping layer, a second linear mapping layer, a normalization layer, an attention calculation layer and a third linear mapping layer. The interaction method of the first interaction feature, the second interaction feature and the memory feature includes: performing feature division on the target feature to obtain the first division feature, the second division feature and the third division feature; performing linear mapping on the first division feature, the second division feature and the third division feature to obtain the first mapping feature, the second mapping feature and the third mapping feature; performing attention weight operation on the first mapping feature and the second mapping feature to obtain the mapping attention feature; performing linear mapping on the mapping attention feature and the scaled third mapping feature to obtain the target interaction feature. Among them, the target feature is a human body feature, an object feature, a memory feature, a first interaction feature and / or a second interaction feature, and the target interaction feature is the first interaction feature, the second interaction feature or the third interaction feature.

[0078] In the first interaction module, the human body features of the same video frame are divided into three groups, namely, the first division feature, the second division feature and the third division feature. The first division feature is linearly mapped by the first linear mapping layer and the second division feature is linearly mapped by the second linear mapping layer to obtain the first mapping feature and the second mapping feature. The third division feature is normalized linearly mapped by the normalization layer to obtain the third mapping feature. The attention calculation layer performs attention weight operation on the first mapping feature and the second mapping feature to obtain the mapping attention feature. After the third mapping feature is scaled to the same dimension as the mapping attention feature, the third linear mapping layer is linearly mapped on the mapping attention feature and the scaled third mapping feature to obtain the target interaction feature. In the second interaction module, the first interaction features of the same video frame and the object features of the same video frame are fused and divided into three groups, namely, the first division features, the second division features and the third division features. The first division features are linearly mapped by the first linear mapping layer and the second division features are linearly mapped by the second linear mapping layer to obtain the first mapping features and the second mapping features. The third division features are normalized linearly mapped by the normalization layer to obtain the third mapping features. The attention calculation layer performs attention weight operation on the first mapping features and the second mapping features to obtain the mapping attention features. After the third mapping features are scaled to the same dimension as the mapping attention features, the third linear mapping layer is linearly mapped on the mapping attention features and the scaled third mapping features to obtain the target interaction features. In the third interaction module, the first division feature, the second division feature and the third division feature are fused and divided into three groups, namely the first division feature, the second division feature and the third division feature. The first division feature is linearly mapped by the first linear mapping layer and the second division feature is linearly mapped by the second linear mapping layer to obtain the first mapping feature and the second mapping feature. The third division feature is normalized linearly mapped by the normalization layer to obtain the third mapping feature. The attention calculation layer performs attention weight operation on the first mapping feature and the second mapping feature to obtain the mapping attention feature. After the third mapping feature is scaled to the same dimension as the mapping attention feature, the third linear mapping layer is linearly mapped on the mapping attention feature and the scaled third mapping feature to obtain the target interaction feature.

[0079] In some embodiments, the above step S203 specifically includes: performing matrix transformation processing on the first interaction feature, the second interaction feature and the third interaction feature to obtain corresponding attention query vectors, attention key vectors and attention value vectors; performing self-attention mechanism processing based on the attention query vector, the attention key vector and the attention value vector to obtain interactive attention features.

[0080] Among them, the self-query vector Q and the key vector K are feature vectors used to calculate the attention weights, and the force value vector V represents the vector of input features. In the self-attention mechanism, the self-attention query vector, the self-attention key vector, and the self-attention value vector are all obtained by performing matrix transformation on the activation feature map, and their feature dimensions are the same.

[0081] Specifically, the self-attention mechanism is a variant of the attention mechanism, which reduces the dependence on external information and is better at capturing the internal correlation of data or features. In the scheme of the present application, the local feature representation is made more diverse mainly by performing self-attention processing on the first interactive feature, the second interactive feature and the third interactive feature. As for the processing process of self-attention, the first interactive feature, the second interactive feature and the third interactive feature can be first subjected to matrix transformation processing to obtain a self-attention query vector, a self-attention key vector and a self-attention value vector, and then the self-attention query vector, the self-attention key vector and the self-attention value vector are subjected to self-attention mechanism processing to obtain interactive attention features. In a specific embodiment, since the feature encoding of the video frame of the video to be detected is realized through the first interactive feature, the second interactive feature and the third interactive feature, the features in the first interactive feature, the second interactive feature and the third interactive feature can be converted into semantic tokens for attention calculation.

[0082] Self-attention specifically satisfies the following formula:

[0083] ,

[0084] Among them, softmax represents the normalized exponential function, Q, K and V represent the self-attention query vector, self-attention key vector and self-attention value vector respectively, and the denominator in the brackets is the scaling numerator.

[0085] Figure 5 is a flowchart of a specific method of step S204 provided in an embodiment of the present application. Figure 5 The method includes but is not limited to steps S501 to S503.

[0086] Step S501, linearly map the interactive attention features to obtain an action feature sequence.

[0087] Step S502: performing feature retrieval in a feature database based on the action feature sequence to obtain a feature retrieval result.

[0088] Step S503: determining the action detection result based on the feature retrieval result.

[0089] In a specific implementation, a feature database can be established based on the purpose of action detection. When performing action detection to obtain action detection results, the interactive attention features can be linearly mapped first to obtain an action feature sequence, and the feature dimension can be reduced by linear mapping. Then, based on the action feature sequence, feature retrieval is performed in the feature database to find feature retrieval results that can match the current action feature sequence, and then the action detection results corresponding to the video to be detected are obtained according to the specific name corresponding to the feature retrieval result.

[0090] The action detection method provided in the above embodiment is implemented by an action detection model, which is obtained by training a preset deep neural network model using an asynchronous memory update algorithm. The action detection model consists of a feature extraction model, a feature interaction model, a self-attention mechanism model and a classification module.

[0091] See also Figure 6 The embodiment of the present application further provides an action detection device, which can implement the above-mentioned action detection method, and the device includes:

[0092] The first module 601 is used to extract features from video frames of the video to be detected to obtain human features, object features and memory features; the memory features include human features in a number of consecutive video frames;

[0093] The second module 602 is used to perform feature interaction on human features, object features and memory features to obtain first interaction features, second interaction features and third interaction features; the first interaction feature is a feature obtained by performing feature interaction on human features of the same video frame, the second interaction feature describes a feature obtained by performing feature interaction on the first interaction feature of the same video frame and the object feature of the same video frame, and the third interaction feature describes a feature obtained by performing feature interaction on the first interaction feature, the second interaction feature and the memory feature;

[0094] The third module 603 is used to perform self-attention mechanism processing on the first interaction feature, the second interaction feature and the third interaction feature to obtain an interaction attention feature;

[0095] The fourth module 604 is used to perform action detection based on the interactive attention feature to obtain the action detection result.

[0096] The specific implementation of the motion detection device is substantially the same as the specific implementation of the above-mentioned motion detection method, and will not be described in detail herein.

[0097] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment.

[0098] Refer to the following Figure 7 An electronic device 700 according to this embodiment of the present disclosure is described. Figure 7The electronic device 700 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0099] like Figure 7 As shown, the electronic device 700 is in the form of a general computing device. The components of the electronic device 700 may include, but are not limited to: at least one processing unit 710, at least one storage unit 720, a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710), a display unit 740, etc.

[0100] The storage unit stores program codes, which can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present disclosure described in the above action detection method section of this specification.

[0101] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 7201 and / or a cache storage unit 7202 , and may further include a read-only storage unit (ROM) 7203 .

[0102] The storage unit 720 may also include a program / utility 7204 having a set (at least one) of program modules 7205, such program modules 7205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0103] Bus 730 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0104] The electronic device 700 may also communicate with one or more external devices 700' (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device that enables the electronic device 700 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 750. Furthermore, the electronic device 700 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 760. The network adapter 760 may communicate with other modules of the electronic device 700 via a bus 730. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0105] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned action detection method is implemented.

[0106] The action detection method, device, equipment and storage medium provided by the embodiments of the present application extract human features, object features and memory features in the video to be detected, make the human features of the same video frame interact with each other to generate a first interaction feature, make the first interaction feature of the same video frame and the object features of the same video frame interact with each other to generate a second interaction feature, make the first interaction feature, the second interaction feature and the memory feature interact with each other to obtain a third interaction feature, and perform action detection based on the first interaction feature, the second interaction feature and the third interaction feature to obtain an action detection result. Since the human features, object features and memory features are subjected to corresponding feature interactions, the interaction features between people in the video frame, the interaction features between people and objects, and the interaction features of people and objects that change over time are fully captured, providing rich semantic clues for action detection, enhancing the ability to understand complex action scenes, and improving the detection accuracy of video action detection.

[0107] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the implementation of the present disclosure.

[0108] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0109] Computer readable storage media may include data signals propagated in baseband or as part of a carrier wave, wherein readable program codes are carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program codes contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0110] Those skilled in the art will appreciate that the above modules can be distributed in the device according to the description of the embodiment, or can be changed accordingly and only used in one or more devices different from the embodiment. The modules of the above embodiments can be combined into one module, or further divided into multiple sub-modules.

[0111] The exemplary embodiments of the present disclosure are specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, configurations or implementations described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent configurations included in the spirit and scope of the appended claims.

Claims

1. A motion detection method, characterized in that: include: Perform feature extraction on the video frames of the video to be detected to obtain human features, object features and memory features; the memory features include human features in a number of consecutive video frames; Performing feature interaction on the human features, the object features and the memory features to obtain a first interaction feature, a second interaction feature and a third interaction feature; the first interaction feature is a feature obtained by performing feature interaction on the human features of the same video frame, the second interaction feature describes a feature obtained by performing feature interaction on the first interaction feature and the object features of the same video frame, and the third interaction feature describes a feature obtained by performing feature interaction on the first interaction feature, the second interaction feature and the memory feature; Performing self-attention mechanism processing on the first interaction feature, the second interaction feature, and the third interaction feature to obtain an interaction attention feature; Performing action detection based on the interactive attention feature to obtain an action detection result; The interaction method of the first interaction feature, the second interaction feature and the memory feature includes: Performing feature division on the target feature to obtain a first division feature, a second division feature and a third division feature; the target feature is the human body feature, the object feature, the memory feature, the first interaction feature and / or the second interaction feature; Performing linear mapping on the first segmentation feature, the second segmentation feature, and the third segmentation feature to obtain a first mapping feature, a second mapping feature, and a third mapping feature; Performing an attention weight operation on the first mapping feature and the second mapping feature to obtain a mapping attention feature; Linearly map the mapping attention feature and the scaled third mapping feature to obtain a target interaction feature; the target interaction feature is the first interaction feature, the second interaction feature or the third interaction feature.

2. The motion detection method according to claim 1, characterized in that: The feature extraction of the video frame of the video to be detected to obtain human features, object features and memory features includes: Disassembling the video to be detected to obtain a plurality of video segments to be detected; Performing edge detection on the human body contour and the object contour in the video frame of the video segment to be detected to obtain a human body contour edge detection result and an object contour edge detection result; Perform feature extraction based on the human body contour edge detection result and the object contour edge detection result to obtain the human body features and the object features; The human body features in a number of consecutive video frames are aggregated to obtain the memory features.

3. The motion detection method according to claim 1, characterized in that: The performing feature interaction on the human body feature, the object feature and the memory feature to obtain a first interaction feature, a second interaction feature and a third interaction feature includes: Performing feature interaction on the human body features of the same video frame to obtain the first interaction features; Performing feature interaction on the first interaction feature of the same video frame and the object feature of the same video frame to obtain the second interaction feature; Perform feature interaction on the first interaction feature, the second interaction feature, and the memory feature to obtain the third interaction feature.

4. The motion detection method according to claim 1, characterized in that: The performing of self-attention mechanism processing on the first interaction feature, the second interaction feature, and the third interaction feature to obtain an interaction attention feature includes: Performing matrix transformation processing on the first interaction feature, the second interaction feature, and the third interaction feature to obtain corresponding attention query vectors, attention key vectors, and attention value vectors; A self-attention mechanism is performed based on the attention query vector, the attention key vector and the attention value vector to obtain the interactive attention feature.

5. The motion detection method according to claim 1, characterized in that: The performing action detection based on the interactive attention feature to obtain an action detection result includes: Linearly mapping the interactive attention features to obtain an action feature sequence; Performing feature retrieval in a feature database based on the action feature sequence to obtain a feature retrieval result; The action detection result is determined based on the feature retrieval result.

6. The motion detection method according to any one of claims 1 to 5, characterized in that: The action detection method is implemented by an action detection model, and the action detection model is obtained by training a preset deep neural network model using an asynchronous memory update algorithm.

7. A motion detection device, characterized in that: include: The first module is used to extract features from the video frames of the video to be detected to obtain human features, object features and memory features; The memory features include human features in a number of consecutive video frames; The second module is used to perform feature interaction on the human features, the object features and the memory features to obtain a first interaction feature, a second interaction feature and a third interaction feature; the first interaction feature is a feature obtained by performing feature interaction on the human features of the same video frame, the second interaction feature describes a feature obtained by performing feature interaction on the first interaction feature and the object features of the same video frame, and the third interaction feature describes a feature obtained by performing feature interaction on the first interaction feature, the second interaction feature and the memory feature; A third module is used to perform self-attention mechanism processing on the first interaction feature, the second interaction feature and the third interaction feature to obtain an interaction attention feature; A fourth module is used to perform action detection based on the interactive attention feature to obtain an action detection result; The interaction method of the first interaction feature, the second interaction feature and the memory feature includes: Performing feature division on the target feature to obtain a first division feature, a second division feature and a third division feature; the target feature is the human body feature, the object feature, the memory feature, the first interaction feature and / or the second interaction feature; Performing linear mapping on the first segmentation feature, the second segmentation feature, and the third segmentation feature to obtain a first mapping feature, a second mapping feature, and a third mapping feature; Performing an attention weight operation on the first mapping feature and the second mapping feature to obtain a mapping attention feature; Linearly map the mapping attention feature and the scaled third mapping feature to obtain a target interaction feature; the target interaction feature is the first interaction feature, the second interaction feature or the third interaction feature.

8. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the action detection method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the action detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Motion recognition method based on feature interactive learning, and terminal device

    WO2022073282A1

  • Video target segmentation method based on space-time decoupling attention mechanism

    WO2024183024A1