Limb conflict behavior identification method and device and storage medium

By fusing multimodal information through dynamic visual sensors, depth cameras, and voice acquisition units, the speed, posture, and emotional characteristics of body movements are extracted, solving the low accuracy problem of traditional methods in complex environments and achieving accurate recognition in high-security scenarios.

CN120766366AActive Publication Date: 2025-10-10LINGYANGE SEMICONDUCTOR, INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511278865.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-10
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Traditional methods for identifying physical conflict behavior have low accuracy in environments such as strong light, occlusion, and poor shooting angles. They find it difficult to distinguish between normal physical behavior and conflict behavior, and the single visual modality is easily interfered with.

Method used

Dynamic visual sensors, depth cameras and voice acquisition units are used to fuse multimodal information. Event spatiotemporal pyramid convolution is used to extract limb movement speed features, three-dimensional posture estimation network extracts limb posture features, speech recognition model extracts emotion features, and comprehensive judgment is made through multimodal attention fusion network.

Benefits of technology

It improves the accuracy and timeliness of identifying physical conflict behaviors, reduces missed reports and false alarms, and is particularly suitable for high-security scenarios such as campuses and public transportation, providing intelligent and automated behavior recognition capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766366A_ABST
    Figure CN120766366A_ABST
Patent Text Reader

Abstract

The invention provides a limb conflict behavior recognition method and device and a storage medium, and the method comprises the steps: arranging a monitoring module at a monitoring place, and the monitoring module at least comprises a dynamic vision sensor, a depth camera and a voice collection unit; according to an event flow collected by a dynamic vision sensor, extracting limb movement speed characteristics by adopting event space-time pyramid convolution; according to the three-dimensional data acquired by the depth camera, extracting limb posture features by adopting a three-dimensional posture estimation network model; according to the voice data collected by the voice collection unit, a voice recognition model and a large language model are adopted to extract emotional features; and inputting the limb movement speed features, the limb posture features and the emotion features into a multi-modal attention fusion network, and determining whether a limb conflict behavior exists or not. According to the method, the limitation of a traditional single vision mode is broken through, multi-dimensional information of dynamic vision, three-dimensional postures and voice emotions is fused, the interference of environmental factors such as strong light and shielding on the single mode is reduced, and the missing report rate and the false report rate are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence and monitoring technology, and in particular to a method, device, and storage medium for identifying physical conflict behavior. Background Art

[0002] With the widespread application of artificial intelligence (AI) technology, behavior recognition systems based on computer vision and speech analysis are playing a vital role in surveillance. Accurately and efficiently identifying behaviors like physical altercations has become a key requirement for smart security systems, particularly in high-risk scenarios like campuses, vehicle cabins, public transportation, and hospitals.

[0003] Traditional methods for identifying physical aggression primarily use RGB cameras to collect visual information and rely on frame-sequence image recognition technologies, such as optical flow analysis, C3D, and I3D motion classification networks. These techniques capture features such as the trajectory of a person's body movements or changes in speed to infer the presence of physical aggression. However, in conditions such as strong lighting, obstructions, and poor camera angles, the collected visual information is susceptible to interference, and image recognition technology struggles to distinguish between normal physical behaviors like hugging, resulting in low accuracy in identifying physical aggression. Summary of the Invention

[0004] The embodiments of the present disclosure provide a method, device, and storage medium for identifying physical conflict behaviors, so as to solve the problems of poor robustness of existing single-modality recognition and low recognition accuracy of physical conflict behaviors.

[0005] In view of the above problems, in a first aspect, an embodiment of the present disclosure provides a method for identifying physical conflict behavior, comprising: Arrange a monitoring module at the monitoring location, wherein the monitoring module includes at least: a dynamic visual sensor, a depth camera and a voice acquisition unit; According to the event stream collected by the dynamic vision sensor, event spatiotemporal pyramid convolution is used to extract limb movement speed features; Extracting limb posture features using a three-dimensional posture estimation network model based on the three-dimensional data collected by the depth camera; Extracting emotional features using a speech recognition model and a large language model based on the speech data collected by the speech collection unit; The limb movement speed feature, the limb posture feature and the emotion feature are input into a multimodal attention fusion network to determine whether there is limb conflict behavior.

[0006] In a second aspect, a device for identifying physical conflict behavior is provided, comprising: A monitoring location arrangement module is used to arrange monitoring modules at the monitoring location, wherein the monitoring modules include at least: a dynamic visual sensor, a depth camera, and a voice acquisition unit; A dynamic visual sensor acquisition module is used to extract limb movement speed features using event spatiotemporal pyramid convolution based on the event stream collected by the dynamic visual sensor; A depth camera acquisition module is used to extract limb posture features using a three-dimensional posture estimation network model based on the three-dimensional data collected by the depth camera; A voice acquisition unit acquisition module, configured to extract emotional features using a voice recognition model and a large language model based on the voice data collected by the voice acquisition unit; The physical conflict behavior determination module is used to input the physical movement speed feature, the physical posture feature and the emotional feature into a multimodal attention fusion network to determine whether there is physical conflict behavior.

[0007] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for identifying physical conflict behavior as described in the first aspect or any possible implementation method in combination with the first aspect are executed.

[0008] The beneficial effects of the embodiments of the present disclosure include: The present disclosure provides a method, device, and storage medium for identifying physical conflict behaviors. The method comprises: deploying a monitoring module at a monitoring location, the monitoring module comprising at least a dynamic visual sensor, a depth camera, and a voice acquisition unit; extracting body movement speed features using event spatiotemporal pyramid convolution based on the event stream collected by the dynamic visual sensor; extracting body posture features using a 3D posture estimation network model based on the 3D data collected by the depth camera; and extracting emotion features using a speech recognition model and a large language model based on the voice data collected by the voice acquisition unit; and inputting the body movement speed features, body posture features, and emotion features into a multimodal attention fusion network to determine whether physical conflict behaviors exist. The method for identifying physical conflict behaviors provided by the disclosed embodiments transcends the limitations of traditional single visual modalities by integrating multi-dimensional information from dynamic vision, 3D posture, and voice emotion. This method reduces interference from environmental factors such as glare and occlusion on a single modality and reduces the rate of missed and false positives. The dynamic visual sensor captures the speed features of fast movements, while the depth camera provides 3D data, enabling accurate distinction between normal behaviors such as hugging and physical conflict. The emotion features complement the semantic and emotional dimensions, and the fusion of multiple features enables a more comprehensive judgment. The multimodal data collection and fusion mechanism significantly improves the accuracy, timeliness and explainability of physical conflict behavior recognition. It is particularly suitable for scenarios requiring a high level of security, such as campuses and public transportation, and provides security systems with intelligent and automated behavior recognition capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1A flowchart of a method for identifying physical conflict behavior provided in an embodiment of the present disclosure; Figure 2 A flow chart for determining limb movement speed characteristics provided by an embodiment of the present disclosure; Figure 3 A flowchart for determining limb posture features provided in an embodiment of the present disclosure; Figure 4 A flowchart for determining whether there is a physical conflict behavior provided in an embodiment of the present disclosure; Figure 5 This is a structural diagram of a device for identifying physical conflict behavior provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0010] The present disclosure provides a method, device, and storage medium for identifying physical conflict behavior. Preferred embodiments of the present disclosure are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are intended only to illustrate and explain the present disclosure and are not intended to limit the present disclosure. Furthermore, the embodiments and features within the embodiments of the present disclosure may be combined with one another unless there is a conflict.

[0011] The present disclosure provides a method for identifying physical conflict behavior. Figure 1 Shown, including: S101. Arrange a monitoring module at a monitoring location, wherein the monitoring module includes at least: a dynamic visual sensor, a depth camera, and a voice acquisition unit; S102, extracting limb movement speed features using event spatiotemporal pyramid convolution based on the event stream collected by the dynamic visual sensor; S103, extracting limb posture features using a three-dimensional posture estimation network model based on the three-dimensional data collected by the depth camera; S104, extracting emotional features using a speech recognition model and a large language model based on the speech data collected by the speech collection unit; S105: Input the limb movement speed features, limb posture features, and emotion features into the multimodal attention fusion network to determine whether there is limb conflict behavior.

[0012] With the widespread adoption of artificial intelligence technology, behavior recognition systems based on computer vision and speech analysis have become increasingly important in the surveillance field. Accurately and efficiently identifying behaviors such as physical altercations has become a critical requirement for smart security systems, particularly in high-risk scenarios such as campuses, vehicle cabins, public transportation, and hospitals. Traditional methods for identifying physical altercations primarily rely on RGB cameras to capture visual information. Their core technology relies on image recognition techniques for frame sequences, such as optical flow analysis and motion classification networks like C3D and I3D. These methods infer the presence of physical altercations by capturing features such as the movement trajectory and speed of human limbs. However, these methods have significant drawbacks: The collected visual information is easily disrupted by strong light, obstructions, or poor camera angles. Furthermore, image recognition technology is incapable of distinguishing between normal physical behaviors, such as hugging, and physical altercations. Traditional cameras often blur images when capturing fast movements, such as punching and shoving, making it difficult to accurately distinguish between hugging and physical altercations. This results in low accuracy in identifying physical altercations. Therefore, there is an urgent need to provide a method for identifying physical conflict behaviors that integrates multimodal data and large-scale model reasoning to solve the problem of low recognition accuracy of physical conflict behaviors.

[0013] In the embodiments of the present disclosure, a monitoring module including a dynamic vision sensor, a depth camera and a voice acquisition unit is arranged in a monitoring place to respectively collect event stream, three-dimensional data and voice data. The dynamic vision sensor (DVS, Dynamic Vision Sensor) can be an event-driven vision sensor inspired by the biological retina. Unlike traditional RGB cameras, the DVS generates event stream by capturing pixel brightness change events, can efficiently record dynamic information of fast body movements, and has strong anti-motion blur capability. The depth camera can be an imaging device that can simultaneously obtain a color image (RGB) of a scene and the distance (Depth) of each pixel to the camera. Unlike traditional RGB cameras that can only record colors, the depth camera additionally outputs a depth map (Depth Map), thereby expanding the 2D picture into 3D spatial information. The depth camera can collect three-dimensional data of the scene, accurately obtain the three-dimensional coordinates and spatial relationship of the human body, and make up for the errors in posture judgment of 2D images. The voice acquisition unit, such as a single microphone or a microphone array, is used to pick up voice signals in the monitoring area, including conversation content, tone, volume and other information, to provide data support for emotion feature extraction. Based on the event stream collected by the dynamic vision sensor, the time and space information of the pixel brightness change are recorded, and the event spatiotemporal pyramid convolution technology is used for feature extraction. The event spatiotemporal pyramid convolution (ESPC, Event-based Spatiotemporal Pyramid Convolution) can effectively capture the speed change (such as the speed of punching) of body movements in the time dimension and the movement range in the space dimension through multi-level convolution operation on the event stream in different time scales and space scales, accurately extract the body movement speed features reflecting the intensity of the movement, and solve the blur problem when the traditional RGB camera captures fast movements. The three-dimensional data collected by the depth camera, including the three-dimensional coordinates of each joint of the human body, are input into a three-dimensional pose estimation network model for processing. Through analysis of the three-dimensional coordinate data, the body posture features such as punching height, body contact angle, distance between people, etc. can be extracted, the spatial interaction relationship of the human body is accurately restored, the behaviors such as “hugging” and “pushing” which are easy to confuse are effectively distinguished, and the errors in posture judgment of 2D images are made up. Based on the voice data obtained by the voice acquisition unit, a speech recognition model is used to transcribe the collected voice data into text, and a large language model is input for context semantic and emotion analysis to identify potential threatening statements or emotion escalation behaviors and extract emotion features. For example, a speech recognition model such as Whisper is used to convert the voice data into a text sequence , which is input into a large language model (LLM) for semantic embedding to obtain emotion features, represented as: ; in, Represents the frozen or fine-tuned encoder of the large language model, input text sequence , The pre-trained weights can be completely frozen or fine-tuned on sentiment data. Freezing means that the parameters are no longer updated, while fine-tuning means that backpropagation continues on downstream tasks. The encoder uses a multi-layer Transformer structure. After the self-attention and position feedforward network, the context vector of each token is obtained, and then the sentence-level vector is aggregated through the pooling strategy. , The dimension is recorded as , Usually equal to the hidden layer dimension of the large language model, for example: 768, 1024, 4096, etc. Input linear mapping layer, composed of weight matrix and the bias vector Work together. The shape of is K×d, where K represents the total number of emotion categories and each row corresponds to the projection vector of an emotion. The dimension is K, providing a learnable offset for each emotion category. K-dimensional logits are obtained through matrix multiplication and addition. These are then fed into the Softmax function, which converts the logits into a probability distribution where all components are non-negative and sum to 1. Each component corresponds to the confidence level of an emotion category. The final output probability vector is the emotional signature of the model's emotional tendency for the entire speech segment.

[0014] Furthermore, a multimodal fusion decision network (such as a cross-modal Transformer or a fusion comparison model) performs joint inference on the output features of each modality to determine whether physical conflict has occurred. The multimodal attention fusion network uses an attention mechanism to automatically learn the weights of different features when determining physical conflict. For example, intense speed features and angry emotion features are given higher weights, thus deeply integrating and reasoning with multimodal features. By comprehensively analyzing a combination of multiple features, including body movement speed, posture, and emotion, it accurately outputs a judgment result on whether physical conflict has occurred, achieving efficient recognition in complex scenarios.

[0015] The embodiments of the present application break through the limitations of the traditional single visual modality and integrate multi-dimensional information of dynamic vision, three-dimensional posture and voice emotion, reducing the interference of environmental factors such as strong light and occlusion on the single modality, and reducing the rate of missed reports and false alarms. The dynamic vision sensor captures the speed characteristics of fast movements, and the depth camera provides three-dimensional data, which can accurately distinguish between normal behaviors such as hugging and physical conflicts; emotional characteristics supplement the semantic and emotional dimensions, and the fusion of multiple features makes the judgment more comprehensive. The multimodal data collection and fusion mechanism significantly improves the accuracy, timeliness and explainability of physical conflict behavior recognition, and is particularly suitable for scenarios requiring a high level of security, such as campuses and public transportation, providing security systems with intelligent and automated behavior recognition capabilities.

[0016] In another embodiment of the present disclosure, in the above step S102, based on the event stream collected by the dynamic vision sensor, event spatiotemporal pyramid convolution is used to extract limb movement speed features, including: Step 1: The event frames in the event stream collected by the dynamic visual sensor are divided into regions of multiple scales in the spatial dimension to obtain multi-scale spatial blocks; Step 2: For each spatial tile, segment the event stream in the time dimension, extract the activation of events in multiple time periods, and obtain multiple spatiotemporal sub-regions; Step 3: For each spatiotemporal sub-region, perform event aggregation convolution operation to extract velocity vector features; Step 4: Fuse the velocity vector features to obtain the limb movement velocity features.

[0017] In the disclosed embodiment, multi-scale spatial tiles are divided spatially, then segmented temporally. Velocity vector features are extracted from each spatiotemporal subregion and fused to obtain velocity features for body movements, providing a key basis for determining the intensity of body movements. Regarding step 1 above, event frames in the event stream captured by the dynamic visual sensor are spatially segmented into regions of multiple scales to obtain multi-scale spatial tiles. For example, spatial tiles are segmented based on the spatial granularity of the original image, 1 / 2 scaled, and 1 / 4 scaled. These spatial tiles of different scales can correspond to both overall body movements and localized subtle movements, ensuring that subsequent velocity features extracted do not miss any motion information within any spatial range, laying the foundation for comprehensively capturing the spatial characteristics of body movements. Regarding step 2 above, the event stream is segmented temporally for each spatial tile. The event stream corresponding to each spatial tile is then divided into several consecutive time periods at regular time intervals. The occurrence of events within each time period is then observed, and the activation of events across multiple time periods, such as the number and distribution of events, is extracted. Combining spatial and temporal information, by capturing event activation within different time periods, can reflect the dynamic changes in limb movements along the temporal dimension, such as the different manifestations of a punching motion at its initiation, midpoint, and end. This results in multiple spatiotemporal subregions. For step 3 above, an event aggregation convolution operation is performed on each spatiotemporal subregion to extract velocity vector features. This aggregation convolution integrates the event information within each spatiotemporal subregion, analyzing the event density (i.e., the number of events occurring per unit time and space) and the event trends, such as transitions from sparse to dense or vice versa. Through this analysis, the velocity direction and magnitude of the limb movement within the region are calculated to form a velocity vector feature. Velocity vector features intuitively and accurately reflect the motion state of a limb within a specific spatiotemporal range. For step 4 above, the velocity vector features are fused to generate a limb movement velocity feature. For example, weighted fusion is used to integrate the velocity vector features extracted from each spatiotemporal subregion. This fusion integrates limb movement velocity information across different spatial scales and time periods to form a comprehensive feature reflecting limb movement velocity. The velocity characteristics of body movements can accurately reflect the intensity of body movements, providing an important speed dimension for subsequent judgments about whether there is physical conflict. Multi-scale spatial partitioning takes into account both overall and local body movements, and temporal segmentation captures changes in movement at different moments, allowing the extracted velocity characteristics to fully reflect the movement state of the body within the scope of time and space. Event aggregation convolution operates on each spatiotemporal sub-region, accurately calculating the speed direction and magnitude of the body movements within that region, making the velocity vector characteristics more accurate. This combination of time and space effectively solves the problem of blurring caused by traditional frame-based cameras when capturing fast movements, and improves the extraction of speed characteristics of fast body movements.

[0018] In another embodiment of the present disclosure, in step 2 above, an event aggregation convolution operation is performed on each spatiotemporal sub-region to extract velocity vector features, including: Step 1: For each spatiotemporal sub-region, perform event aggregation convolution operation to obtain the number of events, the displacement between events, and the time interval between events; Step 2: Determine the velocity vector characteristics based on the number of events, the displacement between events, and the time interval between events. The formula is: ; in, represents the velocity vector feature, Indicates the number of events, represents the displacement between events, Indicates the time interval between events.

[0019] In the embodiment of the present disclosure, the event aggregation convolution operation is used to extract three key parameters: the number of events, the displacement between events, and the time interval between events from each spatiotemporal sub-region; then, based on these three parameters, the velocity vector feature is calculated to accurately quantify the velocity state of the limb movement in the spatiotemporal sub-region. For the above step 1, the data point of each event includes the pixel Coordinates, pixels Coordinates, timestamps and polarity, polarity reflects the brightness change of the pixel, expressed as , . For each spatiotemporal sub-region, parameter extraction is completed through event aggregation convolution operation. Event aggregation convolution can be a convolution method specifically used to process dynamic visual sensor event streams, and performs sliding window aggregation on discrete brightness change events in spatiotemporal sub-regions. The number of events, the displacement between events, and the time interval between events are obtained. The number of events represents the total number of events contained in the statistical window, reflecting the activity level of the action in the sub-region. The displacement between events represents the calculation of the difference between two adjacent events in the window. Direction and The coordinate difference of the direction reflects the spatial movement distance of the limb movement. The time interval between events means calculating the difference between the timestamps of two adjacent events, which reflects the time rhythm of the action. The convolution operation is used to achieve efficient parameter extraction, ensuring that the action details of each spatiotemporal sub-region are accurately captured. For the above step 2, Indicates the For adjacent events The instantaneous velocity component in the direction. Indicates the For adjacent events The instantaneous velocity component in the direction. By formula: ; The obtained speed vector feature integrates the motion information of events in the spatio-temporal sub-region, can reflect the speed of the action, and can also reflect the direction (positive and negative of the component) of the action, thereby providing a quantitative basis for judging the intensity of the limb action. The event aggregation convolution operation simultaneously obtains the number of events, the displacement between events, and the time interval, covers the number, spatial change, and time span of the action, provides a complete data basis for speed calculation, and avoids the deviation of the speed feature caused by the lack of parameters. The speed components of multiple events are integrated by the average formula, which weakens the accidental error of a single event, and the obtained speed vector feature can better reflect the real speed trend of the limb action. The speed vector feature is directly related to the spatial displacement and time interval of the event, has clear physical meaning, and is convenient for weight distribution and logical reasoning of the speed vector feature in subsequent multi-modal fusion. As shown in Figure 2 Figure 2 A flowchart for determining the limb action speed feature is shown in FIG. 8, including the following steps: S201, dividing the event frame in the event stream collected by the dynamic vision sensor into multiple scale regions in the spatial dimension, to obtain multiple scale spatial tiles; S202, for each spatial tile, segmenting the event stream in the time dimension, extracting the activation of events in multiple time periods, to obtain multiple spatio-temporal sub-regions; S203, for each spatio-temporal sub-region, performing an event aggregation convolution operation to obtain the number of events, the displacement between events, and the time interval between events; S204, determining the speed vector feature according to the number of events, the displacement between events, and the time interval between events; S205, fusing the speed vector feature to obtain the limb action speed feature; and the flow ends.

[0020] In another embodiment of the present disclosure, in step S103, the limb pose feature is extracted from the three-dimensional data collected by the depth camera using a three-dimensional pose estimation network model, including: Step 1, extracting the position information of the joint node from the three-dimensional data collected by the depth camera using a three-dimensional pose estimation network model; Step 2, constructing a limb pose time sequence feature sequence according to the position information of the joint node; Step 3, inputting the limb pose time sequence feature sequence into a time convolution network model to model and classify the limb action, and determining the limb pose feature.

[0021] ​​In the disclosed embodiments, a 3D pose estimation network model is used to extract joint position information from 3D data. This joint position information is used to construct a temporal feature sequence of limb poses. This temporal feature sequence is then fed into a temporal convolutional network model. Through modeling and classification, limb pose features are determined, enabling accurate characterization of the limb's spatial pose and dynamic changes. For step 1 above, a 3D pose estimation network model (3DPoseNet) is used to process the 3D data captured by a depth camera. By learning features from the 3D data, the 3D pose estimation network model automatically identifies human joints, such as the shoulder, elbow, and wrist, and outputs the 3D coordinates (x, y, z) of each joint. For example, when identifying a punching motion, the 3D position coordinates of the wrist and elbow joints can be precisely located, reflecting the spatial extension of the arm. This eliminates the ambiguity of joint position in 2D images and provides a precise spatial coordinate basis for subsequent pose analysis. For step 2 above, a temporal feature sequence of limb poses is constructed. Based on the extracted joint position information, the joint coordinates of consecutive frames are concatenated in chronological order to form a temporal feature sequence of limb poses. Temporal feature sequences of limb postures can intuitively reflect dynamic changes in limb posture. For example, in limb contact movements, the displacement of the chest from back to front and the change in angle of the elbow from flexion to extension all produce specific numerical variations in the temporal sequence. Converting static spatial postures into dynamic temporal sequences provides data support for capturing the continuity of movements. For step 3 above, a temporal convolutional network (TCN) model models and classifies limb movements. The temporal feature sequence of limb postures is input into a temporal convolutional network (TCN) model. Through multi-layer convolution operations, the modeling and classification of the temporal feature sequence of limb postures is performed, ultimately outputting limb posture features. Using dilated convolution kernels, the TCN can expand the temporal receptive field without increasing parameters, effectively capturing long-range temporal dependencies in the temporal sequence, such as the complete temporal logic of "hand raise-acceleration-strike" in a punching motion. By learning different action sequences, such as the slow and gentle "hug" sequence and the violent and rapid "push" sequence, the model can classify the movement patterns of postures and convert them into distinctive body posture features, providing a key basis for judging physical conflict. Relying on the three-dimensional data of the depth camera, it can accurately obtain the three-dimensional coordinates of joints compared to 2D images, avoiding posture misjudgments caused by occlusion and overlap in a two-dimensional perspective, and laying a precise foundation for subsequent feature extraction. By constructing a temporal feature sequence of body postures, the static joint positions are converted into dynamic action trajectories, which can fully record the changes in body posture over time, compensating for the limitations of posture information at a single moment.Temporal convolutional networks are good at capturing local and global dependencies in time series data, can effectively distinguish similar postures, and improve the classification accuracy of different limb behaviors through the temporal evolution pattern of movements.

[0022] In yet another embodiment of the present disclosure, the limb posture temporal feature sequence includes: a sequence of velocity and angle changes between joints; In step 2 above, based on the position information of the joint points, a temporal feature sequence of the limb posture is constructed, including: According to the position information of the joint points, the velocity and angle change sequence between joints is constructed, and the formula is expressed as: ; ; ; in, represents the set of joint node positions, express The position of the joint nodes, Indicates the position coordinates of the joint node, Indicates time, Indicates the time difference, Indicates the joint nodes, Indicates the joint nodes, Represents a joint node and joint nodes The change in the angle between Represents the inverse cosine function.

[0023] In the disclosed embodiment, based on the position information of the joint points, the inter-joint velocity is obtained by calculating the difference in joint position at adjacent time points, and the angle change between the joints is calculated using the vector dot product and the arc cosine function. Finally, these velocity and angle changes are arranged in chronological order to form a sequence of inter-joint velocity and angle changes that reflects the dynamic relationship of joint movement, thereby depicting the movement trend of the joints and the law of spatial angle change during limb movements. Through the formula: ; By calculating vector dot products and modulus lengths, the spatial angular relationships between joints can be accurately quantified. This avoids the interference of distance on angle judgment and enhances the robustness of the features. The velocity between joints directly reflects the speed of movement, while the angle change reflects the spatial posture transformation of the joints. The combination of the two can better capture the dynamic characteristics of the action, providing a key basis for distinguishing between "hugging" (slow, small angle changes) and "pushing" (fast, large angle changes). Recording the sequence of velocity and angle changes between joints in a time series format fully preserves the evolution of the limb movement from start to finish, providing rich dynamic features for the subsequent temporal convolutional network to capture the long-term dependencies of the action.

[0024] In yet another embodiment of the present disclosure, the limb posture temporal feature sequence includes: an elbow and wrist motion speed sequence and a punching motion sequence; In step 2 above, based on the position information of the joint points, a temporal feature sequence of the limb posture is constructed, including: Step 1: Based on the position information of the joint points, construct the elbow and wrist motion speed sequence, which is expressed as: ; ; in, Indicates the position coordinates of the elbow or wrist joint node, Indicates time, Indicates time difference; Step 2: When the elbow and wrist motion speed represented by the elbow and wrist motion speed sequence is greater than a first threshold and the speed direction is toward the target person, determine the punching motion sequence according to the position coordinates of the elbow and wrist joint nodes.

[0025] In the disclosed embodiment, the motion speed is calculated using the position information of the joint points to form an elbow and wrist motion speed sequence. The punching motion sequence is then filtered out from the elbow and wrist motion speed sequence based on the speed threshold and direction conditions. Regarding step 1 above, the instantaneous speed of the elbow and wrist joints is calculated to form a time series sequence reflecting their dynamic changes, thereby obtaining the elbow and wrist motion speed sequence. The formula is: ; ; in, Indicates the position coordinates of the elbow or wrist joint node, Indicates time, Represents the time difference. For step 2 above, based on the elbow and wrist speed sequences, conditional screening is performed to extract time sequences that match the characteristics of a punch. The elbow and wrist speeds are greater than a preset first threshold. This threshold is set based on the speed difference between normal and violent movements. For example, a punch is typically much faster than a normal hand swing. For example, if the elbow and wrist speed is 3 m / s, the first threshold is 2 m / s. The speed is directed toward the target person. For example, if the angle between the elbow and wrist speed directions and the target person's relative position is less than a preset angle threshold, the action is identified as a punch. For example, the preset angle threshold is 90 degrees. The punch sequence is then derived based on the position coordinates of the corresponding elbow and wrist joint nodes. Focusing on the elbow and wrist reduces irrelevant information interference compared to full-body joint features and improves sensitivity to conflicting movements. For example, during a punch, the velocity changes of the elbow and wrist are far more significant than those of other joints. This targeted feature can improve recognition efficiency. Speed ​​changes are recorded in a time series format, fully preserving the dynamic process of the punching action. Combined with direction changes, this provides a complete basis for subsequent recognition of the punching action and improves the temporal accuracy of action recognition.

[0026] In yet another embodiment of the present disclosure, the limb posture temporal feature sequence includes: a distance change sequence between target person's chest positions and a limb contact action sequence; In step 2 above, based on the position information of the joint points, a temporal feature sequence of the limb posture is constructed, including: Step 1: Based on the position information of the joint points, construct the distance change sequence between the target person's chest positions. The formula is expressed as: ; ; in, Indicates the distance change between the target person's chest position, Represents the position coordinates of the chest joint node, Indicates time, and Represents the first target person and the second target person respectively; Step 2: When the instantaneous change rate of the distance between the target person's chest positions represented by the distance change sequence between the target person's chest positions is less than the second threshold, and the distance between the target person's chest positions is less than the third threshold, determine the limb contact action sequence based on the position coordinates of the chest joint nodes.

[0027] In the embodiments of the present disclosure, the distance between the chest positions of the first target person and the second target person is calculated based on the chest position information of the target person, and a sequence of distance changes over time is formed, and then a sequence of limb contact actions is filtered from the sequence of distance changes based on threshold conditions of distance size and change rate. The distance between the chest positions of the first target person and the second target person is calculated respectively, and then arranged in time sequence to form a sequence of distance changes between the chest positions of the target persons. For example, when two people approach from a distance, the distance between their chest positions will gradually decrease, when they push and touch each other, the distance between their chest positions may fluctuate rapidly within a small range; when they separate, the distance between their chest positions will gradually increase. The sequence of distance changes between the chest positions of the target persons directly reflects the dynamic relationship between the relative positions of the two people. According to the sequence of distance changes between the chest positions of the target persons, a sequence of time sequences that meet the characteristics of limb contact is extracted through conditional filtering to obtain a sequence of limb contact actions. For example, the instantaneous change rate of the distance between the chest positions of the target persons in the sequence of distance changes between the chest positions of the target persons is less than a second threshold, and the distance between the chest positions of the target persons is less than a third threshold, which is expressed by the formula: ; and ; wherein represents the instantaneous change rate of the distance between the chest positions of the target persons, represents the second threshold, and the negative sign “-” indicates that the speed direction is that the chest positions of the target persons are approaching each other, represents the distance between the chest positions of the target persons, represents the third threshold, for example, , For the “push and touch” behavior, the time sequence is: , ; , ; , ; , ; . For the “hug” behavior, the time sequence is: , ; , ; , ; , ; , ; . The corresponding physical contact action sequence is obtained according to the position coordinates of the chest joint nodes. Focusing on the core area of ​​the human body, the chest, its distance change can intuitively reflect the relative position relationship between characters (such as close, far away), and can more stably reflect the overall interaction state than the limbs and other parts, providing core spatial features for judging physical conflicts such as "pushing and bumping". The dual conditions of "distance less than the third threshold" (close enough in space) and "instantaneous change rate less than the second threshold" can effectively distinguish between "hugging" and "pushing and bumping" to reduce misjudgment. The constructed physical contact action sequence can accurately capture the close interaction state between characters, and provide key spatial and dynamic basis for the identification of contact behaviors such as "pushing and bumping" in physical conflicts. Figure 3 As shown, Figure 3 The flowchart for determining limb posture characteristics includes the following steps: S301, extracting the position information of the joint points using a 3D posture estimation network model based on the 3D data collected by the depth camera; S302: Constructing a limb posture temporal feature sequence based on the position information of the joint points; the limb posture temporal feature sequence may include: a velocity and angle change sequence between joints, a velocity sequence of elbow and wrist movements, a punching action sequence, a distance change sequence between the target person's chest positions, and a limb contact action sequence; S303: Input the temporal feature sequence of the limb posture into the temporal convolutional network model to model and classify the limb movements and determine the limb posture features; the process ends.

[0028] In another embodiment of the present disclosure, Figure 4 As shown, in the above step S105, the body movement speed feature, the body posture feature and the emotion feature are input into the multimodal attention fusion network to determine whether there is a physical conflict behavior, including: S401: After concatenating or weighted summing the limb movement speed features, limb posture features, and emotion features to unify their dimensions, the features are input into a multi-head sub-attention structure to obtain a cross-modal fusion unified feature vector. S402: Inputting the cross-modal fusion unified feature vector into a fully connected classifier to obtain a probability score of whether physical conflict behavior occurs; The method also includes: Step 3: Use weighted cross entropy loss to train the multimodal attention fusion network.

[0029] In the embodiments of the present disclosure, after unifying the limb motion speed feature, the limb posture feature and the emotion feature in the same dimension, a cross-modal fusion unified feature vector is generated through a multi-head sub-attention structure; then the cross-modal fusion unified feature vector is input into a fully connected classifier to obtain a probability score of the existence of limb conflict behavior; and a weighted cross-entropy loss is used to train the network to optimize the model performance. Through the deep fusion of multi-modal features and the optimization of model training, accurate judgment of limb conflict behavior is realized. For the above step S401, since the dimensions of the limb motion speed feature, the limb posture feature and the emotion feature may be different, the dimensions need to be unified through splicing or weighted summation to ensure that the features can be input into the attention network. The unified dimension features are input into a multi-head sub-attention structure, which includes multiple parallel sub-attention heads, each of which focuses on learning the correlation between different modal features. The features are weighted and fused by calculating the attention weights between the features, and finally a cross-modal fusion unified feature vector is output. For example, after the limb motion speed feature, the limb posture feature and the emotion feature are unified in the same dimension through splicing or weighted summation, they are represented as 、 、 respectively, and are input into the multi-head sub-attention structure to obtain a cross-modal fusion unified feature vector represented as: ); For the above step S402, the fully connected classifier is composed of a multi-layer neural network, the input is the cross-modal fusion unified feature vector, and the cross-modal fusion unified feature vector is mapped to a probability score in the interval [0, 1] through linear transformation and activation function (such as ReLU, Sigmoid). The closer the score is to 1, the higher the possibility of the existence of limb conflict behavior; the closer to 0, the higher the possibility of normal behavior. The cross-modal fusion unified feature vector is input into the fully connected classifier to obtain the probability score of the existence of limb conflict behavior, which is represented by the formula: ); wherein, and represent the weights and biases of the fully connected classifier, represents mapping to [0, 1] to obtain the probability score of limb conflict behavior. For the above step 3, the weighted cross-entropy loss pays more attention to the prediction error of the limb conflict behavior sample by introducing a weight parameter. The multi-modal attention fusion network is trained using the weighted cross-entropy loss, which is represented by the formula: ; wherein, represents the true label, represents the output probability of the multi-modal attention fusion network, and Represent the weights of positive and negative samples, respectively. During the training phase of the multimodal attention fusion network, the output probability of the multimodal attention fusion network and the true label are substituted into the weighted cross-entropy loss formula, and the network parameters are updated through backpropagation to minimize the loss value. Ultimately, the multimodal attention fusion network significantly improves its ability to recognize a small number of conflicting samples while learning a large number of normal samples. The multi-head sub-attention structure can automatically learn the weights of different modal features, avoiding information redundancy caused by simple splicing, enhancing the correlation between features, and solving the information island problem in traditional modal fusion. The fully connected classifier outputs a probability score based on the unified fused feature vector, reducing omissions or false positives caused by misjudgment of a single feature. The weighted cross-entropy loss addresses the sample imbalance problem by assigning higher weights to minority samples, allowing the model to focus more on key conflicting samples during training and improving recognition stability in complex scenarios.

[0030] Based on the same disclosed concept, the embodiments of the present disclosure also provide a device for identifying physical conflict behavior. Since the principles of the problems solved by these devices are similar to those of the aforementioned methods for identifying physical conflict behavior, the implementation of the device can refer to the implementation of the aforementioned methods, and the repeated parts will not be repeated.

[0031] The present disclosure provides a device for identifying physical conflict behaviors. Figure 5 Shown, including: A monitoring location arrangement module 501 is configured to arrange monitoring modules at the monitoring location, wherein the monitoring modules include at least a dynamic visual sensor, a depth camera, and a voice acquisition unit; A dynamic vision sensor acquisition module 502 is configured to extract limb movement speed features using event spatiotemporal pyramid convolution based on the event stream acquired by the dynamic vision sensor; A depth camera acquisition module 503 is configured to extract limb posture features using a three-dimensional posture estimation network model based on the three-dimensional data acquired by the depth camera; The speech acquisition unit acquisition module 504 is used to extract emotional features using a speech recognition model and a large language model based on the speech data collected by the speech acquisition unit; The physical conflict behavior determination module 505 is used to input the physical movement speed feature, the physical posture feature and the emotional feature into the multimodal attention fusion network to determine whether there is physical conflict behavior.

[0032] In another embodiment of the present disclosure, the dynamic vision sensor acquisition module 502 is configured to divide event frames in the event stream acquired by the dynamic vision sensor into regions of multiple scales in a spatial dimension to obtain multi-scale spatial tiles; For each spatial tile, the event flow is segmented in the time dimension, and the activation of events in multiple time periods is extracted to obtain multiple spatiotemporal sub-regions; For each spatiotemporal sub-region, an event aggregation convolution operation is performed to extract velocity vector features; The velocity vector features are fused to obtain limb movement velocity features.

[0033] In another embodiment of the present disclosure, the dynamic visual sensor acquisition module 502 is configured to perform an event aggregation convolution operation on each spatiotemporal sub-region to obtain the number of events, the displacement between events, and the time interval between events; The velocity vector characteristics are determined according to the number of events, the displacement between events, and the time interval between events. The formula is expressed as follows: ; in, represents the velocity vector feature, Indicates the number of events, represents the displacement between events, Indicates the time interval between events.

[0034] In another embodiment of the present disclosure, the depth camera acquisition module 503 is configured to extract position information of joint points using a three-dimensional pose estimation network model based on the three-dimensional data acquired by the depth camera; Constructing a temporal feature sequence of limb postures according to the position information of the joint points; The limb posture temporal feature sequence is input into a temporal convolutional network model to model and classify limb movements and determine limb posture features.

[0035] In yet another embodiment of the present disclosure, the limb posture temporal feature sequence includes: a sequence of velocity and angle changes between joints; The depth camera acquisition module 503 is used to construct a sequence of joint velocity and angle changes based on the position information of the joint points, which can be expressed as follows: ; ; ; in, represents the set of joint node positions, express The position of the joint nodes, Indicates the position coordinates of the joint node, Indicates time, Indicates the time difference, Indicates the joint nodes, Indicates the joint nodes, Represents a joint node and joint nodes The change in the angle between Represents the inverse cosine function.

[0036] In another embodiment of the present disclosure, the limb posture temporal feature sequence includes: an elbow and wrist motion speed sequence and a punching motion sequence; The depth camera acquisition module 503 is used to construct the elbow and wrist motion speed sequence based on the position information of the joint points, which is expressed as: ; ; in, Indicates the position coordinates of the elbow or wrist joint node, Indicates time, Indicates time difference; When the elbow and wrist movement speed represented by the elbow and wrist movement speed sequence is greater than a first threshold and the speed direction is toward the target person, a punching action sequence is determined according to the position coordinates of the elbow and wrist joint nodes.

[0037] In yet another embodiment of the present disclosure, the limb posture temporal feature sequence includes: a distance change sequence between target person's chest positions and a limb contact action sequence; The depth camera acquisition module 503 is used to construct a distance change sequence between the target person's chest positions based on the position information of the joint points, which is expressed as follows: ; ; in, Indicates the distance change between the target person's chest position, Represents the position coordinates of the chest joint node, Indicates time, and Represents the first target person and the second target person respectively; When the instantaneous change rate of the distance between the target person's chest positions represented by the distance change sequence between the target person's chest positions is less than a second threshold, and the distance between the target person's chest positions is less than a third threshold, the limb contact action sequence is determined according to the position coordinates of the chest joint nodes.

[0038] In another embodiment of the present disclosure, the physical conflict behavior determination module 505 is configured to concatenate or weighted-sum the physical movement speed feature, the physical posture feature, and the emotional feature to unify their dimensions, and then input the concatenated features into a multi-head sub-attention structure to obtain a cross-modal fusion unified feature vector. Inputting the cross-modal fusion unified feature vector into a fully connected classifier to obtain a probability score of whether physical conflict behavior occurs; The physical conflict behavior determination module 505 is further configured to: The multimodal attention fusion network is trained using weighted cross entropy loss.

[0039] Based on the same disclosed concept, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for identifying physical conflict behavior as described in any of the above embodiments are executed.

[0040] Through the above description of the embodiments, those skilled in the art will clearly understand that the embodiments of the present disclosure can be implemented through hardware or through software plus the necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in the various embodiments of the present disclosure.

[0041] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes in the accompanying drawings are not necessarily required for implementing the present disclosure.

[0042] Those skilled in the art will appreciate that the modules in the devices of the embodiments may be distributed in the devices of the embodiments as described in the embodiments, or may be located in one or more devices different from the embodiments with corresponding changes. The modules of the above embodiments may be combined into one module or further split into multiple submodules.

[0043] The serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.

[0044] Obviously, those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.

Claims

1. A method for identifying physical conflict behavior, characterized in that: include: Arrange a monitoring module at the monitoring location, wherein the monitoring module includes at least: a dynamic visual sensor, a depth camera and a voice acquisition unit; According to the event stream collected by the dynamic vision sensor, event spatiotemporal pyramid convolution is used to extract limb movement speed features; Extracting limb posture features using a three-dimensional posture estimation network model based on the three-dimensional data collected by the depth camera; Extracting emotional features using a speech recognition model and a large language model based on the speech data collected by the speech collection unit; The limb movement speed feature, the limb posture feature and the emotion feature are input into a multimodal attention fusion network to determine whether there is limb conflict behavior.

2. The method according to claim 1, wherein The method of extracting limb movement speed features using event spatiotemporal pyramid convolution based on the event stream collected by the dynamic vision sensor includes: Dividing event frames in the event stream collected by the dynamic vision sensor into regions of multiple scales in the spatial dimension to obtain multi-scale spatial blocks; For each spatial tile, the event flow is segmented in the time dimension, and the activation of events in multiple time periods is extracted to obtain multiple spatiotemporal sub-regions; For each spatiotemporal sub-region, an event aggregation convolution operation is performed to extract velocity vector features; The velocity vector features are fused to obtain limb movement velocity features.

3. The method according to claim 2, wherein The event aggregation convolution operation is performed on each spatiotemporal sub-region to extract velocity vector features, including: For each spatiotemporal subregion, an event aggregation convolution operation is performed to obtain the number of events, the displacement between events, and the time interval between events; The velocity vector characteristics are determined according to the number of events, the displacement between events, and the time interval between events. The formula is expressed as follows: ; in, represents the velocity vector feature, Indicates the number of events, represents the displacement between events, Indicates the time interval between events.

4. The method according to claim 1, wherein The method of extracting limb posture features using a three-dimensional posture estimation network model based on the three-dimensional data collected by the depth camera includes: Extracting the position information of the joint points using a three-dimensional posture estimation network model based on the three-dimensional data collected by the depth camera; Constructing a temporal feature sequence of limb postures according to the position information of the joint points; The limb posture temporal feature sequence is input into a temporal convolutional network model to model and classify limb movements and determine limb posture features.

5. The method according to claim 4, wherein The limb posture temporal feature sequence includes: a sequence of velocity and angle changes between joints; The step of constructing a temporal feature sequence of limb postures based on the position information of the joints includes: According to the position information of the joint points, the velocity and angle change sequence between joints is constructed, and the formula is expressed as: ; ; ; in, represents the set of joint node positions, express The position of the joint nodes, Indicates the position coordinates of the joint node, Indicates time, Indicates the time difference, Indicates the joint nodes, Indicates the joint nodes, Represents a joint node and joint nodes The change in the angle between Represents the inverse cosine function.

6. The method according to claim 4, wherein The limb posture temporal feature sequence includes: elbow and wrist movement speed sequence and punching movement sequence; The step of constructing a temporal feature sequence of limb postures based on the position information of the joints includes: According to the position information of the joint points, the elbow and wrist motion speed sequence is constructed, and the formula is expressed as: ; ; in, Indicates the position coordinates of the elbow or wrist joint node, Indicates time, Indicates time difference; When the elbow and wrist movement speed represented by the elbow and wrist movement speed sequence is greater than a first threshold and the speed direction is toward the target person, a punching action sequence is determined according to the position coordinates of the elbow and wrist joint nodes.

7. The method according to claim 4, wherein The limb posture temporal feature sequence includes: a distance change sequence between the target person's chest positions and a limb contact action sequence; The step of constructing a temporal feature sequence of limb postures based on the position information of the joints includes: According to the position information of the joint points, a distance change sequence between the chest positions of the target person is constructed, and the formula is expressed as: ; ; in, Indicates the distance change between the target person's chest position, Represents the position coordinates of the chest joint node, Indicates time, and Represents the first target person and the second target person respectively; When the instantaneous change rate of the distance between the target person's chest positions represented by the distance change sequence between the target person's chest positions is less than a second threshold, and the distance between the target person's chest positions is less than a third threshold, the limb contact action sequence is determined according to the position coordinates of the chest joint nodes.

8. The method according to claim 1, wherein Inputting the limb movement speed feature, the limb posture feature, and the emotion feature into a multimodal attention fusion network to determine whether there is a physical conflict behavior includes: The limb movement speed feature, the limb posture feature, and the emotion feature are concatenated or weighted summed to unify the dimensions, and then input into a multi-head sub-attention structure to obtain a cross-modal fusion unified feature vector; Inputting the cross-modal fusion unified feature vector into a fully connected classifier to obtain a probability score of whether physical conflict behavior occurs; The method further comprises: The multimodal attention fusion network is trained using weighted cross entropy loss.

9. A device for identifying physical conflict behavior, characterized in that: include: A monitoring location arrangement module is used to arrange monitoring modules at the monitoring location, wherein the monitoring modules include at least: a dynamic visual sensor, a depth camera, and a voice acquisition unit; A dynamic visual sensor acquisition module is used to extract limb movement speed features using event spatiotemporal pyramid convolution based on the event stream collected by the dynamic visual sensor; A depth camera acquisition module is used to extract limb posture features using a three-dimensional posture estimation network model based on the three-dimensional data collected by the depth camera; A voice acquisition unit acquisition module, configured to extract emotional features using a voice recognition model and a large language model based on the voice data collected by the voice acquisition unit; The physical conflict behavior determination module is used to input the physical movement speed feature, the physical posture feature and the emotional feature into a multimodal attention fusion network to determine whether there is physical conflict behavior.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for identifying physical conflict behavior according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Emotion recognition method based on large model and related device

    CN119904901A

  • Exercise rehabilitation evaluation method and system based on limb posture and emotion recognition

    CN120340110A

  • Emotion recognition method and device based on artificial intelligence

    CN120472518A

  • Multimodal heterogeneous feature fusion-based compact video event description method

    WO2023050295A1