Limb conflict detection and early warning method based on video vibration image analysis
Through the method based on video vibration image analysis, the vibration image data and bone feature data of the target object in the video are extracted, and the risk assessment of physical conflict is combined with these data, which solves the problem of real-time detection and early warning of physical conflict events in the prior art, and achieves efficient and accurate safety monitoring and early warning.
Patent Information
- Application Number
- CN202510000580.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art is difficult to realize real-time detection and early warning of physical conflict events, especially in scenarios involving subtle movements and posture changes, and there is a lack of effective detection and early warning solutions.
The early warning method of limb conflict detection based on video vibration image analysis is used. By obtaining the video files to be detected, decomposed into image sequences, the vibration image feature extraction model and bone feature extraction model are used, the vibration image data and bone feature data are extracted, and the body conflict detection model is used to evaluate the limb conflict risk level of the target object.
Real-time detection and early warning of physical conflict events is achieved, the efficiency and accuracy of safety monitoring is improved, potential physical conflict events can be identified more accurately, and safety warnings are provided strongly.
Smart Images

Figure CN119920008A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to data processing technology, and in particular to a physical conflict detection and early warning method based on video vibration image analysis. Background Art
[0002] With the increasing demand for public safety, rapid detection and early warning of potential physical conflict incidents have become an important technical challenge. Traditional security monitoring methods mainly rely on manual monitoring and post-event analysis, which is not only inefficient but also prone to missing key information and unable to achieve real-time early warning of physical conflict incidents. In addition, although traditional image processing methods can automatically identify abnormal behaviors to a certain extent, due to the lack of in-depth analysis of subtle human movements and posture changes, it is often difficult to accurately judge the risk level of physical conflict.
[0003] In recent years, with the rapid development of computer vision and deep learning technologies, intelligent analysis technology based on video data has shown great potential in the field of security monitoring. However, most of the existing intelligent analysis technologies based on video focus on understanding macroscopic scenes and identifying overall behaviors. For scenes such as physical conflicts involving subtle movements and posture changes, there is still no effective detection and warning solution in the existing technology. Summary of the invention
[0004] The present application provides a physical conflict detection and early warning method based on video vibration image analysis, which is used to detect and predict physical conflicts, thereby identifying potential physical conflict incidents, and providing strong support for security monitoring and early warning.
[0005] In a first aspect, the present application provides a method for detecting and warning physical conflict based on video vibration image analysis, comprising:
[0006] Acquire a video file to be detected, and decompose the video file to be detected into a sequence of images to be detected, wherein the number of frames of the video file to be detected is greater than a preset frame number threshold, and each image to be detected in the sequence of images to be detected includes a target object to be detected;
[0007] Extracting vibration image data from the image sequence to be detected using a preset vibration image feature extraction model, wherein the vibration image data includes vibration frequency data and amplitude data;
[0008] Extracting bone feature data from the image sequence to be detected by using a preset bone feature extraction model, wherein the bone feature data includes bone key point position data and bone posture data;
[0009] The physical conflict risk level of the target object to be detected is detected by using a preset physical conflict detection model and according to the vibration image data and the bone feature data.
[0010] In the above scheme, by obtaining the video file to be detected and decomposing it into the image sequence to be detected, and ensuring that the number of frames of the video file to be detected is greater than the preset frame number threshold, the sufficiency and continuity of the data are guaranteed. At the same time, each image to be detected contains the target object to be detected, that is, the potential conflict participant, which further narrows the scope of analysis and improves the detection efficiency. Then, the vibration image data is extracted from the image sequence to be detected using the preset vibration image feature extraction model. The vibration image data includes vibration frequency data and amplitude data, which can reflect the small movements and vibrations of the target object in the video. By analyzing these vibration data, it can be preliminarily determined whether the target object has violent movements or abnormal behaviors, thereby providing clues for the detection of physical conflicts. Then, the preset bone feature extraction model is further used to extract bone feature data from the image sequence to be detected. The bone feature data includes bone key point position data and bone posture data, which can depict the body posture and motion trajectory of the target object. By comparing the bone feature data between different frames, the body posture changes of the target object can be monitored in real time, and then it can be determined whether it has attacking or defensive physical movements. Finally, the preset physical conflict detection model is used to comprehensively evaluate the physical conflict risk level of the target object to be detected based on the vibration image data and bone feature data, thereby combining vibration information and bone posture information to achieve comprehensive and accurate detection of physical conflicts. Through the preset risk feature level mapping table, the calculated physical conflict risk feature value can be converted into an intuitive risk level, thereby identifying potential physical conflict events, providing strong support for security monitoring and early warning.
[0011] Optionally, the extracting vibration image data from the image sequence to be detected by using a preset vibration image feature extraction model includes:
[0012] Determine the motion vector of each pixel point according to the image sequence to be detected to form a time domain motion vector field;
[0013] Determine a corresponding frequency domain motion vector field according to the time domain motion vector field;
[0014] The vibration image data is extracted according to the frequency domain motion vector field.
[0015] In the above scheme, the motion vector of each pixel is determined according to the image sequence to be detected to form a time domain motion vector field, so as to capture the dynamic behavior of the target object in the video by calculating the displacement change of the pixel points between adjacent frames. The construction of the time domain motion vector field provides basic data for the subsequent frequency domain analysis, ensuring the continuity and accuracy of vibration feature extraction. Then, the corresponding frequency domain motion vector field is determined according to the time domain motion vector field, and the time domain signal is converted into a frequency domain signal, thereby revealing the distribution characteristics of the vibration signal at different frequencies. The construction of the frequency domain motion vector field makes the analysis of vibration characteristics more in-depth and detailed, and can effectively distinguish different types of vibration modes. Finally, the vibration image data, including vibration frequency data and amplitude data, is extracted according to the frequency domain motion vector field. Through frequency domain analysis, the frequency and amplitude of the vibration signal are calculated, thereby reflecting the vibration intensity and frequency characteristics of the target object in the video. The accurate extraction of vibration image data provides reliable data support for the input of the subsequent physical conflict detection model, ensuring the accuracy and reliability of the detection results.
[0016] Optionally, the extracting vibration image data from the image sequence to be detected by using a preset vibration image feature extraction model includes:
[0017] Determining the motion vector of each pixel point according to a first image to be detected and a second image to be detected in the sequence of images to be detected to form a time domain motion vector field, wherein the first image to be detected and the second image to be detected are any adjacent images to be detected in the sequence of images to be detected;
[0018] Determine a corresponding frequency domain motion vector field according to the time domain motion vector field;
[0019] The vibration image data is extracted according to the frequency domain motion vector field.
[0020] In the above scheme, the motion vector of each pixel is determined according to any two adjacent frames of the image sequence to be detected (i.e., the first image to be detected and the second image to be detected), and then the time domain motion vector field is constructed. This step calculates the motion direction and speed of each pixel by comparing the position changes of the pixels between adjacent frames, so as to accurately capture the tiny movements of the target object. This analysis method based on adjacent frames ensures that the construction of the time domain motion vector field is both comprehensive and accurate, laying a solid foundation for the subsequent frequency domain analysis. Then, the time domain motion vector field is efficiently converted into the frequency domain motion vector field. This process reveals the distribution law of the vibration signal at different frequencies through frequency domain analysis, making the identification of vibration characteristics more in-depth and detailed. The construction of the frequency domain motion vector field not only helps to distinguish different types of vibration modes, but also provides rich frequency information for the subsequent extraction of vibration image data, enhancing the depth and breadth of the analysis. Finally, based on the frequency domain motion vector field, the vibration image data, including vibration frequency data and amplitude data, is extracted to accurately reflect the vibration intensity and frequency characteristics of the target object in the video by calculating the frequency component and amplitude component in the frequency domain signal. The accurate extraction of vibration image data provides direct and effective input data for the subsequent physical conflict detection model, ensuring the accuracy and reliability of the detection results. At the same time, since the vibration image data is directly related to the actual vibration of the target object, it can more realistically reflect the dynamic characteristics of potential physical conflicts.
[0021] Optionally, the extracting bone feature data from the image sequence to be detected by using a preset bone feature extraction model includes:
[0022] Determine a corresponding feature heat map according to each frame of the image to be detected in the sequence of images to be detected, so as to form a feature heat map sequence;
[0023] Determine the position of the skeleton key point according to the feature heat map sequence;
[0024] Determine the bone vector according to the position of each bone key point;
[0025] Determine the bone length according to the bone vector;
[0026] Determine the bone angle between adjacent bone vectors according to each bone vector;
[0027] The bone feature data is determined according to the positions of the bone key points, the bone vectors, the bone lengths and the bone angles.
[0028] In the above scheme, the corresponding feature heat map is determined according to each frame of the image sequence to be detected, so as to construct a feature heat map sequence. The heat map value on the feature heat map reflects the probability that the corresponding bone key point appears at the position. This probability distribution representation method effectively enhances the accuracy and robustness of key point detection. Through the feature heat map sequence, the position changes of the bone key points of the target object in different frames can be determined, which provides intuitive and rich information for subsequent analysis. Then, the position of the bone key points is determined according to the feature heat map sequence. This process realizes the accurate extraction of the key point position by calculating the extreme value or weighted average of the heat map value. Accurate positioning of the bone key points not only provides a reliable basis for the subsequent steps, but also significantly improves the accuracy and efficiency of the overall algorithm. After determining the position of the bone key points, the bone vectors between each key point are further calculated. The bone vector reflects the connection relationship and directionality between adjacent key points and is an important parameter for describing the bone structure. By scientifically calculating the bone vector, the bone morphology and motion characteristics of the target object can be more accurately described. Then, the bone length is determined according to the bone vector, that is, the straight-line distance between adjacent key points is calculated. This step achieves accurate acquisition of bone length through measurement and calculation. The accurate measurement of bone length provides key data support for subsequent analysis, which helps to more accurately evaluate the body size and range of motion of the target object. In addition to bone length, the bone angle between adjacent bone vectors is also determined based on each bone vector. The bone angle reflects the relative position and angle relationship between bones and is an important indicator for describing the complexity of bone structure. By accurately calculating the bone angle, we can have a deeper understanding of the bone structure and movement pattern of the target object. Finally, based on parameters such as the position of bone key points, bone vectors, bone length, and bone angle, the bone feature data is comprehensively extracted. These data not only contain the bone structure information of the target object, but also reflect its movement characteristics and dynamic changes. By comprehensively extracting bone feature data, comprehensive and effective input information is provided for the training and prediction of subsequent limb conflict detection models.
[0029] Optionally, determining a corresponding feature heat map according to each frame of the image to be detected in the sequence of images to be detected includes:
[0030] Using a preset key point detection model, each frame of the image to be detected in the image sequence to be detected is mapped to the corresponding feature heat map. The preset key point detection model is a model based on a convolutional neural network, and the convolutional neural network includes a convolutional layer, a pooling layer, an activation function and a fully connected layer connected in sequence.
[0031] In the above scheme, the convolution layer, as the first processing unit of the convolutional neural network, slides the convolution kernel on the input image and performs dot product operations to effectively extract local features in the image. These local features contain basic information such as the edge and texture of the image, providing rich data support for subsequent processing. In the key point detection task, the convolution layer can capture the local image pattern related to the skeleton key points, laying the foundation for the subsequent key point positioning. The pooling layer follows the convolution layer and achieves feature dimensionality reduction by downsampling the feature map output by the convolution layer. This step not only reduces the amount of calculation, but also enhances the robustness of the model to image translation, rotation and other transformations. In key point detection, the pooling layer helps the model better cope with the slight position changes of the target object in the video frame and improves the stability of key point detection. Then, the activation function is introduced into the convolutional neural network to introduce nonlinear factors, so that the model can learn more complex feature representations. In the key point detection task, the activation function enables the model to more accurately capture the feature information related to the skeleton key points by enhancing the expression ability of the feature map. Next, the fully connected layer, as the output layer of the convolutional neural network, integrates the features extracted by the previous layers and outputs the final key point detection results. In the fully connected layer, through the operation of the weight matrix and the bias term, the model can map the feature map to the key point coordinate space to achieve accurate positioning of the key points. This step makes full use of the feature information extracted by the previous layers to ensure the accuracy and reliability of key point detection. Finally, through the processing flow of the above convolutional neural network, each frame of the image to be detected is mapped to the corresponding feature heat map. The heat map value on the feature heat map intuitively reflects the probability of the corresponding skeletal key point appearing at that position, without the need for manual intervention, which significantly improves the efficiency of key point detection.
[0032] Optionally, the using a preset physical conflict detection model and detecting the physical conflict risk level of the target object to be detected according to the vibration image data and the bone feature data includes:
[0033] Determine a vibration characteristic value according to the vibration frequency data and the amplitude data in the vibration image data;
[0034] Determine a bone structure vector diagram according to the bone key point positions, the bone vectors, the bone lengths and the bone angles in the bone feature data;
[0035] Determine the bone dynamic feature value according to the bone structure vector diagram corresponding to each frame of the image to be detected;
[0036] Determine a physical conflict risk characteristic value according to the vibration characteristic value and the bone dynamic characteristic value;
[0037] The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
[0038] In the above scheme, the vibration characteristic value is determined according to the vibration frequency data and amplitude data in the vibration image data, and the dynamic behavior characteristics of the target object in the video are reflected by comprehensively considering the vibration frequency and amplitude of the target object. The accurate calculation of the vibration characteristic value provides key data support for subsequent analysis, which helps to more accurately evaluate the possibility of physical conflict. Then, the bone structure vector map is constructed according to the information such as the position of the key points of the bones, the bone vector, the bone length and the bone angle in the bone feature data. The bone structure vector map intuitively displays the bone structure and its dynamic changes of the target object, providing intuitive visual support for subsequent analysis. By constructing the bone structure vector map, the method can more deeply understand the body posture and movement pattern of the target object, providing an important basis for the assessment of the risk of physical conflict. After constructing the bone structure vector map, the bone dynamic characteristic value is further determined according to the bone structure vector map corresponding to each frame of the image to be detected. The bone dynamic characteristic value reflects the changes in the bone structure of the target object in continuous video frames, which is an important parameter for assessing the risk of physical conflict. By extracting the bone dynamic characteristic value, the method can capture the subtle changes of the target object during the movement, so as to more accurately judge whether it has the risk of physical conflict. Subsequently, the physical conflict risk characteristic value is determined based on the vibration characteristic value and the bone dynamic characteristic value. This step achieves a comprehensive assessment of the physical conflict risk by comprehensively considering the vibration behavior and bone structure changes of the target object. The calculation process of the physical conflict risk characteristic value takes into account both the dynamic behavior characteristics of the target object and the changes in its body posture and movement pattern, thereby ensuring the accuracy and reliability of the evaluation results. Finally, the physical conflict risk level of the target object to be detected is determined based on the physical conflict risk characteristic value and the preset risk characteristic level mapping table. The preset risk characteristic level mapping table establishes a corresponding relationship between the physical conflict risk characteristic value and the risk level through prior training and verification. Through the table lookup operation, the physical conflict risk level of the target object can be quickly and accurately determined, providing a strong basis for subsequent early warning and intervention.
[0039] Optionally, the using a preset physical conflict detection model and detecting the physical conflict risk level of the target object to be detected according to the vibration image data and the bone feature data includes:
[0040] fusing the vibration image data and the bone feature data to form a fused limb feature vector;
[0041] Outputting a physical conflict risk feature value according to the fused physical feature vector;
[0042] The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
[0043] In the above scheme, a fused limb feature vector is formed by fusing vibration image data with bone feature data. This process not only considers the dynamic vibration characteristics of the target object (such as frequency and amplitude), but also comprehensively considers its bone structure characteristics (such as the position of key bone points, bone vectors, bone length and bone angle). This comprehensive data fusion method enables the fused limb feature vector to more comprehensively reflect the physical state and action pattern of the target object, providing a richer and more accurate information basis for subsequent risk assessment. On the basis of data fusion, the physical conflict risk feature value is further extracted from the fused limb feature vector. This process uses a specific algorithm or model to efficiently process the fused feature vector and extract key features directly related to the physical conflict risk. This feature extraction method not only improves the processing efficiency, but also ensures that the extracted features are highly representative and discriminative, providing strong support for subsequent risk level determination. According to the physical conflict risk feature value and the preset risk feature level mapping table, the physical conflict risk level of the target object to be detected is determined. The preset risk feature level mapping table is obtained through a large amount of training data and actual case analysis, and an accurate correspondence between the physical conflict risk feature value and the risk level is established. Therefore, by looking up the table, the physical conflict risk level of the target object can be quickly and accurately determined, providing a scientific basis for subsequent early warning and intervention measures.
[0044] In addition, since the entire risk assessment process is based on video vibration image analysis, it is highly real-time. Once the target object's abnormal behavior or potential physical conflict risk appears in the video, data processing and analysis can be performed immediately, and the risk level assessment results can be quickly output. This real-time early warning mechanism helps to take timely intervention measures to prevent the occurrence or escalation of physical conflict incidents, thereby ensuring public safety.
[0045] In a second aspect, the present application provides a physical conflict detection and early warning device based on video vibration image analysis, comprising:
[0046] An acquisition module is used to acquire a video file to be detected and decompose the video file to be detected into a sequence of images to be detected, wherein the number of frames of the video file to be detected is greater than a preset frame number threshold, and each image to be detected in the sequence of images to be detected includes a target object to be detected;
[0047] An extraction module, used to extract vibration image data from the image sequence to be detected by using a preset vibration image feature extraction model, wherein the vibration image data includes vibration frequency data and amplitude data;
[0048] The extraction module is used to extract bone feature data from the image sequence to be detected by using a preset bone feature extraction model, wherein the bone feature data includes bone key point position data and bone posture data;
[0049] The processing module is used to use a preset physical conflict detection model to detect the physical conflict risk level of the target object to be detected according to the vibration image data and the bone feature data.
[0050] Optionally, the extraction module is specifically used to:
[0051] Determine the motion vector of each pixel point according to the image sequence to be detected to form a time domain motion vector field;
[0052] Determine a corresponding frequency domain motion vector field according to the time domain motion vector field;
[0053] The vibration image data is extracted according to the frequency domain motion vector field.
[0054] Optionally, the extraction module is specifically used to:
[0055] Determining the motion vector of each pixel point according to a first image to be detected and a second image to be detected in the sequence of images to be detected to form a time domain motion vector field, wherein the first image to be detected and the second image to be detected are any adjacent images to be detected in the sequence of images to be detected;
[0056] Determine a corresponding frequency domain motion vector field according to the time domain motion vector field;
[0057] The vibration image data is extracted according to the frequency domain motion vector field.
[0058] Optionally, the extraction module is specifically used to:
[0059] Determine a corresponding feature heat map according to each frame of the image to be detected in the sequence of images to be detected, so as to form a feature heat map sequence;
[0060] Determine the position of the skeleton key point according to the feature heat map sequence;
[0061] Determine the bone vector according to the position of each bone key point;
[0062] Determine the bone length according to the bone vector;
[0063] Determine the bone angle between adjacent bone vectors according to each bone vector;
[0064] The bone feature data is determined according to the positions of the bone key points, the bone vectors, the bone lengths and the bone angles.
[0065] Optionally, the extraction module is specifically used to:
[0066] Using a preset key point detection model, each frame of the image to be detected in the image sequence to be detected is mapped to the corresponding feature heat map. The preset key point detection model is a model based on a convolutional neural network, and the convolutional neural network includes a convolutional layer, a pooling layer, an activation function and a fully connected layer connected in sequence.
[0067] Optionally, the processing module is specifically used to:
[0068] Determine a vibration characteristic value according to the vibration frequency data and the amplitude data in the vibration image data;
[0069] Determine a bone structure vector diagram according to the bone key point positions, the bone vectors, the bone lengths and the bone angles in the bone feature data;
[0070] Determine the bone dynamic feature value according to the bone structure vector diagram corresponding to each frame of the image to be detected;
[0071] Determine a physical conflict risk characteristic value according to the vibration characteristic value and the bone dynamic characteristic value;
[0072] The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
[0073] Optionally, the processing module is specifically used to:
[0074] fusing the vibration image data and the bone feature data to form a fused limb feature vector;
[0075] Outputting a physical conflict risk feature value according to the fused physical feature vector;
[0076] The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
[0077] In a third aspect, the present application provides an electronic device, including:
[0078] processor; and,
[0079] A memory, configured to store executable instructions of the processor;
[0080] The processor is configured to perform any possible method described in the first aspect by executing the executable instructions.
[0081] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement any possible method described in the first aspect.
[0082] The physical conflict detection and early warning method based on video vibration image analysis provided by the present application obtains a video file to be detected and decomposes the video file to be detected into an image sequence to be detected, then uses a preset vibration image feature extraction model to extract vibration image data from the image sequence to be detected, uses a preset bone feature extraction model to extract bone feature data from the image sequence to be detected, uses a preset physical conflict detection model, and detects the physical conflict risk level of the target object to be detected based on the vibration image data and the bone feature data, thereby detecting and predicting physical conflicts to identify potential physical conflict incidents, providing strong support for security monitoring and early warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0084] Figure 1 is a flow chart of a method for detecting and warning physical conflict based on video vibration image analysis according to an exemplary embodiment of the present application;
[0085] Figure 2 is a flow chart of a method for detecting and warning physical conflict based on video vibration image analysis according to another exemplary embodiment of the present application;
[0086] Figure 3 is a structural schematic diagram of a physical conflict detection and early warning device based on video vibration image analysis according to an exemplary embodiment of the present application;
[0087] Figure 4 It is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present application.
[0088] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0089] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0090] Figure 1 FIG. 1 is a flow chart of a method for detecting and warning physical conflicts based on video vibration image analysis according to an exemplary embodiment of the present application. Figure 1 As shown, the method provided in this embodiment includes:
[0091] S101, obtaining a video file to be detected, and decomposing the video file to be detected into a sequence of images to be detected.
[0092] In this step, the video file to be detected may be obtained and decomposed into a sequence of images to be detected, wherein the number of frames of the video file to be detected is greater than a preset frame number threshold, and each image to be detected in the sequence of images to be detected includes a target object to be detected.
[0093] Specifically, the above video files usually come from security monitoring systems, cameras in public places or other video recording devices. In order to ensure the accuracy and effectiveness of subsequent processing, the frame number of the acquired video file to be detected must be greater than the preset frame number threshold, which is usually set according to actual needs and technical conditions to ensure that the video file contains enough information for analysis and to ensure the continuity of the movement changes of the target object to be detected.
[0094] After acquiring the video file, the system decomposes it into a series of continuous image frames to form a sequence of images to be detected. Each image frame to be detected contains the target object to be detected, which may be a person, an animal or other entity to be monitored. The decomposition process is usually implemented through video processing software or algorithms to ensure the clarity and integrity of each image frame, laying the foundation for subsequent feature extraction and analysis.
[0095] S102: Extract vibration image data from the image sequence to be detected using a preset vibration image feature extraction model.
[0096] In this step, a preset vibration image feature extraction model is used to extract vibration image data from the image sequence to be detected, wherein the vibration image data includes vibration frequency data and amplitude data.
[0097] Based on the image sequence to be detected, the system uses a preset vibration image feature extraction model to extract vibration image data from the image. This model is built based on deep learning or image processing technology and can accurately identify and analyze vibration information in the image.
[0098] Vibration image data mainly includes frequency data and amplitude data. Frequency data reflects the frequency of the target object's vibration, that is, the number of vibrations per unit time; amplitude data reflects the intensity or amplitude of the vibration. In order to extract this data, the model first pre-processes the image sequence, such as denoising, contrast enhancement, etc., to improve the accuracy of the analysis. Then, by applying a specific model, the motion vector of each pixel can be calculated, and the time domain and frequency domain motion vector fields can be formed. Finally, the frequency and amplitude data are extracted from these vector fields to provide key information for subsequent physical conflict detection.
[0099] In a possible specific implementation, Formula 1 may be used to determine the motion vector of each pixel point according to the image sequence to be detected, so as to form a time domain motion vector field E(x, y, t), where Formula 1 is:
[0100] Formula 1 is:
[0101]
[0102] Among them, x is the horizontal coordinate of the pixel, y is the vertical coordinate of the pixel, and t is the time;
[0103] Using formula 2, and determining the corresponding frequency domain motion vector field E(u,v,ω) according to the time domain motion vector field E(x,y,t), formula 2 is:
[0104]
[0105] Wherein, u is the first spatial frequency, which is used to characterize the rate of change of the pixel point in the horizontal direction, v is the second spatial frequency, which is used to characterize the rate of change of the pixel point in the vertical direction, ω is the time frequency, and i is the imaginary unit;
[0106] The vibration image data is extracted according to the frequency domain motion vector field E(u,v,ω) using Formula 3, where Formula 3 is:
[0107]
[0108] Among them, f(u,v,ω) is the frequency data, and A(u,v,ω) is the amplitude data.
[0109] In the above scheme, using formula 1, the motion vector in the time dimension can be calculated according to each pixel in the image sequence to be detected, and then the time domain motion vector field can be constructed. This step accurately captures the dynamic changes of each pixel in the image by calculating the change rate of the pixel in the horizontal and vertical directions over time. The construction of the time domain motion vector field provides a basis for subsequent frequency domain analysis, ensuring the accuracy and reliability of vibration image data extraction.
[0110] Through Formula 2, the system converts the time domain motion vector field into the frequency domain motion vector field. This conversion process maps the information in the time domain and space domain to the frequency domain, thereby revealing the characteristics of the vibration signal in the frequency domain. The first spatial frequency and the second spatial frequency represent the rate of change of the pixel point in the horizontal and vertical directions, respectively, while the time frequency reflects the time periodicity of the vibration signal. The construction of the frequency domain motion vector field enables the system to analyze the vibration signal more deeply and extract key features such as frequency and amplitude.
[0111] Using Formula 3, the system directly extracts frequency data and amplitude data from the frequency domain motion vector field. Frequency data reflects the frequency characteristics of the vibration signal, that is, the speed of the vibration; while amplitude data reflects the intensity or magnitude of the vibration signal. Through this process, the system can efficiently extract concise and effective vibration image data from complex image sequences, providing key information for subsequent physical conflict detection.
[0112] The vibration image data extracted through the above steps provides rich input information for the physical conflict detection model. The frequency and amplitude characteristics in the vibration image data can reflect the dynamic behavior pattern of the target object. Combined with the bone feature data, it can more accurately determine whether the target object is at risk of physical conflict. In addition, since the extraction process of vibration image data is based on efficient mathematical algorithms and computing technology, it can ensure the real-time nature of physical conflict detection, allowing the system to respond to potential conflict events in the first place.
[0113] In addition, the extraction of vibration image data through the above algorithm avoids the problems of reliance on empirical parameters and manual intervention that may exist in traditional methods. This enables the system to maintain high robustness and adaptability when facing different scenarios and complex environments. At the same time, due to the high degree of automation and intelligence in the vibration image data extraction process, the system's dependence on manual operation and maintenance is reduced, and the overall performance and reliability of the system are improved.
[0114] In summary, the use of a preset vibration image feature extraction model to extract vibration image data from the image sequence to be detected plays a key role in the physical conflict detection and early warning method based on video vibration image analysis. By accurately constructing the motion vector field, realizing frequency domain conversion and feature extraction, and efficiently extracting vibration image data, the accuracy and real-time performance of physical conflict detection are significantly improved, and the robustness and adaptability of the system are enhanced.
[0115] S103, using a preset bone feature extraction model to extract bone feature data from the image sequence to be detected.
[0116] In this step, a preset bone feature extraction model is used to extract bone feature data from the image sequence to be detected, wherein the bone feature data includes bone key point position data and bone posture data.
[0117] Specifically, in addition to the vibration image data, the system also uses a preset bone feature extraction model to extract bone feature data from the image sequence to be detected. This model is also based on deep learning or image processing technology, and can accurately identify the key point positions of human or animal bones in the image and the bone posture (i.e. the direction and angle of the bones in space).
[0118] Skeleton feature data mainly includes skeleton key point position data and skeleton posture data. Skeleton key point position data identifies the position of key nodes on human or animal skeletons, such as joints, endpoints, etc.; skeleton posture data describes the relative position and angle relationship between these key points, reflecting the overall shape and motion state of the skeleton. In order to extract this data, the model usually performs feature extraction and key point detection on the image, and constructs a skeleton structure vector map by calculating the bone vector and determining the bone angle. This process not only provides the static features of the skeleton, but also reveals the dynamic changes of the skeleton through the analysis of continuous frames.
[0119] S104: Using a preset physical conflict detection model, and based on the vibration image data and the bone feature data, detecting the physical conflict risk level of the target object to be detected.
[0120] Specifically, after obtaining the vibration image data and bone feature data, the preset physical conflict detection model is used to detect the physical conflict risk level of the target object to be detected based on these data. This model is built based on machine learning or deep learning technology, which can learn and understand the complex relationship between vibration image and bone feature data, and judge whether the target object has the risk of physical conflict based on this.
[0121] The detection process usually includes the following steps: first, determine the vibration characteristic value based on the frequency and amplitude data in the vibration image data; second, determine the bone structure vector map based on the key point position, bone vector, bone length and bone angle in the bone feature data, and further extract the bone dynamic characteristic value; then, determine the physical conflict risk characteristic value based on the vibration characteristic value and the bone dynamic characteristic value; finally, determine the physical conflict risk level of the target object based on the physical conflict risk characteristic value and the preset risk characteristic level mapping table. This process not only takes into account the dynamic vibration characteristics of the target object, but also comprehensively considers the changes in its bone structure and movement pattern, thereby achieving an accurate assessment of the risk of physical conflict.
[0122] In a possible implementation, the vibration image data and the bone feature data may be fused to form a fused limb feature vector F;
[0123] The physical conflict risk characteristic value R is output according to the fused limb characteristic vector, wherein the physical conflict risk characteristic value R can be determined according to the following formula:
[0124]
[0125] Among them, F i is the i-th feature in the fused limb feature vector F, f i is the i-th feature extraction function, ω i is the weight value of the i-th feature in the weight matrix; K is the number of features in the fused limb feature vector;
[0126] The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
[0127] In the above scheme, the vibration image data is first fused with the bone feature data to form a fused limb feature vector. This step effectively combines the information about the dynamic behavior of the target object in the vibration image data (such as vibration frequency and amplitude) and the information about the body structure of the target object in the bone feature data (such as the position of key bone points and bone posture). By fusing these two different dimensions of data, the system can more comprehensively understand the limb movement state of the target object, providing a richer and more accurate information basis for subsequent risk level detection.
[0128] Moreover, when calculating the physical conflict risk feature value using the above formula, the system considers the importance of each feature in the fused physical feature vector, and assigns a corresponding weight value to each feature through the weight matrix. This feature weighting mechanism enables the system to pay more attention to those features that have an important impact on the judgment of physical conflict risk, thereby improving the accuracy of risk level detection. At the same time, feature weighting also enables the system to adapt to the needs of different scenarios and optimize the detection effect by adjusting the weight value. Then, the linearly weighted feature value is mapped to the set interval to obtain the physical conflict risk feature value. This nonlinear mapping mechanism enhances the expressiveness of the model and enables the system to handle complex physical conflict detection tasks more flexibly. In addition, based on the calculated physical conflict risk feature value, the system determines the physical conflict risk level of the target object to be detected by comparing it with the preset risk feature level mapping table. This process realizes the quantitative assessment of physical conflict risk, enables the system to display the detection results more intuitively, and provides strong support for subsequent safety management and emergency response. At the same time, risk level mapping also enables the system to flexibly set risk level thresholds and corresponding countermeasures according to the safety requirements in different scenarios.
[0129] In summary, the specific technical effects of physical conflict risk level detection in the physical conflict detection and early warning method based on video vibration image analysis are reflected in the features of improving information comprehensiveness through feature fusion, enhancing detection accuracy through feature weighting, enhancing model expression ability through nonlinear mapping, achieving quantitative evaluation through risk level mapping, and improving system real-time and intelligent level. These technical effects work together in the physical conflict detection and early warning process, enabling the system to more accurately judge the physical conflict risk level of the target object, providing strong support for public safety monitoring and security management.
[0130] In this embodiment, by acquiring a video file to be detected and decomposing the video file to be detected into an image sequence to be detected, then using a preset vibration image feature extraction model to extract vibration image data from the image sequence to be detected, using a preset bone feature extraction model to extract bone feature data from the image sequence to be detected, using a preset physical conflict detection model, and detecting the physical conflict risk level of the target object to be detected based on the vibration image data and the bone feature data, physical conflicts are detected and predicted to identify potential physical conflict incidents, thereby providing strong support for security monitoring and early warning.
[0131] Figure 2 FIG. 1 is a flow chart of a method for detecting and warning physical conflicts based on video vibration image analysis according to another exemplary embodiment of the present application. Figure 2 As shown, the physical conflict detection and early warning method based on video vibration image analysis provided in this embodiment includes:
[0132] S201, obtaining a video file to be detected, and decomposing the video file to be detected into a sequence of images to be detected.
[0133] In this step, the video file to be detected may be obtained and decomposed into a sequence of images to be detected, wherein the number of frames of the video file to be detected is greater than a preset frame number threshold, and each image to be detected in the sequence of images to be detected includes a target object to be detected.
[0134] Specifically, the above video files usually come from security monitoring systems, cameras in public places or other video recording devices. In order to ensure the accuracy and effectiveness of subsequent processing, the frame number of the acquired video file to be detected must be greater than the preset frame number threshold, which is usually set according to actual needs and technical conditions to ensure that the video file contains enough information for analysis and to ensure the continuity of the movement changes of the target object to be detected.
[0135] After acquiring the video file, the system decomposes it into a series of continuous image frames to form a sequence of images to be detected. Each image frame to be detected contains the target object to be detected, which may be a person, an animal or other entity to be monitored. The decomposition process is usually implemented through video processing software or algorithms to ensure the clarity and integrity of each image frame, laying the foundation for subsequent feature extraction and analysis.
[0136] S202 : determining a motion vector of each pixel point according to a first image to be detected and a second image to be detected in a sequence of images to be detected, so as to form a time domain motion vector field.
[0137] Specifically, using Formula 4, the motion vector of each pixel is determined according to the first image to be detected and the second image to be detected in the image sequence to form a time domain motion vector field E(x, y, t), and the first image to be detected and the second image to be detected are any adjacent images to be detected in the image sequence to be detected, wherein Formula 4 is:
[0138]
[0139] Where I(x, y, t) is the pixel value of the pixel point in the first image to be detected at time t, I(x ′ ,y ′ ,t+τ) is the pixel value of the corresponding pixel point in the second image to be detected at time t+τ, and τ is the time difference between the first image to be detected and the second image to be detected.
[0140] Using Formula 4, the system can calculate the motion vector of each pixel based on any two adjacent frames in the image sequence to be detected (the first image to be detected and the second image to be detected), thereby constructing a time domain motion vector field. This step accurately reflects the dynamic changes of pixels in the image sequence by calculating the difference in pixel values of corresponding pixels in the two frames. The introduction of time difference ensures that the motion changes between adjacent frames are calculated, providing accurate basic data for subsequent frequency domain analysis.
[0141] S203: Determine a corresponding frequency domain motion vector field according to the time domain motion vector field.
[0142] Specifically, using Formula 5, and determining the corresponding frequency domain motion vector field E(u,v,ω) according to the time domain motion vector field E(x,y,t), Formula 5 is:
[0143]
[0144] Among them, M is the number of pixels in the horizontal direction, N is the number of pixels in the vertical direction, u is the first spatial frequency, which is used to characterize the change rate of the pixel point in the horizontal direction, v is the second spatial frequency, which is used to characterize the change rate of the pixel point in the vertical direction, ω is the time frequency, and i is the imaginary unit.
[0145] Through formula 5, the system converts the time domain motion vector field into the frequency domain motion vector field. This process maps the information in the space domain and time domain to the frequency domain, revealing the characteristics of the vibration signal in the frequency domain. The introduction of the number of pixels in the horizontal and vertical directions ensures the accuracy of the transformation, while the first spatial frequency, the second spatial frequency and the time frequency characterize the rate of change of the vibration signal in the horizontal, vertical and time dimensions respectively. The construction of the frequency domain motion vector field facilitates the subsequent extraction of vibration frequency and amplitude data.
[0146] S204: Extract vibration image data according to the frequency domain motion vector field.
[0147] Specifically, the vibration image data is extracted using Formula 6 according to the frequency domain motion vector field E(u,v,ω), where Formula 6 is:
[0148]
[0149] Among them, f(u,v,ω) is the frequency data, A(u,v,ω) is the amplitude data, Re[E(u,v,ω)] is the real part of E(u,v,ω), and Im[E(u,v,ω)] is the imaginary part of E(u,v,ω).
[0150] Using Formula 6, the system directly extracts the frequency data and amplitude data from the frequency domain motion vector field. The frequency data reflects the frequency characteristics of the vibration signal, that is, the speed of the vibration; the amplitude data reflects the intensity or amplitude of the vibration signal. This process calculates the real and imaginary parts of the frequency domain motion vector field and directly obtains the frequency and amplitude by combining the formula, thus achieving efficient extraction of vibration image data.
[0151] The vibration image data extracted through S202-S204 provides key information for subsequent physical conflict detection. The frequency and amplitude characteristics in the vibration image data can reflect the dynamic behavior pattern of the target object, especially in the physical conflict scene, the changes of these characteristics are often more obvious. Therefore, using these features for physical conflict detection can significantly improve the accuracy of detection and reduce false positives and false negatives.
[0152] In summary, the use of a preset vibration image feature extraction model to extract vibration image data from the image sequence to be detected plays a key role in the physical conflict detection and early warning method based on video vibration image analysis. By accurately constructing the time domain motion vector field, efficiently converting it to the frequency domain motion vector field, and directly extracting vibration image data, the accuracy and real-time performance of physical conflict detection are significantly improved, and the robustness of the system is enhanced.
[0153] S205 . Determine a corresponding feature heat map according to each frame of the image to be detected in the image sequence to form a feature heat map sequence.
[0154] Determine the corresponding feature heat map H according to each frame of the image to be detected in the image sequence to be detected t , to form a feature heat map sequence, where the feature heat map H t The first dimension is used to represent the number of channels, each channel corresponds to a skeleton key point, the second dimension is used to represent the height, and the third dimension is used to represent the width. t The heat map value on is used to characterize the probability of the corresponding bone key point appearing at the corresponding position. t The generation of can be based on mature heat map regression technology, which converts the key point positions in the image into heat maps and then lets the neural network learn how to predict these heat maps. A heat map is a two-dimensional array in which each element represents the probability or confidence of the existence of a key point at the corresponding position. By regressing these heat maps, the network can learn the spatial distribution of key points, thereby achieving high-precision key point detection and posture estimation.
[0155] In addition, in another possible implementation, a preset key point detection model may be used to map each frame of the image to be detected in the image sequence to a corresponding feature heat map H t, the preset key point detection model is a model based on a convolutional neural network. The convolutional neural network includes a convolutional layer, a pooling layer, an activation function, and a fully connected layer connected in sequence. The convolutional layer performs a convolution operation using formula 10, which is:
[0156]
[0157] in, is the output value at position (i, j) in the feature map of the mth layer, is the weight of the m-th convolution kernel at position (p,q), is the input value at position (i+p,j+q) in the feature map of the m-1th layer, P is the height of the convolution kernel, Q is the width of the convolution kernel, and b m is the bias term of the mth layer;
[0158] The pooling layer uses formula 11 for pooling operation, which is:
[0159]
[0160] Among them, R ij is the area covered by the pooling window at position (i, j), is the input value at position (p,q) in the feature map of the m-1th layer;
[0161] The activation function includes formula 12, which is:
[0162]
[0163] in, is the input value at position (i, j) in the feature map of the m-1th layer;
[0164] The fully connected layer includes formula 13, which is:
[0165]
[0166] Among them, N m-1 is the number of neurons in the m-1th layer, is the weight of the mth fully connected layer, is the bias term of the mth fully connected layer.
[0167] Through the above steps, in the process of mapping each frame of the image to be detected in the sequence of images to be detected to the corresponding feature heat map using the preset key point detection model, the convolution layer, pooling layer, activation function and fully connected layer each play an important role. The convolution layer efficiently extracts local features, the pooling layer reduces the dimension and enhances robustness, the activation function introduces nonlinear mapping and feature selection mechanism, and the fully connected layer integrates features and makes decisions. The collaborative work of these components enables the model to accurately and efficiently generate feature heat maps, providing strong support for subsequent key point positioning and physical conflict detection and early warning.
[0168] S206: Determine the positions of the key points of the bones, the bone vectors, the bone lengths and the bone angles according to the feature heat map sequence, and then determine the bone feature data.
[0169] Specifically, using formula 7, the positions of the key points of the skeleton are determined according to the feature heat map sequence. Wherein, Formula 7 is:
[0170]
[0171] in, is the feature heat map H corresponding to the t-th frame of the image to be detected t The heat map value of the kth key point in, is the integral of the entire feature heat map area corresponding to the t-th frame of the image to be detected;
[0172] Determine the bone vector based on the position of each bone key point in, is the bone vector between the i-th bone key point and the j-th bone key point;
[0173] Using formula 8, and according to the bone vector Determining bone length Wherein, formula 8 is:
[0174]
[0175] Using formula 9, determine the bone angle between adjacent bone vectors based on each bone vector Wherein, formula 9 is:
[0176]
[0177] in, is the first bone vector, is the second bone vector, the first bone vector and the second bone vector are vectors corresponding to adjacent bones;
[0178] The bone feature data is determined based on the positions of the bone key points, the bone vectors, the bone lengths and the bone angles.
[0179] Through the above steps, the corresponding feature heat map is determined according to each frame of the image to be detected in the image sequence to be detected. The first dimension of the feature heat map represents the number of channels, and each channel corresponds to a skeleton key point; the second dimension and the third dimension represent the height and width respectively. The heat map value is used to represent the probability of the corresponding skeleton key point appearing at the corresponding position. By generating a feature heat map, the position probability distribution of each skeleton key point in the image can be accurately represented, which provides a solid foundation for the subsequent key point positioning. The multi-dimensional representation method of the feature heat map ensures the integrity and accuracy of the information and improves the accuracy and robustness of the key point positioning. Formula 7 is used to determine the position of the skeleton key point according to the feature heat map sequence. This formula finds the position with the largest heat map value, that is, the key point position, by integrating and normalizing the feature heat map. Through Formula 7, the position of each skeleton key point can be accurately calculated, avoiding the positioning error caused by noise or occlusion in the traditional method. This method has high accuracy and stability and is suitable for skeleton key point positioning in various complex scenes.
[0180] According to the position of each bone key point, the bone vector is determined. The bone vector represents the relative position relationship between two key points, and the bone length reflects the actual physical size of the bone. By calculating the bone vector and bone length, the geometric structure and spatial relationship of the human skeleton can be accurately described. This information is crucial for subsequent limb conflict detection because it can reflect the movement state and deformation of the skeleton.
[0181] Formula 9 is used to determine the bone angle based on adjacent bone vectors. This formula calculates the dot product and modulus of two vectors to obtain the angle between them. The bone angle is an important parameter that reflects the posture and movement of the human body. By calculating the bone angle, the relative movement trend and change law between bones can be captured, providing richer feature information for physical conflict detection.
[0182] The bone feature data is determined based on the position of the key points of the bones, the bone vector, the bone length and the bone angle. These data comprehensively reflect the geometric structure, spatial relationship and motion state of the human skeleton. By comprehensively applying these bone feature data, a comprehensive description and analysis of human posture and movement can be achieved. In the physical conflict detection and early warning, these data can provide more accurate and reliable input information for the algorithm, improving the accuracy and real-time performance of the detection. At the same time, this method also has strong generalization ability and can adapt to the needs of human posture detection in different scenes and environments.
[0183] S207 : Determine the vibration characteristic value according to the vibration frequency data and the amplitude data in the vibration image data.
[0184] Using formula 14, the vibration characteristic value F is determined based on the vibration frequency data and amplitude data in the vibration image data. v , where Formula 14 is:
[0185]
[0186] Among them, G is the number of frames of the video file to be detected, H is the number of skeleton key points, To calibrate the vibration frequency, is the calibrated amplitude, f is the frequency in the frequency data, A is the amplitude in the amplitude data, α is the preset frequency weight value, and β is the preset amplitude weight value.
[0187] S208, determining a skeleton structure vector diagram according to the skeleton key point positions, skeleton vectors, skeleton lengths, and skeleton angles in the skeleton feature data.
[0188] The bone structure vector diagram is determined based on the bone key point positions, bone vectors, bone lengths, and bone angles in the bone feature data.
[0189] S209: Determine the bone dynamic feature value according to the bone structure vector diagram corresponding to each frame of the image to be detected.
[0190] Using formula 15, the bone dynamic feature value F is determined according to the bone structure vector map corresponding to each frame of the image to be detected. s , where Formula 15 is:
[0191]
[0192] Among them, S i+1,j -S i,j is the displacement vector of the jth key point between the vector bone structure vector maps corresponding to the images to be detected in adjacent frames.
[0193] S210: Determine a physical conflict risk characteristic value according to the vibration characteristic value and the bone dynamic characteristic value.
[0194] Using formula 16 and according to the vibration characteristic value F v And the dynamic eigenvalue F of the skeleton s Determine the physical conflict risk characteristic value R, where formula 16 is:
[0195]
[0196] Among them, μ is the first weight value, is the second weight value, γ is the third weight value, S i,j -S i,j+1 is the distance between two adjacent key points in the vector skeleton structure vector map corresponding to the image to be detected in the same frame, and θ is the threshold parameter.
[0197] S211 . Determine the physical conflict risk level of the target object to be detected according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
[0198] In the above scheme, the vibration characteristic value is calculated using Formula 14 according to the frequency data and amplitude data in the vibration image data, combined with the calibrated frequency and calibrated amplitude, and the preset frequency weight value and amplitude weight value β. The vibration characteristic value comprehensively considers the two key vibration parameters of frequency and amplitude, and ensures the reasonable contribution of different parameters to the final result through weight adjustment. This method can more accurately reflect the actual vibration state of the target object and provide a reliable basis for the subsequent assessment of the risk level of physical conflict. By accurately calculating the vibration characteristic value, the system can more sensitively capture potential precursors to physical conflict and improve the accuracy and timeliness of early warning.
[0199] Then, according to the bone key point positions, bone vectors, bone lengths, and bone angles in the bone feature data, a bone structure vector map is constructed. In addition, the displacement vectors of the key points between the bone structure vector maps of adjacent frames are calculated using Formula 15 to obtain the bone dynamic eigenvalues. The bone structure vector map intuitively displays the spatial layout and movement trends of the bones, providing a geometric basis for analyzing physical conflicts. The bone dynamic eigenvalues effectively capture the dynamic behavior characteristics of the target object by quantifying the changes in the bone structure between adjacent frames. This method can accurately reflect the slight changes in the bone structure, which is of great significance for identifying the uncoordinated movement patterns in the early stages of physical conflicts.
[0200] Next, using formula 16, the physical conflict risk characteristic value is calculated by combining the vibration characteristic value, the bone dynamic characteristic value and the change in the distance between adjacent key points in the same frame. The formula also includes weight values and threshold parameters to ensure the comprehensiveness and accuracy of the evaluation. The physical conflict risk characteristic value is a comprehensive indicator that comprehensively considers the vibration characteristics, bone dynamic characteristics and the relative position changes between key points, and can comprehensively reflect the risk level of physical conflict of the target object. By adjusting the weight values and threshold parameters, the system can perform customized evaluations according to different scenarios and needs, thereby improving the pertinence and practicality of the early warning. In addition, this method also takes into account the mutual influence between key points in the same frame, further enhancing the accuracy and robustness of the evaluation.
[0201] In summary, by calculating vibration eigenvalues, constructing bone structure vector diagrams, extracting bone dynamic eigenvalues, and comprehensively evaluating physical conflict risk eigenvalues, a comprehensive and accurate assessment of the risk level of physical conflict is achieved. This not only improves the accuracy and timeliness of early warning, but also provides strong support for subsequent intervention measures, which is of great significance for maintaining public safety and social order.
[0202] In addition, using Formula 16, and according to the vibration characteristic value F v And the dynamic eigenvalue F of the skeleton s Before determining the physical conflict risk characteristic value R, it also includes:
[0203] Using formula 17, the first weight value μ and the second weight value are determined according to the preset training database. The third weight value γ and the threshold parameter θ, the preset training database includes physical conflict samples and non-physical conflict samples, formula 17 is:
[0204]
[0205] Where t is the number of iterations, δ is the learning rate, and L is the loss function. When the change in the loss function L is less than the preset threshold or the number of iterations reaches the preset threshold, the iteration is stopped and the first weight value μ and the second weight value are determined. The third weight value γ and the threshold parameter θ, c i is the physical conflict risk feature training value, Label the value of the physical conflict risk feature.
[0206] In the above scheme, formula 17 is used to perform iterative optimization based on a preset training database (including physical conflict samples and non-physical conflict samples) to determine the first weight value, the second weight value, the third weight value and the threshold parameter. Formula 17 gradually adjusts the weight value and the threshold parameter through the gradient descent method to minimize the loss function. This process ensures the accuracy of the weight value and the threshold parameter, so that it can better reflect the impact of vibration eigenvalues, bone dynamic eigenvalues and changes in distance between key points on the physical conflict risk eigenvalue. Through iterative optimization, the system can automatically learn and adapt to the data distribution in different scenarios, and improve the generalization ability and robustness of the evaluation.
[0207] The first weight value, the second weight value, and the third weight value respectively represent the contribution of the vibration eigenvalue, the skeletal dynamic eigenvalue, and the distance change between key points to the physical conflict risk eigenvalue; the threshold parameter is used to adjust the sensitivity of the impact of the distance change between key points on the assessment results. By precisely adjusting these parameters, the system can more accurately balance the contribution of different features to the physical conflict risk assessment and improve the accuracy and reliability of the assessment. At the same time, the adjustment of the threshold parameters can also enable the system to adapt to the needs of different scenarios and improve the flexibility and practicality of the assessment.
[0208] In summary, the process of using formula 17 and determining the weight value and threshold parameter according to the preset training database is of great significance for improving the performance of the physical conflict detection and early warning method based on video vibration image analysis. By optimizing the weight value and threshold parameter, the system can more accurately assess the risk of physical conflict and provide strong support for timely intervention and prevention of physical conflict.
[0209] Figure 3 FIG. 1 is a schematic diagram of a physical conflict detection and early warning device based on video vibration image analysis according to an exemplary embodiment of the present application. Figure 3 As shown, the physical conflict detection and early warning device 300 based on video vibration image analysis provided in this embodiment includes:
[0210] An acquisition module 310 is used to acquire a video file to be detected and decompose the video file to be detected into a sequence of images to be detected, wherein the number of frames of the video file to be detected is greater than a preset frame number threshold, and each image to be detected in the sequence of images to be detected includes a target object to be detected;
[0211] An extraction module 320 is used to extract vibration image data from the image sequence to be detected using a preset vibration image feature extraction model, wherein the vibration image data includes vibration frequency data and amplitude data;
[0212] The extraction module 320 is used to extract bone feature data from the image sequence to be detected using a preset bone feature extraction model, wherein the bone feature data includes bone key point position data and bone posture data;
[0213] The processing module 330 is used to detect the physical conflict risk level of the target object to be detected according to the vibration image data and the bone feature data by using a preset physical conflict detection model.
[0214] Optionally, the extraction module 320 is specifically used to:
[0215] Determine the motion vector of each pixel point according to the image sequence to be detected to form a time domain motion vector field;
[0216] Determine a corresponding frequency domain motion vector field according to the time domain motion vector field;
[0217] The vibration image data is extracted according to the frequency domain motion vector field.
[0218] Optionally, the extraction module 320 is specifically used to:
[0219] Determining the motion vector of each pixel point according to a first image to be detected and a second image to be detected in the sequence of images to be detected to form a time domain motion vector field, wherein the first image to be detected and the second image to be detected are any adjacent images to be detected in the sequence of images to be detected;
[0220] Determine a corresponding frequency domain motion vector field according to the time domain motion vector field;
[0221] The vibration image data is extracted according to the frequency domain motion vector field.
[0222] Optionally, the extraction module 320 is specifically used to:
[0223] Determine a corresponding feature heat map according to each frame of the image to be detected in the sequence of images to be detected, so as to form a feature heat map sequence;
[0224] Determine the position of the skeleton key point according to the feature heat map sequence;
[0225] Determine the bone vector according to the position of each bone key point;
[0226] Determine the bone length according to the bone vector;
[0227] Determine the bone angle between adjacent bone vectors according to each bone vector;
[0228] The bone feature data is determined according to the positions of the bone key points, the bone vectors, the bone lengths and the bone angles.
[0229] Optionally, the extraction module 320 is specifically used to:
[0230] Using a preset key point detection model, each frame of the image to be detected in the image sequence to be detected is mapped to the corresponding feature heat map. The preset key point detection model is a model based on a convolutional neural network, and the convolutional neural network includes a convolutional layer, a pooling layer, an activation function and a fully connected layer connected in sequence.
[0231] Optionally, the processing module 330 is specifically configured to:
[0232] Determine a vibration characteristic value according to the vibration frequency data and the amplitude data in the vibration image data;
[0233] Determine a bone structure vector diagram according to the bone key point positions, the bone vectors, the bone lengths and the bone angles in the bone feature data;
[0234] Determine the bone dynamic feature value according to the bone structure vector diagram corresponding to each frame of the image to be detected;
[0235] Determine a physical conflict risk characteristic value according to the vibration characteristic value and the bone dynamic characteristic value;
[0236] The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
[0237] Optionally, the processing module 330 is specifically configured to:
[0238] fusing the vibration image data and the bone feature data to form a fused limb feature vector;
[0239] Outputting a physical conflict risk feature value according to the fused physical feature vector;
[0240] The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
[0241] Figure 4 is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present application. Figure 4 As shown, an electronic device 400 provided in this embodiment includes: a processor 401 and a memory 402; wherein:
[0242] The memory 402 is used to store computer programs, and the memory may also be a flash memory.
[0243] The processor 401 is used to execute the execution instructions stored in the memory to implement each step in the above method. For details, please refer to the relevant description in the above method embodiment.
[0244] Optionally, the memory 402 may be independent or integrated with the processor 401 .
[0245] When the memory 402 is a device independent of the processor 401, the electronic device 400 may further include:
[0246] The bus 403 is used to connect the memory 402 and the processor 401 .
[0247] This embodiment further provides a readable storage medium, in which a computer program is stored. When at least one processor of an electronic device executes the computer program, the electronic device executes the methods provided in the above-mentioned various implementation modes.
[0248] This embodiment also provides a program product, which includes a computer program stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device implements the methods provided in the above various embodiments.
[0249] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.
[0250] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A physical conflict detection and early warning method based on video vibration image analysis, characterized in that: include: Acquire a video file to be detected, and decompose the video file to be detected into a sequence of images to be detected, wherein the number of frames of the video file to be detected is greater than a preset frame number threshold, and each image to be detected in the sequence of images to be detected includes a target object to be detected; Extracting vibration image data from the image sequence to be detected using a preset vibration image feature extraction model, wherein the vibration image data includes vibration frequency data and amplitude data; Extracting bone feature data from the image sequence to be detected by using a preset bone feature extraction model, wherein the bone feature data includes bone key point position data and bone posture data; The physical conflict risk level of the target object to be detected is detected by using a preset physical conflict detection model and according to the vibration image data and the bone feature data.
2. The method for detecting and warning physical conflict based on video vibration image analysis according to claim 1, characterized in that: The step of extracting vibration image data from the image sequence to be detected by using a preset vibration image feature extraction model includes: Determine the motion vector of each pixel point according to the image sequence to be detected to form a time domain motion vector field; Determine a corresponding frequency domain motion vector field according to the time domain motion vector field; The vibration image data is extracted according to the frequency domain motion vector field.
3. The method for detecting and warning physical conflict based on video vibration image analysis according to claim 2, characterized in that: The step of extracting vibration image data from the image sequence to be detected by using a preset vibration image feature extraction model includes: Determining the motion vector of each pixel point according to a first image to be detected and a second image to be detected in the sequence of images to be detected to form a time domain motion vector field, wherein the first image to be detected and the second image to be detected are any adjacent images to be detected in the sequence of images to be detected; Determine a corresponding frequency domain motion vector field according to the time domain motion vector field; The vibration image data is extracted according to the frequency domain motion vector field.
4. The method for detecting and warning physical conflict based on video vibration image analysis according to claim 1, characterized in that: The method of extracting bone feature data from the image sequence to be detected by using a preset bone feature extraction model includes: Determine a corresponding feature heat map according to each frame of the image to be detected in the sequence of images to be detected, so as to form a feature heat map sequence; Determine the position of the skeleton key point according to the feature heat map sequence; Determine the bone vector according to the position of each bone key point; Determine the bone length according to the bone vector; Determine the bone angle between adjacent bone vectors according to each bone vector; The bone feature data is determined according to the positions of the bone key points, the bone vectors, the bone lengths and the bone angles.
5. The method for detecting and warning physical conflict based on video vibration image analysis according to claim 4, characterized in that: The step of determining a corresponding feature heat map according to each frame of the image to be detected in the sequence of images to be detected includes: Using a preset key point detection model, each frame of the image to be detected in the image sequence to be detected is mapped to the corresponding feature heat map. The preset key point detection model is a model based on a convolutional neural network, and the convolutional neural network includes a convolutional layer, a pooling layer, an activation function and a fully connected layer connected in sequence.
6. The method for detecting and warning physical conflict based on video vibration image analysis according to claim 4 or 5, characterized in that: The method of using a preset physical conflict detection model and detecting the physical conflict risk level of the target object to be detected according to the vibration image data and the bone feature data includes: Determine a vibration characteristic value according to the vibration frequency data and the amplitude data in the vibration image data; Determine a bone structure vector diagram according to the bone key point positions, the bone vectors, the bone lengths and the bone angles in the bone feature data; Determine the bone dynamic feature value according to the bone structure vector diagram corresponding to each frame of the image to be detected; Determine a physical conflict risk characteristic value according to the vibration characteristic value and the bone dynamic characteristic value; The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
7. The method for detecting and warning physical conflict based on video vibration image analysis according to any one of claims 1 to 5, characterized in that: The method of using a preset physical conflict detection model and detecting the physical conflict risk level of the target object to be detected according to the vibration image data and the bone feature data includes: fusing the vibration image data and the bone feature data to form a fused limb feature vector; Outputting a physical conflict risk feature value according to the fused physical feature vector; The physical conflict risk level of the target object to be detected is determined according to the physical conflict risk characteristic value R and a preset risk characteristic level mapping table.
8. A physical conflict detection and early warning device based on video vibration image analysis, characterized in that: include: An acquisition module is used to acquire a video file to be detected and decompose the video file to be detected into a sequence of images to be detected, wherein the number of frames of the video file to be detected is greater than a preset frame number threshold, and each image to be detected in the sequence of images to be detected includes a target object to be detected; An extraction module, used to extract vibration image data from the image sequence to be detected by using a preset vibration image feature extraction model, wherein the vibration image data includes vibration frequency data and amplitude data; The extraction module is used to extract bone feature data from the image sequence to be detected by using a preset bone feature extraction model, wherein the bone feature data includes bone key point position data and bone posture data; The processing module is used to use a preset physical conflict detection model to detect the physical conflict risk level of the target object to be detected according to the vibration image data and the bone feature data.
9. An electronic device, characterized in that: include: processor; as well as, A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.