Violence event detection method for audio and video after three-dimensional reconstruction based on multiple cameras
Through the three-dimensional reconstruction of multi-camera and multi-modal data fusion methods, the problem of low detection accuracy of brute-events in occlusion and complex scenarios is solved, and high-precision and robust brute-force detection are achieved.
Patent Information
- Application Number
- CN202510702113.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has low detection accuracy for violent incidents in handling occlusion and complex scenarios, and single camera and single mode detection methods have problems of resource waste and false detection.
Multi-camera three-dimensional reconstruction combined with audio and video data, video and audio data are processed through graph convolution networks and two-dimensional convolution neural networks, multi-dimensional violence detection is performed, and violent behavior is judged by adaptive adjustment of weights and thresholds.
It improves detection accuracy in occlusion and complex scenarios, reduces false detection, enhances robustness, realizes the distinction between voiceprint characteristics of different people, and improves the accuracy and adaptability of detection.
Smart Images

Figure CN120343208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a method for detecting violent events in audio and video after three-dimensional reconstruction based on multiple cameras. Background Art
[0002] The combination of a single camera and a human body tracking model, or neural network learning without audio processing: A single camera brings convenience, but at the same time brings the inability to resolve occlusions. When the occlusion exceeds a certain degree, the detection becomes meaningless. And the method of simply placing multiple cameras to run multiple single cameras is too wasteful of hardware resources, lacking the technology for integrating multiple cameras. For the detection of the voices of different people, using neural network learning without distinguishing voiceprints will cause problems in specific recognition. For example, when two people are singing loudly in chorus, those without voiceprint detection will identify it as a single person shouting loudly, resulting in false detection. After detecting the voiceprint, it can better handle the situation of distinguishing the voices of different people, making the neural network model more perfect and more robust. Summary of the Invention
[0003] The main object of the present invention is to provide a method for detecting violent events in audio and video after three-dimensional reconstruction based on multiple cameras, so as to solve the problems existing in the existing violent behavior detection and three-dimensional reconstruction technology of human body key points, such as low detection accuracy, poor adaptability to complex scenes, insufficient reconstruction accuracy, and dependence on manual intervention.
[0004] To solve the above technical problems, the technical solution adopted by the present invention is: A method for detecting violent events in audio and video after three-dimensional reconstruction based on multiple cameras, the method comprising: S1. Obtain video data through multiple cameras and perform time series alignment in combination with audio data; S2. Perform three-dimensional reconstruction of human body key points based on the video data to generate three-dimensional coordinate data; S3. Extract features from the three-dimensional coordinate data and the audio data respectively, and construct detection models for video and audio; S4. Output the probabilities of violent behaviors in the video and audio through the detection models; Set the initial weights of the probabilities of the video and audio according to the scene characteristics, and obtain the fused violent probability through adaptive adjustment; S5. Judge the fused violent probability through a preset threshold to determine whether a violent behavior has occurred.
[0005] In a preferred solution, in step S1, multiple cameras are used to cover the detection area to obtain video data; the multiple cameras are arranged in a triangular state to include the positions of the violent detection area to the greatest extent; Obtain audio data through the audio acquisition device built into the camera; Use timestamps to align the time of the video data and the audio data to ensure the consistency of data from different sources on the time axis; specifically, use absolute timestamps to correspond the video information captured by different cameras with the audio information captured by the microphone; For the video data, adjust the position of the camera to form a stereoscopic coverage perspective and obtain visual information from multiple angles; For the audio data, record environmental sound information to assist subsequent processing.
[0006] In the preferred solution, in step S2, calculate the internal parameter matrix and the external parameter matrix of the camera through a calibration method; Adopt a key point detection algorithm to extract two-dimensional key point coordinates from the video data; Match the two-dimensional key point coordinates through geometric constraints between multiple cameras; Construct a projection equation system according to the matching result, and use an optimization algorithm to solve for the three-dimensional coordinate data of the human body key points; Perform error correction on the three-dimensional coordinate data to ensure the reconstruction accuracy.
[0007] In the preferred solution, the specific steps are as follows: Calculate the internal parameter matrix and the external parameter matrix of the camera through a calibration method; Specifically set a black and white checkerboard or a circular array, use Zhang Zhengyou's calibration method to calculate the internal parameter matrix of each camera, and solve the rotation matrix and the translation vector by arranging known three-dimensional coordinate feature points and combining the least squares method. If the internal and external parameter errors are large, calculate again until the error is reduced below the threshold; For internal parameter calibration, each camera takes at least 10 checkerboard images from different angles and positions, uses the built-in library of matlab to input the grid size to extract corner points, and calculates the reprojection error through the camera model. The calculation formula is: ; Use the library of matlab to calculate the internal and external parameter matrices. The expression of the internal parameter matrix is: ; Among them, and respectively represent the focal lengths of the camera in the horizontal and vertical directions, which are used to describe the scaling degree of the camera's imaging of an object; and are the principal point coordinates, representing the intersection position of the optical axis and the image plane in the image plane, and the units are all pixels; The external parameter calibration obtains the projection equation by combining the conversion relationship between the world coordinate system and the camera coordinate system with the internal parameter matrix: ; Among them, (u, v) is the two-dimensional coordinate of point P on the image plane, is the aforementioned intrinsic matrix, the rotation matrix R is a 3×3 orthogonal matrix used to describe the rotation angle of the camera coordinate system relative to the world coordinate system; T is the translation vector used to represent the position of the camera in the world coordinate system, They are the corresponding coordinates of point P in the camera coordinate system; A key point detection algorithm is used to extract two-dimensional key point coordinates from the video data; specifically, openpose is used to obtain some key point coordinates; The coordinates of the two-dimensional key points are matched by geometric constraints between multiple cameras; the key points are matched by epipolar constraints, and at least one pair of epipolar lines must be matched for each point of at least three cameras, and then two-by-two verifications are performed, and finally the corresponding matching pairs are output after successful verification; Constructing a projection equation group according to the matching results, and using an optimization algorithm to solve and obtain the three-dimensional coordinate data of key points of the human body; The projection equations corresponding to at least three cameras are combined to form an overdetermined system of equations, and the overdetermined system of equations is solved using the least square method to obtain the predicted two-dimensional coordinates of the projection; Assume the observed two-dimensional image coordinates are ( Corresponding to three cameras respectively), the predicted coordinates calculated according to the projection equation are: , then the error function is: ; Then the error function is about , , Find the partial derivative, set it to 0, and substitute it to get the coordinates of the three-dimensional point The estimated value of; the Levenberg-Marquardt algorithm is used for iterative optimization, by solving the equation To calculate the coordinate update amount ; Error correction is performed on the three-dimensional coordinate data to ensure reconstruction accuracy.
[0008] In the preferred solution, in step S3, the three-dimensional coordinate data and the audio data are respectively subjected to feature extraction to construct a video and audio detection model, including: Input the three-dimensional coordinate data as node features into a graph convolutional network model, and construct edge features based on human skeleton relationships; Fuse the node features and edge features, and extract the behavior features in the video through convolution operations; Perform voiceprint discrimination and spectral feature conversion on the audio data to generate spectrogram data; Input the spectrogram data into a two-dimensional convolutional neural network model, and extract the abnormal features in the audio through multi-layer convolution and pooling operations.
[0009] In the preferred solution, it is characterized in that the specific steps include: Input the three-dimensional coordinate data as node features into the graph convolutional network model, and construct edge features according to the human body bone relationship; Add the contact speed to the input matrix. The calculation method of the contact speed is: For key point pairs of key parts where the distance is less than the threshold Calculate the speed between adjacent frames; Suppose at the frame, individual , The hand key point coordinates are respectively , At the frame, individual , The hand key point coordinates become , If the frame rate is fps then the time interval is , and the displacement is m / s; Fuse the node features and edge features, and extract the behavior features in the video through convolution operations; Use the graph convolution formula for convolution processing, then use the Relu activation function for activation, and perform convolution again to extract information; Perform voiceprint discrimination and spectral feature conversion on the audio data to generate spectrogram data; First, perform voiceprint discrimination on the audio data, and then perform Mel scale conversion operation to convert the linear frequency to Mel frequency The formula is: , construct a Mel filter bank containing 40 - 80 triangular filters, and filter the spectrum of the audio signal , the output of the th Mel filter is calculated as follows: ; Perform logarithmic transformation on the output of the Mel filter bank to obtain the Mel spectrogram. The formula is: ; where , which is used to avoid zero values in logarithmic operations; Input the spectrogram data into a two-dimensional convolutional neural network model, and extract abnormal features in the audio through multi-layer convolution and pooling operations; the network architecture of the two-dimensional convolutional neural network model includes a convolutional layer, a max pooling layer, a fully connected layer, and an output layer. The output formula of the convolutional layer is: , the pooling layer uses max pooling, and the output formula is: , the fully connected layer integrates the features extracted by the convolutional layer and the pooling layer, and the output layer outputs the probabilities of violence and normal situations through the Softmax function. The formula is: ; Among them, is the i-th element of the output vector of the fully connected layer, K is the number of categories, is the probability belonging to the i-th category.
[0010] In the preferred solution, in step S4, the three-dimensional coordinate data is processed through a graph convolutional network model to output the probability value of violent behavior in the video; The audio data is processed through a two-dimensional convolutional neural network model to output the probability value of violent behavior in the audio; Normalize the video probability value and the audio probability value to ensure the consistency of the probability distribution; Record the detection results of different modality data according to the probability value, providing a basis for subsequent fusion.
[0011] In the preferred solution, in step S4, set the initial weight values of the video probability and the audio probability according to the scene noise level; Optimize and adjust the initial weight values through a loss function, and dynamically update the weights using the gradient descent method; Perform weighted summation on the video probability and the audio probability according to the adjusted weights to obtain the fused violence probability; Continuously monitor the fused violence probability and record the change trend during the weight adjustment process.
[0012] In the preferred solution, the specific steps are: Set the initial weight values of the video probability and the audio probability according to the scene noise level; Optimize and adjust the initial weight values through a loss function, and dynamically update the weights using the gradient descent method; select the cross-entropy loss function as the optimization target, and calculate the gradients of the loss function with respect to the weights and through backpropagation. The weight update formula is: , ; Weight the video probability and the audio probability according to the adjusted weights, and sum them to obtain the fused violence probability; the fused violence probability ; Compare it with a set violence threshold If p> , it is determined that a violent act has occurred, thus realizing the detection of violent acts under multi-modal data fusion.
[0013] In a preferred solution, in step S5, the fused violence probability is judged by a preset threshold to determine whether a violent act has occurred, including: Use the threshold moving method to determine the initial violence probability threshold; Evaluate the performance of the initial violence probability threshold through verification data and adjust the threshold range; for each value, use it as the current violence threshold , use the trained model to predict the verification set data, and calculate the corresponding evaluation metrics; Including accuracy, recall rate, and F1 value. Specifically: accuracy , recall rate , , where ; Compare the fused violence probability with the adjusted threshold; If the fused violence probability is greater than the threshold, it is determined as a violent act; If the fused violence probability is less than or equal to the threshold, it is determined as a normal situation; Record the determination result and trigger the corresponding response mechanism.
[0014] The present invention provides a method for detecting violent events in audio and video after three-dimensional reconstruction based on multiple cameras. Compared with general violent detection and recognition methods, it has good recognition characteristics for occlusion situations and situations of the same pose but different behaviors. Moreover, for the voices of different people at the same time, it can roughly distinguish them, which is beneficial for training the 2DCNN model. In general, due to three-dimensional reconstruction, this model performs feature point recognition, reducing the error of point coordinates, and adding a one-dimensional speed relationship, improving the recognition accuracy of the model. In the case of occlusion and multi-person chatting, we have a three-dimensional observation ability and a voiceprint unique audio detection method, leading existing violent recognition methods in multiple aspects in terms of robustness and detectability. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The following further describes the present invention in conjunction with the drawings and embodiments: Figure 1 It is a histogram comparing the data of the model in this article of the present invention with other models; Figure 2It is a schematic diagram of the general implementation method of the text model of the present invention; Figure 3 It is a schematic diagram of the multi-modal combination weight setting method of the present invention; Figure 4 It is a schematic diagram of the method of the video module part of the present invention; Figure 5 It is a schematic diagram of the detection method of the audio module part of the present invention; Figure 6 It is a schematic diagram of the detection method of the audio module part of the present invention; Figure 7 It is a schematic diagram of the specific implementation method inside the improved GCN model of the present invention; Figure 8 It is a schematic diagram of two people who are easily misdetected as being beaten under occlusion in the present invention; Figure 9 It is a detection and prediction diagram A for the other two sides in the present invention; Figure 10 It is a detection and prediction diagram B for the other two sides in the present invention.
[0016] Figure 11 It is a schematic diagram of using a calibration board in the present invention Specific implementation manner Example 1 As Figures 1-11 shown, a violent event detection method for audio and video after 3D reconstruction based on multiple cameras, the method includes: S1. Obtain video data through multiple cameras and perform time series alignment in combination with audio data; S2. Perform 3D reconstruction of human key points based on the video data to generate 3D coordinate data; S3. Extract features from the 3D coordinate data and the audio data respectively, and construct detection models for video and audio; S4. Output the probabilities of violent behaviors in the video and audio through the detection models; Set the initial weights of the video and audio probabilities according to the scene characteristics, and obtain the fused violent probability through adaptive adjustment; S5. Judge the fused violent probability through a preset threshold to determine whether a violent behavior occurs.
[0017] Step 1: Alignment of input video and audio time series - Place cameras at 3 positions that can contain each other and can maximally contain the violent detection area, do not place them unilaterally, it is best to arrange them in a triangular state, and then use the microphones built in the cameras to receive audio information. Now we need to align the time. The absolute time stamp can be used to correspond the video information captured by different cameras and the audio information captured by the microphones.
[0018] Step 2: 3D reconstruction of human key points by multiple cameras Set up a black and white chessboard or circular array according to the current specific situation, and then use Zhang Zhengyou's calibration method to calculate the intrinsic parameter matrix of each camera, and solve the rotation matrix and translation vector (extrinsic parameter) by arranging known 3D coordinate feature points and combining the least squares method. In this way, the internal and external parameters are calibrated. If a large error is found in the internal and external parameters, consider recalculating until the error is reduced below the threshold.
[0019] Step 3: Prepare the data for violent behavior detection by using openpose to obtain the coordinates of some key points. Then, match the feature points based on the polar line matching method between the two cameras. Then, build the projection equation based on the calibration parameters and the matched two-dimensional key point coordinates. Combine multiple camera equations, solve them using the least squares method, and iterate and optimize them through the Levenberg-Marquardt algorithm to obtain high-precision three-dimensional coordinates of the key points of the human body.
[0020] Audio data processing: First, the audio data is distinguished by voiceprint and then converted to Mel scale, and finally the Mel spectrum is obtained, which is then used as the input of the 2DCNN model. Voiceprint is a way for humans to distinguish sounds. This allows the model to understand the source of sounds between different people and specifically identify violent and abnormal sounds.
[0021] Step 4: Construct a model to detect violent behavior - GCN inputs the coordinates of the three-dimensional human key points as nodes, and then specifies the construction of the edges between the nodes based on the connection of the human skeleton. The input features are speed, acceleration, confidence, relative position and other features. These features are input into the feature matrix. The aggregation and complexity of features are necessary, because more judgments are required for phenomena such as hugging and beating, which have a lot of physical contact but completely different results, because the same posture and different speeds will bring different results. Moreover, more features can reduce the occurrence of underfitting problems, that is, reduce the phenomenon of biased training of more classified data sets in the binary classification problem in this method. Then follow the repeated operations of convolution, relu activation, and convolution to gradually increase the aggregation between information, and finally output the probability of violence in video processing.
[0022] The input audio is identified using the voiceprint method and then processed using the Mel-spectrogram to distinguish the voiceprint features and information of different people. It is then input into the 2DCNN neural network, where it passes through a convolutional layer, a pooling layer, and a fully connected layer. It is then normalized using the Softmax function and the probability of violent behavior in the audio is output.
[0023] Set the initial weights of the video and audio model output probabilities according to the scenario. When there is a lot of noise, the audio accuracy is low, so the audio weight is set lower. Otherwise, it is increased. The cross entropy loss function is used to optimize the weight. Stop the operation when the final result has a high accuracy rate and determine the weight.
[0024] The same information is input for the determined weights to finally obtain the fused violence probability. The violence probability threshold is optimized by the threshold moving method and the loss function-based optimization method.
[0025] Step 5: Identification of violent behavior: Once the violent behavior is confirmed, the three cameras will start recording the current video and audio information and give cloud prompts. The staff can use the camera's built-in microphone to prevent the violence from further deteriorating.
[0026] In the preferred embodiment, in step S1, multiple cameras are used to cover the detection area to obtain video data; the multiple cameras are arranged in a triangle to cover the location of the violence detection area to the greatest extent; Get audio data through the camera's built-in audio acquisition device; Using timestamps to time-align the video data and the audio data to ensure consistency of data from different sources on the time axis; specifically using absolute timestamps to correspond video information captured by different cameras with audio information captured by microphones; With respect to the video data, adjusting the camera position to form a stereoscopic coverage perspective and obtain multi-angle visual information; For the audio data, environmental sound information is recorded to assist subsequent processing.
[0027] In the preferred embodiment, in step S2, the intrinsic parameter matrix and the extrinsic parameter matrix of the camera are calculated by a calibration method; Extracting two-dimensional key point coordinates from the video data using a key point detection algorithm; Matching the coordinates of the two-dimensional key points through geometric constraints between multiple cameras; Constructing a projection equation group according to the matching results, and using an optimization algorithm to solve and obtain the three-dimensional coordinate data of key points of the human body; Error correction is performed on the three-dimensional coordinate data to ensure reconstruction accuracy.
[0028] In the preferred embodiment, the specific steps are: Calculate the camera's intrinsic and extrinsic matrix through calibration methods; Specifically, a black and white chessboard or circular array is set up, and the Zhang Zhengyou calibration method is used to calculate the intrinsic parameter matrix of each camera. The rotation matrix and translation vector are solved by arranging known three-dimensional coordinate feature points and combining the least squares method. If the internal and external parameter errors are large, they are calculated again until the error is reduced below the threshold; The intrinsic calibration uses each camera to shoot at least 10 chessboard images from different angles and positions, uses the built-in library of Matlab to input the square size to extract the corner points, and calculates the reprojection error through the camera model. The calculation formula is: ; Use the Matlab library to calculate the internal and external parameter matrix, the internal parameter matrix The expression is: ; in, and Respectively represent the focal length of the camera in the horizontal and vertical directions, and are used to describe the degree of zoom of the camera's imaging of an object; and is the principal point coordinate, representing the intersection of the optical axis and the image plane in the image plane, The units are all pixels; The external parameter calibration obtains the projection equation by combining the conversion relationship between the world coordinate system and the camera coordinate system with the internal parameter matrix: ; Among them, (u, v) is the two-dimensional coordinate of point P on the image plane, is the aforementioned intrinsic matrix, the rotation matrix R is a 3×3 orthogonal matrix used to describe the rotation angle of the camera coordinate system relative to the world coordinate system; T is the translation vector used to represent the position of the camera in the world coordinate system, They are the corresponding coordinates of point P in the camera coordinate system; A key point detection algorithm is used to extract two-dimensional key point coordinates from the video data; specifically, openpose is used to obtain some key point coordinates; The coordinates of the two-dimensional key points are matched by geometric constraints between multiple cameras; the key points are matched by epipolar constraints, and at least one pair of epipolar lines must be matched for each point of at least three cameras, and then two-by-two verifications are performed, and finally the corresponding matching pairs are output after successful verification; Constructing a projection equation group according to the matching results, and using an optimization algorithm to solve and obtain the three-dimensional coordinate data of key points of the human body; The projection equations corresponding to at least three cameras are combined to form an overdetermined system of equations, and the overdetermined system of equations is solved using the least square method to obtain the predicted two-dimensional coordinates of the projection; Assume the observed two-dimensional image coordinates are ( Corresponding to three cameras respectively), the predicted coordinates calculated according to the projection equation are: , then the error function is: ; Then, take the partial derivatives of the error function with respect to , , , and set the partial derivatives to 0, and substitute to obtain the estimated values of the coordinates of the three-dimensional points ; Use the Levenberg-Marquardt algorithm for iterative optimization, and calculate the update amount of the coordinates by solving the equation ; Perform error correction on the three-dimensional coordinate data to ensure the reconstruction accuracy.
[0029] In the preferred solution, in step S3, the three-dimensional coordinate data and the audio data are respectively subjected to feature extraction to construct a detection model for video and audio, including: Take the three-dimensional coordinate data as node features and input them into the graph convolutional network model, and construct edge features according to the human body bone relationship; Perform fusion processing on the node features and edge features, and extract the behavior features in the video through convolutional operations; Perform voiceprint discrimination and spectral feature conversion on the audio data to generate spectrogram data; Input the spectrogram data into a two-dimensional convolutional neural network model, and extract the abnormal features in the audio through multi-layer convolutional and pooling operations.
[0030] In the preferred solution, the features are: the specific steps include: Take the three-dimensional coordinate data as node features and input them into the graph convolutional network model, and construct edge features according to the human body bone relationship; Add the contact speed to the input matrix, and the calculation method of the contact speed is: For the key part key point pairs with a distance less than the threshold , calculate the speed between adjacent frames; Suppose at the th frame, the individual , The hand key point coordinates are respectively , , and at the th frame, the individual , The hand key point coordinates become , , if the frame rate is fps then the time interval is , and the displacement is m / s; Perform fusion processing on the node features and edge features, and extract the behavior features in the video through convolutional operations; Use the graph convolution formula for convolutional processing, and then use the Relu activation function for activation, and perform convolution again to extract information; Perform voiceprint differentiation and spectral feature conversion on the audio data to generate spectrogram data; first perform voiceprint differentiation on the audio data, and then perform Mel-scale conversion operation to convert linear frequency to Mel frequency The formula is: , construct a Mel filter bank containing 40 - 80 triangular filters, and perform spectral analysis on the audio signal , for the output of the Mel filter calculate as follows: ; Perform logarithmic transformation on the output of the Mel filter bank to obtain the Mel spectrogram, and the formula is: ; where is used to avoid zero values in logarithmic operations; Input the spectrogram data into a two-dimensional convolutional neural network model, and extract abnormal features in the audio through multi-layer convolution and pooling operations; the network architecture of the two-dimensional convolutional neural network model includes a convolutional layer, a max pooling layer, a fully connected layer, and an output layer. The output formula of the convolutional layer is: , the pooling layer uses max pooling, and the output formula is: , the fully connected layer integrates the features extracted by the convolutional layer and the pooling layer, and the output layer outputs the probabilities of violence and normal situations through the Softmax function. The formula is: ; where, is the i-th element of the output vector of the fully connected layer, K is the number of categories, is the probability belonging to the i-th category.
[0031] In the preferred solution, in step S4, process the three-dimensional coordinate data through a graph convolutional network model to output the probability value of violent behavior in the video; Process the audio data through a two-dimensional convolutional neural network model to output the probability value of violent behavior in the audio; Perform normalization processing on the video probability value and the audio probability value to ensure the consistency of the probability distribution; Record the detection results of different modal data according to the probability value to provide a basis for subsequent fusion.
[0032] In the preferred solution, in step S4, set the initial weight values of the video probability and the audio probability according to the scene noise level; Optimize and adjust the initial weight values through a loss function, and dynamically update the weights using the gradient descent method; Weight the video probability and the audio probability according to the adjusted weights, and sum them up to obtain the fused violence probability; Continuously monitor the fused violence probability and record the change trend during the weight adjustment process.
[0033] In the preferred solution, the specific steps are as follows: Set the initial weight values of the video probability and the audio probability according to the scene noise level; Optimize and adjust the initial weight values through a loss function, and dynamically update the weights using the gradient descent method; select the cross-entropy loss function As the optimization objective, calculate the gradient of the loss function with respect to the weights and through backpropagation. The weight update formula is: , ; Weight the video probability and the audio probability according to the adjusted weights, and sum them up to obtain the fused violence probability; the fused violence probability ; Compare it with the set violence threshold . If p > , it is determined that a violent behavior has occurred, thus realizing the detection of violent behaviors under multi-modal data fusion.
[0034] In the preferred solution, in step S5, judge whether a violent behavior has occurred by comparing the fused violence probability with a preset threshold, including: Determine the initial violence probability threshold using the threshold moving method; Evaluate the performance of the initial violence probability threshold through verification data and adjust the threshold range; for each value, use it as the current violence threshold , use the trained model to predict the verification set data, and calculate the corresponding evaluation metrics; Including accuracy, recall rate, and F1 value. Specifically: accuracy , recall rate , , where ; Compare the fused violence probability with the adjusted threshold; If the fused violence probability is greater than the threshold, it is determined that a violent behavior has occurred; If the fused violence probability is less than or equal to the threshold, it is determined to be a normal situation; Record the judgment result and trigger the corresponding response mechanism.
[0035] Embodiment 2 Further illustrate in combination with Embodiment 1. For exampleFigures 1-11 The structure shown. Step 1: Place cameras in the venue where violence needs to be detected. Try to include the most important detection areas with all three cameras, and place them separately at 120° + 120° + 120° to maximize detection.
[0036] 2. Prepare a clear black-and-white chessboard. Do not use a calibration board with a high reflectivity such as an acrylic board. The size and spacing of each square must be fixed. At the same time, sample in a place with sufficient light to ensure clear imaging and accurate focusing within the camera's shooting range, so that the image contains as many chessboard corner points as possible.
[0037] Step 2: Camera calibration 1. Intrinsic parameter calibration (1) Use each camera to capture chessboard images from different angles and positions. At least ensure that the entire calibration board is in the image, and the number of captured images is at least 10 or more.
[0038] (2) Use the built-in library of Matlab to input the square size to extract corner points. For each set of corner points, there is a set of known three-dimensional world coordinates and two-dimensional image coordinates , and then through the camera model, calculate the reprojection error. The calculation formula is: ; Among them, represents the reprojection error of the th corner point. , : The actual two-dimensional image coordinates of the th corner point.
[0039] , : The reprojected two-dimensional image coordinates of the th corner point.
[0040] The purpose of calculating this error is to use the non-linear least squares method to reduce the error and solve the internal and external parameters.
[0041] Use the Matlab library to calculate the internal and external parameter matrices.
[0042] The expression of the intrinsic parameter matrix K is: ; Among them, and respectively represent the focal lengths of the camera in the horizontal and vertical directions, which are used to describe the scaling degree of the camera's imaging of objects; and are the principal point coordinates, representing the intersection position of the optical axis and the image plane in the image plane, The units are all pixels; 2. Extrinsic calibration: Let the point in the world coordinate system be P(X,Y,Z) The coordinates in the camera coordinate system are Pc( ), and its transformation relationship is: ; Combined with the intrinsic matrix K, the camera coordinate system coordinates are then transformed into image coordinates to obtain the projection equation: ; Among them, (u, v) are the two-dimensional coordinates of point P on the image plane, is the aforementioned intrinsic matrix. The rotation matrix R is a 3×3 orthogonal matrix used to describe the rotation angle of the camera coordinate system relative to the world coordinate system; T is the translation vector used to represent the position of the camera in the world coordinate system, are the corresponding coordinates of point P in the camera coordinate system respectively; Step 3: Feature point matching: The input includes three camera identifiers and the key points of the corresponding cameras extracted by OpenPose.
[0043] The method of epipolar constraint is used for key point matching: First, explain the epipolar constraint method for two cameras (name the two cameras as Camera 1 and Camera 2): For the point = , ), we map it to Camera 2. The epipolar line equation in Camera 2 is: ; The distance formula from the point in Camera 2 to the above epipolar line: ; Then, perform matching and consider 3 pixels as the maximum allowable distance For each point of the three cameras, at least 1 pair of epipolar lines (the result of rounding down 3 / 2) needs to be matched, and then pairwise verification is performed.
[0044] Finally, if the verification is successful, the corresponding matching pairs are output.
[0045] Step 4: Calculate the projection equation of the three-dimensional coordinates: ; Among them, s is the scale factor, indicating the scaling ratio of the distance from the three-dimensional point to the camera optical center on the image plane. Solve for the three-dimensional coordinates: Combine the projection equations corresponding to the three cameras to form an overdetermined system of equations. Use the least squares method to solve this overdetermined system of equations to obtain the predicted two-dimensional coordinates of the projection. Let the observed two-dimensional image coordinates be , where correspond to three cameras respectively, and the predicted coordinates calculated according to the projection equation are , then the error function is: ; Then, take the partial derivatives of the error function with respect to X, Y, and Z, and set the partial derivatives to 0, and substitute to obtain the estimated values of the coordinates (X, Y, Z) of the three-dimensional points. The Levenberg-Marquardt algorithm is used for iterative optimization. This algorithm gradually reduces the value of the error function E by continuously adjusting the three-dimensional coordinate values. By solving the equation to calculate the coordinate update amount ∆x, and then calculate the error to judge whether it is feasible. If it is feasible, use it; if it is not feasible, re-substitute ∆x and calculate the error until the error meets the expectation. Among them, λ is the damping parameter, e is the error vector, ∆x is the coordinate update amount, and J is the Jacobian determinant of the three-dimensional coordinates.
[0046] Then, through the processing of three-dimensional human key point data and the construction of the GCN (Graph Convolutional Network) model, the detection of fighting scenes is realized. The specific operation steps are as follows: I. Data Preparation The processing of three-dimensional human key point data obtains the coordinate information of human key points from the three-dimensional reconstruction results. Use openpose to select key points, and each key point contains 3 coordinate values (x, y, z).
[0047] Dataset combination combines the key point data of different individuals Calculation of contact speed and distance judgment Distance calculation: For the key point pairs of key parts of different individuals, use the Euclidean distance formula to calculate their distances in each frame. The distance ; It is necessary to specify the threshold according to the specific situation: The initial threshold setting should be compared with the specific environmental situation. The weight of the audio input needs to be lowered in noisy locations, and the weight of the video input needs to be increased in places with many people but quiet.
[0048] For the key point pairs of key parts with a distance less than the threshold , calculate the speed between adjacent frames; Suppose at the th frame, individual , The hand key point coordinates are respectively , , at the th frame, individual , The hand key point coordinates become , , if the frame rate is fps then the time interval is , the displacement is m / s; II. GCN Input Layer (1) Input the 3D human key point coordinates and contact speed as node features; Edge construction: Connect the corresponding edges according to the set bone relationships. The initial edge feature is 0, and the distance relationship between the edges is judged. If the distance is close to a certain value, the value of the edge feature is increased.
[0049] (2) Fusion operation: Fuse the updated edge features with the node features: That is, map the edge features (contact speed) to the same dimension as the node features through linear transformation and then add them.
[0050] Graph convolution operation: Perform convolution processing on the fused features. The graph convolution formula is , and then use the Relu activation function for activation and perform convolution again to extract information. Among them, Hl is the node feature matrix of the l-th layer, is the adjacency matrix with self-connection added, D is the degree matrix of Aˆ, is the weight matrix of the l-th layer, and is the activation function (such as ReLU). Assume is an n×3 matrix (n is the number of nodes), A is an n×n adjacency matrix, I is an n×n identity matrix, D is a diagonal matrix, and its diagonal elements are the sum of the elements in each row of Aˆ, is a 3×k matrix (k is the node feature dimension of the next layer).
[0051] (1) Fully connected layer: Convert the output node feature vector into a vector with a fixed length. This flattening process facilitates normalization.
[0052] (2) Probability output: Normalize the output of the fully connected layer through the Softmax function to obtain the probability distributions of fighting and normal situations. The formula is , where is the i-th element of the output vector of the fully connected layer, and K = 2.
[0053] III. Define the loss function and optimizer: The binary cross-entropy loss function is selected as the loss function. The formula is , where y is the true label (0 represents normal, 1 represents fighting), and p is the probability that the model predicts as the positive class (fighting). Due to its exponential loss calculation, the cross-entropy model has very good effects.
[0054] Application of 2DCNN and Mel method in violent behavior detection In the violent behavior detection based on multi-modal data fusion, the Mel method is used to extract audio features in the audio data processing part, and the 2DCNN model is used for feature analysis and behavior judgment. The specific content is as follows: Mel-frequency cepstral coefficient (MFCC) feature extraction. The Mel scale is a frequency scale designed based on the auditory characteristics of the human ear. The human ear's perception of sounds at different frequencies is not a linear relationship, and the Mel scale can better simulate this non-linear perception. The formula for converting linear frequency f to Mel frequency m is: ; where is the linear frequency (Hz), and m is the corresponding Mel frequency value. The Mel-frequency cepstral coefficient (MFCC) feature can be combined with a 2D-CNN model to achieve high accuracy in audio recognition. This conversion makes the Mel frequency change relatively slowly in the low-frequency band, corresponding to the human ear's more sensitive perception of low-frequency sounds; in the high-frequency band, the Mel frequency changes faster, which is in line with the human ear's relatively rough perception of high-frequency sounds. Through this conversion, the frequency components that are more important to human hearing can be retained, redundant information can be removed, and a foundation for subsequent feature extraction can be laid.
[0055] Mel filter bank design. A Mel filter bank consisting of 40 - 80 triangular filters is constructed. Each filter has a specific center frequency and bandwidth in the Mel frequency domain. For the spectrum X(k) of an audio signal (k represents the frequency index, and N is the spectrum length), the output M(m) of the m-th Mel filter is calculated as follows, with the center frequencies evenly distributed on the Mel scale: ; where is the center frequency of the m-th filter. The Mel filter bank filters the audio spectrum and converts the audio signal into a Mel spectrum. Since the design of the filter bank conforms to the auditory characteristics of the human ear, it can highlight the frequency components that the human ear is sensitive to, effectively reduce the data dimension, reduce the computational amount, and at the same time enhance the representativeness of the features.
[0056] Logarithmic energy calculation. A logarithmic transformation is performed on the output of the Mel filter bank to obtain a Mel spectrogram, and the formula is: ; where , which is used to avoid zero values in the logarithmic operation. The logarithmic transformation makes the amplitude change of the Mel spectrogram more in line with the human ear's perception characteristics, further enhancing the weight of the low-frequency part, while compressing the dynamic range of the high-frequency part. This not only makes the features more stable but also improves the model's adaptability to different audio intensities, enhances the robustness of the features, and makes the extracted features more suitable for subsequent model processing.
[0057] Table 1: 2DCNN network architecture
[0058] Constructing a 2DCNN model generally includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer.
[0059] The input layer takes the Mel spectrogram as the input of the 2D CNN. The Mel spectrogram is similar to a two-dimensional image in form. Its width corresponds to the time frames, recording the changes of audio over time; its height corresponds to the Mel frequency channels, reflecting the information of different frequency components; the number of channels is generally equal to the number of Mel filters. This input method of converting audio features into image form makes good use of the powerful ability of 2D CNN in image feature extraction, enabling 2D CNN to effectively analyze audio features.
[0060] The convolutional layer uses multiple convolutional layers to extract features from the Mel spectrogram. The convolutional layer extracts local features through the convolution operation of the convolution kernel and the input data. Assuming the input feature map is X, the convolution kernel is W, and the bias is b, then the output Y of the convolutional layer is: ; where * represents the convolution operation, and σ is the activation function (such as ReLU). Convolution kernels of different sizes and numbers can extract features of different scales and frequencies in the Mel spectrogram. Since we are observing violent behavior audio, we need larger convolution kernels to observe macroscopic features, which can improve the correct recognition rate of the model.
[0061] The pooling layer inserts a pooling layer between convolutional layers to reduce the data dimension, reduce the computational amount, and improve the robustness of the model. We adopt max pooling, taking the maximum value within the pooling window of size p×q as the output. Assuming the input feature map is X and the output feature map is Y, then: ; where 0≤m<p, 0≤n<q. The pooling operation retains the main features and removes redundant information through downsampling, while making the model more adaptable to minor changes and local differences in audio features, then reducing the computational burden of the network and preventing overfitting.
[0062] After passing through multiple convolutional layers and pooling layers, the fully connected layer and the output layer flatten the feature map and input it into the fully connected layer. The fully connected layer maps the features to a specific dimension through matrix multiplication, and the fully connected layer integrates the features extracted by the convolutional layer and the pooling layer. Finally, the probabilities of violent and normal situations are output through the Softmax function, and the formula is: ; where is the i-th element of the output vector of the fully connected layer, K is the number of categories (here K = 2, representing violent and normal respectively), is the probability belonging to the i-th category. The fully connected layer integrates the features extracted by the previous layers, and the Softmax function converts the output into a probability distribution, turning the abstract model prediction into numerical analysis, which is convenient for the final weight calculation.
[0063] In the detection of violent behavior through multi-modal data fusion, after the audio data is processed by the above-mentioned 2DCNN and Mel methods, the probability of violent behavior in the audio is output. ; The video data is processed through three-dimensional reconstruction, OpenPose, and GCN models to output the probability of violent behavior in the video. . Set the initial weights according to the scene classification. For example, in a noisy environment, the weight of the video data = 0.8, and the weight of the audio data = 0.2; In a quiet environment, = 0.3, = 0.7.
[0064] Use an adaptive algorithm (such as a gradient descent-based method) to adjust the weights. With the cross-entropy loss function: ; ( is the true label) as the optimization objective, calculate the gradients of the loss function with respect to the weights and through backpropagation, and update the weights: ; ; Among them, is the learning rate. The fused probability of violence is , compare it with the set violence threshold . If p > , it is determined that a violent behavior has occurred, thus realizing the detection of violent behavior under multi-modal data fusion.
[0065] In a violent behavior detection system based on multi-modal data fusion, the video data is processed by the GCN model to output the probability of violence , and the audio data is processed by the 2DCNN model to output the probability of violence p2. To give full play to the advantages of the two modal data, it is necessary to reasonably set the weights and perform adaptive adjustment. The specific steps are as follows: Set the initial weights. According to the characteristics of different scenarios, assign initial weights to the detection results of GCN and 2DCNN. The scene characteristics and human perception habits determine the reliability differences of different modal data in detecting violent behavior.
[0066] Noisy environment: In scenarios with complex background sounds such as construction sites and markets, the characteristics of violent behavior in the audio are easily masked by environmental noise. At this time, the video data is relatively more reliable. Set the weight of the output probability of the video GCN = 0.8, and the weight of the output probability of the audio 2DCNN = 0.2. This weight assignment highlights the dominant role of video data in capturing the visual features of violent behavior in a noisy environment.
[0067] Quiet environment: In quiet scenes such as meeting rooms and libraries, the sounds of violent behavior (quarreling sounds, crashing sounds) in the audio will be more obvious, and the audio data is more valuable for reference. Set = 0.3, = 0.7 emphasizes the importance of audio data in detecting violent behavior in a quiet environment.
[0068] This initial weight setting based on scene classification is formulated according to the perceptual characteristics of human beings for visual and auditory information in different environments, as well as the differences in the manifestation forms of violent behavior in different environments, providing a reasonable starting point for subsequent weight optimization.
[0069] An adaptive algorithm based on gradient descent is used to dynamically adjust the weights to adapt to the actual data and model prediction situation, further optimizing the fusion effect.
[0070] Select the cross-entropy loss function L as the optimization objective, and the formula is: ; Among them, is the true label ( corresponding to the video modality, corresponding to the audio modality, with a value of 0 or 1, representing whether it is violent behavior), and pi is the predicted probability after fusion.
[0071] The fused violent probability p is obtained by weighted summation of the output probabilities of GCN and 2DCNN, and the formula is: ; Calculate the gradients of the loss function L with respect to the weights and through the backpropagation algorithm, and then update the weights according to the gradient descent method. The formula is: ; ; Among them, α is the learning rate, which is a preset hyperparameter used to control the step size of weight update.
[0072] Compare the fused violent probability p with the set violent threshold to determine whether a violent behavior has occurred.
[0073] The threshold moving method is a method of gradually changing the value of the violent threshold and observing the performance of the model on the validation set to find a better threshold. The specific operation process is as follows: First, determine an initial threshold range, for example, from 0.1 to 0.9, and take values at a certain step size (such as 0.05).
[0074] For each value, use it as the current brute - force threshold , use the trained model to predict the validation set data, and calculate the corresponding evaluation metrics, such as accuracy, recall, F1 - score, etc.
[0075] Accuracy , where TP represents the number of samples correctly predicted as violent behavior, TN represents the number of samples correctly predicted as normal conditions, FP represents the number of samples wrongly predicted as violent behavior, and FN represents the number of samples wrongly predicted as normal conditions.
[0076] Recall 、 , where ; Optimize the brute - force threshold based on the loss function , which is to incorporate the threshold into the optimization process of the loss function and adjust it together with the weights 、 . The specific approach is to expand the original loss function L into a new loss function L′ that includes the threshold Pthresh. To construct the new loss function, a penalty mechanism in different situations needs to be considered.
[0077] When the predicted probability p is greater than the threshold , but the true label is normal (i.e., wrongly predicted as violent behavior), a certain penalty is given; when the predicted probability p is less than or equal to the threshold , but the true label is violent (i.e., wrongly predicted as normal), a corresponding penalty is also given. Assuming the penalty coefficients are and , then the new loss function L′ can be expressed as: ; where, and represent whether the condition holds, with 1 for holding and 0 for not holding. Update these three parameters simultaneously according to the gradient descent method: ; ; ; where, and are the learning rates of the weight and the threshold respectively, is the loss function , are two different weights, The threshold for backpropagation. If p > , it is determined that a violent act has occurred; if p ≤ , it is determined to be a normal situation.
[0078] In summary, through reasonable initial weight setting, adaptive weight adjustment, and scientific violent threshold setting and optimization methods, the detection results of GCN and 2DCNN can be effectively integrated, improving the accuracy and reliability of violent behavior detection.
[0079] The system initializes the weights of modal data in different environments and continuously optimizes the weights and thresholds using an adaptive algorithm and loss function, thus effectively integrating video and audio information and providing a reliable and flexible solution for violent behavior detection.
[0080] Accuracy rate:
[0081] Recall rate:
[0082] F1 value:
[0083] Tested on the public datasets ViolentCrowd and Hockfights: Video data: Includes 2000 normal / violent segment video data Audio data: Includes 1500 annotated violent audios.
[0084] Table 2: Performance comparison of each method on the test set (%):
[0085] The accuracy rate is 5.9% higher than the best baseline (HUNet). The recall rate is 6.8% higher. The F1-score is 6.3% higher.
[0086] The multi-modal violent detection method proposed by the present invention has the following significant advantages compared with the prior art: Through the dual-modal analysis of fusing video GCN and audio 2DCNN, it achieves an accuracy rate of 93.5% on the public test set, which is 5 - 7 percentage points higher than traditional single-modal methods (such as HUNet 87.6%, 3DCNN 86.5%).
[0087] The video module uses multi-camera three-dimensional reconstruction (accuracy 2.1mm) and still maintains an accuracy rate of 91.2% in the 40% occlusion scenario. The audio module processes through a Mel filter bank and achieves a recall rate of 89.7% in the -5dB low signal-to-noise ratio environment The processing latency of the optimized hybrid architecture is <200ms, which is three times faster than similar Transformer solutions and meets the requirements of real-time monitoring.
[0088] The following shows the point cloud detected by OpenPose: In Figure 8 when two people are fighting and there is occlusion, it is impossible to judge the distance of the contact part, which will result in misjudgment. However, Figure 9 and Figure 10 it can be detected and predicted from the other two sides, and the contact distance of the nearest point can be judged from multiple dimensions to avoid misjudging non-violent actions such as rough play and hugging.
[0089] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present invention.
Claims
1. A method for detecting violent events in audio and video after 3D reconstruction based on multiple cameras, characterized in that: The method includes: S1, obtain video data through multiple cameras and combine it with audio data for time series alignment; S2, performing three-dimensional reconstruction of key points of the human body based on the video data to generate three-dimensional coordinate data; S3, extracting features from the three-dimensional coordinate data and the audio data respectively, and constructing video and audio detection models; S4. Outputting the probability of violent behavior in the video and audio through the detection model; Setting the initial weights of the video and audio probabilities according to scene characteristics, and obtaining the fused violence probability through adaptive adjustment; S5. Judging the fused violence probability by a preset threshold value to determine whether a violent behavior occurs.
2. The violent event detection method for the audio and video after three-dimensional reconstruction based on multiple cameras according to claim 1, characterized in that: In step S1, multiple cameras are used to cover the detection area to obtain video data; the multiple cameras are arranged in a triangle to cover the violence detection area to the greatest extent; Get audio data through the camera's built-in audio acquisition device; Using timestamps to time-align the video data and the audio data to ensure consistency of data from different sources on the time axis; Specifically, absolute timestamps are used to match video information captured by different cameras with audio information captured by microphones; With respect to the video data, adjusting the camera position to form a stereoscopic coverage perspective and obtain multi-angle visual information; For the audio data, environmental sound information is recorded to assist subsequent processing.
3. The violent event detection method for the audio and video after 3D reconstruction based on multiple cameras according to claim 1, characterized in that: In step S2, the intrinsic parameter matrix and extrinsic parameter matrix of the camera are calculated by a calibration method; Extracting two-dimensional key point coordinates from the video data using a key point detection algorithm; Matching the coordinates of the two-dimensional key points through geometric constraints between multiple cameras; Constructing a projection equation group according to the matching results, and using an optimization algorithm to solve and obtain the three-dimensional coordinate data of key points of the human body; Error correction is performed on the three-dimensional coordinate data to ensure reconstruction accuracy.
4. The violent event detection method for the audio and video after 3D reconstruction based on multiple cameras according to claim 3, characterized in that: Specific steps for: Calculate the camera's intrinsic and extrinsic matrix through calibration methods; Specifically, a black and white chessboard or circular array is set up, and the Zhang Zhengyou calibration method is used to calculate the intrinsic parameter matrix of each camera. The rotation matrix and translation vector are solved by arranging known three-dimensional coordinate feature points and combining the least squares method. If the internal and external parameter errors are large, they are calculated again until the error is reduced below the threshold; The intrinsic calibration uses each camera to shoot at least 10 chessboard images from different angles and positions, uses the built-in library of Matlab to input the square size to extract the corner points, and calculates the reprojection error through the camera model. The calculation formula is: ; Use the libraries of Matlab to calculate the internal and external parameter matrices. The expression of the internal parameter matrix is as follows: ; Among them, and respectively represent the focal lengths of the camera in the horizontal and vertical directions, and are used to describe the scaling degree of the object imaging by the camera; and are the principal point coordinates, representing the intersection position of the optical axis and the image plane in the image plane, and the unit is pixel; The external parameter calibration obtains the projection equation by combining the conversion relationship between the world coordinate system and the camera coordinate system with the internal parameter matrix: ; where \((u, v)\) are the two-dimensional coordinates of point \(P\) on the image plane, is the aforementioned intrinsic parameter matrix, the rotation matrix \(R\) is a \(3\times3\) orthogonal matrix used to describe the rotation angle of the camera coordinate system relative to the world coordinate system; \(T\) is the translation vector used to represent the position of the camera in the world coordinate system, are the correspondences of the coordinates of point \(P\) in the camera coordinate system respectively; A key point detection algorithm is used to extract two-dimensional key point coordinates from the video data; specifically, openpose is used to obtain some key point coordinates; The coordinates of the two-dimensional key points are matched by geometric constraints between multiple cameras; the key points are matched by epipolar constraints, and at least one pair of epipolar lines must be matched for each point of at least three cameras, and then two-by-two verifications are performed, and finally the corresponding matching pairs are output after successful verification; Constructing a projection equation group according to the matching results, and using an optimization algorithm to solve and obtain the three-dimensional coordinate data of key points of the human body; The projection equations corresponding to at least three cameras are combined to form an overdetermined system of equations, and the overdetermined system of equations is solved using the least square method to obtain the predicted two-dimensional coordinates of the projection; Let the observed two-dimensional image coordinates be , where correspond to three cameras respectively, and the predicted coordinates calculated according to the projection equation are , then the error function is: ; Then, take the partial derivatives of the error function with respect to , , , set the partial derivatives to 0, and substitute to obtain the estimated values of the coordinates of the three-dimensional points ; Use the Levenberg-Marquardt algorithm for iterative optimization and calculate the update amount of the coordinates by solving the equation ; Error correction is performed on the three-dimensional coordinate data to ensure reconstruction accuracy.
5. The violent event detection method for audio and video after 3D reconstruction based on multiple cameras according to claim 1, characterized in that: In step S3, the three-dimensional coordinate data and the audio data are respectively subjected to feature extraction to construct a video and audio detection model, including: Input the three-dimensional coordinate data as node features into a graph convolutional network model, and construct edge features based on human skeleton relationships; The node features and edge features are fused and processed, and the behavior features in the video are extracted through convolution operation; Performing voiceprint differentiation and spectrum feature conversion on the audio data to generate spectrum graph data; The spectrogram data is input into a two-dimensional convolutional neural network model, and abnormal features in the audio are extracted through multi-layer convolution and pooling operations.
6. The violent event detection method for the audio and video after 3D reconstruction based on multiple cameras according to claim 5, characterized in that: The specific steps include: Input the three-dimensional coordinate data as node features into a graph convolutional network model, and construct edge features based on human skeleton relationships; The input matrix plus the contact velocity, the contact velocity is calculated as: For key part key point pairs with a distance less than the threshold , calculate the velocity between adjacent frames; Suppose at the th frame, the individual , hand key point coordinates are respectively , . At the th frame, the individual , hand key point coordinates become , . If the frame rate is fps , then the time interval is , and the displacement is m / s; Fuse the node features and edge features, and extract the behavior features in the video through convolution operations; adopt the graph convolution formula Perform convolution processing, then use the Relu activation function for activation, and perform convolution again to extract information; Perform voiceprint discrimination and spectral feature conversion on the audio data to generate spectrogram data; first perform voiceprint discrimination on the audio data, and then perform Mel-scale conversion operation to convert linear frequency to Mel frequency The formula is: , construct a Mel filter bank containing 40 - 80 triangular filters, and perform spectral analysis on the audio signal , the output of the th Mel filter is calculated as follows: ; The output of the Mel filter bank is logarithmically transformed to obtain the Mel spectrum graph, the formula is: ; wherein , which is used to avoid zero values in logarithmic operations; Input the spectrogram data into a two-dimensional convolutional neural network model, and extract abnormal features in the audio through multi-layer convolution and pooling operations; the network architecture of the two-dimensional convolutional neural network model includes a convolutional layer, a max pooling layer, a fully connected layer, and an output layer. The output formula of the convolutional layer is: , the pooling layer uses max pooling, and the output formula is: , the fully connected layer integrates the features extracted by the convolutional layer and the pooling layer, and the output layer outputs the probabilities of violence and normal conditions through the Softmax function. The formula is: ; wherein, is the i-th element of the output vector of the fully connected layer, K is the number of categories, is the probability of belonging to the i-th category.
7. The method for detecting violent events of the audio and video after 3D reconstruction based on multiple cameras according to claim 1, wherein: In step S4, the three-dimensional coordinate data is processed by a graph convolutional network model to output a probability value of violent behavior in the video; Processing the audio data through a two-dimensional convolutional neural network model to output a probability value of violent behavior in the audio; Normalizing the video probability value and the audio probability value to ensure consistency of probability distribution; The detection results of different modal data are recorded according to the probability values to provide a basis for subsequent fusion.
8. The violent event detection method for the audio and video after 3D reconstruction based on multiple cameras according to claim 7, characterized in that: In step S4, initial weight values of the video probability and the audio probability are set according to the scene noise level; The initial weight value is optimized and adjusted by a loss function, and the weight is dynamically updated by a gradient descent method; Performing a weighted summation of the video probability and the audio probability according to the adjusted weight to obtain a fused violence probability; The fused violence probability is continuously monitored, and the changing trend during the weight adjustment process is recorded.
9. The method for detecting violent events of the audio and video after 3D reconstruction based on multiple cameras according to claim 8, characterized in that: Specific steps for: Setting initial weight values of the video probability and the audio probability according to the scene noise level; Optimize and adjust the initial weight value through a loss function, and dynamically update the weight using the gradient descent method; select the cross-entropy loss function as the optimization objective, and calculate the gradient of the loss function with respect to the weight and through backpropagation. The weight update formula is: , ; Weighted sum the video probability and the audio probability according to the adjusted weights to obtain the fused violence probability; the fused violence probability ; Compare it with the set violence threshold If p > , it is determined that a violent act has occurred, thus realizing the detection of violent acts under multi-modal data fusion.
10. The method for detecting violent events of the audio and video after three-dimensional reconstruction based on multiple cameras according to claim 1, characterized in that: In step S5, the fused violence probability is judged by a preset threshold to determine whether a violent behavior occurs, including: The initial violence probability threshold is determined using the threshold shift method; Perform performance evaluation on the initial brute-force probability threshold using verification data and adjust the threshold range; for each value, use it as the current brute-force threshold , use the trained model to predict the verification set data and calculate the corresponding evaluation metrics; Including accuracy, recall rate, and F1 value, specifically: accuracy , recall rate , , where ; comparing the fused violence probability with the adjusted threshold; If the fused violence probability is greater than the threshold, it is determined to be a violent behavior; If the fused violence probability is less than or equal to the threshold, it is determined to be a normal situation; Record the judgment results and trigger the corresponding response mechanism.
Citation Information
Patent Citations
Behavior recognition method based on ensemble learning method fused with time attention graph convolution
CN114708649A
Video image violent behavior detection model and detection method
CN116189286A
Violent behavior recognition method and system based on convolutional neural network
CN118262408A
Model training method and device, information determination method and device, equipment and computer readable storage medium
CN119048957A
Violent behavior detection and tracking method based on audio and video
CN119541042A