Tracking identification method based on audio and video feature joint judgment
By combining audio processing and video feature analysis, the problem of accuracy and false alarm rate in fighting recognition in complex environments has been solved, achieving high-precision fighting event recognition.
Patent Information
- Application Number
- CN202511143011.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies are easily affected by lighting, occlusion, and changes in perspective when identifying fights in complex environments, resulting in decreased recognition accuracy, high false alarm rate, and inability to fully reflect the whole picture of the event.
By analyzing sound features in real time through audio processing, and combining video human detection and posture recognition algorithms, a deep learning model is used to jointly judge audio and video features to identify fighting events.
It improves the accuracy of fight detection, reduces the false alarm rate, and comprehensively reflects the whole picture of the event.
Smart Images

Figure CN120877781A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and audio signal processing technology, and in particular to a method for fighting recognition based on joint judgment of audio and video features. Background Technology
[0002] Fights frequently occur in public places such as shopping malls, schools, and streets. Timely identification and prevention of fights are crucial for maintaining social order and ensuring personal safety. Current methods for fighting detection rely on feature analysis of video images, including human posture and movement trajectories.
[0003] However, existing methods for identifying fights are susceptible to factors such as lighting, occlusion, and changes in perspective in complex environments, which can lead to decreased recognition accuracy. Furthermore, while they can accurately capture the audio features of fights, they cannot fully reflect the overall picture of the event, resulting in a high false alarm rate. Summary of the Invention
[0004] The purpose of this invention is to provide a fight recognition method based on joint judgment of audio and video features. It aims to solve the technical problems in the prior art that when recognizing fights, the recognition accuracy is easily affected by factors such as lighting, occlusion and changes in viewing angle in complex environments, which leads to a decrease in recognition accuracy. Furthermore, it is difficult to accurately capture the audio features of fights, and therefore cannot fully reflect the whole picture of the event, resulting in a high false alarm rate.
[0005] To achieve the above objectives, the present invention employs a fight recognition method based on joint judgment of audio and video features, comprising the following steps:
[0006] By analyzing the sounds in the monitored environment in real time through audio processing, it can be determined whether there are any noisy or unusual sounds caused by fighting.
[0007] Perform abnormal sound detection;
[0008] Upon detecting unusual noises, the system captures the current camera feed and applies human detection and posture recognition algorithms to identify multiple individuals within the image.
[0009] The judgment is made by combining audio and video features.
[0010] In the step of analyzing the sound in the monitored environment in real time through audio processing to determine whether there are any noisy or unusual sounds caused by fighting:
[0011] Extract specific features from the audio signal and perform noise reduction and smoothing on the original audio signal. Specific features include frequency, volume, and duration.
[0012] Perform a Fourier transform on the processed signal to obtain the spectrum, calculate the energy distribution of the spectrum, and select the frequency range where the energy is concentrated.
[0013] The volume is obtained by calculating the root mean square value of the signal;
[0014] Based on the set energy threshold, find the start and end times when the signal energy exceeds the threshold, and obtain the duration;
[0015] Compare whether the features meet the preset minimum threshold for fighting audio features to determine whether there are any noisy or abnormal noises.
[0016] In the abnormal sound detection step:
[0017] Identify unusual sounds such as sudden screams, impacts, and fighting sounds in audio recordings as a preliminary basis for judging fighting incidents;
[0018] The audio signal is divided into multiple short-time segments called frames by performing a short-time Fourier transform on the audio signal. The power spectrum of each frame is extracted to represent the energy distribution of the signal at different frequencies. A Mel filter bank is applied to perform a nonlinear transformation on the frequency, and the logarithm of the output of each filter is taken to generate the final log-Mel spectrum.
[0019] A suitable threshold is determined by training log-Mel spectrum data using a Gaussian mixture model and analyzing the abnormal score distribution of normal sounds.
[0020] Once real-time log-Mel spectrogram data is obtained, the anomaly score of the log-Mel spectrum is calculated, and a threshold is used to determine whether a new audio sample is an anomalous sound.
[0021] Among the steps, in the process of detecting unusual noises, capturing the current camera view, and applying human detection and pose recognition algorithms to identify multiple people in the image:
[0022] By using video processing technology, the camera captures real-time images at a frequency of at least 5 frames per second, and preprocesses the images to improve the accuracy of subsequent processing.
[0023] Image quality is improved by denoising the image, and rotation is used to reduce possible errors in subsequent processing.
[0024] Deep learning models based on self-training;
[0025] The optimal model was obtained by training with a mixture of publicly available online datasets and real-world scene datasets. This model was then used to detect people in the footage and continuously track their movements, enabling real-time monitoring of each person's behavior.
[0026] We collected publicly available online datasets and real-world scene datasets related to fights and brawls. We used personnel detection for initial screening, used IOU calculation to filter out data with non-overlapping personnel coordinates, and trained a deep learning model by self-labeling the remaining datasets to detect aggressive physical contact behavior of multiple people in the scene.
[0027] When aggressive physical contact is detected, the system counts the aggressive behavior and determines whether the number of recorded aggressive behaviors has reached the specified threshold. If the number exceeds the specified threshold, it is determined to be a fight.
[0028] In the steps of a deep learning model based on self-training:
[0029] A deep learning model consists of an input segment, a backbone network, a neck network, and a detection head, which is composed of multiple convolutional neural network layers.
[0030] In the step of training using a combination of publicly available online datasets and real-world scene datasets:
[0031] The dataset contains over 100,000 images. During training, a subset training method is used to gradually expand the dataset from a small number to a larger number to obtain phased models. During the training process, subsequent training is continuously adjusted based on the problems exhibited by the phased models to obtain the optimal model.
[0032] The process includes collecting publicly available online datasets and real-world scene datasets related to fights, performing initial screening using personnel detection, filtering data with non-overlapping personnel coordinates using IOU calculation, and training a deep learning model with relevant self-annotations on the remaining dataset to detect aggressive physical contact behavior by multiple individuals in the scene.
[0033] A deep learning model consists of an input segment, a backbone network, a neck network, and a detection head. It is composed of multiple convolutional neural network layers. During training, a subset training method is used to gradually expand the model from a small number to a larger number to obtain a stage model. During the training process, subsequent training is continuously adjusted based on the problems shown by the stage model to obtain the optimal model.
[0034] In the step of jointly judging audio and video features:
[0035] By combining the noise and unusual sounds in the audio features and the aggressive physical contact in the video features, relevant information about the fight can be fully captured.
[0036] By adding unified time information to the audio and video data stream units respectively, the audio and video data are made to correspond in time.
[0037] The collected audio and image data are processed by the algorithm model. When the audio and image recognition results are combined, the time information is aligned. The audio and image results with consistent time information are subjected to AND logic calculation to obtain the final judgment result.
[0038] Among them, in the steps of processing the collected audio and image data by the algorithm model, aligning the time information when combining the audio and image recognition results, and performing AND logic calculation on the audio and image results with consistent time information to obtain the final judgment result:
[0039] Audio features provide direct evidence of fighting, while video features provide a visual representation of aggressive physical contact between multiple people.
[0040] This invention discloses a fight identification method based on joint judgment of audio and video features. First, it analyzes the sound in the monitored environment in real time through audio processing to determine if there are any unusual noises or disturbances caused by a fight, and performs abnormal sound detection. Upon detection of unusual noises, it captures the current camera frame and applies human detection and posture recognition algorithms to identify multiple people in the frame. Then, it combines audio and video features for joint judgment. This method, combining audio and video features for joint judgment, can more comprehensively capture relevant information about fight events. Audio features provide direct evidence of unusual noises, while video features provide a visual demonstration of aggressive physical contact. Furthermore, the detection of continuous aggressive physical contact between multiple people improves the accuracy of fight identification, comprehensively reflects the entire event, and reduces the false alarm rate. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart of the steps of the fight recognition method based on joint judgment of audio and video features of the present invention.
[0043] Figure 2 This is a flowchart of steps S100 of the present invention.
[0044] Figure 3 This is a flowchart of steps S200 of the present invention.
[0045] Figure 4 This is a flowchart of steps S300 of the present invention.
[0046] Figure 5This is a flowchart of steps S400 of the present invention. Detailed Implementation
[0047] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0048] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0049] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0050] Please see Figures 1-5 This invention provides a method for identifying fights based on joint judgment of audio and video features, comprising the following steps:
[0051] S100: Through audio processing, it performs real-time analysis of the sound in the monitored environment to determine whether there are any noisy or unusual sounds caused by fighting.
[0052] In this embodiment, audio processing is used to analyze the sounds in the monitored environment in real time to determine whether there are any unusual noises caused by fighting. The specific process is as follows:
[0053] S101: Extract specific features of the audio signal and perform noise reduction and smoothing on the original audio signal. Specific features include frequency, volume, and duration.
[0054] S102: Perform Fourier transform on the processed signal to obtain the spectrum, calculate the energy distribution of the spectrum, and select the frequency range where the energy is concentrated.
[0055] S103: The volume is obtained by calculating the root mean square value of the signal;
[0056] S104: Based on the set energy threshold, find the start and end times when the signal energy exceeds the threshold, and obtain the duration;
[0057] S105: Compare whether the features meet the preset minimum threshold for fighting audio features to determine whether there are any noisy or abnormal noises.
[0058] In the above process, specific features of the audio signal are extracted, and the original audio signal is denoised and smoothed. The specific features include frequency, volume, and duration. Then, the processed signal is subjected to Fourier transform to obtain the spectrum, the energy distribution of the spectrum is calculated, the frequency range where the energy is concentrated is selected, and the volume is obtained by calculating the root mean square value of the signal. According to the set energy threshold, the start and end times when the signal energy exceeds the threshold are found to obtain the duration. Then, the features (frequency, volume, duration) are compared to see if they meet the preset minimum threshold for fighting audio features to determine whether there is any noise or abnormal sound.
[0059] S200: Perform abnormal sound detection.
[0060] In this embodiment, abnormal sound detection is performed, and the specific process is as follows:
[0061] S201: Identify unusual sounds such as sudden screams, impacts, and fighting sounds in audio as a preliminary basis for judging a fight.
[0062] S202: Perform a short-time Fourier transform on the audio signal to divide the audio into multiple short-time segments, which are called frames. Perform a Fourier transform on each frame to calculate its spectral information and extract the power spectrum of each frame to represent the energy distribution of the signal at different frequencies.
[0063] S203: Apply Mel filter bank to perform nonlinear transformation on the frequency, and take the logarithm of the output of each filter to generate the final log-Mel spectrum.
[0064] S204: Use a Gaussian mixture model to train on log-Mel spectrum data and determine a suitable threshold by analyzing the abnormal score distribution of normal sounds;
[0065] S205: Once real-time log-Mel spectrogram data is obtained, calculate the anomaly score of the log-Mel spectrogram and use a threshold to determine whether a new audio sample is an anomalous sound.
[0066] In the above process, abnormal sounds such as sudden screams, impacts, and fighting sounds are identified in the audio as a preliminary basis for judging fighting events. By performing a short-time Fourier transform on the audio signal, the audio is divided into multiple short-time segments, which are called frames. A Fourier transform is performed on each frame to calculate its spectral information. The power spectrum of each frame is extracted to represent the energy distribution of the signal at different frequencies. A Mel filter bank is applied to perform a nonlinear transformation on the frequency, and the logarithm of each filter output is taken to generate the final log-Mel spectrum. The log-Mel spectrum data is trained in advance using a Gaussian mixture model, and an appropriate threshold is determined by analyzing the abnormal score distribution of normal sounds. When real-time log-Mel spectrum data is obtained, the abnormal score of the log-Mel spectrum is calculated, and the threshold is used to determine whether a new audio sample is an abnormal sound.
[0067] S300: Upon detecting unusual noises, it captures the current camera image and applies human detection and posture recognition algorithms to identify multiple people in the image.
[0068] In this embodiment, upon detecting unusual noises, the current camera view is captured, and human detection and pose recognition algorithms are applied to identify multiple people in the view. The specific process is as follows:
[0069] S301: Through video processing technology, it maintains a frequency of at least 5 frames per second, captures real-time images from the camera, and preprocesses the images to improve the accuracy of subsequent processing.
[0070] S302: Improve image quality by denoising the image and use rotation to reduce possible errors in subsequent processing;
[0071] S303: Deep learning model based on self-training;
[0072] S304: The optimal model is obtained by training a mixture of publicly available online datasets and real-world scene datasets. This model is then used to detect people in the scene and continuously track their movement trajectories, enabling real-time monitoring of each person's behavior.
[0073] S305: Collect publicly available online datasets and real-world scene datasets related to fighting and brawling, use personnel detection for initial screening, use IOU calculation to filter out data with non-overlapping personnel coordinates, and perform relevant self-labeling training on the remaining dataset to obtain a deep learning model, which is used to detect aggressive physical contact behavior of multiple people in the scene.
[0074] S306: When aggressive physical contact is detected, the system counts the aggressive behavior and determines whether the number of recorded aggressive behaviors has reached the specified threshold based on the set number threshold. If the number exceeds the specified number, it is determined to be a fighting event.
[0075] In the above process, video processing technology is used to maintain a frequency of at least 5 frames per second, capturing real-time images from the camera across the entire screen. Preprocessing is performed on the images to improve the accuracy of subsequent processing. Image denoising is used to improve image quality, and rotation is employed to reduce potential errors in subsequent processing. The system is based on a self-trained deep learning model, which includes an input segment, a backbone network, a neck network, and a detection head, composed of multiple convolutional neural network layers. The model is then trained using a mixture of publicly available online datasets and real-world scene datasets, with a total dataset exceeding 100,000 images, to obtain the optimal model. This model is then used to detect people in the footage and continuously track their movement trajectories, providing real-time monitoring of each person's behavior. Based on the personnel detection and tracking achieved above, publicly available online datasets and real-world scene datasets related to fights are collected to enable... Initial screening is performed using personnel detection. IOU calculation is used to filter out data with non-overlapping personnel coordinates, leaving a dataset of at least 100,000 people. The remaining dataset is then used for self-labeling and training to obtain a deep learning model. This model includes an input segment, a backbone network, a neck network, and a detection head, composed of multiple convolutional neural network layers. During training, a subset training approach is used to gradually expand the model from a small number to a larger number, resulting in phased models. During training, subsequent training is continuously adjusted based on the problems exhibited by the phased models to obtain the optimal model, which is used to detect aggressive physical contact behavior among multiple people in the scene. When aggressive physical contact is detected, the system counts the aggressive behavior and determines whether the number of recorded aggressive behaviors has reached the specified threshold. If the number exceeds the specified threshold, it is determined to be a fighting event.
[0076] S400: Combines audio and video features for judgment.
[0077] In this embodiment, audio and video features are combined for judgment. The specific process is as follows:
[0078] S401: By combining the noise and unusual sounds in the audio features and the aggressive physical contact in the video features, it comprehensively captures relevant information about fighting events;
[0079] S402: By adding unified time information to the audio and video data stream units respectively, the audio and video data are made to correspond in time.
[0080] S403: The acquired audio and image data are processed by the algorithm model. When merging the audio and image recognition results, the time information is aligned. The audio and image results with consistent time information are ANDed to obtain the final judgment result.
[0081] In the above process, by combining the noise and abnormal sounds in the audio features and the aggressive physical contact in the video features, the system can comprehensively capture relevant information about fighting events. By adding unified time information (millisecond-level timestamps) to the audio and video data stream units respectively, the audio and video data are made to correspond in time. At the same time, the collected audio and image data are processed by the algorithm model. When the audio and image recognition results are combined, the time information is aligned. The audio and image results with consistent time information are ANDed to obtain the final judgment result. Among them, the audio features provide direct evidence of fighting, while the video features provide an intuitive display of aggressive physical contact between multiple people. This multimodal information fusion method significantly improves the accuracy of fight recognition and reduces the false alarm rate.
[0082] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0083] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for fighting identification based on joint judgment of audio and video features, characterized in that, Includes the following steps: By analyzing the sounds in the monitored environment in real time through audio processing, it can be determined whether there are any noisy or unusual sounds caused by fighting. Perform abnormal sound detection; Upon detecting unusual noises, the system captures the current camera feed and applies human detection and posture recognition algorithms to identify multiple individuals within the image. The judgment is made by combining audio and video features.
2. The fight recognition method based on joint judgment of audio and video features as described in claim 1, characterized in that, In the step of analyzing the sound in the monitored environment in real time through audio processing to determine whether there are any noisy or unusual sounds caused by fighting: Extract specific features from the audio signal and perform noise reduction and smoothing on the original audio signal. Specific features include frequency, volume, and duration. Perform a Fourier transform on the processed signal to obtain the spectrum, calculate the energy distribution of the spectrum, and select the frequency range where the energy is concentrated. The volume is obtained by calculating the root mean square value of the signal; Based on the set energy threshold, find the start and end times when the signal energy exceeds the threshold, and obtain the duration; Compare whether the features meet the preset minimum threshold for fighting audio features to determine whether there are any noisy or abnormal noises.
3. The fight recognition method based on joint judgment of audio and video features as described in claim 1, characterized in that, In the steps of abnormal sound detection: Identify unusual sounds such as sudden screams, impacts, and fighting sounds in audio recordings as a preliminary basis for judging fighting incidents; The audio signal is divided into multiple short-time segments called frames by performing a short-time Fourier transform on the audio signal. The power spectrum of each frame is extracted to represent the energy distribution of the signal at different frequencies. By applying a Mel filter bank, a nonlinear transformation of the frequency is performed, and the logarithm of the output of each filter is taken to generate the final log-Mel spectrum. A suitable threshold is determined by training log-Mel spectrum data using a Gaussian mixture model and analyzing the abnormal score distribution of normal sounds. Once real-time log-Mel spectrogram data is obtained, the anomaly score of the log-Mel spectrum is calculated, and a threshold is used to determine whether a new audio sample is an anomalous sound.
4. The fight recognition method based on joint judgment of audio and video features as described in claim 1, characterized in that, In the steps of detecting unusual noises, capturing the current camera view, and applying human detection and pose recognition algorithms to identify multiple people in the image: By using video processing technology, the camera captures real-time images at a frequency of at least 5 frames per second, and preprocesses the images to improve the accuracy of subsequent processing. Image quality is improved by denoising the image, and rotation is used to reduce possible errors in subsequent processing. Deep learning models based on self-training; The optimal model was obtained by training with a mixture of publicly available online datasets and real-world scene datasets. This model was then used to detect people in the footage and continuously track their movements, enabling real-time monitoring of each person's behavior. We collected publicly available online datasets and real-world scene datasets related to fights and brawls. We used personnel detection for initial screening, used IOU calculation to filter out data with non-overlapping personnel coordinates, and trained a deep learning model by self-labeling the remaining datasets to detect aggressive physical contact behavior of multiple people in the scene. When aggressive physical contact is detected, the system counts the aggressive behavior and determines whether the number of recorded aggressive behaviors has reached the specified threshold. If the number exceeds the specified threshold, it is determined to be a fight.
5. The fight recognition method based on joint judgment of audio and video features as described in claim 4, characterized in that, In the steps of a self-trained deep learning model: A deep learning model consists of an input segment, a backbone network, a neck network, and a detection head, which is composed of multiple convolutional neural network layers.
6. The fight recognition method based on joint judgment of audio and video features as described in claim 5, characterized in that, In the step of training using a combination of publicly available online datasets and real-world scene datasets: The dataset contains over 100,000 images. During training, a subset training method is used to gradually expand the dataset from a small number to a larger number to obtain phased models. During the training process, subsequent training is continuously adjusted based on the problems exhibited by the phased models to obtain the optimal model.
7. The fight recognition method based on joint judgment of audio and video features as described in claim 6, characterized in that, In the process of collecting publicly available online datasets and real-world scene datasets related to fights, using personnel detection for initial screening, using IOU calculation to filter out data with non-overlapping personnel coordinates, and then training a deep learning model with relevant self-labeled annotations on the remaining dataset to detect aggressive physical contact behavior among multiple people in the scene: A deep learning model consists of an input segment, a backbone network, a neck network, and a detection head. It is composed of multiple convolutional neural network layers. During training, a subset training method is used to gradually expand the model from a small number to a larger number to obtain a stage model. During the training process, subsequent training is continuously adjusted based on the problems shown by the stage model to obtain the optimal model.
8. The fight recognition method based on joint judgment of audio and video features as described in claim 1, characterized in that, In the step of combining audio and video features for judgment: By combining the noise and unusual sounds in the audio features and the aggressive physical contact in the video features, relevant information about the fight can be fully captured. By adding unified time information to the audio and video data stream units respectively, the audio and video data are made to correspond in time. The collected audio and image data are processed by the algorithm model. When the audio and image recognition results are combined, the time information is aligned. The audio and image results with consistent time information are subjected to AND logic calculation to obtain the final judgment result.
9. The fight recognition method based on joint judgment of audio and video features as described in claim 8, characterized in that, In the process of feeding the collected audio and image data to the algorithm model for processing, aligning the time information when combining the audio and image recognition results, and performing AND logic calculations on audio and image results with consistent time information to obtain the final judgment result: Audio features provide direct evidence of fighting, while video features provide a visual representation of aggressive physical contact between multiple people.