Audio and video stream multi-mode real-time anomaly detection method and system based on intelligent feature tracking

Through the multi-modal real-time abnormality detection method of audio and video streams, combined with video and audio feature point analysis, the real-time and accuracy problems of the existing monitoring system are solved, and timely capture and complete data storage of abnormal events are achieved.

CN120259688AActive Publication Date: 2025-07-04BROAD VISION (XIAMEN) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510734076.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The existing monitoring systems rely on manual monitoring or simple motion detection algorithms, resulting in insufficient real-time performance, low detection accuracy, inability to fully reflect on the on-site situation, and difficult data management, and difficulty in capturing abnormal changes and saving key data in a timely manner.

Method used

The multi-modal real-time anomaly detection method of audio and video streams based on intelligent feature tracking is used to determine video anomaly through grayscale conversion, feature point extraction and global alignment, combining feature point displacement and local grayscale differences, and combining audio energy ratio, spectrum center difference and Frobenius norm difference to determine audio and video multi-modal comprehensive judgment.

Benefits of technology

It improves the adaptability and accuracy of abnormal detection, can timely capture subtle changes in the video, enhances the comprehensiveness and reliability of real-time monitoring, and ensures the complete storage of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259688A_ABST
    Figure CN120259688A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of audio and video processing, and discloses an audio and video stream multi-mode real-time anomaly detection method and system based on intelligent feature tracking, and the method comprises the steps: carrying out the gray-scale map conversion, feature point extraction and global alignment, precise matching of a current frame and a historical frame, and judgment of abnormal feature points through the combination of feature point displacement and local gray scale difference; and subtle changes in the video can be effectively captured. Based on a dynamic threshold value, video abnormity is judged by combining the number of abnormal, disappearing and newly-added feature points, and the adaptability and accuracy of detection are enhanced. Meanwhile, sampling parameters are set, framing and STFT processing are carried out, a Mel filter bank is used for extracting a logarithmic Mel spectrogram, and through multi-index comparison of the energy ratio, the spectrum center difference value and the Frobenius norm difference, the audio abnormity is comprehensively analyzed. Audio and video are combined in multiple modes, comprehensive judgment is carried out from different dimensions, the comprehensiveness and reliability of anomaly detection are improved, and the method is suitable for scenes such as real-time monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio - video processing, and relates to a multi - modal real - time anomaly detection method and system for audio - video streams based on intelligent feature tracking. Background Art

[0002] With the continuous improvement of social security requirements and the acceleration of the urbanization process, video surveillance systems have been widely used in the fields of public security, intelligent transportation, industrial monitoring, etc. However, existing surveillance systems mainly rely on manual monitoring or simple motion detection algorithms such as frame difference method, background subtraction method, etc., which expose the following obvious deficiencies in practical applications: Manual monitoring is limited by manpower and working conditions and cannot achieve all - weather and full - time monitoring; at the same time, simple motion detection algorithms have a lag in responding to emergencies in complex environments and are difficult to capture early abnormal changes in the picture in a timely manner. Methods based on simple differences are easily affected by factors such as light changes and background noise, resulting in a high false alarm rate and a large missed alarm rate; at the same time, single - modal (only video or only audio) detection has limitations and cannot comprehensively reflect the on - site situation. And existing systems lack an efficient video data caching and storage mechanism, making it difficult to quickly locate key time periods and image data after an abnormal event occurs, thus affecting the analysis and evidence collection in the later stage of the accident. Therefore, there is an urgent need for an advanced monitoring technical solution that integrates visual and auditory multi - modal information and has intelligent feature tracking, global motion compensation, and real - time data backtracking, so as to accurately capture tiny changes and give early warnings at the budding stage of abnormal events, while ensuring the complete preservation of post - event data and providing reliable support for safety disposal and subsequent investigations. Summary of the Invention

[0003] The purpose of the present invention is to solve the problems of insufficient real - time performance, low detection accuracy, inability to comprehensively reflect the on - site situation, and difficult data management in the prior art, and provide a multi - modal real - time anomaly detection method and system for audio - video streams based on intelligent feature tracking.

[0004] To achieve the above - mentioned purpose, the present invention adopts the following technical solutions:

[0005] A multi - modal real - time anomaly detection method for audio - video streams based on intelligent feature tracking includes:

[0006] Collect and store audio - video stream data, and select the video stream data of the current frame and any historical frame video stream data stored for conversion to obtain grayscale images;

[0007] Extract feature points from the grayscale images corresponding to the current - frame video stream data and the historical - frame video stream data respectively, and globally align the grayscale image and feature points corresponding to the historical - frame video stream data, and then match the current frame and the globally - aligned historical - frame image;

[0008] Extract local regions from the current frame and the aligned historical frame images, and determine whether each feature point is an abnormal feature point based on the displacement of feature points and the average grayscale difference of local regions at different times;

[0009] Based on the relationship between the number of abnormal feature points, the sum of disappeared feature points and newly added feature points, and a preset dynamic threshold, determine whether the selected video stream data is abnormal;

[0010] Set audio sampling parameters and extract audio data from the selected audio-visual stream data;

[0011] Perform frame splitting and STFT processing on the audio data to obtain an audio spectrum, and extract and convert it into a logarithmic mel spectrogram with the help of a mel filter bank;

[0012] Based on the set time interval , obtain the energy ratio index and spectral centroid difference between the current audio segment and the previous second audio segment; and based on the logarithmic mel spectrogram, obtain the Frobenius norm difference of the logarithmic mel spectrogram;

[0013] Compare the energy ratio index, spectral centroid difference, and Frobenius norm difference with the preset energy ratio threshold, spectral centroid difference threshold, and Frobenius norm difference threshold respectively to determine whether the extracted audio is abnormal sound.

[0014] A further improvement of the present invention lies in:

[0015] Further, perform feature point extraction on the grayscale images corresponding to the current frame video stream data and the historical frame video stream data respectively, and perform global alignment on the grayscale image and feature points corresponding to the historical frame video stream data, and then match the current frame and the globally aligned historical frame image. Specifically:

[0016] Convert the video stream data of the current frame and the video stream data of the historical frame into grayscale images;

[0017] Extract a set of feature points in the current frame based on the corner detection algorithm;

[0018]

[0019] Among them, is the th feature point of the current frame, is the total number of feature points;

[0020] Perform global alignment on the historical frame and its corresponding feature points, so that the current frame is aligned with the historical frame Match in the same coordinate system;

[0021]

[0022]

[0023] wherein, is the th feature point of the historical frame; is the alignment transformation matrix; is the aligned historical frame, is the th feature point of the aligned historical frame.

[0024] Furthermore, extract local regions from the current frame and the aligned historical frame image, and judge whether each feature point is an abnormal feature point based on the displacement of the feature point and the average gray level difference of the local region at different times, specifically:

[0025] The displacement of the feature point at different times is:

[0026]

[0027]

[0028] wherein, is the displacement value of the feature point at different times, is the absolute value of;

[0029]

[0030]

[0031]

[0032] wherein, Generally take the diagonal length of the image or the upper bound set according to the device field of view; is the normalized value of the feature point;

[0033] For each feature point, extract a local Patch of size from the current frame and the aligned historical frame, normalize the gray value to and then calculate:

[0034]

[0035] wherein, represents the average gray level difference of the local region; Indicates the neighborhood image frame centered on the th feature point of the current frame, Indicates the neighborhood image frame centered on the th feature point of the historical frame;

[0036] For each feature point Determine whether it is an abnormal feature point. When the displacement value of the normalized feature point is greater than the displacement anomaly threshold, or the average gray level difference in the local area is greater than the local difference threshold, the current feature point is an abnormal feature point. Specifically:

[0037]

[0038] Among them, is the displacement anomaly threshold, is the local difference threshold; Indicates the abnormal feature point flag bit. When it is 1, it is an abnormal point. When it is 0, it is a non-abnormal point.

[0039] Furthermore, based on the relationship between the number of abnormal feature points, the sum of disappeared feature points and newly added feature points and a preset dynamic threshold, determine whether the selected video stream data is abnormal. Specifically:

[0040] Count the number of abnormal feature points. Specifically:

[0041]

[0042] Among them, is the number of abnormal feature points;

[0043] Calculate the comprehensive total number of anomalies. Specifically:

[0044]

[0045] Among them, is the number of disappeared feature points, that is, the number of feature points that existed in the historical frame but were not matched in the current frame; is the number of newly added feature points, that is, the number of feature points that appeared in the current frame but did not exist in the historical frame;

[0046] Compare the obtained comprehensive total number of anomalies with the preset dynamic threshold. When the comprehensive total number of anomalies is greater than the preset dynamic threshold, determine whether the selected video stream data is abnormal;

[0047] The preset dynamic threshold is

[0048]

[0049] Among them, is the total number of all feature points in the current frame; represents the probability that a single feature point in the current scene is misjudged as abnormal; is the compensation coefficient.

[0050] Further, set the audio sampling parameters, and extract the audio data from the selected audio-visual stream data. Specifically:

[0051] The audio sampling parameters include: sampling rate, quantization bit depth, number of channels, data format;

[0052] Divide the collected audio data to obtain several data segments;

[0053] The length of the data segment is processed with a 1-second time window, and each segment contains 16,000 sampling data;

[0054] The data per second constitutes a time window,

[0055]

[0056] where, represents the th second of the audio sampling data vector, represents the rd th sampling point value in the th second,

[0057] Further, perform frame division and STFT processing on the audio data to obtain the audio spectrum, and extract and convert it into a logarithmic mel spectrogram with the help of a mel filter bank. Specifically:

[0058] Divide the audio data into frames:

[0059]

[0060] where, is the sample index in each frame, indicating which sampling point within the frame, and the range is where, is the length of each frame;

[0061] is the frame number, indicating which frame it is currently, and the range is where is the total number of frames;

[0062] is the frame shift, indicating the interval between the starting points of adjacent frames, in the unit of the number of sampling points;

[0063] Weighted Hamming window:

[0064]

[0065] Among them, is a mathematical symbol;

[0066] Calculate the STFT for each frame:

[0067]

[0068] Among them, is the imaginary unit;

[0069] Use the Mel filter bank to map the STFT to the Mel scale:

[0070]

[0071] Among them, represents the weight of the th Mel filter corresponding to the th frequency point; represents the number of Mel filters;

[0072] Then perform a logarithmic transformation, specifically:

[0073]

[0074] Among them, is a positive constant used to avoid computation problems in logarithmic operations.

[0075] Furthermore, based on the set time interval , obtain the energy ratio index and spectral centroid difference between the current audio segment and the previous -second audio segment, specifically:

[0076] The energy ratio index between the current audio segment and the previous -second audio segment, specifically:

[0077] The root mean square energy of the current audio segment is:

[0078]

[0079] Among them, N is the number of sampled data;

[0080] The root mean square energy of the current segment and the root mean square energy of the previous segment The ratio is

[0081]

[0082] Among them, is a positive constant used to avoid a zero denominator;

[0083] The audio is subjected to FFT to obtain the frequency spectrum , and calculate the center of the frequency spectrum, specifically:

[0084]

[0085] The difference in the center of the frequency spectrum between the current audio segment and the previous second audio segment is:

[0086]

[0087] Based on the logarithmic mel spectrogram, obtain the Frobenius norm difference of the logarithmic mel spectrogram, specifically:

[0088] The current time and the previous second, the Frobenius norm difference of the logarithmic mel spectrogram is , expressed as the Frobenius norm difference of the logarithmic mel spectrogram is

[0089]

[0090] Among them, is the logarithmic mel spectrogram at the current time , is the logarithmic mel spectrogram of the previous second.

[0091] An audio-visual stream multi-modal real-time anomaly detection system based on intelligent feature tracking, including:

[0092] An acquisition module, which acquires and stores audio-visual stream data, and selects the video stream data of the current frame and any historical frame video stream data stored for conversion to obtain a grayscale image;

[0093] A global alignment module, which extracts feature points from the grayscale images corresponding to the video stream data of the current frame and the historical frame video stream data respectively, and globally aligns the grayscale image and feature points corresponding to the historical frame video stream data, and then matches the current frame and the globally aligned historical frame image;

[0094] A first judgment module, which extracts local regions from the current frame and the aligned historical frame image, and judges whether each feature point is an abnormal feature point based on the displacement of the feature points and the average grayscale difference of the local regions at different times;

[0095] A second judgment module, which judges whether the selected video stream data is abnormal based on the relationship between the sum of the number of abnormal feature points, the disappeared feature points and the newly added feature points and a preset dynamic threshold;

[0096] An extraction module, which sets audio sampling parameters and extracts audio data from the selected audio-visual stream data;

[0097] A conversion module, which frames and performs STFT processing on the audio data to obtain an audio spectrum, and extracts and converts it into a logarithmic mel spectrogram by means of a mel filter bank;

[0098] An acquisition module, which, based on a set time interval , acquires the energy ratio index and the spectral centroid difference between the current audio segment and the previous -second audio segment; and based on the logarithmic mel spectrogram, acquires the Frobenius norm difference of the logarithmic mel spectrogram;

[0099] A third judgment module, which respectively compares the energy ratio index, the spectral centroid difference, and the Frobenius norm difference with preset energy ratio thresholds, spectral centroid difference thresholds, and Frobenius norm difference thresholds to judge whether the extracted audio is abnormal sound.

[0100] Compared with the prior art, the present invention has the following beneficial effects:

[0101] Through grayscale conversion, feature point extraction and global alignment, the present invention accurately matches the current frame and the historical frame, combines the displacement of feature points and local grayscale differences to judge abnormal feature points, and can effectively capture subtle changes in the video. Based on dynamic thresholds, combined with the number of abnormal, disappeared, and newly added feature points to judge video anomalies, the adaptability and accuracy of detection are enhanced. At the same time, sampling parameters are set and framed and STFT processed, and a logarithmic mel spectrogram is extracted by means of a mel filter bank. Through multi-index comparison of energy ratio, spectral centroid difference, and Frobenius norm difference, audio anomalies are comprehensively analyzed. The combination of audio-visual multi-modalities comprehensively judges from different dimensions, improves the comprehensiveness and reliability of anomaly detection, and is applicable to scenarios such as real-time monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0103] Figure 1 It is a schematic flowchart of the audio-visual stream multi-modal real-time anomaly detection method based on intelligent feature tracking of the present invention;

[0104] Figure 2Schematic diagram of the audio - video stream multi - modal real - time anomaly detection system based on intelligent feature tracking of the present invention. Detailed implementation manners

[0105] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0106] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0107] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0108] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed during use, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0109] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0110] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set", "installed", "connected", "connected to" are understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0111] The present invention will be further described in detail below with reference to the accompanying drawings:

[0112] See Figure 1 , the present invention discloses a multi-modal real-time anomaly detection method for audio-visual streams based on intelligent feature tracking, including:

[0113] S101, collect and store audio-visual stream data, and select the video stream data of the current frame and any historical frame video stream data stored for conversion to obtain a grayscale image;

[0114] S102, extract feature points from the grayscale images corresponding to the video stream data of the current frame and the historical frame video stream data respectively, and globally align the grayscale image and feature points corresponding to the historical frame video stream data, and then match the current frame and the globally aligned historical frame image;

[0115] Convert the video stream data of the current frame and the video stream data of the historical frame into grayscale images;

[0116] Extract a set of feature points in the current frame based on the corner detection algorithm;

[0117]

[0118] Among them, is the th feature point of the current frame, is the total number of feature points;

[0119] Globally align the historical frame and its corresponding feature points, so that the current frame and the historical frame are matched in the same coordinate system;

[0120]

[0121]

[0122] Among them, is the th feature point of the historical frame; is the alignment transformation matrix; is the globally aligned historical frame, is the th feature point of the globally aligned historical frame.

[0123] S103. Extract local regions from the current frame and the aligned historical frame images, and determine whether each feature point is an abnormal feature point based on the displacement of the feature points and the average gray-scale difference of the local regions at different times;

[0124] The displacement of the feature points at different times is:

[0125]

[0126]

[0127] Among them, is the displacement value of the feature points at different times, is the absolute value of;

[0128] To eliminate the influence of the field of view and image size of different devices, the displacement of the feature points at different times is normalized, specifically:

[0129]

[0130]

[0131] Among them, Generally, the length of the image diagonal or the upper bound set according to the device field of view is taken; is the normalized value of the feature point;

[0132] For each feature point, a local Patch of size is extracted from the current frame and the aligned historical frame, and the gray-scale value is normalized to and then calculated:

[0133]

[0134] Among them, represents the average gray-scale difference of the local region; represents the neighborhood image frame centered on the th feature point in the current frame, represents the neighborhood image frame centered on the th feature point in the historical frame;

[0135] For each feature point it is determined whether it is an abnormal feature point. When the displacement value of the normalized feature point is greater than the displacement anomaly threshold, or the average gray-scale difference of the local region is greater than the local difference threshold, the current feature point is an abnormal feature point. Specifically:

[0136]

[0137] Among them, is the displacement anomaly threshold, is the local difference threshold; represents the anomaly feature point flag bit, when it is 1, it is an anomaly point, when it is 0, it is a non - anomaly point.

[0138] The displacement anomaly threshold is set to 0.05 - 0.15 within the normalization range, and exceeding the threshold is considered motion anomaly. The local difference threshold takes a decimal value in the normalized grayscale range, and the specific value depends on the noise level and scene complexity.

[0139] S104, based on the relationship between the sum of the number of anomaly feature points, disappearing feature points and newly added feature points and a preset dynamic threshold, determine whether the selected video stream data is abnormal;

[0140] Count the number of anomaly feature points, specifically:

[0141]

[0142] Among them, is the number of anomaly feature points;

[0143] The comprehensive total anomaly count, specifically:

[0144]

[0145] Among them, is the number of disappearing feature points, that is, the number of feature points that existed in the historical frame but were not matched in the current frame; is the number of newly added feature points, that is, the number of feature points that appear in the current frame but did not exist in the historical frame;

[0146] Compare the obtained comprehensive total anomaly count with the preset dynamic threshold. When the comprehensive total anomaly count is greater than the preset dynamic threshold, then determine whether the selected video stream data is abnormal;

[0147] The preset dynamic threshold is

[0148]

[0149] Among them, is the total number of all feature points in the current frame; represents the probability that a single feature point in the current scene is misjudged as abnormal; is the compensation coefficient.

[0150] The acquisition method is as follows: Collect a large number of labeled video samples, including normal and abnormal situations. Conduct statistical analysis on the feature points extracted from each frame, and calculate the proportion of misjudged feature points caused by factors such as illumination changes, noise, and camera jitter under normal conditions. For example: After multiple groups of experiments, if it is statistically found that a certain number of feature points are extracted in a frame, and the average proportion of misjudged feature points is 2% - 5%, then The value of

[0151] is between 0.02 and 0.05. If the value is too low, it may underestimate the misjudgment caused by noise, resulting in a high false alarm rate of the system in a noisy environment; if the value is too high, it may lead to over-sensitivity in anomaly determination, misjudging normal fluctuations as anomalies. Therefore, it is generally recommended to obtain an intermediate value through offline experiments and then adaptively adjust according to the actual scenario.

[0152] The reference value range of

[0153] is usually recommended to be between 5 and 15, and the specific value depends on the resolution of the video collected by the system, the number distribution of feature points, and the experimental statistical results.

[0154] The audio sampling parameters include: sampling rate, quantization bit depth, number of channels, and data format;

[0155] Divide the collected audio data to obtain several data segments;

[0156] The length of the data segment is processed with a 1-second time window, and each segment contains 16,000 sampling data; the data per second constitutes a time window, specifically:

[0157]

[0158] where represents the audio sampling data vector at the th second, represents the value of the th sampling point in the th second, is a set of integers, and the value range is within .

[0159] S106, Frame and perform STFT processing on the audio data to obtain an audio spectrum, and extract and convert it into a logarithmic mel spectrogram with the help of a mel filter bank;

[0160] The audio data is segmented into frames, specifically as follows:

[0161]

[0162] Among them, is the sample index in each frame, indicating the position of the sampling point within the frame, and the range is where, is the length of each frame;

[0163] is the frame number, indicating which frame it is currently, and the range is where is the total number of frames;

[0164] is the frame shift, indicating the interval between the starting points of adjacent frames, with the unit being the number of sampling points;

[0165] Weighted Hamming window:

[0166]

[0167] Among them, is a mathematical symbol;

[0168] Calculate the STFT for each frame:

[0169]

[0170] Among them, is the imaginary unit;

[0171] Use the Mel filter bank to map the STFT to the Mel scale:

[0172]

[0173] Among them, represents the weight of the th Mel filter corresponding to the th frequency point; represents the number of Mel filters;

[0174] Then perform logarithmic transformation, specifically as follows:

[0175]

[0176] Among them, is a positive constant used to avoid computation problems in logarithmic operations.

[0177] S107. Based on the set time interval , obtain the current audio segment and the previous The energy ratio index and the spectral centroid difference of the second audio segment; and based on the logarithmic Mel spectrogram, obtain the Frobenius norm difference of the logarithmic Mel spectrogram;

[0178] Based on the set time interval , obtain the energy ratio index and the spectral centroid difference between the current audio segment and the previous second audio segment, specifically:

[0179] The energy ratio index between the current audio segment and the previous second audio segment is specifically:

[0180] The root mean square energy of the current audio segment is:

[0181]

[0182] where N is the number of sampling data;

[0183] The root mean square energy of the current segment and the root mean square energy of the previous segment The ratio is

[0184]

[0185] where is a positive constant used to avoid a zero denominator;

[0186] Perform FFT on the audio to obtain the spectrum , calculate the spectral centroid, specifically:

[0187]

[0188] The spectral centroid difference between the current audio segment and the previous second audio segment is:

[0189]

[0190] The Frobenius norm difference of the logarithmic Mel spectrogram is obtained based on the logarithmic Mel spectrogram, specifically:

[0191] The current time and the previous The Frobenius norm difference of the logarithmic Mel spectrogram between seconds is , expressed as

[0192]

[0193] where is the logarithmic Mel spectrogram at the current time , is the previous Logarithmic Mel spectrogram in seconds; Frobenius norm difference between logarithmic Mel spectrograms It is used to measure the change in the spectral structure between the current and historical windows.

[0194] S108, respectively compare the energy ratio index, spectral centroid difference, and Frobenius norm difference with the preset energy ratio threshold, spectral centroid difference threshold, and Frobenius norm difference threshold to determine whether the extracted audio is abnormal sound.

[0195] If When, it indicates that the energy of the current sound segment has increased significantly; if It is regarded as an obvious change in the spectral distribution; when any index meets the condition, for example Or And Is greater than the preset threshold, and this state lasts for more than the preset minimum abnormal duration Such as 1 - 2 seconds, the system determines it as abnormal sound and triggers an alarm.

[0196] See Figure 2 This invention discloses a multi-modal real-time abnormal detection system for audio and video streams based on intelligent feature tracking, including:

[0197] Conversion module, the conversion module collects and stores audio and video stream data, and selects the video stream data of the current frame and any historical frame video stream data stored for conversion to obtain a grayscale image;

[0198] Global alignment module, the global alignment module extracts feature points from the grayscale images corresponding to the current frame video stream data and the historical frame video stream data respectively, and performs global alignment on the grayscale image and feature points corresponding to the historical frame video stream data, and then matches the current frame and the globally aligned historical frame image;

[0199] First judgment module, the first judgment module extracts local regions from the current frame and the aligned historical frame image, and judges whether each feature point is an abnormal feature point based on the displacement of the feature points and the average grayscale difference of the local regions at different times;

[0200] Second judgment module, the second judgment module judges whether the selected video stream data is abnormal based on the relationship between the sum of the number of abnormal feature points, disappearing feature points, and newly added feature points and the preset dynamic threshold;

[0201] Extraction module, the extraction module sets audio sampling parameters and extracts the audio data in the selected audio and video stream data;

[0202] A conversion module that frames and performs STFT processing on audio data to obtain an audio spectrum, and extracts and converts it into a logarithmic mel spectrogram by means of a mel filter bank;

[0203] An acquisition module that, based on a set time interval , acquires the energy ratio index and the spectral centroid difference between the current audio segment and the previous second audio segment; and based on the logarithmic mel spectrogram, acquires the Frobenius norm difference of the logarithmic mel spectrogram;

[0204] A third judgment module that respectively compares the energy ratio index, the spectral centroid difference, and the Frobenius norm difference with preset energy ratio thresholds, spectral centroid difference thresholds, and Frobenius norm difference thresholds to judge whether the extracted audio is abnormal sound.

[0205] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An audio-visual stream multi-modal real-time anomaly detection method based on intelligent feature tracking, characterized in that including: Collect and store audio-visual stream data, select the video stream data of the current frame and any historical frame video stream data stored, perform conversion to obtain grayscale images; Extract feature points from the grayscale images corresponding to the current frame video stream data and the historical frame video stream data respectively, perform global alignment on the grayscale image and feature points corresponding to the historical frame video stream data, and then match the current frame and the historical frame image after global alignment; Extract local regions from the current frame and the historical frame image after alignment, and judge whether each feature point is an abnormal feature point based on the displacement of the feature points at different times and the average grayscale difference of the local regions; Judge whether the selected video stream data is abnormal based on the relationship between the number of abnormal feature points, the sum of disappeared feature points and newly added feature points and a preset dynamic threshold; Set audio sampling parameters and extract audio data from the selected audio-visual stream data; Perform frame segmentation and STFT processing on the audio data to obtain an audio spectrum, and extract and convert it into a logarithmic mel spectrogram with the help of a mel filter bank; Based on a set time interval , obtain the energy ratio index and the spectral centroid difference between the current audio segment and the previous -second audio segment; and based on the log Mel spectrogram, obtain the Frobenius norm difference of the log Mel spectrogram; Compare the energy ratio index, the spectral center difference and the Frobenius norm difference with the preset energy ratio threshold, spectral center difference threshold and Frobenius norm difference threshold respectively to judge whether the extracted audio is abnormal sound.

2. The method for multi-modal real-time anomaly detection of audio and video streams based on intelligent feature tracking according to claim 1, wherein, The extracting feature points from the grayscale images corresponding to the current frame video stream data and the historical frame video stream data respectively, performing global alignment on the grayscale image and feature points corresponding to the historical frame video stream data, and then matching the current frame and the historical frame image after global alignment is specifically as follows: Convert the video stream data of the current frame and the video stream data of the historical frame into grayscale images; Extract a set of feature points based on the corner detection algorithm in the current frame from it; ; Among them, is the th feature point of the current frame, is the total number of feature points; For historical frames and their corresponding feature points are globally aligned so that the current frame matches the historical frame in the same coordinate system; ; ; Among them, is the th feature point of the historical frame; is the alignment transformation matrix; is the aligned historical frame, is the th feature point of the aligned historical frame.

3. The method for multi-modal real-time anomaly detection of audio and video streams based on intelligent feature tracking according to claim 2, wherein The extracting local regions from the current frame and the historical frame image after alignment, and judging whether each feature point is an abnormal feature point based on the displacement of the feature points at different times and the average grayscale difference of the local regions is specifically as follows: The displacement of the feature points at different times is: ; ; Among them, is the displacement value of the feature point at different times, is the absolute value of; To eliminate the influence of the field of view and image size of different devices, normalize the displacement of the feature points at different times, specifically as follows: ; ; Among them, generally take the length of the image diagonal or the upper bound set according to the device field of view; is the normalized value of the feature point; For each feature point, local patches of size are extracted from the current frame and the aligned historical frames, and after normalizing the grayscale values to the following calculations are performed: ; Among them, represents the average gray-scale difference of this local area; represents the neighborhood image frame centered on the th feature point in the current frame, represents the neighborhood image frame centered on the th feature point in the historical frame; For each feature point It is determined whether it is an abnormal feature point. When the displacement value of the normalized feature point is greater than the displacement anomaly threshold, or the average gray level difference in the local area is greater than the local difference threshold, the current feature point is an abnormal feature point. Specifically: ; Among them, is the displacement anomaly threshold, is the local difference threshold; represents the flag bit of the abnormal feature point. When it is 1, it is an abnormal point. When it is 0, it is a non-abnormal point.

4. The method for multi-modal real-time anomaly detection of audio and video streams based on intelligent feature tracking according to claim 3, wherein The judging whether the selected video stream data is abnormal based on the relationship between the number of abnormal feature points, the sum of disappeared feature points and newly added feature points and a preset dynamic threshold is specifically as follows: Count the number of abnormal feature points, specifically as follows: ; Among them, is the number of abnormal feature points; Integrate the total number of abnormalities, specifically as follows: ; Wherein, is the number of disappearing feature points, that is, the number of feature points that existed in the historical frame but were not matched in the current frame; is the number of newly added feature points, that is, the number of feature points that appear in the current frame but did not exist in the historical frame; Compare the obtained integrated total number of abnormalities with the preset dynamic threshold. When the integrated total number of abnormalities is greater than the preset dynamic threshold, judge whether the selected video stream data is abnormal; The preset dynamic threshold is ; Among them, is the total number of all feature points in the current frame; represents the probability that a single feature point is misjudged as abnormal in the current scene; is the compensation coefficient.

5. The method for multi-modal real-time anomaly detection of audio and video streams based on intelligent feature tracking according to claim 4, characterized in that The setting audio sampling parameters and extracting audio data from the selected audio-visual stream data is specifically as follows: The audio sampling parameters include: sampling rate, quantization bit depth, number of channels, data format; Divide the collected audio data to obtain several data segments; The length of the data segment is processed with 1 second as a time window, and each segment contains 16000 sampling data; The data per second constitutes a time window, ; in, Indicates A vector of audio sample data of seconds. Indicates Seconds The value of the sampling point, is a set of integers.

6. The method for multi-modal real-time anomaly detection of audio and video streams based on intelligent feature tracking according to claim 5, characterized in that, The performing frame segmentation and STFT processing on the audio data to obtain an audio spectrum, and extracting and converting it into a logarithmic mel spectrogram with the help of a mel filter bank is specifically as follows: Segment the audio data by frame: ; Among them, is the sample index in each frame, indicating the number of the sampling point within the frame, and the range is , where is the length of each frame; is the frame number, indicating which frame it is currently, with the range being , where is the total number of frames; It is the frame shift, representing the interval between the starting points of adjacent frames, with the unit of the number of sampling points; Weighted Hamming window: ; Among them, is a mathematical symbol; Calculate the STFT of each frame: ; wherein, is the imaginary unit; Mapping the STFT to the Mel scale using a Mel filter bank: ; Among them, represents the weight of the th Mel filter corresponding to the th frequency point; represents the number of Mel filters; Then performing a logarithmic transformation, specifically: ; Among them, is a positive constant used to avoid computation problems in logarithmic operations.

7. The method for multi-modal real-time anomaly detection of audio and video streams based on intelligent feature tracking according to claim 6, wherein Based on the set time interval , obtain the energy ratio index and the spectral centroid difference between the current audio segment and the previous -second audio segment, specifically: The energy ratio index of the current audio segment and the previous second audio segment is specifically as follows: The root mean square energy of the current audio segment is: ; Where N is the number of sampled data; Root mean square energy of the current segment Ratio with root mean square energy of the previous segment is ; wherein, is a positive constant used to avoid a zero denominator; Perform FFT on the audio to obtain the frequency spectrum , calculate the center of the frequency spectrum, specifically: ; The spectral center difference between the current audio segment and the previous second audio segment is: ; Based on the logarithmic Mel spectrogram, obtaining the Frobenius norm difference of the logarithmic Mel spectrogram, specifically: Current time and the previous log Mel spectrogram Frobenius norm difference between seconds is and is expressed as ; Among them, is the logarithmic Mel spectrogram at the current time , and is the logarithmic Mel spectrogram of the previous seconds.

8. An audio-video stream multi-modal real-time anomaly detection system based on intelligent feature tracking, characterized in that, Including: An acquisition module that acquires and stores audio-visual stream data, and selects the video stream data of the current frame and any historical frame video stream data stored for conversion to obtain a grayscale image; A global alignment module that extracts feature points from the grayscale images corresponding to the current frame video stream data and the historical frame video stream data respectively, and globally aligns the grayscale image and feature points corresponding to the historical frame video stream data, and then matches the current frame and the globally aligned historical frame image; A first judgment module that extracts local regions from the current frame and the aligned historical frame image, and judges whether each feature point is an abnormal feature point based on the displacement of the feature points at different times and the average grayscale difference of the local regions; A second judgment module that judges whether the selected video stream data is abnormal based on the relationship between the sum of the number of abnormal feature points, disappearing feature points and newly added feature points and a preset dynamic threshold; An extraction module that sets audio sampling parameters and extracts audio data from the selected audio-visual stream data; A conversion module that frames and performs STFT processing on the audio data to obtain an audio spectrum, and extracts and converts it into a logarithmic Mel spectrogram with the help of a Mel filter bank; An acquisition module, the acquisition module based on a set time interval , acquire the energy ratio index and the spectral centroid difference between the current audio segment and the previous second audio segment; and based on the logarithmic mel spectrogram, acquire the Frobenius norm difference of the logarithmic mel spectrogram; A third judgment module that compares the energy ratio index, the spectral center difference and the Frobenius norm difference with the preset energy ratio threshold, spectral center difference threshold and Frobenius norm difference threshold respectively to judge whether the extracted audio is abnormal sound.

Citation Information

Patent Citations

  • Method for detecting abnormal points during video multi-target tracking

    CN107273801A

  • Equipment abnormal sound detection method and device and inspection robot

    CN115206341A

  • Abnormality identification method and device

    CN118799812A

  • Cloud edge multi-mode abnormal event detection system based on dynamic decision-making mechanism

    CN119494068A

  • Equipment abnormal behavior identification method and system based on audio and video streams

    CN119723191A

Cited By

  • Multi-role collaborative customer service agent method

    CN120911508A

  • Video anomaly detection method based on Mask grid stability

    CN121259736A