Modal fusion intelligent video abnormal behavior monitoring method

Through the multimodal data acquisition and hierarchical structure anomaly behavior analysis model, the problems of insufficient frame rate and insufficient audio information utilization in dynamic scenarios are solved, and more accurate abnormal behavior judgment and resource optimization are achieved.

CN120340129AInactive Publication Date: 2025-07-18ZHEJIANG COLLEGE OF CONSTR
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510399958.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional intelligent video abnormal behavior monitoring systems have insufficient frame rate in dynamic scenarios, resulting in unsuccessful behavior details, wasted resources in static scenarios, insufficient audio information utilization, and the model cannot effectively integrate multimodal data, resulting in misjudgment or misjudgment.

Method used

Multimodal data acquisition, intelligent frame rate adjustment of video cameras, combined with audio and environmental state data, deep fusion analysis is performed through an abnormal behavior analysis model of hierarchical structure, and multimodal data is processed using bilateral filtering, short-time Fourier transform and bidirectional long-term short-term memory network to ensure data quality and time synchronization.

Benefits of technology

It improves the accuracy of abnormal behavior judgments, reduces the rate of misjudgment and misjudgment, optimizes resource utilization, and enhances adaptability and information comprehensiveness to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340129A_ABST
    Figure CN120340129A_ABST
Patent Text Reader

Abstract

The invention discloses a modal fusion intelligent video abnormal behavior monitoring method, and relates to the technical field of intelligent security monitoring, and the method comprises the following steps: collecting video modal data through a video camera, collecting sound modal data through an audio sensor, and obtaining environment state modal data through an environment state sensor; denoising processing is carried out on video data, noise reduction and filtering operation is carried out on sound data, and normalization processing is carried out on environment state data, so that the data availability is improved. By constructing an abnormal behavior analysis model and adopting a unique layered structure, compared with a traditional simple model, the complex incidence relation between multi-modal data can be effectively processed, and in a scene where personnel gather and are accompanied with noisy sound, the model can integrate personnel actions and postures in a video and sound features in an audio frequency, so that the model can be used for analyzing the abnormal behavior. Whether abnormal behaviors such as noisy and fighting occur or not is accurately judged, the accuracy of abnormal behavior judgment is greatly improved, and the misjudgment rate and the missed judgment rate are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent security monitoring, and specifically to an intelligent video abnormal behavior monitoring method based on modal fusion. Background Art

[0002] Intelligent video abnormal behavior monitoring based on modal fusion is dedicated to integrating multiple data modalities such as video, audio, and environmental status, accurately identifying and warning abnormal behaviors in the scene, and providing strong support for fields such as security prevention and public order maintenance.

[0003] In traditional intelligent video abnormal behavior monitoring solutions, first, in the video data acquisition link, most devices use fixed parameter settings. For example, the camera frame rate is constant. Whether the scene is a static warehouse corner or a dynamic crowded activity place, the pictures are collected at the same frequency. This results in the inability to fully capture behavior details due to insufficient frame rate in dynamic scenes. For example, abnormal movements that quickly intersperse in the crowd are easily overlooked; while in static scenes, too high a frame rate causes waste of storage resources.

[0004] In terms of the utilization of audio data, traditional monitoring often only focuses on the presence or absence of sound, and insufficiently mines the rich information carried by the sound. For example, in public areas, abnormal shouts, impacts, etc., their frequency distribution, volume change, and duration and other characteristics all contain key clues for judging abnormal behaviors. However, traditional technologies lack effective means for extracting and analyzing these audio characteristics and cannot combine audio information with video pictures to comprehensively judge abnormal situations in the scene.

[0005] From the perspective of model construction, traditional abnormal behavior analysis model structures are simple, mostly classification models based on shallow neural networks or single features. These models are difficult to handle the complex correlation relationships between multi-modal data. For example, when there is a crowd gathering in the video picture and there is noisy sound in the audio at the same time, traditional models cannot effectively integrate these two modal information to accurately judge whether an abnormal event has occurred, and are prone to false positives or false negatives. Summary of the Invention

[0006] To solve the above technical problems, an intelligent video abnormal behavior monitoring method based on modal fusion is provided, and this technical solution solves the above problems.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] An intelligent video abnormal behavior monitoring method based on modal fusion includes the following steps:

[0009] Collect video modal data using a video camera, and at the same time collect sound modal data with the help of an audio sensor and obtain environmental status modal data through an environmental status sensor;

[0010] Denoise the video data, perform noise reduction and filtering operations on the audio data, and normalize the environmental state data to improve the data availability;

[0011] Extract the visual features of the key frame images from the video data, obtain the acoustic features from the audio data; extract the change trend features of the environmental parameters from the environmental state data;

[0012] Input the multi-modal features into the abnormal behavior analysis model. The abnormal behavior analysis model is designed with a hierarchical structure to achieve the deep fusion and analysis of multi-modal features, so as to judge whether there is an abnormal behavior

[0013] Preferably, the video camera has an intelligent frame rate adjustment function, which automatically adjusts the frame rate according to the dynamic degree of the scene. The adjustment formula is: frame rate = base frame rate + dynamic degree coefficient × scene change amount, where the scene change amount is calculated through the difference between adjacent video frames.

[0014] Preferably, the bilateral filtering algorithm is used for denoising the video data. This algorithm takes into account both the pixel spatial distance and the pixel value difference while filtering. The filtering formula is:

[0015]

[0016] Among them, g(i,j) is the filtered pixel value, f(m,n) is the original pixel value, and w(i,j,m,n) is the weighting coefficient.

[0017] Preferably, when extracting features from the audio data, the short-time Fourier transform is used. The formula is:

[0018]

[0019] Among them, x(m) is the audio time-domain signal, w(m) is the window function, and N is the number of transform points.

[0020] Preferably, in the hierarchical structure of the abnormal behavior analysis model, the input layer groups and inputs features according to the modal categories. The middle hidden layer uses a bidirectional long short-term memory network module to learn the long-term dependence relationship of the feature sequence through forward and backward propagation. The unit state update formula is:

[0021] i t =σ(W ii x t +b ii +W hi h t-1 +b hi )

[0022] f t =σ(W ifx t +b if +W hf h t-1 +b hf )o t =σ(W io x t +b io +W ho h t-1 +b ho )

[0023]

[0024] h t =o t ⊙tanh(C t )

[0025] Among them, i t 、f t 、o t are the values of the input gate, forget gate, and output gate respectively, C t is the candidate memory cell and memory cell, h t is the hidden state, W and v are weights and biases, σ is the activation function, and ⊙ is element-wise multiplication.

[0026] Preferably, in the preprocessing of the environmental state data, linear interpolation is performed on the temperature data to fill in the missing values, and the interpolation formula is:

[0027]

[0028] Among them, T is the interpolated temperature value, T1 and T2 are adjacent known temperature values, t1 and t2 are the corresponding times, and t is the time point to be interpolated.

[0029] Preferably, when the abnormal behavior analysis model determines that there is an abnormal behavior, the alarm system is immediately triggered, and the alarm information is sent to the security personnel terminal in an encrypted form through the network communication module.

[0030] Preferably, the abnormal behavior analysis model is periodically optimized and updated through newly collected multimodal data, and the optimization algorithm used is the Adagrad algorithm. The model parameter update formula is:

[0031]

[0032] In the formula, θ j is the model parameter, η is the learning rate, G ij is the diagonal element of the sum of squared gradients, ∈ is a small constant to prevent division by zero, is the gradient of the loss function with respect to the parameter θ j .

[0033] Preferably, to ensure the time synchronization of multi-modal data, precise timestamps are added to each modal data during data acquisition, and the data is aligned according to the timestamps during the data preprocessing stage.

[0034] Preferably, when extracting key frames from video data, a key frame extraction algorithm based on motion estimation is adopted. Whether to extract key frames is judged by calculating the motion vectors between video frames. The block matching algorithm is used for motion vector calculation, and the matching error calculation formula is:

[0035]

[0036] where MAD(x,y) is the matching error, f1 and f2 are adjacent video frames, (x,y) is the position within the search window, and M and N are the sizes of the blocks.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] 1. This method collects multi-modal data through video cameras, audio sensors, and environmental state sensors, making up for the deficiencies of traditional single video modality. The intelligent frame rate adjustment function of the video camera can automatically adjust the frame rate according to the dynamic degree of the scene, and can capture details more accurately in complex scenes, solving the problem that traditional fixed frame rates cannot adapt to complex scenes.

[0039] 2. The constructed abnormal behavior analysis model adopts a unique hierarchical structure, and the middle hidden layer uses a bidirectional long short-term memory network module. By learning the long-term dependence relationship of the feature sequence through forward and backward propagation (a series of unit state update formulas), compared with traditional simple models, it can effectively process the complex correlation relationships between multi-modal data. In scenes where people gather and there is noise, the model can comprehensively consider the actions and postures of people in the video and the sound characteristics in the audio to accurately judge whether abnormal behaviors such as quarrels and fights occur, greatly improving the accuracy of abnormal behavior judgment and reducing the false positive and false negative rates.

[0040] 3. By applying the short-time Fourier transform to extract frequency features from sound data, it changes the traditional simple way of only focusing on the presence or absence of sound, and can deeply analyze the subtle changes in the audio. For example, from the frequency distribution of abnormal shouts, the emotions of the shouters and the possible types of events can be judged. Combining the audio information with the video images provides a richer basis for abnormal behavior judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and other obvious variations can be conceived by those skilled in the art.

[0043] Referring to Figure 1 as shown, the intelligent video abnormal behavior monitoring method based on modal fusion includes the following steps:

[0044] Using a video camera to collect video modal data, at the same time collecting sound modal data with the help of an audio sensor and obtaining environmental state modal data through an environmental state sensor;

[0045] Performing denoising processing on the video data, performing noise reduction and filtering operations on the sound data, and normalizing the environmental state data to improve the data availability;

[0046] Extracting the visual features of key frame images from the video data, obtaining acoustic features from the sound data; extracting the change trend features of environmental parameters from the environmental state data;

[0047] Inputting the multi-modal features into an abnormal behavior analysis model, the abnormal behavior analysis model is designed through a hierarchical structure to achieve deep fusion and analysis of multi-modal features, thereby judging whether there is an abnormal behavior.

[0048] Specifically, first, multi-modal data is collected using a variety of devices. The video camera is responsible for obtaining video modal data, the audio sensor collects sound modal data, and the environmental state sensor collects environmental state modal data such as temperature and humidity; after collection, targeted preprocessing is performed on different types of data respectively, denoising the video data, noise reduction and filtering the sound data, and normalizing the environmental state data to improve the data quality and availability; then key features are extracted from various types of data, visual features of key frame images are extracted from the video data, such as color, texture, shape, etc.; acoustic features are obtained from the sound data, such as frequency, amplitude, etc.; the change trend features of environmental parameters are extracted from the environmental state data. Finally, these multi-modal features are input into a specially designed abnormal behavior analysis model, which realizes the deep fusion and analysis of multi-modal features through a unique hierarchical structure, and then judges whether there is an abnormal behavior in the scene. The collection of multi-modal data enables the monitoring system to obtain information from multiple dimensions, greatly improving the comprehensiveness and accuracy of information compared with single-modal data. The targeted preprocessing of different data effectively improves the data quality and provides a reliable basis for subsequent feature extraction and analysis. The fusion analysis of multi-modal features gives full play to the advantages of each modal data, complements each other, and can more accurately identify abnormal behaviors, reducing the probability of misjudgment and missed judgment.

[0049] The video camera has an intelligent frame rate adjustment function, which automatically adjusts the frame rate according to the dynamic degree of the scene. The adjustment formula is: Frame rate = Base frame rate + Dynamic degree coefficient × Scene change amount, where the scene change amount is calculated from the difference between adjacent video frames.

[0050] Specifically, the system calculates the scene change amount by calculating the difference between adjacent video frames, and then dynamically adjusts the frame rate according to the adjustment formula "Frame rate = Base frame rate + Dynamic degree coefficient × Scene change amount". When the scene change amount is large, that is, the scene is relatively dynamic, such as at the sports event site or crowded activity places, the frame rate is increased; while in static scenes with small scene change amounts, such as the quiet corner of a warehouse or an empty corridor, the frame rate is decreased. This intelligent frame rate adjustment method avoids the problem of missing key behavior details due to insufficient frame rate in dynamic scenes in traditional fixed frame rates, and at the same time prevents waste of storage resources caused by too high frame rate in static scenes, improves the efficiency and quality of video acquisition, and optimizes the resource utilization of the entire monitoring system.

[0051] The denoising of video data uses the bilateral filtering algorithm, which takes into account both the pixel spatial distance and the pixel value difference while filtering. The filtering formula is:

[0052]

[0053] where g(i,j) is the filtered pixel value, f(m,n) is the original pixel value, and w(i,j,m,n) is the weighting coefficient.

[0054] Specifically, when the bilateral filtering algorithm denoises video data, it will consider both the spatial distance between pixels and the difference in pixel values at the same time. For each pixel point, the algorithm calculates the weighting coefficient according to the spatial distance between the surrounding pixel points and it and the similarity of pixel values, and performs weighted summation on the surrounding pixels through these weighting coefficients to obtain the filtered pixel value. In this way, while removing noise, it can retain the edge and detail information of the image to the greatest extent. Compared with traditional simple filtering algorithms, the bilateral filtering algorithm can more effectively process video images in complex noise environments, retain the details of the image well while removing noise, make the subsequent extraction of image features more accurate, provide a clearer and more reliable image basis for the judgment of abnormal behaviors, and improve the adaptability of the monitoring system to complex scenes.

[0055] When extracting features from sound data, the short-time Fourier transform is used, and the formula is:

[0056]

[0057] where x(m) is the sound time-domain signal, w(n) is the window function, and N is the number of transform points.

[0058] Specifically, when extracting features from sound data, the short-time Fourier transform is used. It divides the time-domain signal of the sound into many short-time segments, performs Fourier transform on each segment, and the window function is used to limit the range of each short-time segment. In this way, the frequency characteristics of the sound at different time points can be obtained. Since the frequency characteristics of the sound change with time, the short-time Fourier transform can capture these dynamic changes, thereby obtaining more comprehensive acoustic information. The traditional processing of sound data is often relatively simple, only focusing on the presence or absence of sound. The short-time Fourier transform can deeply explore the frequency characteristics of the sound, enabling the monitoring system to use sound information to assist in judging abnormal behaviors. For example, by analyzing the frequency changes of abnormal sounds, the source and nature of the sound can be judged, and combined with the video image, the abnormal situations in the scene can be more accurately identified, enriching the basis for judging abnormal behaviors.

[0059] In the hierarchical structure of the abnormal behavior analysis model, the input layer groups and inputs features according to modal categories, and the middle hidden layer uses a bidirectional long short-term memory network module to learn the long-term dependence relationship of the feature sequence through forward and backward propagation. Its unit state update formula is:

[0060] i t = σ(W ii x t + b ii + W hi h t-1 + b hi )

[0061] f t = σ(W if x t + b if + W hf h t-1 + b hf )o t = σ(W io x t + b io + W ho h t-1 + b ho )

[0062]

[0063] h t = o t ⊙ tanh(C t )

[0064] Among them, i t , f t , o t are the values of the input gate, forget gate, and output gate respectively, C tFor candidate memory units and memory units, h t is the hidden state, W and b are weights and biases, σ is the activation function, and ⊙ is element-wise multiplication.

[0065] Specifically, the input layer groups and inputs the feature data of different modalities by category, facilitating the model to separately process and preliminarily integrate data from different sources. The middle hidden layer uses a bidirectional long short-term memory network module, which learns the long-term dependencies of the feature sequence through forward and backward propagation. The values of the input gate, forget gate, and output gate control the input, retention, and output of information. The candidate memory units and memory units store and update information, and the hidden state transmits information at different time steps. Through the update mechanism of these unit states, the model can learn the complex time series relationships and long-term dependency information between multi-modal features. Compared with traditional simple model structures, the application of this hierarchical structure and bidirectional long short-term memory network module enables the model to better handle the complex associations between multi-modal data, fully exploit the information of different modal data in the time dimension, thereby improving the accuracy and reliability of abnormal behavior judgment, especially being more prominent when dealing with multi-modal data in complex scenarios.

[0066] In the preprocessing of the environmental state data, linear interpolation is performed on the temperature data to fill in the missing values. The interpolation formula is:

[0067]

[0068] where T is the interpolated temperature value, T1 and T2 are adjacent known temperature values, t1 and t2 are the corresponding times, and t is the time point for which interpolation is required.

[0069] Specifically, based on the known adjacent temperature values and the corresponding times, the temperature at the missing value time point is estimated through a linear relationship. Assuming that the temperature values corresponding to the known times t1 and t2 are T1 and T2 respectively, for the time point t for which interpolation is required, the corresponding temperature value T is calculated using the formula, making the temperature data sequence more complete. Complete environmental state data is crucial for accurately judging abnormal behavior. Filling in the missing values of temperature data through linear interpolation ensures the continuity and integrity of the environmental state data, enabling the monitoring system to more accurately utilize the environmental state information when comprehensively analyzing multi-modal data, improving the comprehensiveness and accuracy of abnormal behavior judgment, and avoiding misjudgment caused by data missing.

[0070] When the abnormal behavior analysis model determines that there is an abnormal behavior, the alarm system is immediately triggered, and the alarm information is sent to the security personnel's terminal in an encrypted form through the network communication module.

[0071] Specifically, timely alarm can enable security personnel to quickly learn about abnormal situations and take corresponding measures, effectively reducing the losses caused by abnormal events. Encrypted transmission ensures the reliability of alarm information, ensuring that the security personnel receive accurate and unmodified information, improving the security and practicality of the entire monitoring system.

[0072] The abnormal behavior analysis model is regularly optimized and updated through newly collected multimodal data. The optimization algorithm adopted is the Adagrad algorithm, and the model parameter update formula is:

[0073]

[0074] In the formula, θ j is the model parameter, η is the learning rate, G ij is the diagonal element of the sum of squared gradients, ∈ is a small constant to prevent division by zero, is the gradient of the loss function with respect to the parameter θ j of.

[0075] Specifically, the Adagrad algorithm adjusts the learning rate according to the cumulative situation of each parameter's past gradients. For parameters that are updated frequently, the learning rate will gradually decrease, while for parameters that are updated less frequently, the learning rate is relatively large. By this way of adaptively adjusting the learning rate, the model can better adapt to new data and continuously optimize its own performance. As time goes by and the scenario changes, new abnormal behavior patterns may appear, and the old model may not be able to accurately identify them. Regularly optimizing and updating with new data and the Adagrad algorithm can enable the model to maintain its adaptability to new situations, continuously improve the accuracy of abnormal behavior judgment, extend the effective service life of the monitoring system, and better meet the needs of practical applications.

[0076] To ensure the time synchronization of multimodal data, precise timestamps are added to each modal data during data collection, and the data is aligned according to the timestamps during the data preprocessing stage.

[0077] Specifically, to ensure the time synchronization of multimodal data, during data collection, precise timestamps are added to video modal data, audio modal data, and environmental state modal data. During the data preprocessing stage, different modal data are aligned according to these timestamps, enabling the data from different sources to accurately correspond in the time dimension. The time synchronization of multimodal data is the basis for accurate data fusion and analysis. Through timestamp alignment, it can be ensured that when comprehensively analyzing multimodal data, the data of different modalities corresponds to the scene information at the same moment, avoiding incorrect data association caused by time asynchronization, improving the accuracy of abnormal behavior judgment, and enabling the monitoring system to more realistically reflect the actual situation of the scene.

[0078] When extracting key frames from video data, a key frame extraction algorithm based on motion estimation is adopted. Whether to extract key frames is judged by calculating the motion vectors between video frames. The block matching algorithm is used for calculating the motion vectors, and the matching error calculation formula is:

[0079]

[0080] where MAD(x,y) is the matching error, f1 and f2 are adjacent video frames, (x,y) is the position within the search window, and M and N are the sizes of the blocks.

[0081] Specifically, the key frame extraction algorithm based on motion estimation is used to calculate the motion vectors between video frames to judge whether to extract key frames. The block matching algorithm is used for calculating the motion vectors. The current video frame is divided into multiple small blocks, and the block that best matches the current small block is searched for within the search window of the adjacent video frame. The motion vector is determined by calculating the matching error. When the matching error exceeds a certain threshold, it is considered that the frame contains significant motion and can be extracted as a key frame. This key frame extraction algorithm can effectively extract the frames containing important motion information in the video, avoid processing a large number of unimportant frames, reduce the data processing volume, and improve the efficiency of video analysis. At the same time, the accurately extracted key frames retain the key behavior information in the video, provide core data for subsequent feature extraction and abnormal behavior judgment, and help to quickly and accurately identify abnormal behaviors.

[0082] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. An intelligent video abnormal behavior monitoring method based on modal fusion, characterized in that, Including the following steps: Collect video modality data using a video camera, simultaneously collect sound modality data with the help of an audio sensor, and obtain environmental state modality data through an environmental state sensor; Perform denoising processing on the video data, perform noise reduction and filtering operations on the sound data, and perform normalization processing on the environmental state data to improve data availability; Extract visual features of key frame images from the video data and obtain acoustic features from the sound data; Extract the change trend features of environmental parameters from the environmental state data; Input the multi-modal features into an abnormal behavior analysis model. The abnormal behavior analysis model is designed with a hierarchical structure to achieve deep fusion and analysis of multi-modal features, thereby determining whether there is abnormal behavior.

2. The intelligent video abnormal behavior monitoring method based on modal fusion according to claim 1, wherein The video camera has an intelligent frame rate adjustment function, which automatically adjusts the frame rate according to the dynamic degree of the scene. The adjustment formula is: frame rate = base frame rate + dynamic degree coefficient × scene change amount, where the scene change amount is calculated by the difference between adjacent video frames.

3. The intelligent video abnormal behavior monitoring method based on modal fusion according to claim 1, wherein, The bilateral filtering algorithm is used for denoising the video data. This algorithm takes into account both the pixel spatial distance and the pixel value difference while filtering. Its filtering formula is: where g(i,j) is the filtered pixel value, f(m,n) is the original pixel value, and w(i,j,m,n) is the weighting coefficient.

4. The intelligent video abnormal behavior monitoring method based on modal fusion according to claim 1, characterized in that, When extracting features from the sound data, the short-time Fourier transform is used. The formula is: where x(m) is the sound time-domain signal, w(n) is the window function, and N is the number of transformation points.

5. The intelligent video abnormal behavior monitoring method based on modal fusion according to claim 1, wherein, In the hierarchical structure of the abnormal behavior analysis model, the input layer groups and inputs features according to modality categories. The middle hidden layer uses a bidirectional long short-term memory network module to learn the long-term dependence relationship of the feature sequence through forward and backward propagation. Its unit state update formula is: i t = σ(W ii x t + b ii + W hi h t-1 + b hi ) f t = σ(W if x t + b if + W hf h t-1 + b hf ) o t = σ(W io x t + b io + W ho h t-1 + b ho ) h t = o t ⊙tanh(C t ) where, i t , f t , o t are the values of the input gate, forget gate, and output gate respectively, C t is the candidate memory cell and memory cell, h t is the hidden state, W and b are the weights and biases, σ is the activation function, and ⊙ is the element-wise multiplication.

6. The intelligent video abnormal behavior monitoring method based on modal fusion according to claim 1, characterized in that In the preprocessing of the environmental state data, linear interpolation is performed on the temperature data to fill in missing values. The interpolation formula is: where T is the interpolated temperature value, T1 and T2 are adjacent known temperature values, t1 and t2 are the corresponding times, and t is the time point where interpolation is required.

7. The intelligent video abnormal behavior monitoring method based on modal fusion according to claim 1, characterized in that When the abnormal behavior analysis model determines that there is abnormal behavior, it immediately triggers the alarm system, and the alarm information is sent to the security personnel terminal in an encrypted form through the network communication module.

8. The intelligent video abnormal behavior monitoring method based on modal fusion according to claim 1, characterized in that The abnormal behavior analysis model is regularly optimized and updated with newly collected multi-modal data. The optimization algorithm used is the Adagrad algorithm, and the model parameter update formula is: where θ j is a model parameter, η is the learning rate, G ij is the diagonal element of the sum of squared gradients, ∈ is a small constant to prevent division by zero, and j is the gradient of the loss function with respect to the parameter θ.

9. The intelligent video abnormal behavior monitoring method with modal fusion according to claim 1, characterized in that To ensure the time synchronization of multi-modal data, accurate timestamps are added to each modality data during data collection, and the data is aligned according to the timestamps during the data preprocessing stage.

10. The intelligent video abnormal behavior monitoring method based on modal fusion according to claim 1, characterized in that When extracting key frames from the video data, a key frame extraction algorithm based on motion estimation is used. Whether to extract key frames is determined by calculating the motion vector between video frames. The motion vector calculation uses the block matching algorithm, and the matching error calculation formula is: where MAD(x,y) is the matching error, f1 and f2 are adjacent video frames, (x,y) is the position within the search window, and M and N are the sizes of the blocks.

Citation Information

Cited By

  • Family safety AI monitoring method and system based on multi-modal algorithm fusion

    CN120976853A

  • A Home Security AI Monitoring Method and System Based on Multimodal Algorithm Fusion

    CN120976853B