Monitoring video playback method and device based on multi-modal model

Through the monitoring video viewing method based on multimodal model, multimodal features are extracted and fused to generate natural language descriptions, and a vector space for cross-modal search is constructed, which solves the problem that natural language query and complex scene analysis cannot be supported in the existing technology, and realizes efficient and accurate video retrieval and viewing.

CN120201147APending Publication Date: 2025-06-24CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510404509.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing surveillance video back viewing technology cannot support natural language-based query and analysis of complex scenarios, and cannot quickly locate related back viewing clips, resulting in low review efficiency and cannot meet users' accurate query and analysis of specific characters, behaviors or scenarios.

Method used

The monitoring video viewing method based on multimodal model is adopted. By sharding the monitoring video and generating a unique index, multimodal features and global dependency features are extracted using CNN, VGGish and Transformer models, fusion generate natural language descriptions, and a vector space for cross-modal search is constructed to store it in a vector database to support fast retrieval.

Benefits of technology

It realizes intelligent retrieval based on natural language, significantly improves the management and viewing efficiency of surveillance videos, supports diverse viewing needs, meets users' accurate query and analysis of specific content, and improves scalability and applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201147A_ABST
    Figure CN120201147A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of monitoring video processing. The invention provides a monitoring video playback method and device based on a multi-modal model. The method comprises the following steps: fragmenting real-time monitoring videos of different service scenes according to specified duration to generate a video fragment sequence; further generating a unique index identifier for each video clip in the video clip sequence, and establishing a mapping relation table between the index identifiers and the video clips; according to the method, CNN and VGGish models are adopted to extract multi-modal features from real-time monitoring videos of different business scenes, a Transform model is adopted to extract global dependency relationship features, the extracted multi-modal features and the global dependency relationship features are fused to obtain fused features, natural language description is further generated to construct a vector space of cross-modal retrieval, and the cross-modal retrieval is realized. Storing the multi-modal features into a vector database to form a storage path of each video clip, wherein the multi-modal features comprise audio features and behavior recognition features; and receiving a query statement input by a user, converting the query statement into vector representation, and performing similarity search in the vector database to obtain a video clip corresponding to a to-be-queried statement. According to the method, retrieval can be carried out more accurately, and the playback efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video processing, and provides a method and device for reviewing surveillance videos based on a multimodal model. Background Art

[0002] With the wide application of surveillance cameras in fields such as home care, public security, traffic management, and commercial venues, the amount of surveillance video data has increased exponentially. Traditional surveillance review methods rely on manual frame-by-frame retrieval, which is inefficient and prone to missing key information. To improve the retrieval efficiency and intelligent level of surveillance videos, intelligent surveillance review technology has emerged.

[0003] Existing methods pre-obtain the initial images of the camera at different rotation angles, compare the real-time surveillance images with the initial images to screen for abnormal events, and visually display the abnormal events and their related information in the form of a progress bar during playback. This method realizes the intelligent extraction and visual display of abnormal events by comparing the initial images of the camera with the real-time surveillance images, enabling users to intuitively and efficiently view the abnormal events in the playback video, and to a certain extent improving the user experience of the video playback function. However, screening for abnormal events by comparing the initial images with the real-time surveillance images may lead to missed or misdetected abnormal events due to the limitations of the quality and coverage of the initial images. In addition, only screening for abnormal events through image differences does not perform semantic understanding of the video content, and cannot support natural language-based queries and complex scenario analysis. Moreover, although visualizing abnormal events through a progress bar improves the review efficiency, users still need to view the video segments frame by frame and cannot quickly locate the review content related to specific semantic descriptions. Additionally, existing technologies mainly focus on abnormal event detection, do not provide a comprehensive understanding and indexing of video content, and are difficult to adapt to diverse review requirements (such as queries for specific people, behaviors, or scenarios), and thus have insufficient scalability.

[0004] Therefore, it is necessary to provide a method and device for reviewing surveillance videos based on a multimodal model to solve the above problems. Summary of the Invention

[0005] The present invention provides a method and device for reviewing surveillance videos based on a multimodal model to solve the technical problems in the prior art that it is impossible to support natural language-based queries and complex scenario review requirements, it is impossible to quickly locate relevant review segments and thus it is necessary to view frame by frame, resulting in low review efficiency, and it is impossible to meet the accurate queries and analysis of users for specific people, behaviors, or scenarios. The technical problems to be solved by the present invention are achieved through the following technical solutions.

[0006] A method for retrieving surveillance videos based on a multi-modal model according to the first aspect of the present invention, the surveillance video retrieval method comprising: slicing real-time surveillance videos of different business scenarios into segments according to a specified duration to generate a sequence of video segments; further generating a unique index identifier for each video segment in the sequence of video segments, and establishing a mapping relationship table between the index identifier and each video segment; extracting multi-modal features from real-time surveillance videos of different business scenarios using a CNN and a VGGish model, extracting global dependency relationship features using a Transformer model, fusing the extracted multi-modal features and global dependency relationship features to obtain final fused features, further generating a natural language description of the final fused features to construct a cross-modal retrieval vector space, and storing it in a vector database to form a storage path for each video segment, the multi-modal features including audio features and behavior recognition features; receiving a query statement input by a user, converting the query statement into a vector representation, and performing a similarity search in the vector database to obtain a video segment corresponding to the query statement to be queried.

[0007] A surveillance video retrieval device based on a multi-modal model according to the second aspect of the present invention, which is the surveillance video retrieval method according to the first aspect of the present invention, the surveillance video retrieval device comprising: a generation processing module for slicing real-time surveillance videos of different business scenarios into segments according to a specified duration to generate a sequence of video segments; an establishment module for further generating a unique index identifier for each video segment in the sequence of video segments and establishing a mapping relationship table between the index identifier and each video segment; an extraction processing module for extracting multi-modal features from real-time surveillance videos of different business scenarios using a CNN and a VGGish model, extracting global dependency relationship features using a Transformer model, fusing the extracted multi-modal features and global dependency relationship features to obtain final fused features, further generating a natural language description of the final fused features to construct a cross-modal retrieval vector space, and storing it in a vector database to form a storage path for each video segment, the multi-modal features including audio features and behavior recognition features; a query retrieval module for receiving a query statement input by a user, converting the query statement into a vector representation, and performing a similarity search in the vector database to obtain a video segment corresponding to the query statement to be queried.

[0008] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the surveillance video retrieval method according to the first aspect of the present invention.

[0009] A fourth aspect of the present invention provides a computer-readable medium storing a computer program, which when executed by a processor implements the method for retrieving a surveillance video based on a multi-modal model described in the first aspect of the present invention.

[0010] The embodiments of the present invention have the following advantages: Compared with the prior art, the present invention proposes a pipeline parallel architecture for video sharding generation, multi-modal model parsing, and vector encoding. Through the GPU video memory sharing mechanism, zero-copy transfer of sharded data is achieved, ensuring that semantic indexing and vectorized storage are completed instantly when the sharding is completed. This effectively solves the serial mode of separate sharding storage and semantic analysis, effectively reducing the end-to-end processing delay and avoiding the storage bandwidth pressure caused by secondary reading of video segments. The multi-modal features such as visual features and audio features are fused to obtain multi-modal features, generate unified semantic vectors, construct a vector space for cross-modal retrieval, and further perform secondary fusion of the multi-modal features and the extracted global dependency relationship features to obtain fusion features. Using the fusion features to represent each video segment can perform more accurate retrieval, effectively realizing the retrieval method of "searching for images by text". Within each shard, an index is constructed based on timestamps to support querying data by time range. Semantic analysis is performed on the data to extract semantic feature vectors, and a semantic vector index is constructed using a vector index algorithm (and using an open-source vector database, such as Milvus). The data objects are stored in the object storage system OSS, the storage address of each data object is recorded, and an index is constructed for quick positioning. A three-level association structure of shard timestamp index, semantic vector index, and object storage address index is established to quickly locate video segments according to natural language descriptions, effectively improving the video retrieval efficiency.

[0011] In addition, using a multi-modal model to deeply analyze the video content, generate detailed semantic descriptions, support natural language-based queries and the need for retrieving in complex scenarios. By storing the video segment index and semantic descriptions in a vector database, users can quickly locate relevant retrieval segments according to the descriptions, avoiding frame-by-frame viewing and significantly improving the retrieval efficiency; providing a comprehensive understanding and index of the video content, supporting diverse retrieval needs, meeting users' precise queries and analyses of specific people, behaviors, or scenarios, and improving the scalability and applicability.

[0012] In addition, the multi-modal model performs semantic understanding on video segments, effectively avoiding the dependence on initial frame comparison, and improving the accuracy and coverage of scene anomaly event detection. Description of the Drawings

[0013] Figure 1 is a flowchart of the steps of an example of the method for retrieving a surveillance video based on a multi-modal model of the present invention; Figure 2It is a schematic diagram of the model principle of the monitoring video playback method based on the multi-modal model of the present invention; Figure 3 It is a structural block diagram of the monitoring video playback device based on the multi-modal model of the present invention; Figure 4 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention; Figure 5 It is a schematic structural diagram of a computer-readable medium according to an embodiment of the present invention. Detailed implementation manners

[0014] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0015] In view of the above problems, the present invention proposes a monitoring video playback method based on a multi-modal model.

[0016] This method abstracts the monitoring video into a structured mathematical model. First, it slices the real-time monitoring video stream to generate video segments in time intervals and establishes a unique index for each segment. Through the multi-modal model, it conducts semantic understanding on the video segments, generates natural language descriptions, and converts them into high-dimensional feature vectors for storage in the vector database. At the same time, the video segments are stored in the object storage system. In the playback query stage, the query statement input by the user is converted into a feature vector, and similarity retrieval is performed to locate the most relevant video segment index. Finally, the corresponding video content is extracted from the object storage, thus realizing a complete technical closed-loop from video slicing, semantic understanding to efficient retrieval. By slicing, indexing the monitoring video, conducting semantic understanding through the multi-modal model, and combining with the vector model and vector database, intelligent retrieval based on natural language descriptions is realized, greatly improving the management and playback efficiency of the monitoring video.

[0017] Embodiment 1 The following will refer to Figure 1 、 Figure 2 to describe the content of the present invention in detail.

[0018] Figure 1 It is a flowchart of the steps of an example of the video monitoring playback method based on the multi-modal model of the present invention.

[0019] As Figure 1 shown, in step S101, the real-time monitoring videos in different business scenarios are sliced according to a specified duration to generate a sequence of video segments.

[0020] In a specific implementation manner, in the traffic monitoring scenario, the real-time traffic monitoring video is sliced according to a specified duration. The specified duration is in the range of 2 seconds to 30 seconds.

[0021] For example, the input is the real-time traffic monitoring video stream V, which is segmented according to the specified duration ∆t (in this embodiment, the specified duration is 5 seconds) to generate a series of video segments {v1, v2,..., v n}, where each video segment v i (where i represents the i-th video segment, i is a positive integer, specifically 1, 2,..., n) corresponds to a time interval, which is represented by the following expression: {v1, v2,..., v n} = Slice(V, ∆t) where V represents the input real-time traffic monitoring video stream; ∆t represents the specified duration; Slice represents the segmentation operation, which slices the real-time traffic monitoring video stream V into multiple segments according to the specified duration ∆t.

[0022] It should be noted that in the present invention, the selection of the specified duration is determined according to specific business requirements and the characteristics of the monitoring scenario. The following factors are considered in the selection of the specified duration: the frequency of event occurrence (if the frequency of event occurrence in the monitoring scenario is high, a shorter segmentation duration is configured to capture more event details), the limitations of storage and computing resources (a shorter segmentation duration will increase the storage and computing overhead, so a suitable duration is selected within the range allowed by the resources), and the granularity of business analysis (the segmentation duration matches the granularity of business analysis. For example, if minute-level behavior needs to be analyzed, the segmentation duration can be configured at the minute level). In addition, for the traffic monitoring scenario, in the traffic monitoring video, it is necessary to capture the driving trajectories and illegal behaviors of vehicles, etc. Using a specified duration of 5 seconds can better capture the short-term behaviors of vehicles, such as running red lights and illegal lane changes. For the mall security monitoring scenario, in the mall security monitoring, it is necessary to monitor the behaviors of customers and abnormal events, etc. Using a specified duration of 10 seconds can capture the walking paths and staying times of customers, etc., and can also detect abnormal behaviors such as staying too long in a timely manner. For the factory production monitoring scenario, in the factory production monitoring video, it is necessary to monitor the operating status of the production line and the operation of equipment, etc. Using a specified duration of 30 seconds can better reflect the operating cycle of the production line and facilitate the analysis of production efficiency. For the home care monitoring scenario, in the home care video, sudden events such as the elderly falling, children being accidentally injured or patients suddenly feeling unwell are involved. These events usually occur between a few seconds and more than ten seconds. Using a specified duration of 10 seconds can capture the complete process of key events while ensuring real-time performance. The above is only described as an optional example and should not be construed as a limitation to the present invention.

[0023] Next, in step S102, a unique index identifier is further generated for each video segment in the video segment sequence, and a mapping relationship table between the index identifier and each video segment is established.

[0024] Specifically, for each video segment generate a unique index , and establish a mapping relationship between the index and the segment. The specific expression is as follows: ; ; wherein, represents the unique index of the i-th video segment , i is a positive integer, specifically 1, 2,..., n; represents the mapping relationship table between the index and the video segment. The shard start timestamp, a timestamp accurate to the millisecond level (such as 20250321145678901), is used to record the start moment of the video segment.

[0025] Next, in step S103, the CNN and VGGish models are used to extract multi-modal features from the real-time monitoring videos of different business scenarios, the Transformer model is used to extract global dependency relationship features, the extracted multi-modal features and global dependency relationship features are fused to obtain fused features, and a natural language description of the final fused features is further generated to construct a vector space for cross-modal retrieval and stored in a vector database to form the storage path of each video segment. The multi-modal features include audio features and behavior recognition features.

[0026] Specifically, multi-modal features are extracted from the real-time monitoring videos of different business scenarios.

[0027] Semantic understanding is performed on each video segment using a multi-modal model to generate a corresponding natural language description. The specific expression is as follows:

[0028] ; wherein, represents the natural language description model used to generate the natural language descriptions of each video segment; represents the video segment 's natural language description.

[0029] Specifically, the multi-modal model combines a convolutional neural network (CNN), an audio feature extraction model (VGGish), and a Transformer architecture to perform semantic understanding on each video segment and generate a corresponding natural language description. The core of the multi-modal model is through CNN (corresponding to Figure 2("ResNet" in it) and VGGish extract local features related to behaviors (such as human actions, facial expressions, glass breaking, background music, etc.) in each video clip, and capture the global dependencies of each video clip through the Transformer mechanism, and finally generate a natural language description. For details, please refer to Figure 2 。

[0030] It should be noted that in this example, the audio feature extraction model uses the VGGish model. The VGGish model is a model constructed based on the VGG network structure. Using the monitoring video data of each application scenario applicable to the present invention, the video data (one frame or multiple frame picture data) annotated with key persons, risk behaviors or abnormal behaviors, facial expressions, key events (such as events like glass breaking, painful facial expressions, etc.) is used to establish a training dataset, so as to obtain an audio feature extraction model after incrementally training the existing VGGish model.

[0031] Preprocess the video clip to be input into the multi-modal model. First, split the video clip into multiple frames, and the pre-trained CNN model for each frame of picture extracts behavior recognition features. Suppose the video clip contains frames of pictures, and the feature vector of each frame of picture is , then the feature representation of the video clip is , where d f represents the dimension of the behavior recognition features extracted by the pre-trained CNN model for each frame of picture, and T represents the time step.

[0032] It should be noted that a convolutional neural network (CNN) model is constructed using ResNet (Residual Neural Network). By introducing residual connections (also known as skip connections), the problems of gradient disappearance and degradation in the training of deep neural networks are solved, and it is particularly suitable for tasks such as image classification, object detection, and feature extraction.

[0033] Use the multi-layer convolution part of the CNN model to extract local features (such as human actions, facial expressions) of each frame of picture. Use a one-dimensional convolutional kernel , where is the convolutional kernel size, is the output feature dimension. The following expression is used to represent the convolution operation:

[0034] ; where, represents the feature vector of the th frame obtained through the one-dimensional convolution operation; represents from the from frame to the features of the frame is the bias term; represents the one-dimensional convolutional kernel of the convolutional part.

[0035] Through multiple layers of convolution (preferably 34 layers of convolution in this example), local features of different scales of each frame of video image can be extracted (that is, various local features related to behavior recognition, that is, behavior recognition features), specifically including the following features: low-level features, such as the outline of an object, color gradient, simple texture; middle-level features, such as eyes, mouth, and hand movements of a face; high-level features, such as more semantic visual information such as "running", "fighting", "opening a door", etc.

[0036] Next, audio features are extracted from each frame of video image. Using a pre-trained audio extraction model (such as the VGGish model) with the input being audio data, after passing through, for example, 4 layers of convolutional neural network and fully connected layers, a 128-dimensional feature vector is output to extract features from the audio in each video segment, obtaining features such as acoustic texture, temporal pattern, semantic encoding, etc. Let the audio segment correspond to the video segment , and the audio model outputs the feature vector , then the feature representation of the audio segment is where d a represents the dimension of each audio feature vector, and the value of d a is in the range of [64, 256]. In this example, it is 128 dimensions.

[0037] Next, multi-modal feature fusion is performed, and the Transformer model is used to extract global dependency features.

[0038] Specifically, the video features (in this example, the behavior recognition features), and the audio features are fused in a multi-modal manner. First, the audio features are time-aligned to ensure the same time step as the video features. Then, the video features and audio features are mapped to the same dimensional space through fully connected layers:

[0039] ; ; where represents the video features obtained after mapping the video features F; represents the audio features obtained after mapping the audio features; and are trainable weight matrices, and is the bias term. The fused features are weighted and summed to obtain the multi-modal fusion features ’:

[0040] ; where M’ represents the multi-modal fusion feature obtained after weighted summation; are learnable weight parameters. During the training process of the model, is used to balance the contributions of the behavior recognition feature and the audio feature; is the interaction term weight, which controls the local feature interaction intensity; is the attention weight, which is used to adjust the global dependence relationship; 、 、 are trainable parameters, which are all included in the chain rule of backpropagation, with a range of [0.2, 0.9], to avoid modal suppression or overfitting caused by extreme weights, and are optimized by the gradient descent method during the training process of the model to ensure that the model can adaptively learn the importance between different modal features; represents the Hadamard product, which is element-wise multiplication; is the lightweight cross-attention, which belongs to the single-head attention mechanism.

[0041] For the above trainable parameters 、 、 values, they are continuously optimized in the chain rule of backpropagation to effectively avoid modal suppression or overfitting caused by extreme weights, and are optimized by the gradient descent method during the training process of the model to ensure that the model can adaptively learn the importance between different modal features.

[0042] Perform positional encoding on the multi-modal fusion feature. While retaining the time order information, obtain the model input for extracting the global dependence relationship feature: ; ; where, represents the model input for extracting the global dependence relationship feature, , is the multi-modal feature at time step , M1 represents the behavior recognition feature, and M2 represents the audio feature; represents the time position information obtained by performing positional encoding on the model input for extracting the global dependence relationship feature , is a learnable position basis matrix, specifically obtained during the model training optimization process, represents the feature dimension of the model input; is the input sequence length of the model input; is a depthwise separable convolutional layer, where is a depthwise convolutional kernel ; is a convolutional operation along the time dimension; is a convolutional kernel with a size of 1 to 6 (k is preferably 3), is a pointwise convolutional layer, is a 1×1 convolutional kernel , represents the feature dimension of the model input.

[0043] It should be noted that the core of the Transformer model is the self-attention mechanism.

[0044] Specifically, for the model input X, first calculate the query (Query), key (Key), and value (Value) corresponding to the model input X:

[0045] ; where , , is a trainable weight matrix.

[0046] Then, use the following expression to calculate the self-attention score of the model input X: ; where represents the self-attention score of the model input X; Q represents the query vector corresponding to the model input X; K represents the key corresponding to the model input X; V represents the value corresponding to the model input X.

[0047] In this example, the multi-head attention mechanism divides the model input X into h subspaces, each head calculates independent attention, and finally concatenates and linearly transforms: ; where represents the multi-head attention mechanism in the Transformer model; is the output transformation matrix, i represents the i-th attention head; h represents the number of attention heads in the multi-head attention mechanism, both i and h are positive integers, i is specifically 1, 2,..., h, and h is 3 to 10; d c represents the dimension of the input features; d v represents the dimension of the value vector in each attention head.

[0048] Furthermore, the multi-modal fusion features are fused with the global dependency features extracted by the Transformer model. The fusion strategy adopts weighted summation and performs non-linear transformation through a fully connected layer to obtain the final fusion features:

[0049] ; where z represents the final fusion feature obtained after the secondary fusion of the multi-modal fusion features and the extracted global dependency features; and are trainable parameters, which can be specifically obtained during the model optimization process. [M; MultiHead(Q, K, V)] represents the concatenation operation. M represents the multi-modal features obtained by weighted fusion of multiple features, and MultiHead(Q, K, V) represents the multi-head attention mechanism in the Transformer model.

[0050] The obtained final fusion feature z is used to generate a natural language description through the Transformer decoder , and is represented by the following expression: ; where the convolution kernel size k is in the range of 2 to 6, and k is preferably 3; the output feature dimension is in the range of [128, 512], is preferably 512; the number of multi-head attention heads h is in the range of [3, 10], and h is preferably 8; the position encoding dimension is the same as the CNN output feature dimension; the number of decoder layers is in the range of [2, 16], and is preferably 4 layers.

[0051] Specifically, video clips in different business scenarios are used to generate natural language descriptions and converted into text vector representations (i.e., high-dimensional feature vectors) to construct a vector space for cross-modal retrieval and stored in a vector database. At the same time, the video clips are stored in an object storage system.

[0052] The video clip is stored in the object storage , and its storage path is generated . It is expressed by the following expression:

[0053] ; where represents the storage path corresponding to the i-th video clip; represents the object storage system (such as AWS S3, a certain cloud OSS, etc.).

[0054] The natural language description is converted into a text vector through a vector model, and the index of the i-th video clip , the i-th eigenvector , the i-th video segment Storage path Stored in the vector database . Expressed by the following expression:

[0055] ; ; Where: Represents a vector model for converting natural language descriptions into text vectors. Represents a vector database, which supports efficient storage and retrieval of eigenvectors. For example, the vector model uses BERT, the input is a natural language description, usually a piece of text or a sentence corresponding to the video segment to be processed, such as "a person wearing red clothes enters the room". The output is an eigenvector corresponding to the video segment to be processed, representing the vectorized representation of the input text, and the dimension of the vector is, for example, 768 dimensions. The vector database uses the open-source vector database Milvus, which associates the shard timestamp index, semantic vector index, and object storage address index to form a one-to-one multi-level index relationship, realizing multi-dimensional retrieval based on time, semantics, and storage path. The user inputs a timestamp or a semantic vector , first find the corresponding video segment index through the timestamp index or semantic vector index, and then find the storage path of the video segment through the object storage address index, and accurately locate the specific video segment at the millisecond level.

[0056] Next, in step S104, receive the query statement input by the user, convert the query statement into a vector representation, and perform a similarity search in the vector database to obtain the video segment corresponding to the query statement to be queried.

[0057] Specifically, the query statement input by the user , if q is "a person wearing red clothes enters the room", use the vector model BERT to convert the query statement q into a query vector (that is, convert it into a vector representation), and perform a similarity search in the vector database (such as Milvus). The specific search algorithm uses Euclidean distance similarity calculation. Specifically, calculate the Euclidean distance, that is, the straight-line distance, between the query vector A of the query statement and each vector in the vector database. The smaller the calculated distance, the higher the similarity. Calculate using the following expression: ; Where, Represents calculating the Euclidean distance, that is, the straight-line distance, between the query vector of the query statement and each vector in the vector database, and are two dimensional vectors, where Ai represents the query vector value on the i-th dimension, and Bi represents the value of the database vector B on the i-th dimension respectively.

[0058] Specifically, find the index of the video segment closest to . Then, according to the index retrieve the corresponding video segment from the object storage .

[0059] The specific expression is as follows: ; ; ; ; Among them, represents searching in the vector database for the feature vector closest to and returning the corresponding index ; represents retrieving the corresponding storage path from the mapping relation table according to the index ; represents retrieving the video segment from the object storage according to the path .

[0060] Integrate the above steps into a complete formula expression: ; Among them, represents the video segment corresponding to the query statement q input by the user; q represents the query statement input by the user; represents retrieving the video segment from the object storage according to the path ; represents retrieving the corresponding storage path from the mapping relation table according to the index ; represents searching in the vector database for the feature vector closest to and returning the corresponding index

[0061] ​​​Compared with the prior art, the present invention proposes a pipeline parallel architecture for video sharding generation, large model parsing, and vector encoding. Through the GPU video memory sharing mechanism, zero-copy transfer of sharded data is achieved, ensuring that semantic indexing and vectorized storage are completed instantly upon sharding completion. This effectively solves the serial mode of separate sharded storage and semantic analysis, significantly reducing the end-to-end processing latency and avoiding the storage bandwidth pressure caused by secondary reading of video segments. Multimodal features such as visual features and audio features are fused to obtain multimodal features, generate unified semantic vectors, construct a vector space for cross-modal retrieval, and further perform secondary fusion of the multimodal features and the extracted global dependency relationship features to obtain fusion features. Using the fusion features to represent each video segment enables more accurate retrieval, effectively realizing the retrieval method of "searching for images by text". Within each shard, an index is constructed based on timestamps to support querying data by time range. Semantic analysis is performed on the data to extract semantic feature vectors, and a semantic vector index is constructed using a vector indexing algorithm (and using an open-source vector database such as Milvus). The data objects are stored in the object storage system OSS, the storage address of each data object is recorded, and an index is constructed for quick positioning. A three-level association structure of shard timestamp index, semantic vector index, and object storage address index is established to quickly locate video segments according to natural language descriptions, effectively improving the video retrieval efficiency.

[0062] In addition, a multimodal model is used to deeply analyze the video content, generate detailed semantic descriptions, support natural language-based queries and the demand for reviewing complex scenarios. By storing the video segment index and semantic descriptions in a vector database, users can quickly locate relevant review segments according to the descriptions, avoiding frame-by-frame viewing and significantly improving the review efficiency; providing a comprehensive understanding and index of the video content, supporting diverse review needs, meeting users' precise queries and analyses of specific people, behaviors, or scenarios, and enhancing the scalability and applicability.

[0063] In addition, the multimodal model performs semantic understanding on video segments, effectively avoiding the dependence on initial frame comparison and improving the accuracy and coverage of scene anomaly event detection.

[0064] Embodiment 2 The following is an embodiment of the device of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the device embodiment of the present invention, please refer to the method embodiment of the present invention.

[0065] Figure 3 is a schematic structural diagram of an example of a monitoring video review device based on the present invention. The following will refer to Figure 3 to describe the monitoring video review device. The monitoring video review device is used to execute the monitoring video review method described in the first aspect of the present invention.

[0066] As shown Figure 3 in FIG. 345, the monitoring video playback device 300 includes a generation processing module 310, an establishment module 320, an extraction processing module 330, and a query and retrieval module 340.

[0067] In a specific embodiment, the generation processing module 310 is configured to slice the real-time monitoring videos of different business scenarios by a specified duration to generate a video segment sequence. The establishment module 320 further generates a unique index identifier for each video segment in the video segment sequence and establishes a mapping relationship table between the index identifier and each video segment. The extraction processing module 330 extracts multi-modal features from the real-time monitoring videos of different business scenarios using a CNN model, extracts global dependency relationship features using a Transformer model, fuses the extracted multi-modal features and global dependency relationship features to obtain a final fusion feature, further generates a natural language description of the final fusion feature to construct a cross-modal retrieval vector space, and stores it in a vector database to form a storage path for each video segment. The multi-modal features include audio features and behavior recognition features. The query and retrieval module 340 is configured to receive a query statement input by a user, convert the query statement into a vector representation, and perform a similarity search in the vector database to obtain a video segment corresponding to the query statement to be queried.

[0068] According to an alternative embodiment, before fusing the extracted multi-modal features and global dependency relationship features to obtain a fusion feature, the following expression is used to perform weighted fusion on the extracted multi-modal features to obtain a multi-modal fusion feature: ; where M’ represents the multi-modal fusion feature obtained after weighted summation; is a learnable weight parameter, and during the training process of the model, is used to balance the contributions of behavior recognition features and audio features; is an interaction term weight that controls the local feature interaction intensity; is an attention weight used to adjust the global dependency relationship; , , are trainable parameters, all of which are included in the chain rule of backpropagation, with a range of [0.2, 0.9], to avoid modal suppression or overfitting caused by extreme weights, and are optimized by gradient descent during the training process of the model to ensure that the model can adaptively learn the importance between different modal features; represents the Hadamard product, which is element-wise multiplication; is a lightweight cross-attention, which belongs to a single-head attention mechanism.

[0069] According to an alternative embodiment, positional encoding is performed on the multimodal fusion features, and while preserving the temporal order information, a model input for extracting global dependency features is obtained: ; ; wherein, represents the model input for extracting global dependency features, , is the multimodal feature at time step , M1 represents the behavior recognition feature, and M2 represents the audio feature; represents the temporal position information obtained by performing positional encoding on the model input for extracting global dependency features; , is a learnable position basis matrix, represents the feature dimension of the model input; is the input sequence length of the model input; is a depthwise separable convolutional layer, wherein, is the depthwise convolutional kernel ; is the convolutional operation along the time dimension; is the convolutional kernel, with a size of 1 to 6, is the pointwise convolutional layer, is the 1×1 convolutional kernel , represents the feature dimension of the model input.

[0070] According to an alternative embodiment, the multimodal fusion features are fused with the global dependency features extracted by the Transformer model to obtain the final fusion features: ; wherein, z represents the final fusion feature obtained by performing secondary fusion on the multimodal features and the extracted global dependency features; and are trainable parameters, [M; MultiHead(Q, K, V)] represents the concatenation operation, M’ represents the multimodal fusion feature obtained by fusing multiple features with weights, and MultiHead(Q, K, V) represents the multi-head attention mechanism in the Transformer model.

[0071] According to an alternative embodiment, the input X is divided into h subspaces by using the multi-head attention mechanism, each head calculates independent attention, and finally they are concatenated and linearly transformed to obtain the output as the global dependency feature: ; wherein, Represents the output of the multi-head attention mechanism; is the output transformation matrix, i represents the i-th attention head, h represents the number of attention heads in the multi-head attention mechanism, both i and h are positive integers, i is specifically 1, 2,..., h, and h is 3 to 10; d c Represents the dimension of the input feature; d v Represents the dimension of the value vector in each attention head.

[0072] According to an alternative implementation, when receiving a query statement input by the user, convert the query statement into a vector representation and perform retrieval using the following expression: ; where represents the video segment corresponding to the query statement q input by the user; q represents the query statement input by the user; represents obtaining the video segment from the object storage according to the path ; ; represents obtaining the corresponding storage path from the mapping relation table according to the index ; ; represents searching in the vector database for the feature vector closest to and returning the corresponding index .

[0073] According to an alternative implementation, generate a unique index identifier, i.e., a shard timestamp index, for each video segment in the video segment sequence, and establish a mapping relation table between the index identifier and each video segment.

[0074] Use the vector model to convert the natural language description of the final fused feature of each video segment into a semantic vector index.

[0075] Associate the object storage address, shard timestamp index, and semantic vector index formed when storing each video segment. When performing video segment retrieval, find the corresponding video segment index through the timestamp index or semantic vector index, and find the storage path of the video segment through the object storage address index to determine the video segment to be retrieved.

[0076] It should be noted that since Figure 3 the monitoring video playback method executed by the monitoring video playback device is substantially the same as the monitoring video playback method in the example of Figure 1 , the description of the same part is omitted.

[0077] Compared with the prior art, the present invention proposes a pipeline parallel architecture for video sharding generation, large model parsing, and vector encoding. Through the GPU video memory sharing mechanism, zero-copy transfer of sharded data is achieved, ensuring that semantic indexing and vectorized storage are completed instantaneously upon completion of sharding. This effectively solves the serial mode of separate sharded storage and semantic analysis, reduces the end-to-end processing delay, and avoids the storage bandwidth pressure caused by secondary reading of video segments. Multimodal features such as visual features and audio features are fused to obtain multimodal features, generate unified semantic vectors, construct a vector space for cross-modal retrieval, and further perform secondary fusion of the multimodal features and the extracted global dependency relationship features to obtain fused features. Using the fused features to represent each video segment enables more accurate retrieval and effectively realizes the retrieval method of "searching for images by text". Within each shard, an index is constructed based on timestamps to support querying data by time range. Semantic analysis is performed on the data to extract semantic feature vectors, and a semantic vector index is constructed using a vector indexing algorithm (and using an open-source vector database such as Milvus). The data objects are stored in the object storage system OSS, the storage address of each data object is recorded, and an index is constructed for quick positioning. A three-level association structure of shard timestamp index, semantic vector index, and object storage address index is established to quickly locate video segments according to natural language descriptions, effectively improving the video retrieval efficiency.

[0078] In addition, a multimodal model is used to deeply analyze the video content to generate detailed semantic descriptions, supporting natural language-based queries and the need for lookback in complex scenarios. By storing the video segment index and semantic description in a vector database, users can quickly locate relevant lookback segments according to the description, avoiding frame-by-frame viewing and significantly improving the lookback efficiency. It provides a comprehensive understanding and index of the video content, supports diverse lookback requirements, meets users' precise queries and analysis of specific people, behaviors, or scenarios, and improves scalability and applicability.

[0079] In addition, the multimodal model performs semantic understanding on video segments, effectively avoiding dependence on initial frame comparison and improving the accuracy and coverage of scene anomaly event detection.

[0080] Figure 4 It is a schematic structural diagram of an embodiment of an electronic device according to the present invention.

[0081] As Figure 4 shown, the electronic device is presented in the form of a general-purpose computing device. The processor can be one or multiple and work collaboratively. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity and can also be the sum of multiple physical devices.

[0082] The memory stores computer-executable programs, typically machine-readable code. The computer-readable programs can be executed by the processor to enable the electronic device to execute the method of the present invention or at least some of the steps in the method.

[0083] The memory includes volatile memory, such as random access storage units (RAM) and / or cache storage units, and may also be non-volatile memory, such as read-only storage units (ROM).

[0084] Optionally, in this embodiment, the electronic device further includes an I / O interface for data exchange between the electronic device and external devices. The I / O interface can represent one or more of several bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the bus structures in a variety of bus structures.

[0085] It should be understood that Figure 4 The electronic device shown is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include a display unit such as a display screen, and some electronic devices also include human-computer interaction elements, such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement the method of the present invention or at least some of the steps of the method, it can be considered as the electronic device covered by the present invention.

[0086] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, as Figure 5 shown, the technical solution according to the embodiment of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several commands to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present invention.

[0087] The software product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0088] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with a command execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0089] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0090] The above-mentioned computer-readable medium carries one or more programs, which, when executed by one of the devices, cause the computer-readable medium to implement the data interaction method of the present disclosure.

[0091] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are only different from this embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.

[0092] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several commands to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.

[0093] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0094] In the above detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, like symbols typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, the drawings, and the claims are not meant to be limiting. Other embodiments can be used and other changes can be made without departing from the spirit or scope of the subject matter presented herein.

[0095] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A surveillance video review method based on a multimodal model, characterized in that: The method comprises: Divide the real-time surveillance videos of different business scenarios into segments according to the specified time length to generate a video segment sequence; Further generating a unique index identifier for each video segment in the video segment sequence, and establishing a mapping relationship table between the index identifier and each video segment; The CNN and VGGish models are used to extract multimodal features from real-time surveillance videos of different business scenarios, and the Transformer model is used to extract global dependency features. The extracted multimodal features and global dependency features are fused to obtain the final fused features, and a natural language description of the final fused features is further generated to construct a vector space for cross-modal retrieval, which is stored in a vector database to form a storage path for each video clip. The multimodal features include audio features and behavior recognition features. A query statement input by a user is received, the query statement is converted into a vector representation, and a similarity search is performed in the vector database to obtain a video clip corresponding to the query statement.

2. The monitoring video review method according to claim 1, characterized in that: include: Before fusing the extracted multimodal features and the global dependency features to obtain the fused features, the following expression is used to perform weighted fusion on the extracted multimodal features to obtain the multimodal fused features: ; Among them, M' represents the multimodal fusion feature obtained after weighted summation; is a learnable weight parameter. During the model training process, for balancing the contribution of behavioral identification features and audio features; is the interaction term weight, which controls the interaction strength of local features; is the attention weight, which is used to adjust the global dependency; , , are trainable parameters, all included in the chain rule of back propagation, ranging from [0.2, 0.9], to avoid extreme weights leading to modal suppression or overfitting, and are optimized by gradient descent during model training to ensure that the model can adaptively learn the importance of different modal features; Represents the Hadamard product, which is element-by-element multiplication; It is a lightweight cross attention and belongs to the single-head attention mechanism.

3. The monitoring video review method according to claim 2, characterized in that: include: The multimodal fusion features are positionally encoded to obtain a model input for extracting global dependency features while retaining the temporal order information: ; ; in, represents the model input for extracting global dependency features, , is the time step Multimodal features, M1 represents behavior recognition features, and M2 represents audio features; represents the temporal position information obtained by positionally encoding the model input for extracting global dependency features; , is the learnable position basis matrix, Represents the feature dimension of the model input; the length of the input sequence fed into the model; is a depth-wise separable convolutional layer, where is the depth convolution kernel ; is the convolution operation along the time dimension; is the convolution kernel, the size is 1 to 6, is the point-wise convolution layer, is a 1×1 convolution kernel , Represents the feature dimension of the model input.

4. The monitoring video review method according to claim 3, characterized in that: include: The multimodal fusion features are fused with the global dependency features extracted by the Transformer model to obtain the final fusion features: ; Wherein, z represents the final fusion feature obtained by secondary fusion of the multimodal feature and the extracted global dependency feature; and are trainable parameters, [M; MultiHead(Q, K, V)] represents the concatenation operation, M' represents the multimodal fusion feature obtained by weighted fusion of multiple features, and MultiHead(Q, K, V) represents the multi-head attention mechanism in the Transformer model.

5. The monitoring video review method according to claim 3, characterized in that: include: The multi-head attention mechanism is used to divide the input X into h subspaces. Each head calculates independent attention, and finally concatenates and linearly transforms to obtain the output as the global dependency feature: ; in, Represents the output of the multi-head attention mechanism; is the output transformation matrix, i represents the i-th attention head, h represents the number of attention heads in the multi-head attention mechanism, i and h are both positive integers, i is 1, 2, ..., h, and h is 3 to 10; d c Indicates the dimension of the input feature; d v Represents the dimension of the value vector in each attention head.

6. The monitoring video review method according to claim 1, characterized in that: include: When receiving a query statement input by a user, the query statement is converted into a vector representation and searched using the following expression: ; in, Indicates the video clip corresponding to the query sentence q input by the user; q represents the query sentence input by the user; Indicates that the path is obtained from the object storage Get video clips ; Indicates that according to the index From the mapping table Get the corresponding storage path ; Represents searching for the same The closest feature direction; Quantity, return the corresponding index .

7. The monitoring video review method according to claim 1, characterized in that: Further including: Generate a unique index identifier, i.e., a fragment timestamp index, for each video segment in the video segment sequence, and establish a mapping relationship table between the index identifier and each video segment; The natural language description of the final fusion features of each video clip is converted into a semantic vector index using a vector model; The object storage address, fragment timestamp index, and semantic vector index formed when each video clip is stored are associated. When retrieving video clips, the corresponding video clip index is found through the timestamp index or semantic vector index, and the storage path of the video clip is found through the object storage address index to determine the video clip to be retrieved.

8. A surveillance video review device based on a multimodal model, characterized in that: It implements the monitoring video review method according to any one of claims 1 to 7, and the monitoring video review device comprises: A generation processing module is used to segment the real-time surveillance videos of different business scenarios into segments according to the specified time length to generate a sequence of video segments; An establishment module is further used to generate a unique index identifier for each video segment in the video segment sequence, and to establish a mapping relationship table between the index identifier and each video segment; The extraction processing module uses CNN and VGGish models to extract multimodal features from real-time monitoring videos of different business scenarios, uses the Transformer model to extract global dependency features, fuses the extracted multimodal features and global dependency features to obtain the final fused features, further generates a natural language description of the final fused features to construct a vector space for cross-modal retrieval, and stores it in a vector database to form a storage path for each video clip. The multimodal features include audio features and behavior recognition features. The query retrieval module is used to receive a query statement input by a user, convert the query statement into a vector representation, and perform a similarity search in the vector database to obtain a video clip corresponding to the query statement.

9. The monitoring video review device according to claim 8, characterized in that: include: Before fusing the extracted multimodal features and the global dependency features to obtain the fused features, the following expression is used to perform weighted fusion on the extracted multimodal features to obtain the multimodal fused features: ; Among them, M' represents the multimodal fusion feature obtained after weighted summation; is a learnable weight parameter. During the model training process, for balancing the contribution of behavioral identification features and audio features; is the interaction term weight, which controls the interaction strength of local features; is the attention weight, which is used to adjust the global dependency; , , are trainable parameters, all included in the chain rule of back propagation, ranging from [0.2, 0.9], to avoid extreme weights leading to modal suppression or overfitting, and are optimized by gradient descent during model training to ensure that the model can adaptively learn the importance of different modal features; Represents the Hadamard product, which is element-by-element multiplication; It is a lightweight cross attention and belongs to the single-head attention mechanism.

10. The surveillance video review device according to claim 8, characterized in that: include: The multimodal fusion features are positionally encoded to obtain a model input for extracting global dependency features while retaining the temporal order information: ; ; in, represents the model input for extracting global dependency features, , is the time step Multimodal features, M1 represents behavior recognition features, and M2 represents audio features; represents the temporal position information obtained by positionally encoding the model input for extracting global dependency features; , is the learnable position basis matrix, Represents the feature dimension of the model input; the length of the input sequence fed into the model; is a depth-wise separable convolutional layer, where is the depth convolution kernel ; is the convolution operation along the time dimension; is the convolution kernel, the size is 1 to 6, is the point-wise convolution layer, is a 1×1 convolution kernel , Represents the feature dimension of the model input.

Citation Information

Cited By

  • Video retrieval method based on retrieval agent

    CN120892603A

  • Video information base construction method, video searching method and related device

    CN121188238A

  • Video monitoring alarm method and device, electronic equipment and storage medium

    CN121582838A

  • Monitoring search method and device based on natural language and storage medium

    CN121681873A