Method for automatically generating monitoring video language data set with controllable subject and behavior deviation

Through deep learning and cross-validation methods, surveillance videos are separated into subject and behavior parts to generate high-quality video language datasets. This solves the problem of lack of detailed description of surveillance video datasets, reduces annotation costs, provides effective training data, and improves multimodal understanding capabilities.

CN120670746AActive Publication Date: 2025-09-19BEIJING UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510662239.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-19
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

Existing surveillance video datasets lack detailed sentence-level text descriptions, which limits the research on multimodal video language understanding. In addition, manual annotation is costly and inefficient, and the quality of machine annotation is not high.

Method used

Using deep learning technology and mathematical cross-validation methods, surveillance videos are separated into two parts: subject and behavior. An enhanced subtitle model is constructed using a target tracking module. Combined with an iterative deviation cross-validation filtering algorithm, a video language dataset with controllable subject and behavior deviations is generated.

Benefits of technology

Automatically generate high-quality video language annotation datasets, reduce manual annotation costs, provide training data with known noise, and improve the training effect of multimodal surveillance video language understanding models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670746A_ABST
    Figure CN120670746A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically generating a monitoring video language data set with controllable subject and behavior deviation, and belongs to the field of computer vision. The invention relates to a method for automatically generating a monitoring video language data set with controllable subject and behavior deviation based on a deep learning technology and a cross validation method in mathematics. Firstly, a target tracking module is used for constructing an enhanced monitoring video subtitle model for labeling and generating a description text of a monitoring video, and the subject deviation degree in the description text is controlled. And filtering the description text by using a data filtering model based on iteration deviation cross validation, controlling the behavior deviation degree in the text description, and finally obtaining a video language data set with controllable main body and deviation. The data set produced by the method has a known subject and behavior deviation degree, so that effective help can be provided for training tasks such as a multi-mode monitoring video language understanding model and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Based on deep learning technology and cross-validation methods in mathematics, the present invention studies a method for automatically generating a surveillance video language dataset with controllable subject and behavior deviations. First, a target tracking module is used to construct an enhanced surveillance video subtitle model, which is used to annotate and generate descriptive text for surveillance videos and control the degree of subject deviation in the descriptive text. Subsequently, a data filtering model based on iterative deviation cross-validation is used to filter the descriptive text to control the degree of behavioral deviation in the text description, and finally a video language dataset with controllable subject and deviation can be obtained. The present invention belongs to the field of computer vision, and specifically relates to deep learning, target detection, cross-validation and other technologies. Background Art

[0002] Intelligent surveillance plays a vital role in public security and has become a cornerstone for maintaining social order and safeguarding people's lives and property. Multimodal video language understanding technology focuses on understanding the relationship between surveillance video and language modalities. It can automatically understand the content of surveillance videos and generate detailed text descriptions. This provides strong support for video language understanding applications in surveillance, such as video question answering, video subtitle generation, and dense video subtitle generation. It can significantly improve the efficiency of surveillance video investigation and analysis in real-world scenarios, providing important technical support for the development of a new generation of intelligent security.

[0003] However, at the data level, research on multimodal video language understanding requires relying on surveillance video datasets with well-documented text descriptions. However, currently available public video text description datasets target general video domains, while most existing surveillance video data is labeled only with abnormal event category labels and time information. Sentence-level text descriptions related to the video content are largely missing, severely limiting research on multimodal surveillance video language understanding. To address this issue, given the high cost and low efficiency of manual annotation and the low quality of automated machine annotation, a combination of manual and machine annotation should be employed to construct and expand surveillance video datasets with well-documented text descriptions. Previous studies have shown that, despite the low quality of machine-annotated data, machine-generated data can still serve as an important source of training data when the deviation from manually annotated data is manageable.

[0004] Therefore, studying how to use machines to automatically generate text annotations for surveillance videos and controlling the deviation between machine and manual annotations is a problem that technicians in this field need to solve. Summary of the Invention

[0005] This invention differs from existing dataset generation methods by utilizing deep learning techniques and mathematical cross-validation methods. This method, targeting the visual characteristics of surveillance videos, divides them into two major visual components: subject and behavior, to investigate controllable deviations. The method then investigates a method for automatically generating a surveillance video language dataset with controllable deviations between subject and behavior. This method divides the primary visual information of surveillance videos into two key aspects: subject and behavior. First, the method proposes integrating a target tracking model as a new branch into the video caption generation model to enhance the existing caption generation model's ability to identify and describe the target subject performing the action, thereby achieving high-quality video caption generation. It then proposes controlling behavioral deviations in the dataset's text descriptions through a low-quality data filtering algorithm, thereby estimating the degree of behavioral deviation. Finally, the overall deviation of the machine-annotated data is estimated based on the subject and behavior deviations. Because the degree of deviation between subject and behavior in the generated dataset is known, it can serve as an important supplementary source of training data for multimodal surveillance video language understanding models.

[0006] The main process of this method is as follows Figure 1 As shown in the figure, it can be divided into the following four steps: generating surveillance video text data with controllable subject deviation, estimating the degree of subject deviation of automatically annotated text data, filtering low-quality behavior deviation data based on iterative deviation cross-validation, and generating a dataset with controllable subject and behavior deviation.

[0007] 1. Generating text data from surveillance videos where the subject deviates from the controllable subject

[0008] The target tracking module is used to build an enhanced surveillance video subtitle generation model to generate descriptive text for surveillance videos. The degree of subject deviation in the descriptive text is controlled to improve the existing subtitle generation model's ability to identify and describe the target subject of the behavior, thereby achieving high-quality video subtitle generation.

[0009] 2. Estimation of Subject Deviation in Automatically Annotated Text Data

[0010] The existing dataset is used to evaluate the subject deviation degree of the text generated by the surveillance video subtitle model. The Euclidean distance between the feature vectors of the subject vocabulary is calculated to obtain the current subject deviation degree. Based on this deviation degree, the subject deviation degree of the automatically annotated text data is estimated, thereby ensuring that the subject deviation of the automatically annotated text is known.

[0011] 3. Low-quality data filtering based on iterative deviation cross-validation

[0012] For the generated video subtitles, a low-quality data filtering model based on iterative deviation cross-validation is used to filter the description text to control the degree of behavioral deviation in the text description, further improving the quality of video subtitle generation.

[0013] 4. Generation of datasets with controllable subject and behavior deviations

[0014] The sample behavior deviation is calculated based on the data in the final generated dataset, and finally a noise-known dataset with controllable subject and behavior deviation is obtained, laying a data foundation for downstream tasks such as multimodal surveillance video language understanding models.

[0015] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects:

[0016] First, traditional manual annotation methods are extremely expensive and time-consuming, and their annotation quality varies widely, making them difficult to apply to training large-scale and effective multimodal surveillance video language understanding models. Furthermore, the vast majority of existing datasets related to video language understanding tasks are general domains, while most existing surveillance video data is labeled only with abnormal event category labels and time information. Sentence-level textual descriptions related to the video content are largely missing, severely limiting research on multimodal surveillance video language understanding.

[0017] To address these issues, the present invention proposes a method for automatically generating a surveillance video language dataset with controllable subject and behavior deviations. This method can automatically generate high-quality video language annotated datasets, saving the cost and time of manual annotation. Furthermore, the present invention can also determine the degree of subject and behavior deviation in the generated dataset, effectively providing a surveillance video language dataset with known noise. This method can therefore effectively assist in the training of multimodal surveillance video language understanding models and other tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Flowchart of the method for automatically generating a surveillance video language dataset with controllable subject and behavior deviations.

[0019] Figure 2 A video subtitle generation model that integrates global and local features into video coding strategies.

[0020] Figure 3 A two-way model for behavioral description quality detection in machine-annotated text data.

[0021] Figure 4 Lightweight temporal network model DETAILED DESCRIPTION

[0022] Based on the above description, the following is a specific implementation process, but the scope of protection of this patent is not limited to this implementation process.

[0023] Step 1: Generate text data from surveillance videos where the subject deviates from the controllable

[0024] Step 1.1: Use the object tracking module to build an enhanced surveillance video captioning model

[0025] The target tracking module is used to build an enhanced surveillance video caption model. A video encoding strategy that integrates global and local features is adopted. The model architecture is shown in the attached figure. Figure 2 As shown in the figure. First, the model performs a fixed sampling on the input original video, extracts several video frame segments, and extracts global video features through the visual encoder. At the same time, a pre-trained target detector is used to identify and locate the object in each video frame, and then the target filter is used to track and analyze the motion trajectory of the detected object in the video frame, thereby further filtering out the objects that have changed in consecutive frames, and masking the other parts outside the object detection frame to obtain an image containing only the tracked target, which is then input into the visual encoder to obtain local video features. Finally, the global video features and local video features output by the visual encoder are fused and input into the multimodal decoder together with the text features to predict the text output. Overall, the subtitle generation model designed in this section can construct an enhanced visual feature representation that focuses on the active target subject by fusing global video features with filtered local features, which is used to generate text annotations with low deviation from the subject description.

[0026] For example, the two basic models SwinBERT and SORT can be selected as the main branch of video subtitles and the branch of target tracking respectively. The model can be trained using the existing video language dataset in the surveillance field. With the enhanced target subject recognition ability, the model can generate more accurate subject-related text descriptions compared to the original SwinBERT. During the training of the model, the enhanced visual features are input into the multimodal decoder together with the text features. The model will use the enhanced visual features and the generated text sequence to predict the next word. The loss function of the model training continues to use the original cross entropy loss, that is, given the correct index of the previous t-1 words And the visual representation ν, we can get the current t-th word For predictions, use p y represents the probability, and the training loss is calculated as shown in formula (1), where l1 is the word length.

[0027]

[0028] Step 1.2: Generate surveillance video text data

[0029] The collected surveillance videos are fed into the aforementioned surveillance video captioning model to obtain descriptive text for the surveillance videos. The video input is the surveillance video data to be annotated. The text input begins with a start symbol, and the model predicts the next text. The model then continues predicting the next text until the end symbol appears. In this way, a descriptive text annotation is generated for each surveillance video data point.

[0030] Step 2: Estimation of subject deviation degree of automatic annotation text data

[0031] Because the target subject categories in various surveillance videos vary little and are relatively easy to identify, the subject deviations of the trained surveillance video captioning model are relatively consistent when generating text annotations for any surveillance video. In other words, the subject deviations of the generated text dataset can be directly estimated from the subject deviations of existing surveillance video language datasets. Specifically, the subject vocabulary automatically annotated by the surveillance video captioning model on existing surveillance video language datasets is first represented in pairs with the subject vocabulary manually annotated. These vocabulary words are then input into BERT to extract feature vectors, and the deviations between these vocabulary feature vectors are calculated.

[0032] To further improve the accuracy and robustness of subject deviation estimation, this paper proposes a subject deviation measurement method based on semantic distribution alignment. Compared with traditional distance calculations between word pairs, this method evaluates the overall deviation of machine-annotated and human-annotated texts in the semantic space from a holistic distribution perspective.

[0033] Specifically, the main vocabulary set in the machine-annotated text data is represented as S m , the main vocabulary set in the manually annotated text data is represented as S a First, a pre-trained language model (such as BERT) is used to map each subject vocabulary into a unified feature vector space to obtain a feature set.

[0034] Then, the machine annotation subject set S is calculated separately m and the manually annotated subject set S a The characteristic mean vector μ m 、μ a and the feature covariance matrix Σ m ,Σ a Based on the mean and covariance information, the semantic distribution alignment deviation ε1 is defined as follows:

[0035]

[0036] Among them, ||μ m -μ a ||2 represents the Euclidean distance between the feature means of the machine-annotated and manually-annotated subjects; Tr(·) represents the trace operation of the matrix, which is used to measure the overall difference between the covariance matrices; The matrix representing the square root of the product of covariance matrices is used to estimate the similarity of feature distributions.

[0037] By comprehensively considering both mean deviation and covariance structure deviation, this method effectively captures the overall distributional consistency differences between machine-annotated and human-annotated text in the semantic space, avoiding the problem of local error amplification caused by matching only single word pairs. Ultimately, the subject deviation assessment results based on ε1 can serve as an important basis for overall quality control of machine-annotated text data generated by surveillance video captioning models.

[0038] Step 3: Low-quality data filtering based on iterative deviation cross-validation

[0039] Step 3.1: Train a two-way model that verifies the quality of machine-standard text behavior

[0040] A two-way model is trained to estimate whether the behavior description in the input machine-annotated text data is accurate. Figure 3 As shown in Figure 3, the model includes an anomaly detection model based on the C3D feature video branch (the model consists of a lightweight temporal network, see step 3.1.1 for specific model details), and an anomaly detection model based on the GPT3 large language model and prompt words, where the prompt words are used to tell the large language model to analyze whether the input text sentence contains abnormal behavior. The training data uses the existing video language dataset in the surveillance field, where the input of the anomaly detection model of the C3D feature video branch is the video in the dataset, and the input of the GPT3 large language model is the description text and prompt words in the dataset. The output results of both sides are to judge whether there is an anomaly, and the cross entropy loss is used to analyze the attached Figure 3 The neural network module can be trained.

[0041] Subsequently, a pair of videos and machine-labeled text descriptions are input into the trained model. By judging whether the outputs of the two branches in anomaly detection are consistent, the accuracy of the automatically labeled behavior description can be determined.

[0042] Step 3.1.1: Lightweight timing network construction

[0043] The video clip features extracted by C3D are input into the lightweight multi-scale time series module for anomaly detection. This part of the time series module is connected by two models, and the input features are respectively passed through the two modules, such as Figure 4As shown in the figure: one is a 1D hole convolution layer (convolution kernel size is 3, expansion rate r = 3), and the other is a lightweight Transformer layer: only a 1-layer Transformer (attention mechanism and feedforward network FFN), in which an improved window attention mechanism is used. The improved attention mechanism is as follows: the time axis is divided into blocks (such as window size = 5), and self-attention is calculated in each window. Cross-window jump connection: add a hop global connection every 2 windows (such as window 1→4→7...) to retain key long-range dependencies. The hole convolution output features and the Transformer output features are fused through residual connections, and the fused output features are then passed through Figure 3 The anomaly prediction head (single-layer MLP) outputs the anomaly probability.

[0044] The specific parameter settings of the model are as follows: the dilated convolution layer uses 256 convolution kernels (number of output channels), the convolution kernel time span is 3, the dilation rate is 3, and the input and output timing alignment is ensured by padding (padding) 6, and is matched with the ReLU activation function; the lightweight Transformer layer projects the input features to 256 dimensions, adopts an improved attention mechanism with a window size of 5, sets 4 attention heads (64 dimensions each), establishes cross-window jump connections every 2 windows (such as window 1→4→7), the hidden layer dimension of the feedforward network is 512 and uses GELU activation, and introduces layer normalization and residual connection; in the feature fusion stage, the dilated convolution and the 256-dimensional features output by the Transformer are added element by element, and finally the anomaly probability is generated through a single-layer MLP prediction head (input 256 dimensions, output 1 dimension, Sigmoid activation). The design achieves lightweightness with a unified dimension (256), complements local perception with Transformer long-range modeling through dilated convolution, and combines cross-window jump connections (global interaction every 15 frames) to balance computational efficiency and capture key timing dependencies, making it suitable for real-time video anomaly detection scenarios.

[0045] Step 3.2: Use cross-validation to filter out data sets that ultimately deviate from controllable values.

[0046] Using an iterative-validation loop, the algorithm repeatedly filters out samples with high behavioral deviations from the machine-annotated data over multiple iterations, retaining samples with low behavioral deviations. These samples are then used to train the two-way detection model initialized in step 3.1 and to test the confidence of the data samples. Specifically, in each iteration, the unfiltered candidate machine annotation set C (initially consisting of all data in the entire machine-annotated dataset) is divided into two disjoint parts. One part serves as the training set to train the two-way detection model until the model converges (i.e., the loss barely decreases), while the other part serves as the test set for sample filtering. During the training of the two-way model, the anomaly detection and temporal network components of the video branch undergo parameter updates, while the parameters of the other components do not need to be updated. During the testing process, the algorithm stores the samples in the test set that are accurately judged by the two-way model into a sample set S1 of controlled behavioral deviation data. Low-confidence samples are removed from this set according to a pre-set removal ratio γ, further improving the cleanliness of S1. Subsequently, the test set and training set are swapped, and the reinitialized two-way model is used for training and testing. The above process is repeated to obtain another sample set S2 with controllable behavior deviation data. After that, S1, S2, and the data that the two-way model judged inaccurately are removed from the candidate machine annotation set C, and S1 and S2 are added to the sample set S. The two-way detection model pre-trained on the manually annotated dataset is then used to test the accuracy ρ of the set S. Subsequently, according to the formula ρ = (1-ε2) 2 +ε2 2 The behavioral deviation ε2 of the data samples in S is calculated from the accuracy ρ. If the deviation is too high (greater than 50%), the low-quality data in the dataset can be filtered more aggressively by adjusting the value of the removal ratio γ.

[0047] Finally, after the deviation requirement is met and multiple iterations are completed, the sample set S is the machine-annotated text dataset output by this step.

[0048] Step 4: Generate a dataset with controllable subject and behavior deviations

[0049] According to the data in the final sample set S, use the formula ρ=(1-ε2) 2 +ε2 2 By calculating the behavioral deviation ε2 of the data sample and combining it with the subject deviation ε1 obtained in step 2, we can obtain the overall deviation array of the dataset as ε = [ε1, ε2]. This dataset has known noise, so it can serve as an important source of training data for multimodal surveillance video language understanding models. Subsequently, weakly supervised neural network training schemes that account for data noise can be designed to enhance multimodal surveillance video language understanding capabilities.

Claims

1. A method for automatically generating a surveillance video language dataset with controllable subject and behavior deviations, characterized by: For any surveillance video set, we first use the target tracking module to build an enhanced surveillance video captioning model, which is used to automatically annotate and generate descriptive text for the surveillance video and control the degree of subject deviation in the descriptive text; then we use a data filtering model based on iterative deviation cross-validation to filter the descriptive text and control the degree of behavioral deviation in the text description. Finally, we can obtain a video language dataset with controllable subject and deviation.

2. The method according to claim 1, wherein: An enhanced surveillance video captioning model is constructed using a target tracking module. The model adopts a video encoding strategy that integrates global and local features. First, the input original video is fixedly sampled to extract several video frame segments, and global video features are extracted through a visual encoder. At the same time, a pre-trained target detector is used to identify and locate the object in each video frame. The target filter is then used to track and analyze the motion trajectory of the detected object in the video frame, thereby further filtering out objects that have changed in consecutive frames and masking the other parts outside the object detection frame to obtain an image containing only the tracked target. The image is then input into the visual encoder to obtain local video features. Finally, the global video features and local video features output by the visual encoder are fused and input into the multimodal decoder together with the text features to predict the text output and generate text annotations with low deviation from the subject description. During the model training process, the existing video language dataset is used; the model uses the visual features enhanced by the target detector and target filter and the generated text sequence to jointly predict the next word; the loss function of the model training continues to use the cross entropy loss, that is, given the correct index of the previous t-1 words And visual representation ν, get the current t-th word The prediction, p y Expressed as a probability, its training loss is calculated as l1 is the word length.

3. The method according to claim 1, wherein: Description text is used to describe the events in the surveillance video.

4. The method according to claim 1, wherein: The subject deviation degree in the description text is expressed in pairs based on the existing video language dataset, where the subject words automatically annotated by the machine and the subject words manually annotated for the same video are represented. These words are then input into BERT to extract feature vectors, and the subject deviation degree measurement method based on semantic distribution alignment is used to calculate the deviation degree, and its value is recorded as ε1; specifically, The main vocabulary set in the machine-annotated text data is represented as S m , the main vocabulary set in the manually annotated text data is represented as S a ; First, use the pre-trained language model to map each subject vocabulary to a unified feature vector space to obtain a feature set; calculate the machine annotation subject set S m and the manually annotated subject set S a The characteristic mean vector μ m 、μ a and the feature covariance matrix Σ m ,Σ a Based on the mean and covariance information, the semantic distribution alignment deviation ε1 is defined as follows:

5. The method according to claim 1, wherein: In each iteration of the cross-validation algorithm, the algorithm first needs to initialize the two-way detection model used to determine whether the behavior description is accurate. Subsequently, the algorithm divides the unfiltered candidate machine annotation set C into two non-overlapping parts, one of which is used as a training set to train the two-way detection model, and the other is used as a test set for sample filtering. During the training of the two-way model, the anomaly detection and temporal network parts of the video branch will undergo parameter updates, while the parameters of other parts do not need to be updated. During the test process, the algorithm stores the samples in the test set that are accurately judged by the two-way model into the sample set S1 of controllable behavior deviation data, and removes the samples with low confidence according to the pre-set removal ratio γ, thereby further improving the cleanliness of S1. Subsequently, the above-mentioned test set and training set are exchanged, and the reinitialized two-way model is used for training and testing to obtain another sample set S2 of controllable behavior deviation data. Remove S1, S2 and the inaccurate data judged by the two-way model from the candidate machine annotation set C, and add S1 and S2 to the sample set S. Then use the pre-trained two-way detection model to test the accuracy ρ of the set S. According to the formula p=(1-ε2) 2 +e2 2 , Calculate the behavioral deviation ε2 of the data samples in S; After each iteration, the behavioral deviation ε2 will be updated with the measured accuracy ρ; after multiple iterations, if the deviation is less than 50%, the sample set S is a machine-annotated text dataset with controllable overall deviation.

6. The method according to claim 5, characterized in that: The dual-path model consists of an anomaly detection model based on C3D features and a temporal network branch, and an anomaly detection model based on the GPT3 large language model and prompt words. A pair of videos and machine-annotated text descriptions are input into the trained model. By determining whether the anomaly detection outputs of the two branches are consistent, the accuracy of the behavior description in the input machine-annotated text data is estimated.

7. The method according to claim 6, characterized in that: The temporal network branches in the anomaly detection model are specifically: The video clip features extracted by C3D are input into a lightweight multi-scale temporal module for anomaly detection. This temporal module is connected by two models. The input features pass through two modules respectively. One is a 1-layer 1D void convolution layer with a convolution kernel size of 3 and a dilation rate of r = 3, and the other is a lightweight Transformer layer: a single-layer Transformer, namely the attention mechanism and the feed-forward network FFN; the void convolution output features and the Transformer output features are fused through residual connections, and then the fused output features are output through the anomaly prediction head to output the anomaly probability.

8. The method according to claim 7, characterized in that The specific attention mechanism is as follows: divide the timeline into blocks, set the window size to 5, and calculate self-attention within each window; cross-window jump connection: add a hop global connection every 2 windows to establish key long-range dependencies.

9. The method according to claim 1, wherein: The subject and deviation controllable video language dataset is characterized by: The sample set S after the data filtering model based on iterative deviation cross-validation is a biased controllable video language dataset. The subject deviation of the machine-annotated text in this dataset is ε1, and the behavior deviation is ε2. The overall deviation array of the dataset can be obtained as ε=[ε1,ε2].

Citation Information

Patent Citations

  • Multi-modal feature fusion video description text generation method

    CN113806587A

  • Multimodal video behavior recognition method based on language-vision contrast learning

    CN117197708A

  • Supervision data generation method based on video semantic structured analysis

    CN119048964A

  • Image content automatic description method based on construction of chinese visual vocabulary list

    WO2021223323A1

  • Multimodal heterogeneous feature fusion-based compact video event description method

    WO2023050295A1