Monitoring video retrieval method and system based on visual large model under fuzzy description

Through the surveillance video retrieval method based on the video-text model, the fuzzy description is processed using Transformer and KAN networks to achieve efficient and accurate video retrieval in fuzzy scenarios, solving the search difficulties in the existing technology, and improving the practicality and accuracy of the security system.

CN120336582AActive Publication Date: 2025-07-18WUHAN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510355552.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-18
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately retrieve security surveillance videos under fuzzy scenarios, which makes manual retrieval time-consuming and labor-intensive and unable to meet the security needs of the big data era.

Method used

The surveillance video retrieval method based on the video-text big model is adopted. By creating a fuzzy behavior description and feature data set, text-to-text big model is trained, text and video feature extraction is used using the Transformer encoder and KAN network, and combined with a mixed parallel dual network training model, the transformation from fuzzy description to specific description and video matching is achieved.

Benefits of technology

It improves the accuracy and adaptability of surveillance video retrieval under fuzzy descriptions, and can quickly and accurately find surveillance videos that meet needs, improving the practicality and feasibility of the security system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336582A_ABST
    Figure CN120336582A_ABST
Patent Text Reader

Abstract

The invention discloses a monitoring video retrieval method and system based on a visual large model under fuzzy description. The method comprises the following steps: step 1, creating a fuzzy behavior description and feature data set, and training a text-to-text large model; 2, setting a text encoder, encoding a text, and generating a text feature vector; step 3, preprocessing the videos, aggregating the proposed features into a KAN network, setting a video encoder, and performing video encoding on a plurality of monitoring videos to form video feature vectors; step 4, based on the video-text matching pair, using a hybrid parallel dual-network parallel mode to train a text video matching model; and 5, analyzing and expanding characters input by the user by using the trained text interpretation large model, and finding a matched video according to new text description. The corresponding video can be searched according to the content described by the user, the application range is wide, and practicability and feasibility are good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of security monitoring and vision large models, and particularly to a method and system for multi-retrieving security monitoring videos based on a video-text large model in a fuzzy description scenario. Background Art

[0002] In modern society, with the increasing importance of security prevention, security issues have gradually become the focus of public attention. As the mainstream monitoring mode today, security monitoring videos are widely used in various places, including urban streets, railway stations, airports, shopping malls, etc. They are an important means to ensure public safety and an indispensable part of modern social management, and are of great significance for maintaining public order, combating crime, and dealing with emergencies.

[0003] At present, the retrieval of monitoring videos still mainly relies on manual labor. The perspectives of monitoring videos are diverse and the content is complex, making the retrieval work time-consuming and laborious, resulting in the inability to efficiently and smoothly complete tasks such as risk investigation and suspect arrest. Relying on manual video monitoring retrieval is difficult to meet the security needs in today's big data era. In recent years, with the continuous development of computer vision and artificial intelligence large model technologies, there have been attempts to apply technologies such as image processing and video-text large models in the field of security monitoring videos, but there is no precedent for video retrieval in fuzzy scenarios. Therefore, it is urgent to propose a new, efficient method for retrieving security monitoring videos in fuzzy scenarios. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for retrieving security monitoring videos based on a video-text large model in a fuzzy scenario, which is simple and practical, has good adaptability to retrievals based on descriptions of different degrees of fuzziness, and has high accuracy.

[0005] The technical solution adopted by the present invention to solve its technical problems is:

[0006] The present invention provides a method for retrieving security monitoring videos based on a video-text large model in a fuzzy scenario, the method comprising the following steps:

[0007] Step 1, create a fuzzy behavior description and feature data set, and train a text-to-text large model for converting fuzzy descriptions into more specific and structured text information;

[0008] Step 2, first preprocess the text information obtained in Step 1, then set a Transformer-based text encoder to encode the preprocessed text to generate a text feature vector;

[0009] Step 3: First, perform video frame sampling and feature extraction, and then based on the KAN network, perform temporal and semantic aggregation to obtain a feature vector representing the semantics of the entire video, and extract key frames;

[0010] Step 4: Train the text-video matching model through a hybrid parallel dual-network parallel mode. The text-video matching model includes a text branch and a video branch. The text branch is the Transformer-based text encoder in Step 2, and the video branch is the KAN network;

[0011] Step 5: Use the trained text-video matching model to parse and expand the user input text, and find the matching video according to the new text description.

[0012] Furthermore, the text-to-text large model in Step 1 is the T5 or BART model.

[0013] Furthermore, in Step 2, the preprocessing of the text information includes:

[0014] Step 201: Perform word segmentation on the input text, and separate it into several words by spaces.

[0015] Step 202: By looking up the embedding matrix of clip, convert each word into a vector of a fixed dimension: X = [e1, e2, …, en], where ei represents the embedding vector of the i-th word.

[0016] Step 203: Use the sine function to mark the position of each word. The formula is as follows:

[0017]

[0018] where: i represents the position of the current word in the sentence, j represents half of the embedding dimension, d is the dimension of the word embedding, and 10000 is the scaling factor.

[0019] Finally, a fixed vector P representing the position will be generated for each position i, and it is added to the word embedding X to obtain the final vector:

[0020] T = x + p.

[0021] Furthermore, in Step 2, the specific implementation method for generating the text feature vector is as follows:

[0022] First, use the multi-head self-attention mechanism to calculate the relevance between words. The specific sub-steps are as follows:

[0023] Step 20411: For each layer of Transformer, the input vector T is linearly transformed to obtain the query Q, key K, and value V:

[0024] Q = TWQ

[0025] K = TW K

[0026] V = TW V

[0027] where W Q ,W K ,W V are all learnable parameter matrices, with shapes of (feature vector dimension, feature dimension per attention head);

[0028] Step 20412, calculate the attention scores:

[0029]

[0030] where, QK T Calculate the dot product of the query and the key to measure the correlation between each pair of words; Perform scaling; to avoid gradient explosion caused by overly large values.

[0031] Step 20413, normalize the attention scores to obtain attention weights, and multiply the value V by the weights to get the final output;

[0032] Step 2042, skip connection: Add the input vector T to the vector T after being processed by the self-attention mechanism in Step 2041, and after layer normalization, obtain the vector T1;

[0033] Step 2043, further process the vector T1 obtained in the previous step using the feed-forward neural network FFN, and the operation formula is as follows:

[0034] FFN(x) = max(0, xw1 + b1)w2 + b2

[0035] where x is the vector T1 obtained in Step 2042, w1 and w2 are weight matrices learned during the training process, and b1 and b2 are updated together during the training process;

[0036] Step 2044, stack multiple layers of Transformer: The Transformer encoder contains several transformer layers. After the first layer of operation is completed, use the output result of the first layer as the input for the second layer of operation, and repeat multiple layers until all layers of operation are completed;

[0037] After being processed by L layers of Transformer operations, a sequence of processed text vectors is obtained, and the feature vector at the end position is used as the final global text feature vector.

[0038] Furthermore, in step 3, uniform sampling is performed at a fixed frame rate r, and a fixed number of frames Ns = frame rate * sampling interval are extracted per second. The embedded features and high-level semantic information of each frame are extracted through an image encoder, namely the VIT model.

[0039] Furthermore, in step 3, the processing process of the KAN network is as follows:

[0040] First, the features extracted from the video are mapped to an intermediate feature space after preprocessing, denoted as where W h and b h are learnable parameters;

[0041] During temporal aggregation, for the inherent temporal dynamics of the video frame sequence, the features at each moment are processed layer by layer through several KAN modules to form an expression of local temporal information; specifically, the output of each layer can be written as where the initial input is the preprocessed feature, l represents the number of layers, and φ l represents the KAN module; a local attention mechanism or a kernel function-based weighting strategy can be embedded in each layer to capture the fine-grained changes between adjacent frames within the local time domain using the B-spline function; after processing, the outputs of each frame in the temporal branch are weighted or averaged to obtain the global temporal aggregation feature

[0042] In the semantic aggregation process, for the high-level semantic information of each frame feature extracted by the image encoder is processed through several KAN modules to obtain another set of features and then hierarchically processed by a set of independent KAN modules, and the output form of each layer is represents an independent KAN module, and then a non-linear transformation is performed through the B-spline activation function. Finally, using a learnable semantic attention or weighted summation mechanism, the processing results of each frame are integrated into the global semantic aggregation feature

[0043] Finally, the global temporal aggregation feature and the global semantic aggregation feature are fused by weighted summation, that is where λ ∈ [0, 1] is a hyperparameter that adjusts the contributions of the two parts, and the global video feature with temporal dynamics and semantic consistency is obtained. After normalization, the feature vector z’ representing the semantics of the entire video is obtained v .

[0044] Furthermore, the formula for extracting key frames in step 3 is:

[0045]

[0046] Among them, F i represents the feature vector of the i-th frame, and the frame with the largest d(i) is the key frame.

[0047] Furthermore, the specific implementation of step 4 is as follows:

[0048] First, calculate the global similarity, that is, the similarity between the video and text vectors obtained in steps 2 and 3. The formula is as follows:

[0049]

[0050] Among them, v and t represent the global feature vectors of the video and text respectively;

[0051] Then calculate the local matching degree based on the key frames: calculate the similarity between the feature vector of each key frame i and the text description, denoted as s i , and assign weights to each key frame based on the similarity. The formula is as follows:

[0052]

[0053] The weight α of each key frame i is automatically adjusted according to its matching degree with the text. The better the matching frame, the greater the weight. Calculate the local total matching degree:

[0054]

[0055] Weighted sum the global similarity and the local matching degree to obtain the final similarity S;

[0056] Then train the model based on the idea of contrastive learning. The contrastive loss function is:

[0057]

[0058] Among them, i′ represents the index of the positive sample pair, j is the index of all sample pairs in the traversal batch, where τ is the temperature parameter;

[0059] During the training process, adopt a hybrid parallel dual-network training method by combining data parallelism and model parallelism. Deploy the text branch and the video branch on different devices respectively to achieve parallel computing of forward propagation, and calculate gradients independently on each branch. Subsequently, adopt a synchronous gradient reduction mechanism to ensure consistent parameter updates among devices, so as to simultaneously transmit the gradients generated by the contrastive loss function to the Transformer-based text encoder and the KAN network during the backpropagation process. Finally, when the value of the contrastive loss function converges to a certain value, the training is completed.

[0060] Furthermore, the specific implementation of step 5 is as follows:

[0061] Input the text into the text-to-text large model in step 1 to obtain a detailed descriptive text with semantic expansion. Subsequently, input the detailed text description into the video-text large model obtained in step 4 to obtain the video with the maximum similarity, that is, the surveillance video that best matches the description.

[0062] The present invention also provides a surveillance video retrieval system based on a vision large model under fuzzy descriptions, including:

[0063] A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a method for retrieving surveillance videos based on a vision large model under fuzzy descriptions as described in the above technical solution.

[0064] Compared with the prior art, the advantages and beneficial effects of the present invention are:

[0065] 1. Retrieving surveillance videos based on a video-text large model, with strong learning ability, having unique advantages in accuracy, wide application range, and being able to retrieve surveillance videos of different types in different scenarios according to the content description text.

[0066] 2. Introducing a text-to-text large model to process the text input by the user, which can parse the fuzzy and inaccurate description text input by the user into a description with rich details and substantial content, significantly improving the retrieval ability of surveillance videos in fuzzy scenarios, and further enhancing the accuracy and practicality of retrieval.

[0067] 3. Using the KAN network for semantic enhancement of videos and Transformer for semantic enhancement of text, and training based on a dual network, the video-text large model has high accuracy, with superiority that is difficult to match by other methods, filling the technical gap in this aspect of video retrieval, and providing new ideas and methods for the security industry in our country.

[0068] In summary, the method of the present invention is simple and practical, has good adaptability for retrieving surveillance videos under fuzzy descriptions, can find the surveillance videos that meet the requirements based on inaccurate and detail-lacking descriptions, and has good practicality and feasibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0070] Figure 1 is the flow chart of the present invention.

[0071] Figure 2 is the schematic diagram of the principle of the multi-head attention mechanism.

[0072] Figure 3 is the schematic diagram of the core idea of the KAN network.

[0073] Figure 4 This is the schematic diagram of the final effect of the present invention. Specific implementation method

[0075] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0076] A fuzzy scenario refers to a situation where the description content is blurred due to the user's memory loss, mainly including scenarios where descriptive words are too general, specific details are missing, abstract or emotional words are used, etc. A video-text large model is a type of large-scale deep learning model that jointly models text information and visual information through multimodal learning technology. It can extract semantic features from natural language descriptions and align them with the feature space of visual content, thereby achieving cross-modal understanding, generation, and retrieval tasks. Such models are usually pre-trained based on large-scale multimodal data and adopt a dual-encoder or joint-encoder architecture, significantly improving performance in scenarios such as image-text alignment, inter-modal retrieval, content classification, and intelligent parsing.

[0077] Accordingly, the present invention has developed a technology for retrieving security surveillance videos based on a video-text large model in a fuzzy scenario. When the user's memory is missing and they cannot accurately and carefully describe the content of the picture, the required surveillance video can be quickly and accurately found through a fuzzy content description, which is of great significance to social security and provides strong support for maintaining social public order and combating crimes.

[0078] As Figure 1 shown, the surveillance video retrieval method based on a video-text large model in a fuzzy description scenario according to an embodiment of the present invention includes the following steps:

[0079] Step 1: Create a dataset of fuzzy behavior descriptions and features, and train a text-to-text large model. To achieve intelligent parsing of fuzzy descriptions, we first need to construct a dataset of behavior descriptions covering different scenarios and train it in combination with video scene features. This dataset contains various possible user input texts, such as "someone is wandering around" and "a suspicious person approaches the restricted area", and at the same time, corresponding standardized behavior descriptions are established to improve the accuracy of text parsing. We use data augmentation methods to expand multiple expressions of the same behavior to adapt to different user expression habits. In addition, we construct a text-to-text large model, which is based on natural language processing technology and can understand and convert fuzzy descriptions into more specific and structured text information as input data for subsequent video retrieval.

[0080] Step 2: Set up a text encoder to encode the text and generate text feature vectors. After training is completed, we use the pre-trained text encoder to vectorize the description text input by the user. The text encoder converts the input text into high-dimensional feature vectors for subsequent calculation of its similarity with the video content. In this process, we adopt deep learning models such as Transformer or BERT variants to enable the text encoding to have stronger semantic understanding capabilities. In addition, we normalize the text encoding results to ensure that text descriptions of different lengths can obtain consistent feature representations during the matching process, thereby improving the matching accuracy.

[0081] Step 3: Set up a video encoder to perform video encoding on several surveillance videos. While completing the text encoding, we need to extract features from the surveillance video data for comparison with the text features. We adopt pre-trained video encoding models such as architectures based on 3D CNN or spatio-temporal Transformer to perform frame-by-frame analysis on the input video and extract key frame features. During the feature extraction process, we perform normalization processing on videos with different frame rates and resolutions to ensure the stability of the features. After encoding, each video segment is mapped to a high-dimensional feature vector for comparison with the text vector in the same feature space.

[0082] Step 4: Use video-text matching pairs to train the text-video matching model based on contrastive learning. To achieve efficient matching between text and video, we use text-video matching pairs to construct a training dataset and adopt contrastive learning methods to optimize the matching model. During training, we find the corresponding video segments for each text description and introduce negative sample contrast loss to enable the model to learn to distinguish relevant videos from irrelevant videos. During the training process, we introduce self-supervised learning strategies and use unlabeled data to further improve the generalization ability of the model, so as to still achieve accurate matching in complex scenarios.

[0083] Step 5: Use the trained large text interpretation model to parse and expand the text input by the user, and find the matching video based on the new text description. After the model training is completed, we use the trained large text interpretation model to parse the vague description input by the user and automatically expand its semantic scope. For example, when the user inputs "Someone is walking strangely", the system will parse it into more specific descriptions such as "Someone is wandering at the same place for a long time" or "Someone repeatedly enters and exits a certain area". Then, based on the expanded text description, we use the video-text matching model to retrieve relevant segments in the surveillance video library. The retrieval results are sorted by similarity and multiple candidate videos are provided for the user or security personnel to review.

[0084] In the above-mentioned surveillance video retrieval method based on a video-text large model in a fuzzy description scenario, the specific process is as follows:

[0085] Step 101, Behavior description and feature collection: Collect possible target behaviors (such as "running"), appearance features (such as "wearing red clothes"), and scene features (such as "dim environment") in security surveillance to generate a description sample set {d1, d2, …, dn}.

[0086] Step 102, Semantic fuzzification: Perform semantic fuzzification processing on the description samples (for example, expand "running" to "moving quickly" and "sneaky" to "repeatedly moving near a location") to form a preliminary fuzzy behavior description data set D = {t1, t2, …, tn}.

[0087] Step 103, Train the text generation model: Use the constructed data set to train a generative text model (such as T5 or BART) so that it can generate more accurate and diverse descriptions from fuzzy descriptions.

[0088] The specific method in Step 2 is as follows:

[0089] Step 201, Tokenize the input text and separate it into several words by spaces.

[0090] Step 202, Convert each word into a vector of a fixed dimension by looking up the embedding matrix of clip: X = [e1, e2, …, en], where ei represents the embedding vector of the i-th word.

[0091] Step 203, Use the sine function to mark the position of each word. The formula is as follows:

[0092]

[0093] where: i represents the position of the current word in the sentence, j represents half of the embedding dimension, d is the dimension of the word embedding, and 10000 is the scaling factor to ensure reasonable frequency changes in different dimensions.

[0094] Finally, a fixed vector P representing the position will be generated for each position i, and it is added to the word embedding X to obtain the final vector:

[0095] T = x + p

[0096] Step 204, Use a multi-layer Transformer encoder to calculate the global features:

[0097] Step 2041, As Figure 2 shown, use the multi-head self-attention mechanism to calculate the correlation between words:

[0098] Step 20411: For each layer of the Transformer, the input vector T is linearly transformed to obtain the query (Q), key (K), and value (V):

[0099] Q = TW Q

[0100] K = TW K

[0101] V = TW V

[0102] where W Q , W K , W V are all learnable parameter matrices with shapes (feature vector dimension, feature dimension of each attention head).

[0103] Step 20412: Calculate the attention scores:

[0104]

[0105] where, QK T Calculate the dot product of the query and the key to measure the correlation between each pair of words; Perform scaling to avoid gradient explosion caused by overly large values.

[0106] Step 20413: Normalize the attention scores to obtain the attention weights. Multiply the value V by the weights to get the final output.

[0107] Step 2042: Skip connection: Add the input vector T to the vector T after being processed by the self-attention mechanism in Step 2041, and then perform layer normalization to obtain the vector T 1 .

[0108] Step 2043: Further process the vector T obtained in the above step using the feed-forward neural network FFN 1 , and the core operation formula is as follows:

[0109] FFN(x) = max(0, xw1 + b1)w2 + b2

[0110] where x is the vector obtained in the previous step, and w1 and w2 are weight matrices learned during the training process. b1 and b2 are usually updated together during the training process to adjust the network output.

[0111] Step 2044: Stack multiple layers of Transformers: The Transformer encoder contains several Transformer layers. After the first layer of operation is completed, use the output result of the first layer as the input for the second layer of operation, and repeat the above steps until all layers of operation are completed.

[0112] Step 205: Extract global text features: After being processed by L layers of Transformer operations, a processed text vector sequence is obtained, and the feature vector at the end position is used as the final global text feature vector.

[0113] The specific method in Step 3 is as follows:

[0114] Step 301: Video frame sampling and feature extraction: Divide the surveillance video into several frame sequences, perform uniform sampling using a fixed frame rate r, and extract a fixed number of frames Ns per second = frame rate * sampling interval. Each frame extracts its embedded features and high-level semantic information through an image encoder (VIT model).

[0115] Step 302: In order to generate a feature vector z that can represent the semantics of the entire video v , based on KAN, perform temporal modeling aggregation and semantic enhancement on the embedded features of all frames :

[0116] First, map the frame embedded features extracted from the video to the intermediate feature space through preprocessing, which can be formally expressed as where W h and b h are learnable parameters. The KAN network uses the theoretical basis of the Kolmogorov - Arnold representation theorem to represent a multivariate continuous function as a combination of a finite number of univariate functions and addition operations. Therefore, in KAN, each layer is composed of univariate functions, and these functions are usually parameterized through B-spline activation functions to switch between coarse-grained and fine-grained grids, achieving flexible approximation of the input features. As Figure 3 shown, the left figure is the activation symbol flowing through the network; the right figure is that the activation function is parameterized as a B-spline and can switch between coarse-grained and fine-grained grids.

[0117] Based on the above principle, during temporal aggregation, for the inherent temporal dynamics of the video frame sequence, through several layers of KAN modules (denoted as ), process the features at each moment layer by layer to form an expression of local temporal information. Specifically, the output of each layer can be written as where the initial input is the preprocessed feature. Due to the continuity of video frames in time, a local attention mechanism or a kernel function-based weighting strategy can also be embedded in each layer to capture the fine-grained changes between adjacent frames within the local time domain using B-spline functions. After processing, perform appropriate weighting or averaging operations on the outputs of each frame in the temporal branch to obtain the global temporal aggregation feature

[0118] Meanwhile, during the semantic aggregation process, for the high-level semantic information of each frame feature extracted by the image encoder, another set of features is obtained through the same processing as above. Then, a set of independent KAN modules (denoted as ) perform hierarchical processing. The output form of each layer is This branch focuses on extracting semantic features within the frame. It performs a non-linear transformation on the input through a B-spline activation function, and finally uses a learnable semantic attention or weighted summation mechanism to integrate the processing results of each frame into global semantic features.

[0119] Finally, the temporal aggregation features and semantic aggregation features are fused by weighted summation, that is, where λ ∈ [0, 1] is a hyperparameter that adjusts the contributions of the two parts, and global video features that comprehensively consider temporal dynamics and semantic consistency are obtained. To facilitate distance calculation or classification processing in subsequent tasks, it is also necessary to normalize the fused features, and finally output the normalized video features.

[0120] This dual-branch aggregation method based on the KAN network not only theoretically relies on the Kolmogorov-Arnold theorem to ensure the approximation ability of multivariate continuous functions, but also in practice uses the B-spline activation function to switch between fine and coarse grids, endowing the model with higher flexibility and expressive power, thereby effectively improving the discriminability and robustness of video feature representation.

[0121] Step 303: Extract key frames:

[0122] Use a 5-frame sliding window. For the frames from the 2nd to the 5th frame within the window, calculate the cosine distance between each frame and the previous frame as a criterion for measuring the difference between frames. The formula is as follows:

[0123]

[0124] where F i represents the feature vector of the i-th frame. The numerator represents the difference between the feature vectors of the two frames, and the denominator is the product of the norms of the respective vectors.

[0125] Within the window, find the frame with the largest d(i) as the key frame with the largest change and the best representation of this window.

[0126] The specific method in Step 4 is as follows:

[0127] First, calculate the global similarity, that is, the similarity between the video and text vectors obtained in Step 2 and Step 3. The formula is as follows:

[0128]

[0129] Among them, v and t respectively represent the global feature vectors of the video and the text.

[0130] Then, calculate the local matching degree based on the key frames. Calculate the similarity between the feature vector corresponding to each key frame i and the text vector, denoted as s i , and assign weights to each key frame based on the similarity. The formula is as follows:

[0131]

[0132] In this way, the weight α of each key frame i will be automatically adjusted according to its matching degree with the text. The better the match, the greater the weight of the frame.

[0133] Calculate the local total matching degree:

[0134]

[0135] Weightedly sum the global similarity and the local matching degree to obtain the final similarity S = λS part +(1 - λ)S whole . λ is determined to be 0.3 through multiple experimental verifications and can be flexibly adjusted according to the actual task requirements.

[0136] Then, train the model based on the idea of contrastive learning. The contrastive loss function is:

[0137]

[0138] Among them, i represents the index of the current positive sample pair (matched video-text pair). In the loss function, each S i corresponds to the similarity of a positive sample pair. j is the index of all sample pairs in the traversed batch (including positive sample pairs and negative sample pairs).

[0139] τ is the temperature parameter, which is used to adjust the smoothness of the distribution, helps to control the discrimination degree of the similarity between different samples during the contrastive process, so that the similarity of the positive sample pair is relatively more significant than that of the negative sample pair.

[0140] During the training process, a hybrid parallel dual-network training method is adopted. By combining data parallelism and model parallelism, the text branch and the video branch are respectively deployed on different devices to achieve parallel computing for forward propagation, and gradients are independently calculated on their respective branches. Subsequently, a synchronous gradient reduction mechanism is used to ensure consistent parameter updates among devices, so that during the backpropagation process, the gradients generated by the contrast loss function are simultaneously transmitted to the Transformer encoder and the KAN network. Specifically, the normalized feature vectors output by the two branches are used to construct positive and negative sample pairs through cosine similarity calculation. The contrast loss function is used to drive the similarity of matching video-text pairs to approach 1, and the similarity of non-matching pairs to approach 0. The value of the contrast loss function converges to a certain minimum value, and during the dynamic update process using optimizers such as Adam or SGD combined with a learning rate scheduling strategy, the model encoder and weighted parameters are continuously adjusted to achieve precise cross-modal semantic alignment and continuous optimization of the overall training effect.

[0141] The specific method in Step 5 is as follows:

[0142] Input the text into the text-to-text large model obtained in Step 1 to get a detailed description text with semantic expansion. Subsequently, input the detailed text description into the video-text large model obtained in Step 4 to get the video with the largest similarity, that is, the surveillance video that best matches the description.

[0143] To verify the effectiveness of the proposed security surveillance video retrieval method based on video-text large models, the present invention conducted a large number of experimental evaluations. In the experiments, Recall@1, Recall@5, Recall@10 (which respectively represent how many queries contain the correct answer among the first, top five, and top ten returned results), and mean average precision (mAP) were used as the main performance indicators, and the average query response latency of the system was also measured. The datasets used in the experiments included surveillance videos with various scenarios and corresponding fuzzy descriptions from users.

[0144] In the comparative experiments, the following two sets of schemes were set up:

[0145] Baseline scheme: Global video feature extraction is adopted, and the fuzzy text input by the user is directly used for matching. This method only relies on the overall semantic representation of the video and the original fuzzy text description, and cannot further supplement or optimize the text semantics. The specific data is as follows:

[0146] Recall@1: 42.0%

[0147] Recall@5: 65.0%

[0148] Recall@10: 75.0%

[0149] Mean average precision (mAP): 38.0%

[0150] Average query response latency: approximately 90 milliseconds

[0151] The solution of the present invention: First, use a generative large model to supplement the semantic information of the fuzzy text input by the user, making the text description more detailed and rich; then combine the global video features with the local features extracted from the key frames, and achieve matching through weighted fusion. This method not only retains the overall semantic information of the video, but also uses local details to further improve the accuracy of matching. The specific data is as follows:

[0152] Recall@1: 60.0%

[0153] Recall@5: 80.0%

[0154] Recall@10: 88.0%

[0155] Mean Average Precision (mAP): 52.0%

[0156] Average query response latency: approximately 100 milliseconds

[0157] As can be seen from the above data, compared with the baseline solution, the solution of the present invention has improved by 18 percentage points on Recall@1 (a relative improvement of approximately 43%), and by 14 percentage points on mAP (a relative improvement of approximately 37%), while the query response time has only increased slightly by about 10 milliseconds, and the impact on the online real-time retrieval system is negligible.

[0158] These experimental data fully prove that the present invention significantly enhances the video retrieval performance under fuzzy descriptions, improves the recall rate and accuracy of the system, and maintains an efficient response speed by using a generative large model to supplement text descriptions and then combining global and local weighted fusion matching.

[0159] On the other hand, the embodiment of the present invention also provides a surveillance video retrieval system based on a vision large model under fuzzy descriptions, including:

[0160] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a surveillance video retrieval method based on a vision large model under fuzzy descriptions as described in the above technical solution.

[0161] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A monitoring video retrieval method based on a large vision model under fuzzy descriptions, characterized in that, It includes the following steps: Step 1: Create a dataset of fuzzy behavior descriptions and features, and train a large text-to-text model to convert fuzzy descriptions into more specific and structured text information; Step 2: First, preprocess the text information obtained in Step 1, and then set up a Transformer-based text encoder to encode the preprocessed text to generate text feature vectors; Step 3: First, perform video frame sampling and feature extraction, and then perform temporal and semantic aggregation based on the KAN network to obtain a feature vector representing the semantics of the entire video and extract key frames; Step 4: Train a text-video matching model through a hybrid parallel dual-network parallel mode. The text-video matching model includes a text branch and a video branch. The text branch is the Transformer-based text encoder in Step 2, and the video branch is the KAN network; Step 5: Use the trained text-video matching model to parse and expand the user input text, and find the matching video according to the new text description.

2. The monitoring video retrieval method based on a large visual model under fuzzy description according to claim 1, characterized in that: The text-to-text large model described in Step 1 is a T5 or BART model.

3. The monitoring video retrieval method based on a large vision model under fuzzy description according to claim 1, characterized in that: In Step 2, the preprocessing of the text information includes: Step 201: Perform word segmentation on the input text, and separate it into several words by spaces; Step 202: Convert each word into a vector of a fixed dimension by looking up the embedding matrix of clip: X = [e1, e2,..., en], where ei represents the embedding vector of the i-th word; Step 203: Use the sine function to mark the position of each word, and the formula is as follows: where: i represents the position of the current word in the sentence, j represents half of the embedding dimension, d is the dimension of the word embedding, and 10000 is the scaling factor; Finally, a fixed vector P representing the position will be generated for each position i, and it is added to the word embedding X to obtain the final vector: T = x + p.

4. The monitoring video retrieval method based on a vision large model under fuzzy description according to claim 1, wherein: In Step 2, the specific implementation method for generating text feature vectors is as follows: First, use the multi-head self-attention mechanism to calculate the relevance between words. The specific sub-steps are as follows: Step 20411: For each layer of the Transformer, the input vector T is linearly transformed to obtain the query Q, key K, and value V; Q = TW Q K = TW K V = TW V Among which W Q , W K , W V are all learnable parameter matrices, with the shape of (feature vector dimension, feature dimension of each attention head); Step 20412: Calculate the attention scores; Among them, QK T Calculate the dot product of the query and the key to measure the relevance between each pair of words; Perform scaling; to avoid gradient explosion caused by overly large values. Step 20413: Normalize the attention scores to obtain attention weights, and multiply the value V by the weights to obtain the final output; Step 2042, skip connection: The input vector T is added to the vector T after being processed by the self-attention mechanism in Step 2041, and after layer normalization, the vector T is obtained 1 ; Step 2043, use the feedforward neural network FFN to further process the vector T obtained in the previous step 1 The calculation formula is as follows: FFN(x) = max(0, xw1 + b1)w2 + b2 where x is the vector T obtained in step 2042 1 , w1 and w2 are weight matrices learned during the training process, and b1 and b2 are updated together during the training process; Step 2044: Stack multiple layers of Transformers. The Transformer encoder contains several transformer layers. After the first layer of operation is completed, the output result of the first layer is used as the input of the second layer of operation, and the process is repeated for multiple layers until all layers of operation are completed; After being processed by L layers of Transformer operations, a sequence of processed text vectors is obtained, and the feature vector at the end position is used as the final global text feature vector.

5. The monitoring video retrieval method based on a vision large model under fuzzy description according to claim 1, characterized in that: In step 3, uniform sampling is performed using a fixed frame rate r, and a fixed number of frames Ns = frame rate * sampling interval are extracted per second. The embedded features and high-level semantic information of each frame are extracted through an image encoder, i.e., the VIT model.

6. The method for retrieving surveillance videos based on a large vision model under fuzzy descriptions as claimed in claim 1, wherein: In step 3, the processing process of the KAN network is as follows: First, the features extracted from the video are mapped to the intermediate feature space after preprocessing, denoted as where W h and b h are learnable parameters; During temporal aggregation, for the inherent temporal dynamics of the video frame sequence, the features at each moment are processed layer by layer through several layers of KAN modules to form an expression of local temporal information. Specifically, the output of each layer can be written as where the initial input is the preprocessed feature, l represents the number of layers, and φ l represents the KAN module. A local attention mechanism or a kernel function-based weighting strategy can be embedded in each layer to capture the fine-grained changes between adjacent frames within the local time domain using B-spline functions. After processing, the outputs of each frame in the temporal branch are weighted or averaged to obtain the global temporal aggregation feature During the semantic aggregation process, for the high-level semantic information of each frame of features extracted by the image encoder After being processed by several layers of KAN modules, another set of features is obtained Then, a set of independent KAN modules performs hierarchical processing, and the output form of each layer is φ l ′ represents an independent KAN module, and then a non-linear transformation is performed through the B-spline activation function. Finally, using a learnable semantic attention or weighted summation mechanism, the processing results of each frame are integrated into the global semantic aggregation feature Finally, the global temporal aggregation feature and the global semantic aggregation feature are fused by weighted summation, that is where λ ∈ [0, 1] is a hyperparameter for adjusting the contributions of the two parts, obtaining the global video feature with temporal dynamics and semantic consistency, and after normalization, the feature vector z’ representing the semantics of the entire video is obtained v .

7. The method for retrieving surveillance videos based on a large vision model under fuzzy descriptions according to claim 1, characterized in that: The formula for extracting key frames in step 3 is: Among them, F i represents the feature vector of the i-th frame, and the frame with the largest d(i) is the key frame.

8. The monitoring video retrieval method based on a large vision model under fuzzy description according to claim 1, wherein: The specific implementation method of step 4 is as follows: First, calculate the global similarity, that is, the similarity between the video and text vectors obtained in step 2 and step 3. The formula is as follows: Among them, v and t represent the global feature vectors of the video and text respectively; Then calculate the local matching degree based on the key frames: calculate the similarity between the feature vector of each key frame i and the text description, denoted as s i , and assign weights to each key frame based on the similarity, and the formula is as follows: The weight α of each key frame i Automatically adjusted according to its matching degree with the text. The better the match, the greater the weight of the frame. Calculate the local total matching degree: The global similarity and the local matching degree are weighted and summed to obtain the final similarity S; Then, based on the idea of contrastive learning, the model is trained. The contrastive loss function is: Among them, i′ represents the index of the positive sample pair, j is the index of all sample pairs in the traversed batch, and τ is the temperature parameter; During the training process, a hybrid parallel dual-network training method is adopted by combining data parallelism and model parallelism. The text branch and the video branch are deployed on different devices respectively to realize parallel computing of forward propagation, and the gradients are calculated independently on their respective branches. Subsequently, a synchronous gradient reduction mechanism is adopted to ensure consistent parameter updates among devices, so that during the backpropagation process, the gradients generated by the contrastive loss function are simultaneously transmitted to the Transformer-based text encoder and the KAN network. Finally, when the value of the contrastive loss function converges to a certain value, the training is completed.

9. The monitoring video retrieval method based on a large vision model under fuzzy description according to claim 1, characterized in that: The specific implementation method of step 5 is as follows: Input the text into the text-to-text large model in step 1 to obtain a detailed description text with semantic expansion. Subsequently, input the detailed text description into the video-text large model obtained in step 4 to obtain the video with the largest similarity, that is, the surveillance video that best matches the description.

10. A surveillance video retrieval system based on a large vision model under fuzzy descriptions, characterized in that, Including: A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a method for retrieving a surveillance video based on a vision large model under fuzzy description as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Video text retrieval method based on pre-training model

    CN116109960A

  • Video text retrieval method based on BEiT-3 multi-mode large model

    CN118377930A

  • Multi-round dialogue processing method and system based on RAG and knowledge graph

    CN118885627A

  • Multilevel semiotic and fuzzy logic user and metadata interface means for interactive multimedia system having cognitive adaptive capability

    US20090132441A1

  • Large language model-based information retrieval for large datasets

    US20250086215A1