Unmanned aerial vehicle video abstract semantic description method and system based on multi-modal large model
Through video segmentation, feature extraction and semantic description technology of multimodal large model, the problem of unsatisfactory abstract accuracy and integrity in drone video data is solved, efficient and accurate video summary generation and semantic description are achieved, and the utilization efficiency of drone video data is improved.
Patent Information
- Application Number
- CN202510463016.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-01
AI Technical Summary
The video summary generated by the existing methods in drone video data is not ideal for correctness and completeness, making it difficult to extract high-value information efficiently and accurately.
The UAV video abstract semantic description method based on multimodal large model is used to generate a UAV video segmentation, image feature extraction, adaptive clustering and semantic description through video segmentation, feature extraction and semantic description.
It realizes automatic, accurate and efficient extraction of core intelligence information from drone video data, improves video data utilization efficiency, enhances semantic understanding depth and context relevance, and reduces processing costs.
Smart Images

Figure CN120411571A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV video data interpretation, and particularly relates to a method and system for semantic description of UAV video summaries based on a multimodal large model. Background Art
[0002] With the rapid progress of UAV and sensor technologies, the amount of video data obtained by UAVs using sensors shows an explosive growth trend. Quickly and accurately screening high-value content or key information from UAV video data can provide support for efficiently understanding video content. However, with the improvement of the performance of UAV detection devices, the obtained video data not only requires a large storage space, but also the difficulty of quickly finding valuable content from a large amount of video data increases, which requires a large amount of manpower and time and is far beyond what can be handled manually. How to intuitively and efficiently access UAV video data, quickly complete a general understanding of the core content of the video data, and obtain the main information contained in the video data has become an urgent problem to be solved for the efficient understanding of UAV video data and has gradually become a research hotspot.
[0003] In order to quickly, intuitively and efficiently obtain high-value content and key information from a large amount of UAV video data, it is necessary to rely on computers to automatically screen and extract the core content of the video and carry out subsequent processing tasks. Therefore, it is urgent to carry out relevant research on UAV video data. Among them, the combination of video summary generation and semantic description technology can present valuable content in long videos in the form of a concise picture list and semantic description, which can provide an effective solution for the understanding of a large number of videos. Currently, it has been applied in the video understanding of the public domain. The main goal of video summary generation and semantic description is to generate a concise and complete video summary by selecting the part of the video data that can represent the information of the video. The generated video summary is usually a set of representative video frames (such as video key frames) or a shorter video formed by stitching key video segments in chronological order; combining semantic description technology to obtain video summary description information can support intelligence interpreters to quickly understand the core content of video data and significantly reduce the burden on intelligence interpreters. It can be seen from this that video summary generation and semantic description have a relatively broad application prospect in the field of UAV video data interpretation.
[0004] However, although great progress has been made in the research on video summary technology for the public domain, the existing methods have problems of lack of temporal information and incomplete feature representation, which easily affect the correctness and integrity of video summaries; at the same time, the video summary technology for UAVs has not received wide attention. How to construct a video summary generation and semantic description technology for UAV video data and automatically, efficiently and accurately extract high-value information from a large amount of video data is one of the actual urgent problems in the field of UAV video interpretation. Summary of the Invention
[0005] To this end, the present invention provides a method and system for semantic description of UAV video summaries based on multi-modal large models, which solves the problem that the correctness and integrity of existing methods for generating UAV video summaries are not ideal.
[0006] According to the design solution provided by the present invention, on the one hand, a method for semantic description of UAV video summaries based on multi-modal large models is provided, including:
[0007] Preprocess the UAV video data to be processed to obtain a number of segmented video frame images, where the preprocessing includes video segment segmentation and video frame extraction;
[0008] Use a preset multi-modal large model to extract image features from the segmented video frame images, where the multi-modal large model uses an image encoder in the vision-language foundation model to encode the input segmented video frame images and extract the corresponding image features;
[0009] Perform adaptive clustering on the extracted image features to obtain the clustering centers of each segmented video, use the frame positions where the clustering centers are located as the frame positions where the video summaries are located, and generate UAV video summaries through the clustering centers;
[0010] Input the video summary into a semantic description model, and use the semantic description model to obtain the scene semantic description of the UAV video summary, where the semantic description model is obtained by fine-tuning the large model using a UAV image semantic description data set.
[0011] As the method for semantic description of UAV video summaries based on multi-modal large models of the present invention, further, preprocessing the UAV video data to be processed includes:
[0012] Construct a video scene segmentation model based on TransNet V2, where the video scene segmentation model includes: a feature extraction part for extracting the spatial features and inter-frame temporal correlation features of the input video frames to capture the local motion changes and global context relationships of the video frames when the video shots are switched, and a classification decision part for judging whether the corresponding video frames are shot boundaries based on the local motion changes and global context relationships of the video frames;
[0013] Train the video scene segmentation model using a mixed data set with annotation labels, where the mixed data set includes scene synthesis video data, scene real video data, and video frame enhancement data, and the video frame enhancement data is to perform enhancement transformation on the video frame images using an image processing method, and the image processing method includes one or more combinations of, but is not limited to, left and right flipping of the frame image, up and down flipping of the frame image, saturation adjustment of the frame image, contrast adjustment of the frame image, brightness adjustment of the frame image, and hue transformation of the frame image;
[0014] Input the UAV video data to be processed into a video scene segmentation model to obtain several segmented videos using the video scene segmentation model.
[0015] As the UAV video abstract semantic description method based on a multimodal large model of the present invention, further, when the feature extraction part extracts the spatial features and inter-frame temporal correlation features of the input video frames, it calculates the cosine similarity between the video frame images based on the image RGB color histogram and the learned features, and uses the cosine similarity to obtain the inter-frame temporal correlation features.
[0016] As the UAV video abstract semantic description method based on a multimodal large model of the present invention, further, video frame extraction includes:
[0017] Set the frame time interval and extract video frame images from the segmented videos based on the frame time interval;
[0018] If the resolution of the extracted video frame images exceeds the threshold, compress the video frame images to obtain segmented video frame images that meet the resolution requirements.
[0019] As the UAV video abstract semantic description method based on a multimodal large model of the present invention, further, using the multimodal large model to extract the image features in the segmented video frame images includes:
[0020] Use the multi-layer stacked ViT encoder in the vision-language foundation model RemoteCLIP and extract the global features of the segmented video frame images through patch embedding, positional encoding, Transformer encoding, and cross-layer fusion.
[0021] As the UAV video abstract semantic description method based on a multimodal large model of the present invention, further, adaptively clustering the extracted image features includes:
[0022] Determine the neighborhood radius and the minimum number of points, and mark all image feature points as unvisited. The neighborhood radius is used to represent the size of the neighborhood range centered on the image feature, and the minimum number of points is used to determine whether the image feature is a core point;
[0023] Randomly select an unvisited image feature point, mark its visited status as visited, and calculate the neighborhood of the image feature point with the visited status marked as visited. If the number of image feature points included in the neighborhood is not less than the minimum number of points, determine that the image feature point is a core point and create a new cluster, and add the core point and all unclustered image feature points within its neighborhood to the created new cluster. If the number of feature points included in the neighborhood is less than the minimum number of points, determine that the image feature point is a boundary point or a noise point;
[0024] Traverse all image feature points in the unvisited state until the access status of all image feature points is marked as visited, and obtain the clustering result of the image feature points.
[0025] As the method for semantic description of UAV video summaries based on multi-modal large models of the present invention, further, the large model is fine-tuned using the UAV image semantic description dataset, including:
[0026] Extract semantic description data from the open-source remote sensing field image semantic description dataset from a specified perspective, and construct a UAV image semantic description dataset including remote sensing images, text semantic description instructions, and semantic description responses, where the specified perspective is the UAV perspective;
[0027] Use LVLM as the large model, inject a low-rank matrix into the attention layer of the model, and fine-tune the large model based on the LoRA fine-tuning algorithm and using the UAV image semantic description dataset to obtain a semantic description model for semantic description in the UAV video summary scenario.
[0028] On the other hand, the present invention also provides a system for semantic description of UAV video summaries based on multi-modal large models, including: a video segmentation module, a feature extraction module, a summary generation module, and a semantic description module, where,
[0029] The video segmentation module is used to preprocess the UAV video data to be processed to obtain several segmented video frame images, and the preprocessing includes video segment segmentation and video frame extraction;
[0030] The feature extraction module is used to extract image features in the segmented video frame images using a preset multi-modal large model. The multi-modal large model encodes the input segmented video frame images using the image encoder in the vision-language base model and extracts the corresponding image features;
[0031] The summary generation module is used to perform adaptive clustering on the extracted image features to obtain the clustering center of each segmented video, use the frame position where the clustering center is located as the frame position of the video summary, and generate a UAV video summary through the clustering center;
[0032] The semantic description module is used to input the video summary into the semantic description model and use the semantic description model to obtain the scene semantic description of the UAV video summary. The semantic description model is obtained by fine-tuning the large model using the UAV image semantic description dataset.
[0033] The beneficial effects of the present invention:
[0034] The present invention extracts core intelligence information from UAV video data automatically, accurately, and efficiently through stages such as video segmentation, image feature extraction, image feature clustering, and video summary semantic description, improving the utilization efficiency of UAV video data; uses a multimodal large model for feature extraction to enhance the image feature representation ability, obtains the frame position where the video summary is located through adaptive clustering of image features, can adapt to the characteristics of UAV video data with varying lengths and diverse scene changes, uses the fine-tuned large model to generate a summary semantic description of the UAV video scene, avoids information loss, enhances the depth of semantic understanding and context relevance, can simulate human visual perception, achieve coherent description of complex time, and reduce costs and improve the processing efficiency of video summary description. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the UAV video summary semantic description process based on a multimodal large model in the embodiment;
[0036] Figure 2 It is a schematic diagram of the UAV video summary semantic description algorithm architecture in the embodiment;
[0037] Figure 3 It is a schematic diagram of the video composition in the embodiment;
[0038] Figure 4 It is a schematic diagram of the video scene segmentation model structure in the embodiment;
[0039] Figure 5 It is a schematic diagram of the video segmentation processing flow in the embodiment;
[0040] Figure 6 It is a schematic diagram of the UAV video data image frame feature extraction process in the embodiment;
[0041] Figure 7 It is a schematic diagram of the adaptive clustering algorithm process in the embodiment;
[0042] Figure 8 It is a schematic diagram of the UAV video summary semantic description model structure in the embodiment;
[0043] [[ID=3⑦]] Figure 9 It is a schematic diagram of the UAV video summary semantic description model fine-tuning process in the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0044] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the drawings and technical solutions.
[0045] Video generally refers to continuous pictures, which are obtained by arranging a series of pictures in an orderly manner on the time axis. When the change rate of pictures exceeds 24 frames per second, the entire video will present a continuous and smooth effect, which is actually the principle of persistence of vision. The smallest constituent unit of a video is a frame. Multiple frames of images form a video shot, multiple video frame shots form a scene, and multiple scenes ultimately form a video. The detailed structure of the video is as shown in Figure 3 Figure 1. Among them, Frame: A frame is the smallest constituent unit of a video, and each frame represents a picture; Shot: A shot is composed of a series of frames with relatively similar meanings. A video can be split into individual shots according to different contents, and the change amplitude of the frame images in the same shot is relatively small. In short, a video is composed of various different shots, and the meanings represented by each shot are different; Scene: A scene is composed of multiple shots, and there is a certain logical sequence between these shots, which can express a certain event plot and is the movement or plot development process of relevant target objects; Video: Essentially, a video is a collection of frames arranged along the time axis according to a specific time axis and is an unstructured data stream.
[0046] A drone video can be composed of multiple short videos through direct splicing, transition splicing, etc., and the splicing is achieved through shot conversion. The abstracts in different drone short videos are relatively independent. Therefore, for long drone videos, in order to reduce the processing pressure, the video segmentation method can be used to find the boundaries of video splicing and obtain independent short videos. This process can be automatically recognized through shot conversion in the video and is already a classic task in video analysis. However, the performance of various traditional automatic shot conversion recognition algorithms is insufficient. For example, when the frame rate of the video is not high and it contains high-speed moving objects, the difference between adjacent frames will be very large, and it may be misjudged as the boundary of video splicing; when the video has a long transition shot, the algorithm performance is also poor.
[0047] Regarding the problems of drone video abstract generation and semantic description, in the embodiments of the present invention, the algorithm architecture is as shown in Figure 2 Figure 2, which mainly includes video segmentation, image feature extraction, image feature clustering, video abstract semantic description, etc. In the embodiments of this case, the method for semantic description of drone video abstracts based on a multi-modal large model is as shown in Figure 1 Figure 3, and specifically includes the following contents:
[0048] S101. Preprocess the drone video data to be processed to obtain several segmented video frame images. The preprocessing includes video segment segmentation and video frame extraction.
[0049] Specifically, the preprocessing of the drone video data to be processed can be designed to include:
[0050] Build a video scene segmentation model based on TransNet V2. The video scene segmentation model includes: a feature extraction part for extracting the spatial features and inter-frame temporal correlation features of the input video frames to capture the local motion changes and global context relationships of the video frames during video shot transitions, and a classification decision part for determining whether the corresponding video frames are shot boundaries based on the local motion changes and global context relationships of the video frames;
[0051] Train the video scene segmentation model using a mixed dataset with annotation labels. The mixed dataset includes scene synthesis video data, scene real video data, and video frame enhancement data. The video frame enhancement data is obtained by performing enhancement transformations on the video frame images using image processing methods, and the image processing methods include, but are not limited to, one or more combinations of left-right flipping of the frame image, up-down flipping of the frame image, saturation adjustment of the frame image, contrast adjustment of the frame image, brightness adjustment of the frame image, and hue transformation of the frame image;
[0052] Input the drone video data to be processed into the video scene segmentation model to obtain several segmented videos using the video scene segmentation model.
[0053] Such as Figure 4As shown, TransNet V2 is built on the basis of the original TransNet concept. The model can quickly and effectively detect shot transitions in videos and can be used for tasks such as video segmentation and video content analysis, showing many significant advantages in the field of video shot transition detection. In the embodiments of this case, the 3D convolution is decomposed into 2D spatial convolution and 1D temporal convolution by means of convolutional kernel decomposition, which not only helps to more accurately learn image features and temporal features respectively, but also can effectively reduce the number of learnable parameters and the risk of overfitting, making the model have stronger adaptability and generalization ability when facing different types of video content. Secondly, by introducing frame similarity features, the cosine similarity of RGB color histograms (RGBHist.Similarites) and learnable feature similarity (Learnable Similarites) are calculated respectively, making full use of the correlation information between frames, providing a richer basis for judging shot transitions, and further improving the accuracy of detection. Furthermore, a multi-classification head design is used. Among them, one head focuses on predicting the key intermediate frames of shot transitions, and the other head is used to optimize the network's understanding of transitions during training. This way of division of labor and cooperation effectively improves the model's overall cognition and judgment ability of shot transitions. During the training process, a strategy of combining large-scale synthetic training data and partial real data is adopted, and combined with methods such as left-right flipping (probability 0.5), up-down flipping (probability 0.1), saturation adjustment, contrast, brightness, and hue adjustment of frame images, so that the model can fully learn various types of shot transition modes and enhance the processing ability of videos in different scenarios. The obtained UAV video data is directly input into the TransNetV2 large model, as Figure 5 shown, and the segmentation result can be obtained. Compared with the result of the video processing software, the effect is better.
[0054] When extracting video frames, set the frame time interval, and extract video frame images from the segmented video based on the frame time interval; if the resolution of the extracted video frame images exceeds the threshold, compress the video frame images to obtain segmented video frame images that meet the resolution requirements.
[0055] S102. Use a preset multi-modal large model to extract the image features in the segmented video frame images. The multi-modal large model uses the image encoder in the vision-language foundation model to encode the input segmented video frame images and extract the corresponding image features.
[0056] Among them, the global features of the segmented video frame images are extracted by using the stacked ViT encoders in the vision-language foundation model RemoteCLIP and through patch embedding, positional encoding, Transformer encoding, and cross-layer fusion.
[0057] The scene changes contained in UAV video data vary greatly, and general supervised or semi-supervised deep learning algorithms cannot handle the changing UAV video data well. RemoteCLIP is the first vision-language foundation model for the remote sensing field, aiming to learn visual features with rich semantics and robust features aligned with text embeddings to enable seamless downstream applications. To address the problem of insufficient pre-training data, RemoteCLIP converts heterogeneous annotations into a unified image caption data format based on box-to-caption (B2C) and mask-to-box (M2B) conversions through data scaling techniques. By further merging UAV images, a pre-training dataset 12× larger than the combination of all available datasets was generated. RemoteCLIP was evaluated on a variety of downstream tasks, including zero-shot image classification, linear probing, k-NN classification, few-shot classification, image-text retrieval, and object counting in remote sensing images. In the embodiments of this case, the process of extracting image frame features from UAV video data based on RemoteCLIP is as Figure 6 shown. By using RemoteCLIP to implement the changes in images and text, through the RemoteCLIP-ViT-B / 32 model and using the image encoding part of RemoteCLIP, the processed video frame images in the segmented video are input into the RemoteCLIP image encoding part frame by frame to obtain the extracted image features, and normalization is performed to obtain the global features of the video frame images.
[0058] S103. Perform adaptive clustering on the extracted image features to obtain the clustering centers of each segmented video, and use the frame position where the clustering center is located as the frame position of the video summary, so as to generate a UAV video summary through the clustering center;
[0059] Specifically, performing adaptive clustering on the extracted image features can be designed to include:
[0060] Determine the neighborhood radius and the minimum number of points, and mark all image feature points as unvisited. The neighborhood radius is used to represent the size of the neighborhood range centered on the image feature, and the minimum number of points is used to determine whether the image feature is a core point;
[0061] Randomly select an unvisited image feature point, mark its visited status as visited, and calculate the neighborhood of the image feature point with the visited status marked as visited. If the number of image feature points contained in the neighborhood is not less than the minimum number of points, then determine that the image feature point is a core point, and create a new cluster, and add the core point and all unclustered image feature points within its neighborhood to the newly created cluster. If the number of feature points contained in the neighborhood is less than the minimum number of points, then determine that the image feature point is a boundary point or a noise point;
[0062] Traverse all image feature points in the unvisited state until the access status of all image feature points is marked as visited, and obtain the clustering result of the image feature points.
[0063] The features extracted from the UAV reconnaissance video frame images by the remote sensing large model can effectively represent the content of the frame images, and the features between similar frame images have extremely high similarity. The similar frame images can be classified by clustering, and the key frames can be determined by obtaining the cluster centers. Therefore, the clustering algorithm can be used to extract key frames. However, there are a large number of frames in the video, and the frame image features are complex. The actual cluster centers are not clear. Therefore, the clustering algorithm needs to be able to perform clustering without setting the cluster centers and have good efficiency in the case of a large amount of feature data.
[0064] To address this problem, in the embodiments of this case, an adaptive clustering of image features is performed based on the DBSCAN clustering algorithm. The DBSCAN clustering algorithm is a density-based clustering algorithm that aims to discover clusters of any shape and is robust to noise points (Outliers). It finds high-density regions in the data space, treats these regions as clusters, and classifies isolated points (points with low density) as noise. Its basic idea is that if there are enough points (more than a threshold minPts) within the neighborhood radius (ε) of a certain point, this region is considered a high-density region and can be expanded into a cluster. A cluster is expanded by density-connected points. Points that cannot be assigned to any cluster are considered noise points.
[0065] Applying the DBSCAN clustering algorithm to image feature clustering and obtaining the video summary through the cluster centers, the algorithm steps can be summarized as including:
[0066] The first step: Initialization. Determine the neighborhood radius and the minimum number of points, as shown in (a) of Figure 7 . The neighborhood radius determines the size of the neighborhood range centered on a certain image feature, and the minimum number of points is used to determine whether an image feature is a core point. Mark all image feature points as unvisited so that it can be known which points have been processed and which have not been processed during the subsequent traversal.
[0067] The second step: Traverse the image feature points. Randomly select an unvisited image feature point, as shown in Figure 7As shown in (b) in [reference], it is marked as visited. Calculate the neighborhood based on the just-marked image feature points. If the number of image feature points contained within the neighborhood of this image feature point is not less than the pre-set minimum number of points, then this image feature point is determined to be a core point. At this time, a new cluster is created, which can be represented by a class identifier. Add this core point and all the points within its neighborhood that have not been assigned to any cluster to this newly created cluster. Recursively search for and add points in this way until no new image feature points can be added to this cluster. If it is not a core point, it is a boundary point or a noise point and is not processed temporarily.
[0068] Step 3: Repeat Step 2 until all these image feature points are visited, as shown in (c) and (d) in [reference]. Figure 7 As shown in (c) and (d) in [reference]. At this time, all these image feature points marked as clusters constitute the final clustering result, and the points that have not been assigned to a cluster are marked as noise points. Thus, the clustering analysis of this entire image based on the DBSCAN algorithm is completed.
[0069] S104. Input the video summary into the semantic description model, and use the semantic description model to obtain the scene semantic description of the UAV video summary. The semantic description model is obtained by fine-tuning a large model using a UAV image semantic description dataset.
[0070] Specifically, fine-tuning the large model using the UAV image semantic description dataset can be designed to include:
[0071] Extract semantic description data from the open-source remote sensing field image semantic description dataset from a specified perspective, and construct a UAV image semantic description dataset containing remote sensing images, text semantic description instructions, and semantic description responses. The specified perspective is the UAV perspective.
[0072] Use LVLM as the large model, inject a low-rank matrix into the attention layer of the model, and fine-tune the large model using the UAV image semantic description dataset based on the LoRA fine-tuning algorithm to obtain a semantic description model for the scene semantic description of the UAV video summary.
[0073] For the obtained video summaries, in the embodiments of this case, the Qwen2-VL model is fine-tuned using a semantic description dataset in the remote sensing field, and the fine-tuned model is used to process the generated drone video summaries to obtain the scene semantic description information of key frames. As a large vision language model (LVLM), Qwen-VL can take images, text, and detection boxes as inputs and text and detection boxes as outputs. The Qwen-VL series of models has powerful performance and achieves the best results under the same general model size in multiple multimodal tasks. It supports multilingual conversations in English, Chinese, etc., end-to-end supports long text recognition of Chinese and English in pictures, and is the first open-source LVLM model with a resolution of 448. Higher resolution can improve fine-grained text recognition, document question answering, and detection box annotation. Its top-notch understanding ability for images of various resolutions and aspect ratios, the ability to understand videos longer than 20 minutes, has multilingual support including Chinese and English, and can process image inputs of any resolution without having to split the image into blocks. It can convert pictures of different sizes into a dynamic number of tokens, with a minimum of only 4 tokens. This design simulates the natural way of human visual perception, ensuring a high degree of consistency between the model input and the original image information, enabling the model to perform image processing more flexibly and efficiently. As Figure 8 shown in the model structure of Qwen2-VL, on the basis of Qwen-VL, a visual encoder and a language model are integrated. For various scaling adjustments, Qwen2-VL implements a vision transformer (VIT), including approximately 675 million parameters, which is good at processing image and video inputs. In terms of language processing, Qwen2-VL uses the more powerful Qwen2 series of language models. The Qwen2-VL-2B-Instruct model can be used in the embodiments of this case.
[0074] To build a semantic description model for drone video summaries, multiple groups of open-source remote sensing field image semantic description datasets (such as NWPU-Caption, LEVIR-CC, UCM captions, CapERA, etc.) can be fused, data from the drone perspective or close to the drone perspective can be extracted from them to build a large-scale drone image semantic description dataset, and Qwen2-VL can be fine-tuned based on the LoRA fine-tuning algorithm to build a multimodal large model UAV-VL for drone image semantic description, such as Figure 9As shown in the figure. LoRA (Low-Rank Adaptation) is an effective technique used in the fine-tuning of large language models (such as GPT-3, BERT, etc.). Its main purpose is to reduce the number of parameters to be updated by introducing low-rank decomposition matrices without significantly changing the original model structure, so as to quickly adapt to specific downstream tasks. LoRA assumes that the weight matrix of the pre-trained language model is W. During fine-tuning, instead of directly updating W, W is decomposed into W = W0 + BA. Where W0 is the original pre-trained weight matrix, and B and A are newly introduced low-rank matrices, and their product represents the update amount to the original weight matrix. The rank of the low-rank matrix is usually a relatively small value, for example, it can be set between 1 and 8 in some experiments. This means that the number of parameters of BA is much less than that of the original weight matrix. Especially in the field of semantic description of drone images, the advantage of parameter efficiency of the LoRA fine-tuning algorithm is more important and prominent. The dataset of drone image semantic description has a large data scale and high computational costs for processing and training. By using the LoRA fine-tuning algorithm, it is possible to improve the training and inference speed of the model without significantly increasing computational resources and storage requirements, and achieve fine-tuning optimization of large-scale vision-language models. As Figure 9 As shown in the figure, during the fine-tuning stage, the original weight matrix W0 of Qwen2-VL is fixed, and only the two low-rank matrices B and A are trained, which can greatly reduce the number of parameters to be updated during the training process. During the inference stage, the updated B and A are recombined with the original weight matrix to obtain the fine-tuned weight matrix W = W0 + BA. Then this fine-tuned model is used to predict new input data. For the input x, the corresponding output h is expressed as h = (W0 + △W)x = W0x + △Wx = W0x + BAX, and d is the rank of the parameter matrix △W.
[0075] Furthermore, based on the above method, the embodiment of the present invention also provides a drone video summary semantic description system based on a multi-modal large model, including: a video segmentation module, a feature extraction module, a summary generation module, and a semantic description module, where,
[0076] The video segmentation module is used to preprocess the to-be-processed drone video data to obtain several segmented video frame images, and the preprocessing includes video segment segmentation and video frame extraction;
[0077] The feature extraction module is used to extract image features in the segmented video frame images by using a preset multi-modal large model. The multi-modal large model uses an image encoder in the vision-language base model to encode the input segmented video frame images and extract the corresponding image features;
[0078] An abstract generation module, which is used to adaptively cluster the extracted image features to obtain the clustering centers of each segmented video, use the frame positions where the clustering centers are located as the frame positions of the video abstract, and generate a UAV video abstract through the clustering centers;
[0079] A semantic description module, which is used to input the video abstract into a semantic description model and use the semantic description model to obtain the scene semantic description of the UAV video abstract. The semantic description model is obtained by fine-tuning a large model using a UAV image semantic description data set.
[0080] Unless otherwise specifically stated, the relative steps, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0081] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method section.
[0082] The units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the various examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0083] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, the various modules / units in the above embodiments can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of the combination of hardware and software.
[0084] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for semantic description of UAV video summaries based on multi-modal large models, characterized in that, Including: Preprocess the drone video data to be processed to obtain several segmented video frame images. The preprocessing includes video segment segmentation and video frame extraction; Use a preset multi-modal large model to extract image features from the segmented video frame images. The multi-modal large model uses the image encoder in the vision-language base model to encode the input segmented video frame images and extract the corresponding image features; Perform adaptive clustering on the extracted image features to obtain the clustering center of each segmented video. Use the frame position where the clustering center is located as the frame position of the video summary, and generate a drone video summary through the clustering center; Input the video summary into the semantic description model, and use the semantic description model to obtain the scene semantic description of the drone video summary. The semantic description model is obtained by fine-tuning the large model using the drone image semantic description dataset.
2. The method for semantic description of UAV video abstract based on multi-modal large model according to claim 1, characterized in that, Preprocess the drone video data to be processed, including: Build a video scene segmentation model based on TransNet V2. The video scene segmentation model includes: a feature extraction part for extracting the spatial features and inter-frame temporal correlation features of the input video frames to capture the local motion changes and global context relationships of the video frames when the video shot is switched, and a classification decision part for judging whether the corresponding video frame is a shot boundary based on the local motion changes and global context relationships of the video frames; Train the video scene segmentation model using a mixed dataset with annotation labels. The mixed dataset includes scene synthetic video data, scene real video data, and video frame enhancement data. The video frame enhancement data is obtained by performing enhancement transformation on the video frame images using image processing methods. The image processing methods include, but are not limited to, one or more combinations of left-right flipping of the frame image, up-down flipping of the frame image, saturation adjustment of the frame image, contrast adjustment of the frame image, brightness adjustment of the frame image, and hue transformation of the frame image; Input the drone video data to be processed into the video scene segmentation model to obtain several segmented videos using the video scene segmentation model.
3. The method for semantic description of UAV video summary based on a multi-modal large model according to claim 2, wherein, When the feature extraction part extracts the spatial features and inter-frame temporal correlation features of the input video frames, calculate the cosine similarity between the video frame images based on the image RGB color histogram and the learned features, and use the cosine similarity to obtain the inter-frame temporal correlation features.
4. The method for semantic description of UAV video abstract based on multi-modal large model according to claim 1 or 2, characterized in that Video frame extraction, including: Set the frame time interval and extract video frame images from the segmented video based on the frame time interval; If the resolution of the extracted video frame image exceeds the threshold, perform compression processing on the video frame image to obtain segmented video frame images that meet the resolution requirements.
5. The method for semantic description of UAV video abstract based on multi-modal large model according to claim 1, wherein, Using the multi-modal large model to extract image features from the segmented video frame images, including: Use the multi-layer stacked ViT encoder in the vision-language base model RemoteCLIP and extract the global features of the segmented video frame images through patch embedding, positional encoding, Transformer encoding, and cross-layer fusion.
6. The method for semantic description of UAV video abstract based on multi-modal large model according to claim 1, wherein, Performing adaptive clustering on the extracted image features, including: Determine the neighborhood radius and the minimum number of points, and mark all image feature points as unvisited. The neighborhood radius is used to represent the size of the neighborhood range centered on the image feature, and the minimum number of points is used to determine whether an image feature is a core point; Randomly select an unvisited image feature point, mark its visited status as visited, and calculate the neighborhood of the image feature point with the visited status marked as visited. If the number of image feature points contained in the neighborhood is not less than the minimum number of points, then determine that the image feature point is a core point, and create a new cluster. Add the core point and all unclustered image feature points within its neighborhood to the newly created cluster. If the number of feature points contained in the neighborhood is less than the minimum number of points, then determine that the image feature point is a boundary point or a noise point; Traverse all unvisited image feature points until the visited status of all image feature points is marked as visited, and obtain the clustering result of the image feature points.
7. The method for semantic description of UAV video abstract based on multi-modal large model according to claim 1, characterized in that, Fine-tune the large model using the UAV image semantic description dataset, including: Extract semantic description data from the open-source remote sensing field image semantic description dataset from a specified perspective, and construct a UAV image semantic description dataset including remote sensing images, text semantic description instructions, and semantic description responses. The specified perspective is the UAV perspective; Use LVLM as the large model, inject a low-rank matrix into the attention layer of the model, and fine-tune the large model using the UAV image semantic description dataset based on the LoRA fine-tuning algorithm to obtain a semantic description model for semantic description in the UAV video summary scenario.
8. An unmanned aerial vehicle video summary semantic description system based on a multimodal large model, characterized in that, Include: a video segmentation module, a feature extraction module, a summary generation module, and a semantic description module. Among them, The video segmentation module is used to preprocess the UAV video data to be processed to obtain several segmented video frame images. The preprocessing includes video segment segmentation and video frame extraction; The feature extraction module is used to extract image features from the segmented video frame images using a preset multimodal large model. The multimodal large model encodes the input segmented video frame images using the image encoder in the vision-language foundation model and extracts the corresponding image features; The summary generation module is used to perform adaptive clustering on the extracted image features to obtain the clustering center of each segmented video, use the frame position where the clustering center is located as the frame position of the video summary, and generate a UAV video summary through the clustering center; The semantic description module is used to input the video summary into the semantic description model and obtain the scene semantic description of the UAV video summary using the semantic description model. The semantic description model is obtained by fine-tuning the large model using the UAV image semantic description dataset.
9. An electronic device, characterized in that, Include: At least one processor, and a memory coupled to the at least one processor; Wherein, the memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, it can implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
Intelligent data acquisition method, equipment, medium and product
CN120781874A
Video perception feature extraction method based on semantic sentence library and similarity time sequence modeling
CN121415332A
A Video-Aware Feature Extraction Method Based on Semantic Database and Temporal Similarity Modeling
CN121415332B