Multi-modal Marine Scenario Video Description Algorithm Based on Instance Segmentation Auxiliary Information
By using Video-Swin-Transformer and Segment Anything networks to extract video features and instance segmentation in marine scene video description, and combining Bert for text feature extraction, the problem of insufficient richness and standardization of marine scene video description in the prior art is solved, and better semantic alignment and description quality are achieved.
Patent Information
- Application Number
- CN202310727600.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-06-19
AI Technical Summary
The prior art is difficult to effectively capture the semantic information of the video in the description of marine scene videos, resulting in insufficient richness and specification of the description, and the different network architectures of the video feature extractor and the text feature extractor affect the semantic alignment between the modals.
Video-Swin-Transformer is used as a video feature extractor, and instance segmentation is combined with Segment Anything network to obtain auxiliary information dictionaries and enhance the correlation between video texts. Use Bert as a text feature extractor to encode video features and text features multimodal interactively to complete semantic alignment and language reconstruction tasks.
It improves the video feature extraction effect, enhances the semantic richness and standardization of marine scene video descriptions, improves the alignment ability of video text descriptions, and the generated descriptions are smoother and conforms to natural language habits.
Smart Images

Figure CN116778382B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-modal marine scene video description algorithm based on instance segmentation auxiliary information, belonging to the cross field of computer vision and natural language processing, and being a downstream task in the multi-modal field. Background Art
[0002] With the popularization of videos in daily life and the increase in the usage volume, the technology of automatically generating video descriptions has gradually become a popular research direction. The task of generating video descriptions can be regarded as transforming the content and plot shown in the video into a text-based description, which can help users understand the video content more quickly and improve the user experience.
[0003] Marine scene video description is a downstream sub-task of the video description task, which is the process of transforming the content and information of marine scene videos into natural language descriptions. The research on marine scene video description aims to develop automated methods to help computers understand and process the content of marine scene videos. Marine scene video description can be applied to multiple fields, such as marine ecological protection, marine resource exploration, marine tourism, marine science popularization, etc. By automatically describing marine scene videos, it is convenient to obtain knowledge about marine ecology, species, geographical information, etc., and improve the understanding and protection of the marine ecological environment. For the research on marine scene video description, it is necessary to combine the knowledge and technologies of multiple fields such as marine science, computer vision, and natural language processing.
[0004] Most of the models for implementing the video description task follow the encoder-decoder architecture. In the early stage, the Convolutional Neural Network (CNN) and the Recurrent Neural Network (RNN) were concatenated to complete the video description task. CNN generally uses 3D networks such as I3D and S3D to extract features from the video, and the extracted video features are fed into the RNN network to generate corresponding description sentences. LSTM networks are often used in RNN. With the emergence and development of the Transformer network, the network model for this task has gradually been dominated by the Transformer. The video feature extraction still uses the S3D network, and the text feature extractor is replaced by Bert. The output result is obtained after fusing the two modal features. Using 3DCNN to extract features from the video cannot well capture the semantic information of the video, cannot capture the changes and events occurring in the video, and at the same time, because more computational effort is introduced in the time dimension, the training and inference costs may be higher. The video feature extractor and the text feature extractor using different network architectures cannot interact well, affecting the semantic alignment between the two modalities. Therefore, the previous work has great limitations in completing the video description task. Summary of the Invention
[0005] The object of the present invention is to provide a multi-modal marine scene video description algorithm based on instance segmentation auxiliary information. This algorithm uses Video-Swin-Transformer as a video feature extractor, reducing the problem of excessive computational complexity, enhancing the correlation between video and text, and at the same time performing instance segmentation on the marine scene to create an auxiliary information dictionary to obtain richer semantic information in the marine scene video, making the text description under the marine scene more abundant and standardized.
[0006] To achieve the above object, the present invention includes the following steps:
[0007] 1. Design and produce a marine scene video description dataset and an image dataset, including 1000 marine videos and 5000 marine images respectively. Each video in the video dataset corresponds to 5 text labels, which describe the content in the video. The image dataset is made by sampling 5 frames from each video in the video dataset;
[0008] 2. Use the SegmentAnything network to segment the foreground instances and background information in the marine images, record and write the foreground information and background information into the auxiliary information dictionary, and send the content of the auxiliary information dictionary into the text encoder to obtain auxiliary information features;
[0009] 3. Use the Video-Swin-Transformer video feature extractor and the Bert text feature extractor respectively to extract features from the video data and text label data;
[0010] 4. Fuse the video features and text label features and send them into a single-stream multi-modal interaction encoder. In the interaction encoder, the video features and text features complete semantic alignment tasks, text masking tasks, and video frame masking tasks, and obtain multi-modal global information features;
[0011] 5. Implement a dual-stream joint video description algorithm for multi-modal global information features and auxiliary information features based on contrast learning, perform joint contrast learning on the multi-modal global information features and auxiliary information features, interactively fuse the dual-stream features, and send them into the language decoder;
[0012] 6. The language decoder is an autoregressive decoder used to convert the dual-stream features into natural language understandable by humans. The language decoder decodes the fused dual-stream features to obtain a description statement, calculates the loss between the obtained description statement and the annotated text label, and completes the language reconstruction task, continuously optimizing the text description ability and effect.
[0013] The beneficial effects of the present invention are:
[0014] 1. Good video feature extraction effect: The present invention uses Video-Swin-Transformer as the video feature extractor, adopts a multi-scale sliding window method to increase the local receptive field, and uses a local attention mechanism to reduce the problem of too large computational complexity of vision transformer. Moreover, it adopts the same network architecture as the text feature extractor, which can better align the semantic information of the two modalities.
[0015] 2. Rich semantic content: The present invention also uses the Segment Anything network to perform instance segmentation on the ocean scene, extracts the key foreground object information and background information in the video, provides richer information for subsequent text descriptions, makes the semantic connections in the video closer, deepens the understanding of the ocean scene content by the network model, and pays attention to the details of the ocean scene to impose scene constraints on the text description, making the text description under the ocean scene more standardized.
[0016] 3. Rich optimization objectives: The present invention sets five optimization objectives (semantic alignment task, text masking task, video frame masking task, auxiliary information contrast learning task, language reconstruction task) to better train the model. The semantic alignment task aligns video features and text features, creating a basis for better interaction between the two. The text masking task improves the model's language understanding ability and context understanding ability. The video frame masking task improves the model's video understanding ability and context understanding ability. The segmentation network is used to extract richer semantic information in the video as auxiliary information for the main network, making the network model pay more attention to the details and content of the ocean scene video. At the same time, the text description is restricted. The language reconstruction task is responsible for autoregressive decoding of the fused features, making the description generated by the model more fluent and in line with our usual speaking habits. Brief Description of the Drawings
[0017] Figure 1 : It is the flowchart of the multi-modal ocean scene video description algorithm based on instance segmentation auxiliary information of the present invention.
[0018] Figure 2 : It is the network model structure diagram of the multi-modal ocean scene video description algorithm based on instance segmentation auxiliary information of the present invention.
[0019] Figure 3 : Network model diagram of Video-Swin-Transformer.
[0020] Figure 4 : Network model diagram of the multi-modal interaction encoder.
[0021] Figure 5 : Network model diagram of the language decoder. Detailed Embodiment
[0022] The flowchart of the present invention is as Figure 1 shown, and the overall network model structure diagram is as Figure 2 shown. The following describes the specific implementation process of the technical solution of the present invention.
[0023] 1. Make a marine scene video dataset, which contains about 1,000 videos. The video content mainly focuses on the sea surface scene, supplemented by the underwater scene. The sea surface scene includes: ship navigation relationship, ship position relationship, maritime traffic, sea surface movement, shore situation, etc.; the underwater scene includes: marine biological activities, seabed terrain situation, marine garbage situation, etc. Divide this dataset into two parts, one part is the video dataset, and the other part is the image dataset. The video dataset contains 1,000 mp4 files. The video dataset is randomly divided into a training set and a test set according to a ratio of 4:1. At the same time, the video names are named in the way of "video + serial number", such as: "video1", "video2". Record the training video names and test video names into the training csv file and the test csv file respectively. Each video corresponds to 5 text descriptions. Store the video names and text descriptions in one-to-one correspondence in a json file. The image dataset is made on the basis of the video dataset. Randomly sample 5 frames from each video in the video dataset and save them as jpg files. At the same time, the image names are named in the way of "image + serial number", such as: "image1", "image2". Store the image names in the image csv file.
[0024] 2. Use the instance segmentation network to operate on the image dataset to extract the auxiliary semantic information of the video. We use the powerful SegmentAnything network to segment the foreground information and background information in our images. When segmenting the foreground, the number and category of the instance bodies in the foreground need to be recorded to form an auxiliary information dictionary. Write the auxiliary information dictionary into the json file of the video dataset. In this way, in the json file, one video corresponds to 5 images, 5 text descriptions and 1 auxiliary information dictionary, such as: "video1 + picture1 + caption1: "two boats are sailing on the sea under the sun" + dict1: {"boat1", "boat2", "sea", "sun"}". Send the auxiliary information dictionary into Bert, and the output is the extracted auxiliary information feature S. The auxiliary information feature is used as the prior knowledge of the marine scene of the model to assist the subsequent marine scene video description work.
[0025] 3. Feature extraction is performed on the video dataset. First, the video data and text data are embedded into a video sequence f and a text sequence t. Then, we use the Video-Swin-Transformer network to extract features from the video sequence f. The network model of Video-Swin-Transformer is as shown in Figure 3 . The Bert language encoder is used to extract features from the text sequence t. The feature extraction formulas for the two modalities are:
[0026] v = VideoSwinTransformer(f) (1)
[0027] w = Bert(t) (2)
[0028] where v is the video feature and w is the text feature.
[0029] 4. The video feature v and the text feature w are fused and fed into a multi-modal interaction encoder. The multi-modal interaction encoder consists of 6 Transformer encoder blocks. Each Transformer encoder block contains a self-attention layer and a feed-forward layer. The network model is as shown in Figure 4 . The fused feature passes through the multi-modal interaction encoder to obtain the output M. M is the multi-modal global information feature. The formula is:
[0030] M = Interact encoder(v:w) (3)
[0031] In the interaction encoder, the video feature and the text feature complete the semantic alignment task. The loss function is:
[0032] P = E (w,v)~P exp(e(w, v)) (4)
[0033] N = E (w,v)~N exp(e(w, v)) (5)
[0034]
[0035] Among them, (w, v) is a video text feature pair, P is the positive sample of video text feature alignment, N is the negative sample of video text feature alignment, and the loss function of semantic alignment is the result obtained by using Noise Contrastive Estimation (NCE) Loss to perform contrastive learning on positive and negative samples. Text masking task: Provide a text sequence containing a special token [MASK] (i.e., the mask), and the words in this text sequence are masked with a probability of 15%, and then let the model predict the original word at the masked position. For example, provide "The ship is at sea [MASK]", and predict the word at the [MASK] position, such as "sailing", "turning", or "colliding", etc. This task allows the model to simultaneously focus on the context information on both sides of [MASK]. The loss function formula for the text masking task is:
[0036]
[0037] where w is the input text feature, v is the input video feature, w m is the masked text feature, D is the entire training set, and p is the probability. Similarly, based on the text masking task, we propose a video frame masking task: The input video frame sequence contains a special token [MASK], and the frames in the video frame sequence are randomly replaced with [MASK] with a probability of 15%, and then use the model to predict the replaced video frames. Since it is very difficult to directly predict the original RGB video frames, we use the method of contrastive learning to enhance the correlation between video frames, and improve the spatial modeling ability of the model by learning the context information of video frames. The loss function formula for the video frame masking task is:
[0038]
[0039]
[0040] where v is the real-valued vector of video features, is the linear output of v, M v is the video part of the output result of the interactive encoder, belongs to M v .
[0041] 5. The multi-modal global information feature M and the auxiliary information feature S are subjected to contrastive learning. If the multi-modal global information feature contains the foreground information and background information in the auxiliary information, and the number and category of the instance subjects in the foreground information can be matched, we set this feature pair as a positive sample; if they cannot be matched, we set it as a negative sample. The NCE Loss is used to perform contrastive learning on the multi-modal global information feature and the auxiliary information feature to standardize the results of the ocean scene video description statements, while enabling the network to obtain richer semantic information and enhancing the alignment ability between the ocean video and the text description. The loss function formula for this contrastive learning is:
[0042]
[0043]
[0044] L CMS = L M2S + L S2M (12)
[0045] where B is the batch size, σ is the learnable temperature parameter, M i and S j are the normalized embeddings of the i-th multi-modal global information feature and the j-th auxiliary information feature.
[0046] 6. After completing the contrastive learning, the multi-modal global information feature M and the auxiliary information feature S are fused and fed into the language decoder to obtain the text description O corresponding to the ocean scene video. The formula for this process is:
[0047] O = Caption decoder(M:S) (13)
[0048] To reconstruct the input text description and enable the model to have the generation ability, we adopt the autoregressive decoder Caption decoder. Caption decoder consists of 3 Transformer decoder blocks. Each Transformer decoder block contains one self-attention layer and one feed-forward layer. Its network model is as Figure 5 shown. Caption decoder decodes the fused features to complete the language reconstruction task. Its loss function is:
[0049]
[0050] where T is the length of the generated text sequence, t is the t-th word, S is the auxiliary information feature, and M is the multi-modal global information feature.
[0051] 7. The loss functions of the five tasks are combined into a total loss function. The total loss function is shown in formula (15). The marine scene video dataset is input into the network model for training on the training set. In each round, the total loss function is calculated, and then the optimizer is used to optimize the entire network. After completing the training phase, it is tested on the test set to evaluate the effect of the network model and the quality and fluency of the output description statements. Finally, according to the test situation, the model is further fine-tuned.
[0052] L Overall = L VLM + L MLM + L MFM + L CMS + L CAP (15)
[0053] It should be noted that the above description is only an embodiment of the present invention, which is only to explain the present invention and does not limit the scope of the patent of the present invention. Modifications that are merely obvious to those skilled in the art within the technical concept of the present invention are also within the protection scope of the present invention.
Claims
1. A multi-modal marine scene video description method based on instance segmentation auxiliary information, characterized in that, it includes the following steps: (1) Design and produce a marine scene video description dataset and an image dataset, including 1000 marine videos and 5000 marine images respectively. Each video in the video dataset corresponds to 5 text labels, which describe the content in the video. The image dataset is made by sampling 5 frames from each video in the video dataset; (2) SegmentAnything is an instance segmentation tool used to extract the features of the marine image set, which helps to obtain richer visual information and assist in the generation of descriptions. Use the SegmentAnything network to segment the foreground instances and background information in the marine images, record and write the foreground information and background information into the auxiliary information dictionary, and send the content of the auxiliary information dictionary into the text encoder to obtain auxiliary information features; (3) Use the Video-Swin-Transformer video feature extractor and the Bert text feature extractor to extract features from the video data and text label data respectively; (4) Fuse the video features and text label features and send them into a single-stream multi-modal interaction encoder. In the interaction encoder, the video features and text features complete semantic alignment tasks, text masking tasks, and video frame masking tasks, and obtain multi-modal global information features; (5) Implement a two-stream joint video description algorithm for multi-modal global information features and auxiliary information features based on contrast learning. Perform joint contrast learning on the multi-modal global information features and auxiliary information features, interactively fuse the two-stream features, and send them into the language decoder; (6) The language decoder is an autoregressive decoder used to convert the two-stream features into natural language that can be understood by humans. The language decoder decodes the fused two-stream features to obtain a description statement, calculates the loss between the obtained description statement and the annotated text label, and completes the language reconstruction task, continuously optimizing the text description ability and effect.
2. The multi-modal marine scene video description method based on instance segmentation auxiliary information as described in claim 1, characterized in that, According to the production of the marine scene video dataset described in step (1), the dataset includes two parts: video and text labels. The video content mainly focuses on the sea surface scene, supplemented by the underwater scene. The sea surface scene includes: ship navigation relationship, maritime traffic, maritime sports, and shore situation; the underwater scene includes: marine biological activities, underwater terrain situation; each video is annotated with 5 text labels; randomly sample 5 frames from each video in the marine scene video dataset as marine scene images, and each video corresponds to 5 images, which are made into an image dataset.
3. The multi-modal marine scene video description method based on instance segmentation auxiliary information as described in claim 1, characterized in that, Implement an ocean scene feature extraction network based on the instance segmentation auxiliary information dictionary. According to the production of the auxiliary information dictionary described in step (2), extract the single-modal auxiliary information features. We use the Segment Anything network to perform instance segmentation on each image in the ocean scene image dataset, record the number and categories of the segmented foreground objects and background regions, and store them as the auxiliary information dictionary. Then, send it into Bert to extract the auxiliary information features, which are used as the prior information of the ocean scene to assist the subsequent text description work.
4. A multimodal ocean scene video description method based on instance segmentation auxiliary information as described in claim 1, characterized in that, According to the feature extraction described in step (3), we use Video-Swin-Transformer to extract features from the ocean scene video dataset and use Bert to extract features from the text labels corresponding to the videos.
5. A multimodal ocean scene video description method based on instance segmentation auxiliary information as described in claim 1, characterized in that, Implement a multimodal global information feature learning network that interacts and fuses ocean scene video features and text features. According to the multimodal interaction coding described in step (4), use the Transformer Encoderblock to fuse the video features and text features and send them into the interaction encoder to obtain multimodal features. In the interaction encoder, the two-modal data complete the semantic alignment task. The loss function formula is: P = E (w,v)~P exp(e(w, v))⑴ N = E (w,v)~N exp(e(w, v))⑵ where (w, v) is the video-text feature pair, P is the positive sample of the video-text feature alignment, N is the negative sample of the video-text feature alignment, and the loss function of the semantic alignment is the result obtained by using the Noise Contrastive Estimation (NCE) Loss to perform contrastive learning on the positive and negative samples; for the text masking task, 15% probability is used to mask the words in the input text label. The loss function formula is: where w is the input text feature, v is the input video feature, and w m is the masked text feature, D is the entire training set, and p is the probability; similar to the text masking task, the video frame masking task masks the frames in the video with a probability of 15%, and its loss function formula is: where v is a real-valued vector of video features, is a linear output of v, and M v is the video part of the output result of the interaction encoder, belongs to M v .
6. A multimodal ocean scene video description method based on instance segmentation auxiliary information as described in claim 1, characterized in that, Implement a two-stream joint video description algorithm for multimodal global information features and auxiliary information features based on contrastive learning. According to the contrastive learning of the multimodal global information features and auxiliary information features described in step (5), if the multimodal global information features contain the foreground information and background information in the auxiliary information features, and the number and categories of the instance objects can be matched, we set it as a positive sample, and set it as a negative sample if they do not match. Use the NCE Loss to perform contrastive learning on the auxiliary information features and multimodal global information features to standardize the results of the ocean scene video description statements. The contrastive learning loss function formula is: L CMS = L M2S + L S2M ⑼ where B is the batch size, σ is the learnable temperature parameter, M i and S j are the normalized embeddings of the i-th multimodal feature and the j-th auxiliary information feature respectively.
7. A multimodal ocean scene video description method based on instance segmentation auxiliary information as described in claim 1, characterized in that, According to the language decoder described in step (6), the Transformer Decoder block is used to decode the result after fusing the auxiliary information feature and the multimodal global information feature to complete the language reconstruction task, and its loss function is as follows: where T is the length of the generated text sequence, t is the t-th word, S is the auxiliary information feature, and M is the multimodal global information feature.
Citation Information
Patent Citations
Retrieval-oriented monitoring video semantic description and inspection modeling method
CN102880692A
Traffic scene video description generation method and device based on multi-modal feature fusion
CN115496134A