Semi-supervised video description generation method with big language model guiding pseudo label enhancement

By constructing a dual-path collaborative framework and a pseudo-label enhancement method guided by a large language model, the problem of semantic bias in video description generation is solved, resulting in more accurate and consistent video descriptions.

CN120932159AActive Publication Date: 2025-11-11XIAN UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511394804.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-11-11
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Existing semi-supervised video description generation methods do not make full use of the prior knowledge of a small number of base descriptions, making it difficult to capture the fine-grained differences in the structured semantic elements of the video, resulting in semantic deviations in the generated video descriptions.

Method used

A dual-path collaborative framework is constructed, including a video-level coarse-grained pseudo-label generation branch and a high-confidence visual tuple generation branch. Pseudo-labels and visual tuples are generated using the CoCap network and the VIS-MiniCPM model, and enhanced by the LLaMA large language model. Finally, the video description is generated through iterative self-training optimization.

Benefits of technology

It enables semantic expression of video content at different granularities, resulting in smoother and more natural video descriptions. It also reduces redundant word interference in fine-grained video frame-level descriptions, improving the accuracy and consistency of the descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932159A_ABST
    Figure CN120932159A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised video description generation method for large language model-guided pseudo-label enhancement, and the method comprises the steps: constructing a two-way collaborative framework which comprises a video-level coarse-grained pseudo-label generation branch and a high-confidence visual tuple generation branch; respectively preprocessing the videos with the base descriptions, inputting the preprocessed videos into a two-way collaborative framework, and generating a video-level coarse-grained pseudo tag and a high-confidence visual tuple; taking the high-confidence visual tuple as a fine-grained visual prompt, inputting the high-confidence visual tuple and the video-level coarse-grained pseudo-tag into an LLaMA large language model together, and completing enhancement of the video-level coarse-grained pseudo-tag under the guidance of the large language model; and through iterative self-training optimization, a final video description is generated. The method provided by the invention solves the problem that the generated video description has semantic deviation due to the fact that the utilization of a small amount of base descriptions does not fully depend on priori knowledge and the fine-grained difference of video structured semantic elements is difficult to fully capture in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and image processing technology, specifically relating to a semi-supervised video description generation method with pseudo-label enhancement guided by a large language model. Background Technology

[0002] With the rapid development of internet technology and the widespread adoption of social media platforms, video has gradually become one of the mainstream forms of information dissemination. Whether in social media, e-commerce platforms, or education and training, the scale of video content generation has increased dramatically. How to quickly extract the desired content from massive amounts of information has become a common challenge for users. Specific groups (such as visually impaired and cognitively impaired individuals) face numerous obstacles in understanding video content, further limiting their ability to access video information. Against this backdrop, users have new demands for understanding video content: how to achieve deep semantic understanding of video data to achieve automatic comprehension and efficient dissemination. To address this need, video description generation methods have emerged.

[0003] Video description generation methods are mainly divided into three categories: fully supervised video description generation methods, unsupervised video description generation methods, and semi-supervised video description generation methods. Fully supervised video description generation methods rely on a large number of base descriptions and use encoder-decoder models to generate natural language descriptions of video content. However, due to the reliance on a large number of manually labeled base descriptions, these methods incur high labor costs and limited applicability. To alleviate this problem, further research has been conducted on unsupervised video description generation methods. These methods mainly use semantic matching and reconstruction mechanisms to generate text descriptions of video content. However, due to the lack of base descriptions semantically aligned with the video content, unsupervised video description generation methods struggle to accurately establish cross-modal semantic mapping relationships between video and text, making it difficult to generate visual descriptions highly consistent with human natural language expressions and video semantic content. Currently, semi-supervised video description generation methods mainly utilize semantic elements from a small number of base descriptions and a large number of unlabeled data features, employing pseudo-label generation methods to achieve relatively accurate video content understanding and rich visual text descriptions, thus overcoming the aforementioned problems. However, existing semi-supervised video description generation methods do not fully utilize prior knowledge in their use of a small number of base descriptions, making it difficult to fully capture the fine-grained differences in the structured semantic elements of the video, which can easily lead to semantic deviations in the generated video descriptions. Summary of the Invention

[0004] The purpose of this invention is to provide a semi-supervised video description generation method guided by a large language model and enhanced by pseudo-labels. This method solves the problem in existing technologies where the use of a small number of base descriptions does not fully leverage prior knowledge, making it difficult to capture the fine-grained differences in the structured semantic elements of the video, resulting in semantic deviations in the generated video descriptions.

[0005] The technical solution adopted in this invention is a semi-supervised video description generation method with pseudo-label enhancement guided by a large language model, which specifically includes the following steps: Step 1: Construct a dual-path collaborative framework, which includes a video-level coarse-grained pseudo-label generation branch and a high-confidence visual tuple generation branch; Step 2: After preprocessing the video with the base description, input it into the dual-path collaborative framework to generate video-level coarse-grained pseudo-labels and high-confidence visual tuples; Step 3: Input the high-confidence visual tuples as fine-grained visual cues, along with the video-level coarse-grained pseudo-labels, into the LLaMA large language model. Under the guidance of the LLaMA large language model, the video-level coarse-grained pseudo-labels are enhanced. Step 4: Optimize and generate the final video description through iterative self-training.

[0006] The invention is further characterized by: In step 1, the video-level coarse-grained pseudo-label generation branch includes the CoCap network model, which is used to process video with a base description into video-level coarse-grained pseudo-labels by dividing the image groups into several independent encoding and decoding. The high-confidence visual tuple generation branch is the VIS-MiniCPM model. The VIS-MiniCPM model adds a visual event information selector to the output of the MiniCPM-V-2 model. The visual event information selector includes an LLaMA large language model and a visual tuple generation module. The MiniCPM-V-2 model is used to process video divided into continuous frames with base descriptions into fine-grained visual descriptions at the video frame level. The LLaMA large language model of the visual event information selector is used to process the fine-grained visual descriptions at the video frame level into a set of sentences. The visual tuple generation module obtains the weighted score set of each type of word in the sentence set based on the TextRank algorithm and word frequency statistics algorithm, and selects the highest-scoring words from each type of word to form a high-confidence visual tuple.

[0007] In step 2, the process of generating video-level coarse-grained pseudo-tags is as follows: Video with base description The consecutive frames are divided into several independently encoded and decoded image groups, as shown below:

[0008] In the formula, Indicates the first n A group of images that are independently encoded and decoded. Each independently encoded and decoded image group starts with an I-frame, followed by P-frames and B-frames, for a total of M frames. Video with base description, divided into several independently encoded and decoded image groups. Input the CoCap network model; the CoCap network model includes an encoder and a decoder, and the processing procedure is as follows: During the encoding phase, each GOP's I-frames are converted into 768-dimensional feature vectors by an I-frame encoder, and then passed through 12 Transformer encoder layers to output contextual semantic features. The I-frame encoder includes a 16×16 convolutional kernel, a 16×16 stride, and a convolutional layer with 768 output channels. Each GOP's P-frames or B-frames are converted into 192-dimensional feature vectors by a motion encoder and then passed through 2 Transformer encoder layers to output motion vectors. The motion encoder includes an 8×8 convolutional kernel, an 8×8 stride, and a convolutional layer with 192 output channels. Simultaneously, each GOP's P-frames or B-frames are converted into 768-dimensional feature vectors by a residual encoder and then passed through 2 Transformer encoder layers to output residual features. The residual encoder includes a 64×64 convolutional kernel, a 64×64 stride, and a convolutional layer with 768 output channels. After the motion features and residual features are processed by dropout, they are combined with the contextual semantic features of the I-frame to form a temporal feature sequence. The temporal feature sequence is then input into the action encoder. The action encoder first encodes the temporal feature sequence through position embedding and frame type embedding. Then, the encoded temporal feature sequence is sequentially passed through the self-attention module for temporal modeling and through the cross-attention module for feature fusion, and finally outputs a 512-dimensional fused feature vector. In the decoding stage, the decoder adopts a two-layer BERT structure, which includes word embedding, positional encoding, and type encoding. Finally, it generates video-level coarse-grained pseudo-tags through linear mapping. .

[0009] The loss function of the CoCap network model is expressed as follows: (1) In the formula, Represents the loss function. Indicates the length of the predicted word sequence. Indicates the current number t The prediction results for each word Indicates the preceding The true sequence of words, Indicates before The true sequence of words and videos with base descriptions Under the conditions, the first t The word prediction is as follows The probability of.

[0010] In step 2, the process of generating high-confidence visual tuples is represented as follows: Step 2.1: Transfer the video with base description Sampling is performed according to frame rate T as a sequence. Each video frame in the sequence The high-confidence visual tuple generation branch, namely the VIS-MiniCPM model, is input sequentially. First, it passes through the MiniCPM-V-2 model to generate the initial text description for each video frame, as shown below: (2) In the formula, This indicates a fine-grained visual description. This refers to the text description of the nth video frame generated by MiniCPM-V-2. This represents the model parameters of MiniCPM-V-2; Step 2.2: Utilize the LLaMA large language model in the visual event information selector for fine-grained visual description. In summary, a general semantic description of the video segment is generated, as follows: (3) In the formula, A general semantic description of a video segment. Represents the LLaMA large language model. These are the model parameters for LLaMA; Subsequently, fine-grained visual descriptions at the video level will be used. and a general semantic description of the corresponding video clip Merge to obtain a set of sentences , means as follows: (4) In the formula, Indicates will and connect; Step 2.3: Set the sentences The input visual tuple generation module obtains a weighted score set of words of each category in the sentence set based on the TextRank algorithm and word frequency statistics algorithm. It then selects the words with the highest scores from each category to form a high-confidence visual tuple.

[0011] Step 2.3 specifically includes the following sub-steps: Step 2.3.1: Use spaCy to process the sentence set Dependency parsing yields all subject, object, action, and context words in the sentence. Each category forms a candidate word set, represented as follows: , , , ; Step 2.3.2: Based on the word frequency algorithm, calculate the candidate words in each candidate word set. In the sentence collection The frequency of occurrence in the following is represented as follows: (5) In the formula, This represents the i-th word in the text. ; Represents the frequency function, if but =1, otherwise 0 Represents a set of sentences The words in m Indicates the total number of words in the text; Meanwhile, the TextRank scores of the candidates in each candidate word set are calculated based on TextRank, and are represented as follows: (6) In the formula, Indicate candidate words TextRank score, The damping coefficient is... Indicates with candidate words In the sentence collection A set of adjacent candidate words with semantic relationships. Indicate candidate words and Edge weights between them; Used to determine in the first Candidate words in a sliding window and Do they appear simultaneously? If in the first... Simultaneous appearance of multiple windows and The value is 1 if the value is 1, otherwise the value is 0. Step 2.3.3: Calculate the score of all words in each candidate word set, as shown below: (7) In the formula, , Indicates the weighting coefficient. ; Step 2.3.4: Extract high-confidence visual tuples from each candidate word set. ,in, , , , These represent the subject, object, action, and environment, respectively, as follows: (8) In the formula, This indicates selecting the word from the candidate word set that maximizes the corresponding score function value.

[0012] The LLaMA large language model includes a byte pair encoding model and a decoder. The decoder consists of 32 Transformer modules, each containing a multi-head self-attention network and a feedforward neural network. The input of each sub-layer is normalized using the RMSNorm function. Step 3 specifically includes the following sub-steps: Step 3.1: The byte-pair encoding model will input... The input matrix X is obtained by decomposing the words into sub-word units and mapping them to the corresponding tokens in the predefined vocabulary. The tokens are then converted into word vectors. Step 3.2: Enhance the input matrix X input decoder: During the enhancement process, the multi-head self-attention network in each Transformer module introduces positional information by encoding the query and key through rotation. The multi-head self-attention network also integrates fine-grained visual cues provided by high-confidence visual tuples into the semantic representation of coarse-grained pseudo-labels. After processing by each Transformer module, the output features are passed to the next Transformer module for further feature transformation, until the final Transformer module outputs a feature vector containing both coarse-grained pseudo-labels and fine-grained visual cues. , The size is [1, 4096], and the calculation process is as follows: (9) (10) In the formula, Indicates the first Semantic features extracted by an attention head in a subspace For multi-head self-attention functions, This represents the rotational position encoding applied to the query and key matrices. , and For learnable query, key, and value parameter matrices, Represents matrix multiplication; This indicates that the features from multiple attention heads are concatenated. To output the transformation matrix, This represents a feedforward neural network. This indicates a normalization operation. This represents the output of the 32nd layer Transformer module; Feature vectors with coarse-grained pseudo-labels and fine-grained visual cues from the video Vocabulary embedding matrix of the LLaMA large language model Perform a product operation, then normalize using the softmax function to obtain the th... Pace prediction words probability distribution , means as follows: (11); Step 3.3: Based on the greedy sampling strategy, the first... Predicting words for steps From probability distribution The word with the highest probability is selected, as shown below: (12) By iteratively performing single-step predictions, the final enhanced pseudo-label text sequence is generated. , means as follows: (13) The output probability of the enhanced pseudo-labeled text sequence is represented as follows: (14) In the formula, T represents the length of the generated pseudo-label text sequence.

[0013] Step 4, combine the video data set with the base description. With the enhanced pseudo-label text sequence video data set Merging them into a training dataset, represented as Repeat steps 1 through 3, using the following overall loss function for iterative self-training: (15) In the formula, Represents the overall loss function. This indicates the number of samples in a dataset with a basis description. This indicates the number of samples in the dataset without a basis description. This indicates that the data in the dataset has a basis description. One sample, Indicates the true label, This indicates the first term in the dataset that is described without basis. d One sample, Represents samples in an unlabeled dataset Corresponding to the semantically enhanced pseudo-tags, Generally refers to the parameters of the semi-supervised method proposed in this invention. These are hyperparameters used to balance the weights of labeled data loss and augmented pseudo-labeled data loss in the total loss function. This represents the number of self-training iterations.

[0014] The beneficial effects of this invention are: This invention presents a semi-supervised video description generation method guided by a large language model and enhanced pseudo-tags. It constructs a dual-path collaborative framework to generate video-level coarse-grained pseudo-tags to capture the global semantics of the video, while simultaneously generating high-confidence visual tuples to supplement the local visual information of the video, achieving semantic expression of video content at different granularities. Within the high-confidence visual tuple branch of the dual-path collaborative framework, a VIS-MiniCPM model based on a visual event information selector is constructed. This model extracts high-confidence visual tuples from the video frame-level fine-grained description. The designed visual event information selector selects visual tuples that match the video theme and key information, reducing redundant word interference in the video frame-level fine-grained description. Finally, the aforementioned high-confidence visual tuples are used as fine-grained visual cues and input along with the video-level coarse-grained pseudo-tags into the LLaMA large language model. Under the guidance of the large language model, the video-level coarse-grained pseudo-tags are enhanced, resulting in enhanced pseudo-tags. Then, using the base description and enhanced pseudo-tags, a more fluent and natural video description is generated through self-training with relatively few base descriptions. Attached Figure Description

[0015] Figure 1 This is a framework diagram of the semi-supervised video description generation method with large language model-guided pseudo-label enhancement according to the present invention; Figure 2 This is a graph showing the training loss reduction of the semi-supervised video description generation method with large language model-guided pseudo-label enhancement based on the present invention on the MSVD dataset; Figure 3 This is a graph showing the training loss reduction of the semi-supervised video description generation method with large language model-guided pseudo-label enhancement based on the present invention on the MSR-VTT dataset; Figure 4 This is a visualization of the semi-supervised video description generation method for large language model-guided pseudo-label enhancement based on the MSVD dataset. Figure 5 This is a visualization of the semi-supervised video description generation method for large language model-guided pseudo-label enhancement based on the MSR-VTT dataset. Detailed Implementation

[0016] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0017] This invention relates to a semi-supervised video description generation method guided by a large language model and enhanced with pseudo-labels, such as... Figure 1 As shown, the specific steps include the following: Step 1: Construct a dual-path collaborative framework, which includes a video-level coarse-grained pseudo-label generation branch and a high-confidence visual tuple generation branch; The video-level coarse-grained pseudo-label generation branch includes the CoCap network model, which is used to process videos with base descriptions that are divided into several independent encoded and decoded image groups into video-level coarse-grained pseudo-labels. The high-confidence visual tuple generation branch uses the VIS-MiniCPM model, an improvement on the MiniCPM-V-2 model. It adds a visual event information selector to the MiniCPM-V-2 model's output, taking video frames as input and outputting visual tuples. The visual event information selector includes an LLaMA large language model and a visual tuple generation module. The MiniCPM-V-2 model consists of a visual encoder, a compression layer, and a large language model. The visual encoder uses a SigLIPSoViT-400m / 14 and employs adaptive visual coding technology to process high-resolution images with different aspect ratios, slicing them for encoding. The compression layer uses a perceiver resampler structure, compressing 1024 visual tokens per slice into 64. The large language model is based on MiniCPM 2.4B, combining the processed visual tokens with text input to generate fine-grained visual descriptions at the video frame level. The visual event information selector is used to process fine-grained visual descriptions at the video frame level into high-confidence visual tuples. Specifically, the LLaMA large language model is used to process fine-grained visual descriptions at the video frame level into generalized semantic descriptions for video segments. The visual tuple generation module obtains a weighted score set of words in the generalized semantic descriptions of video segments and fine-grained visual descriptions at the video frame level based on the TextRank algorithm and word frequency statistics algorithm. The words with the highest scores in each category are selected to form high-confidence visual tuples.

[0018] Step 2: After preprocessing the video with the base description, input it into the dual-path collaborative framework to generate video-level coarse-grained pseudo-labels and high-confidence visual tuples; The process of generating video-level coarse-grained pseudo-tags is as follows: Video with base description The consecutive frames are divided into several independently encoded and decoded image groups, as shown below:

[0019] In the formula, Indicates the first n A group of images that are independently encoded and decoded. Each independently encoded and decoded image group starts with an I-frame, followed by P-frames and B-frames, for a total of M frames. Video with base description, divided into several independently encoded and decoded image groups. Input the CoCap network model; the CoCap network model is the network model in the paper Shen Y, Gu X, Xu K, et al. Accurate and fast compressed video captioning[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023: 15558-15567.doi: 10.1109 / ICCV51070.2023.01426., specifically including an encoder and a decoder, and the processing procedure is as follows: During the encoding phase, each GOP's I-frame is converted into a 768-dimensional feature vector by an I-frame encoder, and then passed through 12 Transformer encoder layers to output contextual semantic features. The I-frame encoder includes a 16×16 convolutional kernel, a 16×16 stride, and a convolutional layer with 768 output channels. Each GOP's P-frame or B-frame is converted into a 192-dimensional feature vector by a motion encoder and then passed through 2 Transformer encoder layers to output motion features. The motion encoder includes an 8×8 convolutional kernel, an 8×8 stride, and a convolutional layer with 192 output channels. Simultaneously, each GOP's P-frame or B-frame is converted into a 768-dimensional feature vector by a residual encoder and then passed through 2 Transformer encoder layers to output residual features. The residual encoder includes a 64×64 convolutional kernel, a 64×64 stride, and a convolutional layer with 768 output channels. After the motion features and residual features are processed by dropout, they are combined with the contextual semantic features of the I-frame to form a temporal feature sequence. The temporal feature sequence is then input into the action encoder. The action encoder first encodes the temporal feature sequence through position embedding and frame type embedding. Then, the encoded temporal feature sequence is sequentially passed through the self-attention module for temporal modeling and through the cross-attention module for feature fusion, and finally outputs a 512-dimensional fused feature vector. In the decoding stage, the decoder adopts a two-layer BERT structure, which includes word embedding, positional encoding, and type encoding. Finally, it generates video-level coarse-grained pseudo-tags through linear mapping. .

[0020] To prevent overfitting, the dropout rate was set to 0.2 on the motion and residual features of the CoCap network model, and to 0.1 on the BERT model.

[0021] The loss function of the CoCap network model is expressed as follows: (1) In the formula, Represents the loss function. Indicates the length of the predicted word sequence. Indicates the current number t The prediction results for each word Indicates the preceding The true sequence of words, Indicates before The true sequence of words and videos with base descriptions Under the conditions, the first t The word prediction is as follows The probability of.

[0022] The generation process of high-confidence visual tuples is represented as follows: Step 2.1: Transfer the video with base description Sampling is performed according to frame rate T as a sequence. Each video frame in the sequence The high-confidence visual tuple generation branch, namely the VIS-MiniCPM model, is input sequentially. First, it passes through the MiniCPM-V-2 model to generate the initial text description for each video frame, as shown below: (2) In the formula, This indicates a fine-grained visual description. This refers to the text description of the nth video frame generated by MiniCPM-V-2. This represents the model parameters of MiniCPM-V-2; Step 2.2: Utilize the LLaMA large language model in the visual event information selector for fine-grained visual description. In summary, a general semantic description of the video segment is generated, as follows: (3) In the formula, A general semantic description of a video segment. Represents the LLaMA large language model. These are the model parameters for LLaMA; Subsequently, fine-grained visual descriptions at the video level will be implemented. and a general semantic description of the corresponding video clip Merge to obtain a set of sentences , means as follows: (4) In the formula, Indicates will and connect; Step 2.3: Set the sentences The input visual tuple generation module obtains a weighted score set of words of each category in the sentence set based on the TextRank algorithm and word frequency statistics algorithm. It then selects the words with the highest scores from each category to form a high-confidence visual tuple.

[0023] Because the above process generates a large number of redundant words and lacks representation of the video theme, and considering the characteristics of word frequency statistics algorithms in representing text themes by measuring word frequency and TextRank in extracting keywords by constructing word co-occurrence graphs, this invention designs a visual tuple generation module based on word frequency statistics and TextRank algorithms to remove redundant words and extract visual tuples from the above process. The specific process is as follows: Step 2.3.1: Use spaCy to process the sentence set Dependency parsing yields four categories of words in the sentence: Subject, Object, Action, and Environment. Each category forms a candidate word set, denoted as follows: , , , ; Step 2.3.2: Based on the word frequency algorithm, calculate the candidate words in each candidate word set. In the sentence collection The frequency of occurrence in the following is represented as follows: (5) In the formula, This represents the i-th word in the text. ; Represents the frequency function, if but =1, otherwise 0 Represents a set of sentences The words in m Indicates the total number of words in the text; Meanwhile, the TextRank scores of the candidates in each candidate word set are calculated based on TextRank, and are represented as follows: (6) In the formula, Indicate candidate words TextRank score, The damping coefficient is... Indicates with candidate words In the sentence collection A set of adjacent candidate words with semantic relationships. Indicate candidate words and Edge weights between them; Used to determine in the first Candidate words in a sliding window and Do they appear simultaneously? If in the first... Simultaneous appearance of multiple windows and The value is 1 if the value is 1, otherwise the value is 0. Step 2.3.3: Calculate the score of all words in each candidate word set, as shown below: (7) In the formula, , Indicates the weighting coefficient. ; Step 2.3.4: Extract high-confidence visual tuples from each candidate word set. ,in, , , , These represent the subject, object, action, and environment, respectively, as follows: (8) In the formula, This indicates selecting the word from the candidate word set that maximizes the corresponding score function value.

[0024] Step 3: Input the high-confidence visual tuples as fine-grained visual cues, along with the video-level coarse-grained pseudo-labels, into the LLaMA large language model. Under the guidance of the LLaMA large language model, the video-level coarse-grained pseudo-labels are enhanced.

[0025] The high-confidence visual tuples obtained in step 2 As a fine-grained visual cue, it enables coarse-grained pseudo-labeling at the video level under the guidance of the LLaMA large language model. The enhancement process generates enhanced pseudo-tags that match the video content. Specifically, it includes the following steps: The LLaMA large language model includes a Byte Pair Encoding (BPE) model and a decoder. The decoder consists of 32 Transformer modules, each containing a multi-head self-attention network and a feedforward neural network. The input to each sub-layer is normalized using the RMSnorm function.

[0026] Step 3.1: The byte-pair encoding model will input... The input matrix X is obtained by decomposing the words into sub-word units and mapping them to the corresponding tokens in the predefined vocabulary. The tokens are then converted into word vectors. Step 3.2: Enhance the input matrix X input decoder: During the enhancement process, the multi-head self-attention network in each Transformer module introduces positional information by encoding the query and key through rotation. The multi-head self-attention network also integrates fine-grained visual cues provided by high-confidence visual tuples into the semantic representation of coarse-grained pseudo-labels. After processing by each Transformer module, the output features are passed to the next Transformer module for further feature transformation, until the final Transformer module outputs a feature vector containing both coarse-grained pseudo-labels and fine-grained visual cues. , The size is [1, 4096], and the calculation process is as follows: (9) (10) In the formula, Indicates the first Semantic features extracted by an attention head in a subspace For multi-head self-attention functions, This represents the rotational position encoding applied to the query and key matrices. , and For learnable query, key, and value parameter matrices, Represents matrix multiplication; This indicates that the features from multiple attention heads are concatenated. To output the transformation matrix, This represents a feedforward neural network. This indicates a normalization operation. This represents the output of the 32nd layer Transformer module; Feature vectors with coarse-grained pseudo-labels and fine-grained visual cues from the video Vocabulary embedding matrix of the LLaMA large language model (dimension) Perform a product operation, then normalize using the softmax function to obtain the first product. Pace prediction words probability distribution , means as follows: (11); Step 3.3: Based on the greedy sampling strategy, the first... Predicting words for steps From probability distribution The word with the highest probability is selected, as shown below: (12) By iteratively performing single-step predictions, the final enhanced pseudo-label text sequence is generated. , means as follows: (13) The output probability of the enhanced pseudo-labeled text sequence is represented as follows: (14) In the formula, T represents the length of the generated pseudo-label text sequence.

[0027] Step 4: Optimize and generate the final video description through iterative self-training.

[0028] A collection of video data with a base description With the enhanced pseudo-label text sequence video data set Merging them into a training dataset, represented as Repeat steps 1 to 3 to continuously optimize the model parameters and processing results throughout the process through self-training, reduce errors, and improve the accuracy and consistency of the description until the final video text description is generated.

[0029] This invention uses the following overall loss function for iterative self-training: (15) In the formula, Represents the overall loss function. This indicates the number of samples in a dataset with a basis description. This indicates the number of samples in the dataset that describes the data without a basis. This indicates that the data in the dataset has a basis description. One sample, Indicates the true label, This indicates the first term in the dataset that is described without basis. d One sample, Represents samples in an unlabeled dataset Corresponding to the semantically enhanced pseudo-tags, Generally refers to the parameters of the semi-supervised method proposed in this invention. These are hyperparameters used to balance the weights of labeled data loss and pseudo-labeled data loss in the total loss function. This represents the number of self-training iterations.

[0030] Example 1 This embodiment provides a semi-supervised video description generation method guided by a large language model and enhanced with pseudo-labels, which specifically includes the following steps: Step 1: Construct a dual-path collaborative framework, which includes a video-level coarse-grained pseudo-label generation branch and a high-confidence visual tuple generation branch; Step 2: After preprocessing the video with the base description, input it into the dual-path collaborative framework to generate video-level coarse-grained pseudo-labels and high-confidence visual tuples; Step 3: Input the high-confidence visual tuples as fine-grained visual cues, along with the video-level coarse-grained pseudo-labels, into the LLaMA large language model. Under the guidance of the LLaMA large language model, the video-level coarse-grained pseudo-labels are enhanced. Step 4: Optimize and generate the final video description through iterative self-training.

[0031] Example 2 Based on Example 1, the video-level coarse-grained pseudo-label generation branch in step 1 includes the CoCap network model, which is used to process video with base descriptions divided into several independently encoded and decoded image groups into video-level coarse-grained pseudo-labels; the high-confidence visual tuple generation branch is the VIS-MiniCPM model, which adds a visual event information selector to the output of the MiniCPM-V-2 model. The visual event information selector includes the LLaMA large language model and a visual tuple generation module; the MiniCPM-V-2 model is used to process video with base descriptions divided into consecutive frames into video frame-level fine-grained visual descriptions; the LLaMA large language model of the visual event information selector is used to process the video frame-level fine-grained visual descriptions into a set of sentences, and the visual tuple generation module obtains the weighted score set of each type of word in the sentence set based on the TextRank algorithm and the word frequency statistics algorithm, and selects the word with the highest score from each type of word to form a high-confidence visual tuple.

[0032] Example 3 Based on Example 2, the process of generating video-level coarse-grained pseudo tags in step 2 is as follows: Video with base description The consecutive frames are divided into several independently encoded and decoded image groups, as shown below:

[0033] In the formula, Indicates the first n A group of images that are independently encoded and decoded. Each independently encoded and decoded image group starts with an I-frame, followed by P-frames and B-frames, for a total of M frames. Video with base description, divided into several independently encoded and decoded image groups. Input the CoCap network model; the CoCap network model includes an encoder and a decoder, and the processing procedure is as follows: During the encoding phase, each GOP's I-frames are converted into 768-dimensional feature vectors by an I-frame encoder, and then passed through 12 Transformer encoder layers to output contextual semantic features. The I-frame encoder includes a 16×16 convolutional kernel, a 16×16 stride, and a convolutional layer with 768 output channels. Each GOP's P-frames or B-frames are converted into 192-dimensional feature vectors by a motion encoder and then passed through 2 Transformer encoder layers to output motion vectors. The motion encoder includes an 8×8 convolutional kernel, an 8×8 stride, and a convolutional layer with 192 output channels. Simultaneously, each GOP's P-frames or B-frames are converted into 768-dimensional feature vectors by a residual encoder and then passed through 2 Transformer encoder layers to output residual features. The residual encoder includes a 64×64 convolutional kernel, a 64×64 stride, and a convolutional layer with 768 output channels. After the motion features and residual features are processed by dropout, they are combined with the contextual semantic features of the I-frame to form a temporal feature sequence. The temporal feature sequence is then input into the action encoder. The action encoder first encodes the temporal feature sequence through position embedding and frame type embedding. Then, the encoded temporal feature sequence is sequentially passed through the self-attention module for temporal modeling and through the cross-attention module for feature fusion, and finally outputs a 512-dimensional fused feature vector. In the decoding stage, the decoder adopts a two-layer BERT structure, which includes word embedding, positional encoding, and type encoding. Finally, it generates video-level coarse-grained pseudo-tags through linear mapping. .

[0034] To prevent overfitting, the dropout rate was set to 0.2 on the motion and residual features of the CoCap network model and 0.1 on the BERT model.

[0035] The loss function of the CoCap network model is expressed as follows: (1) In the formula, Represents the loss function. Indicates the length of the predicted word sequence. Indicates the current number t The prediction results for each word Indicates the preceding The true sequence of words, Indicates before The true sequence of words and videos with base descriptions Under the conditions, the first t The word prediction is as follows The probability of.

[0036] Example 4 Based on Example 3, the process of generating high-confidence visual tuples in step 2 is as follows: Step 2.1: Transfer the video with base description Sampling is performed according to frame rate T as a sequence. Each video frame in the sequence The high-confidence visual tuple generation branch, namely the VIS-MiniCPM model, is input sequentially. First, it passes through the MiniCPM-V-2 model to generate the initial text description for each video frame, as shown below: (2) In the formula, This indicates a fine-grained visual description. This refers to the text description of the nth video frame generated by MiniCPM-V-2. This represents the model parameters of MiniCPM-V-2; Step 2.2: Utilize the LLaMA large language model in the visual event information selector for fine-grained visual description. In summary, a general semantic description of the video segment is generated, as follows: (3) In the formula, A general semantic description of a video segment. Represents the LLaMA large language model. These are the model parameters for LLaMA; Subsequently, fine-grained visual descriptions at the video level will be used. and a general semantic description of the corresponding video clip Merge to obtain a set of sentences , means as follows: (4) In the formula, Indicates will and connect; Step 2.3: Set the sentences The input visual tuple generation module obtains a weighted score set of words of each category in the sentence set based on the TextRank algorithm and word frequency statistics algorithm. It then selects the words with the highest scores from each category to form a high-confidence visual tuple.

[0037] Step 2.3.1: Use spaCy to process the sentence set Dependency parsing yields all subject, object, action, and context words in the sentence. Each category forms a candidate word set, represented as follows: , , , ; Step 2.3.2: Based on the word frequency algorithm, calculate the candidate words in each candidate word set. In the sentence collection The frequency of occurrence in the text is represented as follows: (5) In the formula, This represents the i-th word in the text. ; Represents the frequency function, if but =1, otherwise 0 Represents a set of sentences The words in m Indicates the total number of words in the text; Meanwhile, the TextRank scores of the candidates in each candidate word set are calculated based on TextRank, and are represented as follows: (6) In the formula, Indicate candidate words TextRank score, The damping coefficient is... Indicates with candidate words In the sentence collection A set of adjacent candidate words with semantic relationships. Indicate candidate words and Edge weights between them; Used to determine in the first Candidate words in a sliding window and Do they appear simultaneously? If in the first... Simultaneous appearance of multiple windows and The value is 1 if the value is 1, otherwise the value is 0. Step 2.3.3: Calculate the score of all words in each candidate word set, as shown below: (7) In the formula, , Indicates the weighting coefficient. This embodiment is set up =0.5; Step 2.3.4: Extract high-confidence visual tuples from each candidate word set. ,in, , , , These represent the subject, object, action, and environment, respectively, as follows: (8) In the formula, This indicates selecting the word from the candidate word set that maximizes the corresponding score function value.

[0038] Example 5 Based on Example 4, the LLaMA large language model includes a byte pair encoding model and a decoder. The decoder consists of 32 Transformer modules, each containing a multi-head self-attention network and a feedforward neural network. The input of each sub-layer is normalized using the RMSNorm function. This example uses an LLaMA large language model with 8B parameters.

[0039] Step 3 specifically includes the following sub-steps: Step 3.1: The byte-pair encoding model will input... The input matrix X is obtained by decomposing the words into sub-word units and mapping them to the corresponding tokens in the predefined vocabulary. The tokens are then converted into word vectors. Step 3.2: Enhance the input matrix X input decoder: During the enhancement process, the multi-head self-attention network in each Transformer module introduces positional information by encoding the query and key through rotation. The multi-head self-attention network also integrates fine-grained visual cues provided by high-confidence visual tuples into the semantic representation of coarse-grained pseudo-labels. After processing by each Transformer module, the output features are passed to the next Transformer module for further feature transformation, until the final Transformer module outputs a feature vector containing both coarse-grained pseudo-labels and fine-grained visual cues. , The size is [1, 4096], and the calculation process is as follows: (9) (10) In the formula, Indicates the first Semantic features extracted by an attention head in a subspace For multi-head self-attention functions, This represents the rotational position encoding applied to the query and key matrices. , and For learnable query, key, and value parameter matrices, Represents matrix multiplication; This indicates that the features from multiple attention heads are concatenated. To output the transformation matrix, This represents a feedforward neural network. This indicates a normalization operation. This represents the output of the 32nd layer Transformer module; Feature vectors with coarse-grained pseudo-labels and fine-grained visual cues from the video Vocabulary embedding matrix of the LLaMA large language model Perform a product operation, then normalize using the softmax function to obtain the th... Pace prediction words probability distribution , means as follows: (11); Step 3.3: Based on the greedy sampling strategy, the first... Predicting words for steps From probability distribution The word with the highest probability is selected, as shown below: (12) By iteratively performing single-step predictions, the final enhanced pseudo-label text sequence is generated. , means as follows: (13) The output probability of the enhanced pseudo-labeled text sequence is represented as follows: (14) In the formula, T represents the length of the generated pseudo-label text sequence.

[0040] Example 6 Based on Example 5, step 4 involves combining the video data set with the base description. With the enhanced pseudo-label text sequence video data set Merging them into a training dataset, represented as Repeat steps 1 through 3, using the following overall loss function for iterative self-training: (15) In the formula, Represents the overall loss function. This indicates the number of samples in a dataset with a basis description. This indicates the number of samples in the dataset without a basis description. This indicates that the data in the dataset has a basis description. One sample, Indicates the true label, This indicates the first term in the dataset that is described without basis. d One sample, Represents samples in an unlabeled dataset Corresponding to the semantically enhanced pseudo-tags, Generally refers to the parameters of the semi-supervised method proposed in this invention. These are hyperparameters used to balance the weights of labeled data loss and pseudo-labeled data loss in the total loss function. This represents the number of self-training iterations.

[0041] The batch size of the semi-supervised video description generation method proposed in this embodiment is set to 2. The number of training iterations is set to 20, and the learning rate is set to... Finally, to avoid overfitting, the regularization rate of the Dropout algorithm was set to 0.2, and the BertAdam optimizer was used for training optimization.

[0042] Simulation Experiment Experiments were conducted using two public datasets, MSVD and MSR-VTT. MSVD, sourced from YouTube, contains 1970 open-domain video clips, each with an average of 40 annotations, divided into 1200 training videos, 100 validation videos, and 670 test videos. MSR-VTT contains approximately 10,000 video clips, each with an average of 20 annotations, divided into 6513 training videos, 497 validation videos, and 2990 test videos. To evaluate the performance of this invention on datasets of different sizes, subsets of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, and 90% were randomly selected from the MSVD and MSR-VTT datasets as labeled data, with the remaining unselected data treated as unlabeled data.

[0043] To verify the performance of this invention, it adopts a set of standard evaluation metrics widely recognized in the field of video description generation, specifically including BLEU@4, METEOR, ROUGE-L, CIDEr, and the overall performance metric OVERALL.

[0044] The calculation formula for the OVERALL index is shown in (16): (16) In the formula, B-4 represents BLEU@4, C represents CIDEr, M represents METEOR, and R represents ROUGE-L. This represents the maximum value of a certain indicator. If the model's results for each indicator approach the maximum value of all model indicators, its overall score approaches 100. An overall score of 100 is achieved if and only if a model has state-of-the-art performance across all indicators.

[0045] Furthermore, this experiment was programmed using Python 3.8 and trained on a Linux environment using the PyTorch 2.0.0 framework. The GPU used in the experiment was an RTX 4090, with 100GB of memory and 100GB of hard disk space. CUDA 11.8 and cuDNN 8.6 were used for acceleration.

[0046] Experimental results: This invention uses 48,000 and 130,260 data points respectively in the MSVD and MSR-VTT training sets as training data. To find the optimal model performance, the learning rate is set to [value missing] during the training phase. , and Experiments were conducted using 80% of the MSVD and MSR-VTT datasets (80% of which have basis descriptions and the remaining 20% ​​do not) as examples. The experimental results are as follows: Figure 2 and Figure 3 As shown, when the learning rate is At this time, the loss of this invention is minimized on both datasets. This is because this invention employs a dual-path collaborative framework to achieve semantic representation of video content at different granularities. At this learning rate, the loss decreases steadily while generating higher quality text. Specifically, as... Figure 2 As shown, in 80% of the MSVD dataset, the model begins to converge smoothly after 20,000 iterations; as Figure 3 As shown, in 80% of the MSR-VTT dataset, the model begins to converge smoothly after 50,000 iterations.

[0047] Table 1 shows the performance comparison of the semi-supervised video description generation method and the comparative method guided by the large language model of the present invention on the MSVD and MSR-VTT datasets and on different fully supervised video description models.

[0048] Table 1

[0049] Comparison of data in Table 1 reveals that, under fully supervised conditions, this invention achieves the best performance in BLEU@4, METEOR, and CIDEr metrics on the MSVD dataset, while the ROUGE-L metric is slightly lower than CoCap(ViT / B16). On the MSR-VTT dataset, this invention achieves the best performance in METEOR and ROUGE-L metrics, while BLEU@4 is slightly lower than the performance of this invention on 80% labeled data, and CIDEr is slightly lower than CoCap(ViT / B16). Although the semi-supervised method proposed in this invention is slightly lower than CoCap(ViT / B16) on some metrics, it still surpasses many fully supervised methods, such as: S2VT, SAAT, TTA, CSA-SR, GRU-EVE, RecNet, aLSTM, TDDF, PickNet, LEAD, VC-STG, and ASGNet.

[0050] Given the current lack of a comprehensive metric for evaluating text description quality, existing evaluation metrics (such as BLEU@4, METEOR, CIDEr, and ROUGE-L) focus on measuring specific aspects of text description. Therefore, this invention uses the overall performance metric OVERALL to balance commonly used metrics for comprehensive evaluation. The comparative experiments in Table 1 show that this invention achieves the best overall performance metric OVERALL on the MSVD and MSR-VTT datasets compared to fully supervised methods such as S2VT, SAAT, TTA, CSA-SR, GRU-EVE, RecNet, aLSTM, TDDF, PickNet, LEAD, VC-STG, CoCap(ViT / B16), and ASGNet.

[0051] Although this invention does not achieve optimal performance across all metrics in fully supervised scenarios, as a semi-supervised learning model, it can still supplement fine-grained visual information in videos through a dual-path collaborative framework and generate semantically consistent video descriptions using enhancement mechanisms. This allows the model to achieve performance similar to or even better than fully supervised methods under semi-supervised conditions. For example, as shown in Table 1, this invention achieved a CIDEr score of 108.5 on the MSVD dataset using only 80% of the labeled data, surpassing many fully supervised methods. Furthermore, even when trained with only 80% or even 50% of the labeled data, this invention still outperforms many fully supervised methods in overall performance metrics. This indicates that even with limited base descriptions, this invention can still generate high-quality video descriptions and effectively compensate for the performance bottleneck caused by insufficient base descriptions, alleviating the high cost of manually annotating base descriptions.

[0052] Table 2 shows the performance comparison of the semi-supervised video description generation method with large language model-guided pseudo-label enhancement of the present invention and the comparison method on different semi-supervised video description models on the MSVD and MSR-VTT datasets.

[0053] Table 2

[0054] The comparative experiments in Table 2 show that, on the MSVD and MSR-VTT datasets, this invention and the semi-supervised method LCABM exhibit better performance on METEOR, ROUGE-L, CIDEr, and the overall score metric OVERALL under different base description annotation ratios (30%, 50%, and 80%). Compared to LCABM, this invention adopts a dual-path collaborative framework, which can fully integrate the rich visual information in video frames when generating coarse-grained descriptions, and further optimize the visual descriptions through enhancement modules. At the same time, the visual selector effectively reduces redundant vocabulary in the fine-grained descriptions at the video frame level, enhances the ability to capture video details, and makes the generated descriptions more accurately reflect the fine-grained information in the video.

[0055] It should be noted that although the BLEU@4 index of this invention is slightly lower than that of LCABM, this is because BLEU@4 is mainly based on n -gram matching focuses on precise overlap between the description and the reference text. However, this invention incorporates a large language model in both the dual-path collaborative framework and the enhancement process. The resulting text is more lexically flexible and includes synonym substitutions, leading to more precise matching with the reference description. n The decrease in -gram matching does not affect its ability to accurately capture semantic information in videos.

[0056] Finally, the proposed invention was used to generate text descriptions on the MSVD and MSR-VTT test sets. For example... Figure 4 and Figure 5 Experimental results show that this method can fully integrate rich visual information in video frames when generating coarse-grained descriptions, effectively alleviating the problem of semantic deviation in video descriptions.

[0057] like Figure 4 and Figure 5 As shown, to verify the impact of the dual-path collaborative framework and the visual selector on model performance, this invention conducts experimental evaluations under an 80% basis description ratio. The experiments include single-path learning, a dual-path collaborative learning strategy using the MiniCPM-V-2 model, and a dual-path collaborative learning strategy using the VIS-MiniCPM model, respectively, to highlight the impact of the dual-path collaborative framework and the VIS-MiniCPM model on model performance. Figure 4As can be seen, when the single-path learning mechanism is selected, the output descriptions suffer from semantic bias, such as generating "riding" and "horse". When the dual-path learning mechanism using the MiniCPM-V-2 model is selected, the model can basically output complete video descriptions. However, when the dual-path collaborative learning strategy of the MiniCPM-V-2 model is selected, redundant words in the video frame descriptions introduce semantic noise into the model, resulting in semantic bias in the generated sentences, such as obtaining descriptions like "in a park" that do not conform to the video content. On the other hand, when the dual-path collaborative learning strategy of the VIS-MiniCPM model is selected, it can help the model generate more accurate descriptions, such as generating "dogs" and "in a field".

[0058] Depend on Figure 5 It is evident that using a single-path learning mechanism cannot improve the quality of generated sentences. For example, the model may predict incorrect descriptive sentences, such as "A man is riding on the water." When a dual-path collaborative learning strategy using the MiniCPM-V-2 model is selected, semantic noise is introduced, leading to a decrease in the quality of generated descriptions, such as generating "watching" and "child." When a dual-path collaborative learning strategy using the VIS-MiniCPM model is selected, the accuracy of the generated descriptive sentences is significantly improved. Ultimately, by Figure 4 and Figure 5 It is evident that when the learning strategy of using the VIS-MiniCPM model in a dual-path collaborative manner is selected, this invention can alleviate the semantic bias problem in video description.

Claims

1. A semi-supervised video description generation method guided by a large language model and augmented with pseudo-labels, characterized in that, Specifically, the steps include the following: Step 1: Construct a dual-path collaborative framework, which includes a video-level coarse-grained pseudo-label generation branch and a high-confidence visual tuple generation branch; Step 2: After preprocessing the video with the base description, input it into the dual-path collaborative framework to generate video-level coarse-grained pseudo-labels and high-confidence visual tuples; Step 3: Input the high-confidence visual tuples as fine-grained visual cues, along with the video-level coarse-grained pseudo-labels, into the LLaMA large language model. Under the guidance of the LLaMA large language model, the video-level coarse-grained pseudo-labels are enhanced. Step 4: Optimize and generate the final video description through iterative self-training.

2. The semi-supervised video description generation method with pseudo-label enhancement guided by a large language model according to claim 1, characterized in that, The video-level coarse-grained pseudo-tag generation branch in step 1 includes the CoCap network model, which is used to process a video with a base description that is divided into several independent encoding and decoding image groups into video-level coarse-grained pseudo-tags. The high-confidence visual tuple generation branch is the VIS-MiniCPM model. The VIS-MiniCPM model adds a visual event information selector to the output of the MiniCPM-V-2 model. The visual event information selector includes an LLaMA large language model and a visual tuple generation module. The MiniCPM-V-2 model is used to process video divided into continuous frames with base descriptions into fine-grained visual descriptions at the video frame level. The LLaMA large language model of the visual event information selector is used to process the fine-grained visual descriptions at the video frame level into a set of sentences. The visual tuple generation module obtains a weighted score set of words of each category in the sentence set based on the TextRank algorithm and word frequency statistics algorithm, and selects the words with the highest scores from each category to form a high-confidence visual tuple.

3. The semi-supervised video description generation method with large language model-guided pseudo-label enhancement according to claim 2, characterized in that, In step 2, the process of generating the video-level coarse-grained pseudo-tags is as follows: Video with base description The consecutive frames are divided into several independently encoded and decoded image groups, as shown below: In the formula, Indicates the first n A group of images that are independently encoded and decoded. Each independently encoded and decoded image group starts with an I-frame, followed by P-frames and B-frames, for a total of M frames. Video with base description, divided into several independently encoded and decoded image groups. Input a CoCap network model; the CoCap network model includes an encoder and a decoder, and the processing procedure is as follows: During the encoding phase, each GOP's I-frame is converted into a 768-dimensional feature vector by an I-frame encoder, and then passed through 12 Transformer encoder layers to output contextual semantic features. The I-frame encoder includes a 16×16 convolutional kernel, a 16×16 stride, and a convolutional layer with 768 output channels. Each GOP's P-frame or B-frame is converted into a 192-dimensional feature vector by a motion encoder and then passed through 2 Transformer encoder layers to output motion vectors. The motion encoder includes an 8×8 convolutional kernel, an 8×8 stride, and a convolutional layer with 192 output channels. Simultaneously, each GOP's P-frame or B-frame is converted into a 768-dimensional feature vector by a residual encoder and then passed through 2 Transformer encoder layers to output residual features. The residual encoder includes a 64×64 convolutional kernel, a 64×64 stride, and a convolutional layer with 768 output channels. After the motion features and residual features are processed by dropout, they are combined with the contextual semantic features of the I-frame to form a temporal feature sequence. The temporal feature sequence is then input into the action encoder. The action encoder first encodes the temporal feature sequence through position embedding and frame type embedding. Then, the encoded temporal feature sequence is sequentially passed through the self-attention module for temporal modeling and through the cross-attention module for feature fusion, and finally outputs a 512-dimensional fused feature vector. During the decoding stage, the decoder employs a two-layer BERT structure, including word embedding, positional encoding, and type encoding, and ultimately generates video-level coarse-grained pseudo-tags through linear mapping. .

4. The semi-supervised video description generation method with pseudo-label enhancement guided by a large language model according to claim 3, characterized in that, The loss function of the CoCap network model is expressed as follows: (1) In the formula, Represents the loss function. Indicates the length of the predicted word sequence. Indicates the current number t The prediction results for each word Indicates the preceding The true sequence of words, Indicates before The true sequence of words and videos with base descriptions Under the conditions, the first t The word prediction is as follows The probability of.

5. The semi-supervised video description generation method with pseudo-label enhancement guided by a large language model according to claim 2, characterized in that, In step 2, the process of generating the high-confidence visual tuple is as follows: Step 2.1: Transfer the video with base description Sampling is performed according to frame rate T as a sequence. Each video frame in the sequence The high-confidence visual tuple generation branch, namely the VIS-MiniCPM model, is input sequentially. First, it passes through the MiniCPM-V-2 model to generate the initial text description for each video frame, as shown below: (2) In the formula, This indicates a fine-grained visual description. This refers to the text description of the nth video frame generated by MiniCPM-V-2. This represents the model parameters of MiniCPM-V-2; Step 2.2: Utilize the LLaMA large language model in the visual event information selector for fine-grained visual description. In summary, a general semantic description of the video segment is generated, as follows: (3) In the formula, A general semantic description of a video segment. Represents the LLaMA large language model. These are the model parameters for LLaMA; Subsequently, fine-grained visual descriptions at the video level will be used. and a general semantic description of the corresponding video clip Merge to obtain a set of sentences , means as follows: (4) In the formula, Indicates will and connect; Step 2.3: Set the sentences The input visual tuple generation module obtains a weighted score set of words of each category in the sentence set based on the TextRank algorithm and word frequency statistics algorithm. It then selects the words with the highest scores from each category to form a high-confidence visual tuple.

6. The semi-supervised video description generation method with pseudo-label enhancement guided by a large language model according to claim 5, characterized in that, Step 2.3 specifically includes the following sub-steps: Step 2.3.1: Use spaCy to process the sentence set Dependency parsing yields all subject, object, action, and context words in the sentence. Each category forms a candidate word set, represented as follows: , , , ; Step 2.3.2: Based on the word frequency algorithm, calculate the candidate words in each candidate word set. In the sentence collection The frequency of occurrence in the text is represented as follows: (5) In the formula, Indicates the first in the text i One word, ; Represents the frequency function, if but =1, otherwise 0 Represents a set of sentences The words in m Indicates the total number of words in the text; Meanwhile, the TextRank scores of the candidates in each candidate word set are calculated based on TextRank, and are represented as follows: (6) In the formula, Indicate candidate words TextRank score, The damping coefficient is... Indicates with candidate words In the sentence collection A set of adjacent candidate words with semantic relationships. Indicate candidate words and Edge weights between them; Used to determine in the first Candidate words in a sliding window and Do they appear simultaneously? If in the first... Simultaneous appearance of multiple windows and The value is 1 if the value is 1, otherwise the value is 0. Step 2.3.3: Calculate the score of all words in each candidate word set, as shown below: (7) In the formula, , Indicates the weighting coefficient. ; Step 2.3.4: Extract high-confidence visual tuples from each candidate word set. ,in, , , , These represent the subject, object, action, and environment, respectively, as follows: (8) In the formula, This indicates selecting the word from the candidate word set that maximizes the corresponding score function value.

7. The semi-supervised video description generation method with pseudo-label enhancement guided by a large language model according to claim 1, characterized in that, The LLaMA large language model includes a byte pair encoding model and a decoder. The decoder consists of 32 Transformer modules. Each Transformer module contains a multi-head self-attention network and a feedforward neural network. The input of each sub-layer is normalized using the RMSNorm function. Step 3 specifically includes the following sub-steps: Step 3.1: The byte-pair encoding model will input... It is broken down into sub-word units, and then the sub-word units are mapped to the corresponding tokens in the preset vocabulary; The tokens are then converted into word vectors to obtain the input matrix X; Step 3.2: Enhance the input matrix X input decoder: During the enhancement process, the multi-head self-attention network in each Transformer module introduces positional information by encoding the rotation position as query and key; A multi-head self-attention network is used to integrate fine-grained visual cues provided by high-confidence visual tuples into the semantic representation of coarse-grained pseudo-labels. After each Transformer module completes its processing, the output features are passed to the next Transformer module for further feature transformation, until the final Transformer module outputs a feature vector containing both coarse-grained pseudo-labels and fine-grained visual cues. , The size is [1, 4096], and the calculation process is as follows: (9) (10) In the formula, Indicates the first Semantic features extracted by an attention head in a subspace For multi-head self-attention functions, This represents the rotational position encoding applied to the query and key matrices. , and For learnable query, key, and value parameter matrices, Represents matrix multiplication; This indicates that the features from multiple attention heads are concatenated. To output the transformation matrix, This represents a feedforward neural network. This indicates a normalization operation. This represents the output of the 32nd layer Transformer module; Feature vectors with coarse-grained pseudo-labels and fine-grained visual cues from the video Vocabulary embedding matrix of the LLaMA large language model Perform a product operation, then normalize using the softmax function to obtain the th... Pace prediction words probability distribution , means as follows: (11); Step 3.3: Based on the greedy sampling strategy, the first... Predicting words for steps From probability distribution The word with the highest probability is selected, as shown below: (12) By iteratively performing single-step predictions, the final enhanced pseudo-label text sequence is generated. , means as follows: (13) The output probability of the enhanced pseudo-labeled text sequence is represented as follows: (14) In the formula, T represents the length of the generated pseudo-label text sequence.

8. The semi-supervised video description generation method with pseudo-label enhancement guided by a large language model according to claim 7, characterized in that, Step 4, combine the video data set with the base description. With the enhanced pseudo-label text sequence video data set Merging them into a training dataset, represented as Repeat steps 1 through 3, using the following overall loss function for iterative self-training: (15) In the formula, Represents the overall loss function. This indicates the number of samples in a dataset with a basis description. This indicates the number of samples in the dataset that describes the data without a basis. This indicates that the data in the dataset has a basis description. One sample, Indicates the true label, This indicates the first term in the dataset that is described without basis. d One sample, Represents samples in an unlabeled dataset Corresponding to the semantically enhanced pseudo-tags, Generally refers to the parameters of the semi-supervised method proposed in this invention. These are hyperparameters used to balance the weights of labeled data loss and pseudo-labeled data loss in the total loss function. This represents the number of self-training iterations.

Citation Information

Patent Citations

  • Multi-modal hierarchical video description generation method and system

    CN117596453A

  • Video description method based on visual context sparse regularization and implicit attention

    CN118397509A

  • Generating Optimized Datasets for Training Machine-Learned Models for Subjective Image Understanding

    US20250201006A1