Video description generation method based on synthetic text and global mixed linear expert

By using a global hybrid linear expert model and an adaptive data generation process, the shortcomings of existing video description generation methods in terms of model structure and data quality are addressed, enabling high-quality and diverse video description generation and improving the model's cross-scene adaptability and description accuracy.

CN121033722APending Publication Date: 2025-11-28ZHEJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511032714.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing video description generation methods lack cross-scene generalization and content specialization capabilities in their model structure, and the quality of training data is insufficient, resulting in inaccurate or inappropriate generated descriptions that are difficult to cover diverse scenarios.

Method used

A global hybrid linear expert model is adopted, replacing all linear layers in the video description generation model with hybrid linear expert layers. Combined with an adaptive structured video description text generation process, experts are dynamically selected to participate in the calculation through a lightweight routing network, and dynamic bias terms are introduced for training to generate high-quality video description text.

Benefits of technology

It achieves high-quality and diverse video description text generation, improves the model's ability to express complex and diverse video content, and has stronger video understanding and description generation capabilities. The generated descriptions are accurate, natural, and context-appropriate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033722A_ABST
    Figure CN121033722A_ABST
Patent Text Reader

Abstract

The invention discloses a video description generation method based on a synthetic text and a global mixed linear expert, and aims to realize high-quality video content automatic description through an innovative model architecture and a data construction strategy. The method is suitable for video auxiliary subtitle systems, education video content abstracts or video content platform intelligent analysis scenes, and high-quality video description texts can be automatically generated. According to the method, through framework innovation, a first whole-network MoE-based Transform framework is realized on a model level, so that a global mixed linear expert model is constructed, and a powerful model basis is constructed for a video description task. According to the model, multi-modal information and diversified content features can be efficiently fused, and good support and expansion capability is provided for subsequent training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of video-text multimodal, and particularly relates to a video description generation method based on synthetic text and global mixed linear experts. BACKGROUND

[0002] With the development of multimedia technology and the surge of online video content, how to let the machine automatically understand and describe the video content has become a research hotspot in the field of multi-modal artificial intelligence. For example, in the application of film and television auxiliary subtitles, educational video abstracts, etc., it is necessary to convert the visual and audio information in the video into accurate text description. This task is usually referred to as video captioning, and the goal is to let the model automatically generate natural language descriptions consistent with the video content. In recent years, video large language models (Video-LLM) have combined video understanding capabilities with the powerful generation capabilities of LLM by using a large amount of video-text paired data for large-scale training, bringing rich knowledge and reasoning capabilities to video description, making the generated description more natural and informative. Although existing video description methods have made significant progress, there are still two bottlenecks: model structure capacity and training data quality.

[0003] Firstly, in terms of model structure, video content is diverse, and a model with a fixed number of activation parameters often has difficulty in fully expressing fine-grained information in different types of videos. Increasing the model size can improve performance, but the computational overhead increases, increasing the training difficulty and inference cost. As a parameter-efficient architecture, Mixture-of-Experts (MoE) selects part of the sub-network for each input token and only activates the selected experts to participate in the calculation, thereby realizing sparse activation. This mechanism allows the total parameter size to be large, but only a portion of the experts is designed for each inference, which can improve training and inference efficiency while maintaining performance. Existing MoE methods mostly integrate experts in the feed-forward network (FFN), such as the Switch Transformer, which replaces the FFN with multiple experts. However, this type of FFN-MoE scheme still has limitations: other modules of the model (such as the linear layer in the attention mechanism) still use dense computation, which cannot fully utilize the advantages of MoE; at the same time, routing easily leads to uneven load of experts, and auxiliary losses need to be introduced to balance the usage rate, increasing the training difficulty and even interfering with the gradient.

[0004] Secondly, in terms of training data, high-quality video description annotations are very scarce. Existing public datasets are limited in size and cover only a single type of content, which results in insufficient description capabilities of the model when it faces new domains or complex scenarios. Manually writing subtitles or descriptions for large-scale videos is not only time-consuming and labor-intensive but also difficult to scale. To alleviate the lack of data, researchers have attempted to use synthetic data to assist training, such as using video metadata, commentary, existing subtitles, or generating pseudo-label text based on video content with the help of pre-trained language models. However, the quality of these data varies greatly, may not be consistent with the video content, and even contains errors and biases. If directly used for training, it can weaken the model's understanding of video semantics, leading to inaccurate or inappropriate descriptions. In addition, different types of videos (such as sports events and interviews) have significant differences in language expression. Existing synthetic methods mostly use unified strategies, lack of class-awareness, and are difficult to cover the key features of various videos. There is a lack of data generation mechanism that is adaptive to video categories, which limits the model's expression ability for diverse scenarios.

[0005] In summary, the existing technology has significant deficiencies in both model structure and data acquisition: on the one hand, the current model architecture lacks a mechanism that simultaneously possesses cross-scene generalization ability and content specialization ability; on the other hand, an automatic, efficient, and high-quality multi-modal pseudo-annotation generation process has not been established. To solve the above problems, it is urgent to propose a video description generation method that combines structural innovation and data strategy innovation to improve the model's expression ability for complex and diverse video content. SUMMARY

[0006] The purpose of the present application is to solve the problems existing in the prior art and provide a video description generation method based on synthetic text and global hybrid linear experts, aiming to achieve automatic description of high-quality video content through innovative model architecture and data construction strategy.

[0007] To achieve the above invention purposes, the present application specifically adopts the following technical solutions:

[0008] In a first aspect, the present application provides a video description generation method based on synthetic text and global hybrid linear experts, comprising the following steps:

[0009] S1: Constructing an adaptive structured video description text generation process, classifying the content type of the obtained original video, and generating high-quality video description text synthesis data based on the type label of the original video content through style adaptive generation and adversarial consistency verification;

[0010] S2: uniformly replace all linear layers in the video description generation model with mixed linear expert layers to form a global mixed linear expert model; wherein each mixed linear expert layer comprises a shared expert, a parallel expert group, and a lightweight routing network;

[0011] S3: use the video description text synthesis data generated in S1 and the corresponding original video as a training data, and form a training data set, and iteratively train the global mixed linear expert model constructed in S2 on the training data set, dynamically select a preset number of experts to participate in calculation through the lightweight routing network, realize sparse activation training, and update the dynamic bias term introduced based on the difference in expert usage frequency at each model parameter update, and finally obtain the optimized global mixed linear expert model parameters;

[0012] S4: input the target video to be described into the trained global mixed linear expert model, extract the spatio-temporal feature representation of the target video through the video encoder, and calculate the expert score through the lightweight routing network of each mixed linear expert layer according to the video feature token input into the current mixed linear expert layer, dynamically select the expert combination most suitable for processing the token, and finally generate a description text consistent with the content of the target video.

[0013] On the basis of the above scheme, each step can be implemented in the following preferred specific manner.

[0014] As a preferred embodiment of the first aspect described above, the specific process of step S1 is as follows:

[0015] S11: extract key frames representing the overall content of the original video from the original video, input the extracted key frames into the InternVideo video base model to extract the spatio-temporal feature representation of the original video as key frame features, output a confidence distribution of predefined categories, and select the category with the highest confidence as the type label of the original video content;

[0016] S12: preset corresponding language templates and style criteria for each type of original video, fuse the theme, main action events and scene background, and dialogue of the original video into key information, fill the key information into the language template to form a prompt word, and input the prompt word into GPT-4o mini to generate a description text conforming to the style criteria;

[0017] S13: use the pre-trained image-text matching model SigLIP2 to calculate the similarity between the key frame features and the generated description text, and use Qwen2.5-VL to determine whether the content of the original video is consistent with the generated description text, filter out the description text with a similarity lower than a preset similarity threshold and the description text inconsistent with the content of the original video, and finally obtain the description text as the video description text synthesis data.

[0018] As a preferred embodiment of the first aspect, for the original video containing multiple action scenes, the original video is detected and segmented by scene using PySceneDetect, the key behavior events in the original video are detected and labeled using the Qwen2.5-VL model, the corresponding action categories are identified and the corresponding time stamps are recorded, and the scene environment and important object information in the original video are extracted using the visual detection model RAM++, and finally the obtained scene, key behavior events, scene environment and important object information are added to the key information.

[0019] As a preferred embodiment of the first aspect, for the original video containing multiple language contents, first, the audio track of the original video is automatically speech recognized using the Whisper model of OpenAI, and the speech or narration therein is converted into text, then the layout analysis and text detection are performed by PaddleOCR to locate the text region and identify the text content, the explicit text information in the original video is obtained and added to the key information.

[0020] As a preferred embodiment of the first aspect, the linear layer in the video description generation model does not include a video encoder and a text token classification head, and includes Q, K, V projection layers in the attention mechanism and an output layer, and dimension lifting projection layers and dimension reduction projection layers in the feedforward neural network.

[0021] As a preferred embodiment of the first aspect, in the hybrid linear expert layer, the weights of the video description generation model linear layer are used as shared expert weights for general basic transformation of all input token features, and the parallel expert group consists of N experts, each of which performs special transformation for a specific input mode; the lightweight routing network itself is a lightweight linear layer, which outputs the selection scores of each expert according to the input token features of the current layer, and the top k experts with the highest selection scores are selected to participate in the calculation of the token features, and the unselected experts do not participate in the calculation.

[0022] As a preferred embodiment of the first aspect, in step S3, the global hybrid linear expert model has the following training process in each round:

[0023] S31: Select data batches from all training data using random sampling or stratified sampling strategy to ensure that each batch contains video description text synthesis data of different types and styles;

[0024] S32: input a batch of original videos and corresponding video description text synthesis data into the global mixed linear expert model, encode the input original video through the video encoder to convert the original video into a token sequence, and initialize the expert use frequency statistics value of each layer lightweight routing network, wherein each token in the token sequence represents the local spatiotemporal feature of the original video;

[0025] S33: in the forward propagation process, for each mixed linear expert layer in the global mixed linear expert model, the lightweight routing network calculates the original score of each expert according to the input token sequence x, and then adds the original score to the dynamic bias term to obtain the final score of the expert;

[0026] The final score of the i-th expert is:

[0027]

[0028] Wherein, represents the transpose of the i-th expert routing parameter; simoid(·) represents the sigmoid function; b i represents the dynamic bias term of the i-th expert;

[0029] S34: select the top k experts with the final score, normalize the final score of the selected experts, and calculate the normalized score;

[0030] S35: multiply the normalized score of the selected experts with the respective weight matrix to obtain the final weight of the selected experts, add the final weight of all selected experts to the shared expert weight as the fusion weight, and multiply the fusion weight with the input token sequence as the weighted output of the mixed linear expert layer;

[0031] S36: after processing by the global mixed linear expert model, the video description text synthesis data is used as the true label, and the cross-entropy loss function is used to calculate the difference between the predicted text and the true label of the global mixed linear expert model;

[0032] S37: calculate the gradient of the loss function on the parameters of the global mixed linear expert model, including the shared expert weight, the weight of the experts in the parallel expert group and the parameters of the lightweight routing network, and update all learnable parameters using the AdamW optimizer;

[0033] S38: after each round of training, the use frequency of each expert in the round is counted, and the average use frequency is calculated based on the use frequency of each expert According to the use frequency f i The dynamic bias term of the i-th expert is dynamically updated according to the difference direction of the average frequency:

[0034]

[0035] wherein sign denotes a sign function that returns -1 when , +1 when , and 0 when ; and λ is a hyperparameter that controls the adjustment magnitude.

[0036] In a second aspect, the present application provides a video description generation system based on synthetic text and global hybrid linear experts, comprising:

[0037] A data generation module is configured to build an adaptive structured video description text generation process, classify the content type of an obtained original video, and perform style adaptive generation and adversarial consistency verification based on the type label of the original video content, so as to generate high-quality video description text synthetic data.

[0038] A model acquisition module is configured to uniformly replace all linear layers in the video description generation model with hybrid linear expert layers to form a global hybrid linear expert model; wherein each hybrid linear expert layer comprises a shared expert, a parallel expert group, and a lightweight routing network.

[0039] A model optimization module is configured to use the video description text synthetic data generated by the data generation module and the corresponding original video as a training data, form a training data set, and iteratively train the global hybrid linear expert model on the training data set. The lightweight routing network dynamically selects a preset number of experts to participate in the calculation, realizes sparse activation training, and synchronously updates the dynamic bias term introduced based on the difference in expert usage frequency during each model parameter update, and finally obtains the optimized global hybrid linear expert model parameters.

[0040] A description text generation module is configured to input a target video to be described into the trained global hybrid linear expert model, extract the spatiotemporal feature representation of the target video through a video encoder, and calculate the expert score of the current hybrid linear expert layer input by the lightweight routing network of each hybrid linear expert layer. The lightweight routing network dynamically selects the most suitable expert combination for processing the token, and finally generates a description text consistent with the content of the target video.

[0041] In a third aspect, the present application provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the video description generation method based on synthetic text and global hybrid linear experts is realized as described in any of the above first aspect.

[0042] In a fourth aspect, the present application provides a computer electronic device, comprising a memory and a processor;

[0043] The memory is configured to store a computer program.

[0044] The processor is configured to implement the method for generating video description based on synthesized text and global mixed linear expert when executing the computer program.

[0045] Compared with the prior art, the present application has the following beneficial effects:

[0046] After multi-stage processing, the present application can generate large-scale high-quality video description text synthesis data with clear structure and various styles, thereby forming a data set. Thanks to the introduction of multi-modal information, compared with traditional methods, the method of the present application has significant improvements in content coverage, language quality and annotation credibility. Combined with the global mixed linear expert model, it adapts different types of data through the expert routing mechanism in training, and unifies the style expression through the shared structure, and the final model has stronger video understanding and description generation capability, and can generate accurate, natural and contextually consistent description text according to the content.

[0047] In addition, the present application has innovation and advantage in model architecture and data construction, breaks through the limitation of existing video description technology in model capacity and data quality, and has wide application value. In terms of data quality, the present application constructs an automatic and multi-modal fusion video description text synthesis data generation process, integrates video classification, OCR, ASR, action analysis, LLM generation and consistency checking, generates multi-type and high-consistency text with quality close to artificial annotation, and is significantly better than traditional methods. In terms of model architecture, the present application proposes the first all-network MoE Transformer architecture, i.e. global mixed linear expert model, which expands the expert mechanism to all linear transformation units, realizes dynamic allocation of parameters on demand, improves the expression and flexibility of the model, and adopts the structure combining shared experts and parallel experts, balances generalization and specialization, and improves the training stability. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 The step flowchart of the method of the present application is shown in the figure;

[0049] Figure 2 The architecture diagram of the global mixed linear expert model of the present application is shown in the figure;

[0050] Figure 3 The flowchart of generating video description text synthesis data of the present application is shown in the figure;

[0051] Figure 4 The system block diagram of the present application is shown in the figure. DETAILED DESCRIPTION

[0052] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.

[0053] like Figure 1 As shown, in a preferred embodiment of the present invention, the video description generation method based on synthesized text and global hybrid linear experts includes the following steps S1 to S4. The specific implementation process of each step will be described in detail below.

[0054] S1: Construct an adaptive structured video description text generation process, classify the acquired original video into content types, and perform style adaptive generation and adversarial consistency verification based on the type tags of the original video content, thereby generating high-quality video description text synthesis data.

[0055] It should be noted that the specific process of step S1 is as follows:

[0056] S11: Extract keyframes representing the overall content of the original video from the original video, input the extracted keyframes into the InternVideo video base model to extract the spatiotemporal feature representation of the original video and use it as keyframe features, output the confidence distribution of predefined categories, and select the category with the highest confidence as the type label of the original video content.

[0057] In this embodiment, the video content is categorized into predefined categories such as sports, news, lifestyle, education, entertainment, games, and assessments.

[0058] S12: Preset corresponding language templates and style guidelines for each type of original video, integrate the theme, main action events, scene background, and dialogue of the original video into key information, fill the key information into the language template to form cue words, and input the cue words into GPT-4o mini to generate descriptive text that conforms to the style guidelines.

[0059] S13: Calculate the similarity between the key frame features and the generated description text using the pre-trained image-text matching model SigLIP2, and use Qwen2.5-VL to determine whether the content of the original video matches the generated description text. Remove the description text with a similarity lower than the preset similarity threshold and the description text that does not match the content of the original video. The final description text is used as the video description text synthesis data.

[0060] In this embodiment, step S13 includes similarity detection and content matching detection. When the similarity is lower than the preset similarity threshold 0.3 or there is a content conflict, the corresponding description text is automatically removed.

[0061] It should be noted that, at the data level, the present application proposes an innovative video description data construction pipeline (i.e., an adaptive structured video description text generation process) for generating large-scale high-quality video description text synthesis data to support video description model training. The construction process adopts a multi-stage, step-by-step refinement strategy to systematically improve the weak links in existing methods, ensuring that the generated text is highly consistent with the video content and the language style is uniform. The pipeline has key features such as multi-modal collaborative processing, video type adaptation, and output structure uniformity and traceability, significantly improving the quality and applicability of the data.

[0062] As shown in Figure 3 , the adaptive structured video description text generation process mainly includes the following key stages:

[0063] Original video content type classification: First, key frames are uniformly extracted from the original video, which represent the overall content of the original video. Then, the InternVideo video base model is used to extract the spatio-temporal feature representation of the original video, output the confidence distribution of the predefined categories, and select the category with the highest confidence as the type label of the original video content, including the category name and hierarchical label (main category and subcategory). This output will provide prior knowledge for the subsequent stages, guiding the style-adaptive generation process to use appropriate tone or professional terminology.

[0064] Further, for videos with rich action scenes, the application deeply analyzes the dynamic content and static scene information of the video, extracts the actions and backgrounds in the original video. Specifically, for an original video containing multiple action scenes, PySceneDetect is used to detect and segment the scenes of the original video, Qwen2.5-VL model is used to detect and mark the key behavior events in the original video, identify the corresponding action categories and record the corresponding time stamps, and use the visual detection model RAM++ to extract the scene environment and important object information in the original video to enrich the description of the background elements. Finally, the obtained scene, key behavior event, scene environment and important object information are added to the key information. This process realizes the effective conversion of visual action information to text description, ensuring accurate content and style consistent with the video type.

[0065] Further, for videos containing a large amount of language content, the application extracts speech content and picture text through speech recognition and character recognition, extracts explicit text information in the video from both audio and visual aspects, and fuses to build a knowledge point outline. Specifically, for an original video containing multiple language contents, first use the Whisper model of OpenAI to perform automatic speech recognition (ASR) on the audio track of the original video, convert the speech or narration in it into text, then use the mature open source OCR engine PaddleOCR in the industry to perform layout analysis and text detection, locate the text area and recognize the text content, obtain the explicit text information in the original video and add it to the key information. This method effectively integrates multi-modal information to generate accurate and clear text that faithfully reflects the explanation content in the video.

[0066] Style adaptive generation: Each type of original video is equipped with a dedicated expression structure. First, sort out key information such as video category and theme, main action events and scene background, and dialogue, etc., then fuse these elements into prompt words, input them into GPT-4o mini, and the model ensures uniform style during generation, while adjusting details according to the context, making the output content natural and contextually appropriate.

[0067] Adversarial consistency check: To ensure the quality of the generated data, the application filters out descriptions that do not match the video content or are redundant. Specifically, a pre-trained image-text matching model SigLIP2 is introduced to calculate the similarity between the key frame features and the generated description text. If the similarity is low, it means that the generated description text may have missed the main theme of the video or has deviations. At the same time, Qwen2.5-VL is used to determine whether the video content and the description text are consistent and whether there are conflicts, further ensuring the correctness of the description through adversarial testing.

[0068] In addition to the above process, the application also designs meta information labeling and unified data structure. After generating each piece of text, relevant meta information is attached, and a unified data structure is constructed, including video ID, type label, time anchor point and confidence score, etc., to facilitate content tracing, data filtering and fine sampling during training.

[0069] S2: all linear layers in the video description generation model are uniformly replaced by mixed linear expert layers to form a global mixed linear expert model; wherein each mixed linear expert layer comprises a shared expert, a parallel expert group and a lightweight routing network.

[0070] It should be noted that in step S2, as shown in Figure 2 The application reconstructs the Transformer model from the architecture level, introduces a global mixed linear expert mechanism, and enables each layer and each submodule to have sparse activation expert routing capability. Specifically, all linear layers in the video description generation model, including the Q, K, V projection layers in the attention mechanism and the output layer, and the dimension lifting projection layer and the dimension reducing projection layer in the feedforward neural network, are uniformly replaced by mixed linear expert structures. Unlike the prior art which only introduces MoE in part of the FFN sublayer, the application realizes the overall specialization of linear computing units and constructs a "full-level MoE" Transformer architecture to form a global mixed linear expert model. Each mixed linear expert layer consists of three parts: a shared expert (Shared Expert), a parallel expert group (Expert Pool) and a lightweight routing network (Router). The weights of the linear layers of the video description generation model are used as shared expert weights for general basic transformation of all input token features; the parallel expert group consists of N experts, each of which performs special transformation for specific input patterns; the lightweight routing network itself is a lightweight linear layer that outputs the selection score or probability of each expert according to the input token features of the current layer, and the top-k experts with the highest selection score or probability are selected to participate in the calculation of the token features, and the unselected experts do not participate in the calculation, thereby realizing sparse activation.

[0071] In this embodiment, the top-k value of expert selection is set to 1, that is, one expert is dynamically selected from the N experts in the parallel expert group to work with the shared expert that is always activated, and the total number N of experts in the parallel expert group is set to 8.

[0072] It should be noted that the shared expert retained in each layer is always effective for all tokens to provide general basic transformation support, while the experts selected by the lightweight routing network are specially processed for specific tokens. This mechanism not only guarantees the cross-task generalization ability of the model, but also realizes personalized modeling of different input features through experts.

[0073] S3: Synthesize the video description text synthesis data generated in S1 and the corresponding original video as a training data, and form a training data set, iteratively train the global mixed linear expert model constructed in S2 on the training data set, dynamically select a preset number of experts to participate in calculation through a lightweight routing network, realize sparse activation training, and update the dynamic bias term introduced based on the difference in expert usage frequency at each model parameter update, and finally obtain the optimized global mixed linear expert model parameters.

[0074] It should be noted that in step S3, the global mixed linear expert model has the following training process each round:

[0075] S31: Data batch preparation. Random sampling or stratified sampling strategy is used to select data batches from all training data, ensuring that each batch contains video description text synthesis data of different types and styles.

[0076] S32: Forward propagation initialization. Input a batch of original videos and corresponding video description text synthesis data into the global mixed linear expert model, encode the input original video through the video encoder to convert the original video into a token sequence, each token in the token sequence represents the local spatio-temporal features of the original video, and initialize the expert usage frequency statistics of each layer of the lightweight routing network.

[0077] S33: Expert selection and routing. In the forward propagation process, for each mixed linear expert layer in the global mixed linear expert model, the lightweight routing network calculates the original score of each expert according to the input token sequence x, and then adds the original score to the dynamic bias term to obtain the final score of the expert.

[0078] In this embodiment, take the ith expert as an example, and its corresponding final score is:

[0079]

[0080] wherein, represents the transpose of the routing parameter of the ith expert; sigmoid(·) represents the sigmoid function; b i represents the dynamic bias term of the ith expert, which is used to realize load balancing without auxiliary loss.

[0081] S34: Expert score normalization. Select the top k experts with the final score, normalize the final score of the selected experts, and calculate the normalized score a i (x) as follows:

[0082]

[0083] where, denotes the expert set consisting of all selected experts; s j (x) denotes the final score of the jth expert.

[0084] In this embodiment, the step S34 ensures that the sum of the final scores of all selected experts is 1, preparing for the subsequent weighted fusion.

[0085] S35: Expert output fusion. Multiply the normalized scores of the selected experts with the respective weight matrices to obtain the final weights of the selected experts, and add the final weights of all selected experts and the shared expert weight W shared to obtain the fusion weight, and multiply the fusion weight with the input token sequence x as the weighted output y of the hybrid linear expert layer:

[0086]

[0087] where, W i is the weight matrix of the ith expert, realizing the organic combination of shared knowledge and special knowledge.

[0088] S36: Loss function calculation. After processing by the global hybrid linear expert model, the video description text synthesis data is taken as the real label, and the cross-entropy loss function in the following form is used to calculate the difference between the text predicted by the global hybrid linear expert model and the real label:

[0089]

[0090] where, z t is the real token at time z t in the real label; P(z t | t <t , x) is the conditional probability predicted by the model.

[0091] In this example, the feature transmission and step-by-step abstraction between multiple layers of the global hybrid linear expert model are performed according to the foregoing process, and each hybrid linear expert layer can select a different expert combination, thereby forming a hierarchical expert collaboration mode. After multiple layers of hybrid linear expert processing, the last layer of the network outputs a feature representation containing rich semantic information, which integrates the global information, local details and different expert knowledge of the video. Then the feature representation of the last layer of the network is input into the text token classification head, and the generation probability distribution of each token in the vocabulary is obtained through linear transformation and softmax activation function, providing a probability basis for sequence generation. The text token classification head here adopts a standard linear layer instead of a hybrid linear expert structure in order to maintain computational efficiency and parameter efficiency. Finally, the description text is generated word by word in a self-recurrent manner, and a sampling strategy is used for decoding to generate the description text.

[0092] S37: Backpropagation update. The gradient of the loss function with respect to the parameters of the global hybrid linear expert model is calculated through the backpropagation algorithm, including the shared expert weights, the weights of the experts in the parallel expert group and the parameters of the lightweight routing network. All learnable parameters are updated using the AdamW optimizer.

[0093] S38: Expert load balancing adjustment. After each round of training, the usage frequency of each expert in that round is counted, and the average usage frequency is calculated based on the usage frequency of each expert. According to the usage frequency f i of the i-th expert, the difference between the average frequency and the usage frequency of the i-th expert is used to dynamically update the dynamic bias term of the i-th expert:

[0094]

[0095] where sign represents the sign function, which returns -1 when to reduce the probability of being selected, returns +1 when to increase the probability of being selected, and returns 0 when to maintain the average frequency. λ is a hyperparameter that controls the adjustment amplitude.

[0096] In this embodiment, the hyperparameter λ is set to 0.001.

[0097] In this embodiment, to address the common issue of uneven expert load in MoE routing, this invention introduces a load balancing strategy without auxiliary loss—a dynamic bias adjustment mechanism. By introducing a dynamic bias term into the routing score, the selection frequency of each expert is balanced without adding an additional loss function. Specifically, after calculating the score for each expert, the routing network assesses whether the expert's usage frequency is too high or too low relative to the average level based on the number of tokens selected in the current iteration round, and makes subtle bias adjustments to the scoring results accordingly. For example, if an expert has been frequently selected recently, a negative bias is added to their score to reduce the probability of them being selected again; conversely, a positive bias is applied to less frequently used experts to increase their participation opportunities. This bias update does not participate in gradient backpropagation, avoiding the introduction of gradient interference, while achieving an effect similar to traditional load balancing loss, promoting more even expert usage.

[0098] Therefore, the dynamic bias adjustment mechanism of this invention ensures a relatively balanced workload for each expert during training, avoiding the training overhead caused by additional regularization terms in previous methods, and achieving "zero-loss" expert scheduling. With this mechanism, the global hybrid linear expert architecture remains stable and efficient during the training phase, allowing each expert to learn fully; during the inference phase, each expert collaborates according to the input features, leveraging their respective strengths.

[0099] In addition, after updating the dynamic bias term for each expert, status monitoring is performed. Training metrics such as the loss value, expert usage distribution, and route selection statistics for the current round are recorded to monitor training progress and model convergence status, ensuring that each expert receives sufficient and balanced training.

[0100] In summary, the global hybrid linear expert model of the present invention has the following significant advantages:

[0101] Global Sparse Activation: Almost all computational units in the entire Transformer network employ the MoE sparse activation mechanism, significantly improving parameter utilization efficiency. Compared to dense models with the same parameter scale, this architecture effectively reduces computational overhead in both pre-training and inference phases.

[0102] It combines sharing and specialization: shared linear experts provide general representation capabilities, ensuring that the model has a consistent basic representation when dealing with different tasks or domains; multiple experts existing in parallel enable the model to automatically adapt to different input features, forming dynamic expert paths.

[0103] Stable training and easy scalability: The introduced unassisted loss balancing routing strategy ensures dynamic balance of expert load. Even with a large number of experts, the training process still converges stably without the need for additional regularization terms. Simultaneously, this mechanism supports flexible expansion of the number of experts. The routing strategy ensures that new experts effectively participate in training, avoiding resource idleness or concentrated load, and providing good elasticity for model scaling.

[0104] S4: Input the target video to be described into the trained global hybrid linear expert model. The video encoder extracts the spatiotemporal feature representation of the target video. The lightweight routing network of each hybrid linear expert layer calculates the expert score based on the video feature token input by the current hybrid linear expert layer, dynamically selects the expert combination most suitable for processing the token, and finally generates a descriptive text that matches the content of the target video.

[0105] In this embodiment, the output is post-processed. Specifically, the generated descriptive text undergoes further grammar checks, duplication removal, and style consistency adjustments to ensure the quality and readability of the output descriptive text, ultimately generating descriptive text that highly matches the target video content and features natural and fluent language.

[0106] To better demonstrate the specific implementation and technical effects of the present invention, the video description generation method based on synthetic text and global hybrid linear expert shown in steps S1 to S4 of the above preferred implementation will be applied to several specific scenarios. The specific implementation process of the method of the present invention is as described above and will not be repeated here.

[0107] Example 1

[0108] In scenarios such as television broadcasting and online video, it is often necessary to automatically generate descriptive subtitles for videos without subtitles or narration to facilitate understanding of the content by visually or hearing impaired users. This embodiment applies the method of the present invention to a video-assisted subtitle generation system, which comprises two parts: training data synthesis and hybrid linear expert model training.

[0109] To synthesize training data, the system first parses the input video, extracting video frames and audio tracks. The audio is processed by the speech recognition module, while the video frames are used for content classification, such as sports, news, and lifestyle. Based on the classification results, the system selects the appropriate sub-process. For example, sports videos undergo action and scene analysis to extract target trajectories and key movements, which are then combined with templates by a large model to generate a description of the event. For landscape or lifestyle videos, keyframes are directly described using a visual model or LLM. The initially generated subtitles are filtered for inconsistencies, such as mentioning people not present in the frame, ensuring consistency between the subtitles and the video. Finally, a descriptive subtitle matching the video duration is generated.

[0110] When training the global hybrid linear expert model, each linear layer adds 8 parallel linear experts in addition to the original parameters as shared experts. One shared expert and one parallel expert are activated at a fixed time, for a total of 2 linear experts. Synthetic data is used to train the newly added linear experts and the routing network, activating different linear experts as needed to generate fluent subtitles that more closely resemble human expression.

[0111] Example 2

[0112] On educational video platforms such as MOOCs and online lectures, it is often necessary to generate content summaries for long teaching videos to help students quickly grasp key points or aid in retrieval. This embodiment demonstrates how to use the method of the present invention to generate an automatic summary for an online educational video.

[0113] First, training data is synthesized, and audio is transcribed using ASR to extract lecture content. Simultaneously, frames are periodically extracted, and OCR is used to recognize the text on the slides. ASR provides conversational long text, while OCR provides structured keywords; the two are combined to extract chapter divisions and core knowledge points, which are then fed into the LLM summarization module. The LLM module writes a third-person, objective summary based on a teaching template, and removes omissions or irrelevant content through consistency checks.

[0114] Subsequently, a global hybrid linear expert model is trained using only synthetic data. Leveraging the multimodal data pipeline integrating ASR and OCR, and the domain-adaptive global hybrid linear expert model of this invention, the system can automatically generate well-structured and linguistically standardized educational video summaries, comprehensively covering key knowledge points and discussion topics in lectures. This significantly improves learning efficiency while effectively reducing manual compilation costs.

[0115] Example 3

[0116] On large video sharing platforms, users upload massive amounts of content daily. The platforms need to automatically analyze these videos for content moderation, tag generation, and search indexing. To handle videos of different genres and generate high-quality text descriptions, this embodiment describes the application of the method of the present invention in an intelligent analysis scenario for video content platforms.

[0117] First, the uploaded videos are categorized by type (e.g., entertainment, news, games, reviews, etc.), and the corresponding processing flow is invoked based on the type:

[0118] Entertainment category: Extract character actions and dialogue to generate lighthearted and humorous descriptions;

[0119] News: Combining ASR and OCR to extract narration and subtitle information, generating a formal summary;

[0120] Game category: Identify key events and voice content to generate player-style commentary;

[0121] Evaluation category: Extract the explanation content and product information to generate a structured evaluation summary.

[0122] The generated text undergoes a consistency check to ensure that the content matches the video and contains no inappropriate information. The final result can be used for search indexing, video description display, or content review.

[0123] Leveraging the global hybrid linear expert model of this invention, the system activates experts on demand based on video content, achieving automatic generation of diverse yet consistent-quality content. For example, a humorous video can output lighthearted captions like "funny yet thrilling," while a news video generates a formal and rigorous summary. The model possesses good scalability and computational efficiency, supporting large-scale deployment scenarios and significantly reducing resource consumption while ensuring output quality.

[0124] The above embodiments fully demonstrate the effectiveness and advancement of the method of the present invention in various scenarios. Through innovation in model architecture (global hybrid linear expert) and data strategy (multimodal adaptive data synthesis), the present invention achieves a dual breakthrough in model performance and application implementation, providing new ideas for the development of video-text multimodal technology, and has significant practical significance and promising industrial application prospects.

[0125] It should also be noted that the video description generation method based on synthetic text and global hybrid linear experts in the above embodiments can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a video description generation system based on synthetic text and global hybrid linear experts, corresponding to the video description generation method based on synthetic text and global hybrid linear experts provided in the above embodiments, such as... Figure 4 As shown, it includes:

[0126] The data generation module is used to build an adaptive structured video description text generation process. It classifies the acquired raw video into content types and performs style adaptive generation and adversarial consistency verification based on the type tags of the raw video content, thereby generating high-quality video description text synthesis data.

[0127] The model acquisition module is used to uniformly replace all linear layers in the video description generation model with hybrid linear expert layers to form a global hybrid linear expert model. Each hybrid linear expert layer consists of three parts: a shared expert, a parallel expert group, and a lightweight routing network.

[0128] The model optimization module is used to take the video description text synthesis data generated in the data generation module and the corresponding original video as training data and form a training dataset. The global mixed linear expert model is iteratively trained on the training dataset. A preset number of experts are dynamically selected to participate in the calculation through a lightweight routing network to achieve sparse activation training. The dynamic bias term introduced based on the difference in the frequency of expert use is updated synchronously with each model parameter update, and finally the optimized global mixed linear expert model parameters are obtained.

[0129] The description text generation module is used to input the target video to be described into the trained global hybrid linear expert model. The video encoder extracts the spatiotemporal feature representation of the target video. The lightweight routing network of each hybrid linear expert layer calculates the expert score based on the video feature token input by the current hybrid linear expert layer, dynamically selects the expert combination that is most suitable for processing the token, and finally generates description text that matches the content of the target video.

[0130] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the video description generation method based on synthetic text and global hybrid linear experts provided in the above embodiments, which includes a memory and a processor;

[0131] The memory is used to store computer programs;

[0132] The processor is configured to implement the video description generation method based on synthetic text and global hybrid linear expert in the above embodiments when executing the computer program.

[0133] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0134] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the video description generation method based on synthetic text and global hybrid linear expert provided in the above embodiments. The storage medium stores a computer program that, when executed by a processor, can implement the video description generation method based on synthetic text and global hybrid linear expert in the above embodiments.

[0135] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.

[0136] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0137] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.

[0138] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A video description generation method based on synthetic text and global hybrid linear expert, characterized in that, Includes the following steps: S1: Construct an adaptive structured video description text generation process, classify the content type of the acquired original video, and perform style adaptive generation and adversarial consistency verification based on the type tags of the original video content, thereby generating high-quality video description text synthesis data. S2: Replace all linear layers in the video description generation model with hybrid linear expert layers to form a global hybrid linear expert model; each hybrid linear expert layer consists of three parts: shared experts, parallel expert groups, and a lightweight routing network. S3: Take the video description text synthesis data generated in S1 and the corresponding original video as training data to form a training dataset. Iteratively train the global hybrid linear expert model constructed in S2 on the training dataset. Dynamically select a preset number of experts to participate in the calculation through a lightweight routing network to achieve sparse activation training. And update the dynamic bias term based on the difference in the frequency of expert use synchronously when the model parameters are updated. Finally, the optimized global hybrid linear expert model parameters are obtained. S4: Input the target video to be described into the trained global hybrid linear expert model. The video encoder extracts the spatiotemporal feature representation of the target video. The lightweight routing network of each hybrid linear expert layer calculates the expert score based on the video feature token input by the current hybrid linear expert layer, dynamically selects the expert combination most suitable for processing the token, and finally generates a descriptive text that matches the content of the target video.

2. The video description generation method based on synthetic text and global hybrid linear expert as described in claim 1, characterized in that, The specific process of step S1 is as follows: S11: Extract keyframes representing the overall content of the original video from the original video, input the extracted keyframes into the InternVideo video basic model to extract the spatiotemporal feature representation of the original video and use it as keyframe features, output the confidence distribution of predefined categories, and select the category with the highest confidence as the type label of the original video content. S12: Preset corresponding language templates and style guidelines for each type of original video, integrate the theme, main action events, scene background, and dialogue of the original video into key information, fill the key information into the language template to form cue words, and input the cue words into GPT-4o mini to generate descriptive text that conforms to the style guidelines; S13: Use the pre-trained image-text matching model SigLIP2 to calculate the similarity between keyframe features and generated descriptive text. At the same time, use Qwen2.5-VL to determine whether the content of the original video matches the generated descriptive text. Descriptive texts with similarity below the preset similarity threshold and descriptive texts that do not match the content of the original video are filtered out. The final descriptive text is used as the video descriptive text synthesis data.

3. The video description generation method based on synthetic text and global hybrid linear expert as described in claim 2, characterized in that, For the original video containing multiple action scenes, PySceneDetect is used to detect and segment the scenes. The Qwen2.5-VL model is used to detect and label key behavioral events in the original video, identify the corresponding action categories and record the corresponding timestamps. At the same time, the visual detection model RAM++ is used to extract scene environment and important object information from the original video. Finally, the obtained scene, key behavioral events, scene environment and important object information are added to the key information.

4. The video description generation method based on synthetic text and global hybrid linear expert as described in claim 2, characterized in that, For original videos containing multiple languages, the OpenAI Whisper model is first used to perform automatic speech recognition on the audio track of the original video, converting the speech or narration into text. Then, PaddleOCR is used for layout analysis and text detection to locate text regions and recognize text content, obtaining explicit text information in the original video and adding it to key information.

5. The video description generation method based on synthetic text and global hybrid linear expert as described in claim 1, characterized in that, The linear layers in the video description generation model do not include the video encoder and the text token classification head. They include the Q, K, and V projection layers and the output layer in the attention mechanism, as well as the up-dimensional projection layer and down-dimensional projection layer in the feedforward neural network.

6. The video description generation method based on synthetic text and global hybrid linear expert as described in claim 1, characterized in that, In the hybrid linear expert layer, the weights of the linear layer of the video description generation model are used as shared expert weights to perform a general basic transformation on the token features of all inputs. The parallel expert group consists of N experts, each of whom performs a special transformation for a specific input pattern. The lightweight routing network itself is a lightweight linear layer. Based on the input token features of the current layer, it outputs the selection score of each expert and selects the top k experts with the highest scores to participate in the calculation of the token features. Experts that are not selected do not participate in the calculation.

7. The video description generation method based on synthetic text and global hybrid linear expert as described in claim 6, characterized in that, In step S3, the training process of the global mixed linear expert model in each round is as follows: S31: Select data batches from all training data using a random sampling or stratified sampling strategy to ensure that each batch contains video description text synthesis data of different types and styles; S32: Input a batch of original videos and corresponding video description text synthetic data into the global hybrid linear expert model, encode the features of the input original videos through the video encoder, convert the original videos into a token sequence, each token in the token sequence represents the local spatiotemporal features of the original video, and initialize the expert usage frequency statistics of each layer of lightweight routing network. S33: During the forward propagation process, for each hybrid linear expert layer in the global hybrid linear expert model, the lightweight routing network calculates the raw score of each expert based on the input token sequence x, and then adds the raw score to the dynamic bias term to obtain the final score of the expert. The final score of the i-th expert is: in, This represents the transpose of the i-th expert routing parameter; sigmoid(·) represents the sigmoid function; b i This represents the dynamic bias term of the i-th expert; S34: Select the top k experts in terms of final score, normalize the final score of the selected experts, and calculate their normalized score. S35: Multiply the normalized score of the selected expert by its respective weight matrix to obtain the final weight of the selected expert. Add the final weights of all selected experts to the weights of the shared experts to obtain the fusion weight. Multiply the fusion weight by the input token sequence to obtain the weighted output of the hybrid linear expert layer. S36: After processing by the global hybrid linear expert model, the synthesized video description text data is used as the real label, and the cross-entropy loss function is used to calculate the difference between the text predicted by the global hybrid linear expert model and the real label. S37: Calculate the gradient of the loss function with respect to the parameters of the global mixed linear expert model using the backpropagation algorithm, including the shared expert weights, the weights of experts in the parallel expert group, and the lightweight routing network parameters, and update all learnable parameters using the AdamW optimizer. S38: After each round of training, the usage frequency of each expert in that round is statistically analyzed, and the average usage frequency is calculated based on the usage frequency of each expert. Based on the usage frequency f of the i-th expert i The dynamic bias term of the i-th expert is dynamically updated based on the direction of the difference from the average frequency: Where sign represents the sign function, when Returns -1 when Returns +1 when It returns 0; λ is a hyperparameter that controls the adjustment range.

8. A video description generation system based on synthetic text and global hybrid linear expert, characterized in that, include: The data generation module is used to build an adaptive structured video description text generation process. It classifies the acquired raw video into content types and performs style adaptive generation and adversarial consistency verification based on the type tags of the raw video content, thereby generating high-quality video description text synthesis data. The model acquisition module is used to uniformly replace all linear layers in the video description generation model with hybrid linear expert layers to form a global hybrid linear expert model. Each hybrid linear expert layer consists of three parts: a shared expert, a parallel expert group, and a lightweight routing network. The model optimization module is used to take the video description text synthesis data generated in the data generation module and the corresponding original video as training data and form a training dataset. The global mixed linear expert model is iteratively trained on the training dataset. A preset number of experts are dynamically selected to participate in the calculation through a lightweight routing network to achieve sparse activation training. The dynamic bias term introduced based on the difference in the frequency of expert use is updated synchronously with each model parameter update, and finally the optimized global mixed linear expert model parameters are obtained. The description text generation module is used to input the target video to be described into the trained global hybrid linear expert model. The video encoder extracts the spatiotemporal feature representation of the target video. The lightweight routing network of each hybrid linear expert layer calculates the expert score based on the video feature token input by the current hybrid linear expert layer, dynamically selects the expert combination that is most suitable for processing the token, and finally generates description text that matches the content of the target video.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the video description generation method based on synthetic text and global hybrid linear expert as described in any one of claims 1 to 7.

10. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the video description generation method based on synthetic text and global hybrid linear expert as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Long-time video generation method based on adaptive video world model and related device

    CN122093637A