Video description generation method, apparatus, device, medium, and program product

By using an event location prediction model and positive and negative mask feature representation, video event locations and descriptions are automatically generated, solving the problem of high annotation costs in dense video description generation and achieving efficient model training and accurate alignment of event locations and descriptions.

CN122640602APending Publication Date: 2026-08-25TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213144.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In existing dense video description generation tasks, the annotation costs for event locations and event descriptions are too high, resulting in low model training efficiency.

Method used

Event locations are automatically generated by an event location prediction model, and event descriptions are generated by combining positive and negative mask feature representations with a first event description model, achieving implicit alignment between event locations and descriptions and reducing reliance on manual annotation.

Benefits of technology

Without relying on manual labeling of location information, the model training efficiency was improved, accurate alignment of event location with description was achieved, and labeling costs were reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122640602A_ABST
    Figure CN122640602A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, medium, and program product for generating video descriptions. The method includes: acquiring the positive mask feature representation and the negative mask feature representation of video frames in sample video data; analyzing the positive mask feature representation using a first event description model to obtain a first event description corresponding to the nth event; and analyzing the negative mask feature representation using the first event description model to obtain second event descriptions corresponding to events other than the nth event. By using the differences between S reference event descriptions and the first and second event descriptions, the accuracy of the location prediction is indirectly evaluated, enabling the event location prediction model to automatically adjust the event location prediction. Thus, the event descriptions generated by the first event description model can accurately correspond to the corresponding event locations. Without providing location labels, effective implicit alignment between event locations and event descriptions is achieved, improving the training efficiency of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method, apparatus, device, medium, and program product for generating video descriptions. Background Technology

[0002] Dense Video Captioning (DVC) is a task that, given a video, generates different descriptions for each of the specific events that occur in the video. The descriptions are brief and general text summarizing the events in the video.

[0003] In related technologies, the DVC task can be understood as a combination of two sub-tasks: 1. Determine the location of events in the video, and 2. Generate text descriptions for each event. For the DVC task, during the model training phase, each video sample needs to be labeled with several real event descriptions and their corresponding start and end times (which can be called event locations) as training labels. For example: "0:00-0:15: The chef is chopping onions and carrots; 0:15-0:30: Water in the pot begins to heat up, producing steam; 0:30-0:45: The chef puts the chopped vegetables into the pot and stir-fries them, with flames rising."

[0004] However, the model training schemes in related technologies require the annotation of event locations and event descriptions, which is too costly and results in low model training efficiency. Summary of the Invention

[0005] This application provides a method, apparatus, device, medium, and program product for generating video descriptions, the technical solution of which is as follows:

[0006] On the one hand, a method for generating video descriptions is provided, the method comprising:

[0007] Obtain reference event descriptions corresponding to S events in the sample video data. The reference event descriptions are used to describe the event content, and S is a positive integer.

[0008] The event location prediction model predicts S event locations in the sample video data, where the nth event location is used to express the time period of the nth event in the sample video data, and 0 < n ≤ S;

[0009] For the nth event location, obtain the positive mask feature representation and the negative mask feature representation of the video frame in the sample video data;

[0010] The first event description corresponding to the nth event is obtained by analyzing the positive mask feature representation using the first event description model; and the second event description corresponding to events other than the nth event is obtained by analyzing the negative mask feature representation using the first event description model.

[0011] The event location prediction model and the first event description model are trained based on S reference event descriptions, the first event description, and the second event description.

[0012] On the other hand, a method for generating a video description is provided, the method comprising:

[0013] Obtain the first video data;

[0014] Predict the event descriptions corresponding to R events in the first video data using the target event description model, where R is a positive integer;

[0015] The target event location prediction model predicts the event locations of the R events in the first video data, where the event location is used to express the time period of the event in the first video data;

[0016] The target event description model and the target event location prediction model are obtained by training sample video data and the reference event description corresponding to the sample video data.

[0017] On the other hand, a video description generation apparatus is provided, the apparatus comprising:

[0018] The first acquisition module is used to acquire reference event descriptions corresponding to S events in the sample video data, wherein the reference event descriptions are used to describe the event content, and S is a positive integer.

[0019] The model training module is used to predict S event locations in the sample video data through the event location prediction model, wherein the nth event location is used to express the time period of the nth event in the sample video data, and 0 < n ≤ S;

[0020] The model training module is used to obtain the positive mask feature representation and the negative mask feature representation of the video frame in the sample video data for the nth event position;

[0021] The model training module is used to analyze the positive mask feature representation through the first event description model to obtain a first event description corresponding to the nth event; and to analyze the inverse mask feature representation through the first event description model to obtain a second event description corresponding to events other than the nth event.

[0022] The model training module is used to train the event location prediction model and the first event description model based on S reference event descriptions, the first event description, and the second event description.

[0023] In some embodiments, the model training module is configured to: generate frame feature representations of video frames in the sample video data; determine the positive mask value and inverse mask value corresponding to the video frame in the sample video data based on the nth event position and the frame position of the video frame in the sample video data; weight the tth frame feature representation with the positive mask value corresponding to the tth video frame to obtain the tth positive mask feature representation, where t is a positive integer; and weight the tth frame feature representation with the inverse mask value corresponding to the tth video frame to obtain the tth inverse mask feature representation.

[0024] In some embodiments, the location of the nth event includes the time period center and the time period width corresponding to the time period of the nth event in the sample video data; the model training module is configured to: determine a first distance between the frame position of the video frame and the time period center; determine the positive mask value of the video frame based on the ratio between the first distance and the time period width; and obtain the difference between a preset value and the positive mask value as the inverse mask value.

[0025] In some embodiments, the model training module is configured to: obtain the number of video frames in the sample video data; and determine a first distance between the frame position of the video frame and the center of the time period based on the ratio between the frame position of the video frame and the number of video frames.

[0026] In some embodiments, the model training module is configured to: obtain the event feature representations corresponding to the S events respectively; input the event feature representations corresponding to the S events respectively, the frame feature representations and the frame position feature representations corresponding to the video frames in the sample video data into the event position prediction model, and output the S event positions of the sample video data.

[0027] In some embodiments, the event location prediction model includes a decoder, a first linear layer, and a second linear layer; the model training module is configured to: input the event feature representations corresponding to the S events, the frame feature representations corresponding to the video frames in the sample video data, and the frame location feature representations into the decoder, and output the decoded feature representations corresponding to the S events; input the decoded feature representations corresponding to the S events into the first linear layer, and output the time period centers corresponding to the S events; input the decoded feature representations corresponding to the S events into the second linear layer, and output the time period widths corresponding to the S events.

[0028] In some embodiments, the S reference event descriptions include a reference event description corresponding to the nth event and reference event descriptions corresponding to events other than the nth event; the model training module is configured to: determine a positive masking loss based on the difference between the reference event description corresponding to the nth event and the first event description; determine an inverse masking loss based on the difference between the reference event descriptions corresponding to events other than the nth event and the second event description; and train the event location prediction model and the first event description model based on the positive masking loss and the inverse masking loss.

[0029] In some embodiments, the model training module is configured to: determine a diversity loss based on the similarity between the positive mask information corresponding to the nth event location and the positive mask information corresponding to events other than the nth event, wherein the positive mask information includes the positive mask value corresponding to the video frame in the sample video data; and train the event location prediction model and the first event description model based on the positive mask loss, the inverse mask loss and the diversity loss.

[0030] In some embodiments, the model training module is configured to: acquire video feature representations of video frames in target sample video data, wherein the target sample video data corresponds to K reference event descriptions for each of the events, where K is a positive integer; analyze the video feature representations using a second event description model to obtain third event descriptions for each of the K events; and train the second event description model based on the differences between the reference event descriptions for each of the K events and the third event descriptions for each of the K events to obtain the first event description model.

[0031] In some embodiments, the model training module is configured to: generate frame feature representations and frame position feature representations of video frames in the target sample video data; and determine the video feature representation of the video frame based on the frame feature representations and the frame position feature representations.

[0032] On the other hand, a video description generation apparatus is provided, the apparatus comprising:

[0033] The second acquisition module is used to acquire the first video data;

[0034] The model prediction module is used to predict the event descriptions corresponding to R events of the first video data through the target event description model, where R is a positive integer;

[0035] The model prediction module is used to predict the event positions of the R events in the first video data through a target event position prediction model, wherein the event position is used to express the time period of the event in the first video data;

[0036] The target event description model and the target event location prediction model are obtained by training sample video data and the reference event description corresponding to the sample video data.

[0037] In some embodiments, the event location includes the time period center and time period width corresponding to the time period of the event in the first video data; the model prediction module is configured to: obtain the event feature representations corresponding to the R events respectively; analyze the event feature representations corresponding to the R events respectively, the frame feature representations corresponding to the video frames in the first video data, and the frame position feature representations through the target event location prediction model to generate the decoding feature representations corresponding to the R events respectively; analyze the decoding feature representations corresponding to the R events respectively through the target event location prediction model to determine the time period center and time period width corresponding to the R events respectively.

[0038] In some embodiments, the model prediction module is configured to: for the j-th event location, obtain the positive mask feature representation of the video frame in the first video data, 0 < j ≤ R; and analyze the positive mask feature representation through the target event description model to obtain the target event description corresponding to the j-th event.

[0039] In some embodiments, the model prediction module is configured to: generate frame feature representations of video frames in the first video data; determine the positive mask value corresponding to the video frame in the first video data based on the j-th event position and the frame position of the video frame in the first video data; and weight the t`-th frame feature representation with the positive mask value corresponding to the t`-th video frame to obtain the t`-th positive mask feature representation, where t` is a positive integer.

[0040] In some embodiments, the j-th event location includes the time period center and time period width corresponding to the time period of the j-th event in the first video data; the model prediction module is configured to: determine a first distance between the frame position of the video frame and the time period center; and determine the positive mask value of the video frame based on the ratio between the first distance and the time period width.

[0041] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the video description generation method described above.

[0042] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the video description generation method described above.

[0043] On the other hand, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the generation method described in any of the above-described video descriptions.

[0044] The beneficial effects of the technical solutions provided in this application include at least the following:

[0045] The event location prediction model automatically generates event locations from sample video data, eliminating the need for manually labeled locations and reducing labeling costs. For each event location, positive and negative mask feature representations are obtained. The first event description model uses the positive mask feature representation to focus on video frame features related to the specified event, thereby generating a first event description corresponding to the specified event. It uses the negative mask feature representation to focus on video frame features related to other events, thereby generating a second event description corresponding to other events. The accuracy of location prediction is indirectly evaluated by comparing the differences between S reference event descriptions and the first and second event descriptions. This allows the event location prediction model to automatically adjust event location predictions, ensuring that the event descriptions generated by the first event description model accurately correspond to the corresponding event locations. In the absence of provided location labels, this achieves effective implicit alignment between event locations and event descriptions, improving the training efficiency of the model. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a schematic diagram of a video description generation method provided in an exemplary embodiment of this application;

[0048] Figure 2 This is a schematic diagram of a computer system provided in an exemplary embodiment of this application;

[0049] Figure 3This is a flowchart of a video description generation method provided in an exemplary embodiment of this application;

[0050] Figure 4 This is a schematic diagram illustrating the construction of mask information provided in an exemplary embodiment of this application;

[0051] Figure 5 This is a flowchart of a method for generating a video description provided in another exemplary embodiment of this application;

[0052] Figure 6 This is a flowchart of a method for generating a video description provided in yet another exemplary embodiment of this application;

[0053] Figure 7 This is a schematic diagram illustrating the generation of a full video description provided in an exemplary embodiment of this application;

[0054] Figure 8 This is a schematic diagram illustrating the generation of a positive mask description provided in an exemplary embodiment of this application;

[0055] Figure 9 This is a schematic diagram illustrating the generation of an inverse mask description provided in an exemplary embodiment of this application;

[0056] Figure 10 This is a flowchart of a video description generation method provided in another exemplary embodiment of this application;

[0057] Figure 11 This is a structural block diagram of a video description generation apparatus provided in an exemplary embodiment of this application;

[0058] Figure 12 This is a structural block diagram of a video description generation apparatus provided in another exemplary embodiment of this application;

[0059] Figure 13 This is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0061] In this application, the terms "first" and "second" are used to distinguish between identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first" and "second", nor is there any limitation on the quantity or execution order.

[0062] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with relevant laws, regulations, and standards.

[0063] First, a brief introduction to the terms used in the embodiments of this application will be given.

[0064] Video Captioning (VC) task: Given a video, generate a brief summary text describing the content of the video. For example: The video shows a person preparing dinner in the kitchen. The corresponding description could be: "A person is preparing dinner in the kitchen."

[0065] Dense Video Captioning (DVC) task: Given a video, for several specific events occurring in the video, the model generates different descriptions corresponding to each event. The DVC task can be understood as a combination of two subtasks: 1. Determine the location of events in the video; 2. Generate a text description for each event. Unlike traditional video caption generation methods (which typically generate only a global description), the DVC task focuses on multiple time segments in the video and generates corresponding text descriptions for each segment. For example, if the video content is a person preparing dinner in a kitchen, the corresponding dense description generation could be: "0:00-0:15: The chef is chopping onions and carrots; 0:15-0:30: Water in the pot begins to heat up, steam is emitted; 0:30-0:45: The chef puts the chopped vegetables into the pot and stir-fries them, flames are rising."

[0066] The Weakly-Supervised Dense Video Captioning (WSDVC) task differs from the DVC task in that it uses different training data. Specifically, in the DVC task, during training, each training video sample contains several event descriptions and corresponding start and end times (referred to as "timestamps" or "locations"). The model uses this fully labeled data for training (i.e., fully supervised). After training, it is tested on unlabeled test video samples to generate several event descriptions and corresponding event locations. For the WSDVC task, the training samples used during training only contain event descriptions without specific event location annotations. For example, if the video content is about someone preparing dinner in a kitchen, a weakly supervised annotation might be: "The chef is chopping onions and carrots. Water in the pot starts to heat up, steam is coming out. The chef puts the chopped vegetables into the pot and stirs them up, flames are rising." The goal of the WSDVC task during the testing phase is the same as the DVC task: to generate several event descriptions and corresponding event locations. Compared to DVC tasks, WSDVC tasks require simpler annotation methods and content, thus having greater potential for practical applications.

[0067] This application proposes an implicit position-description alignment paradigm based on complementary masks, aiming to address the lack of supervision over event position in WSDVC tasks. For illustrative examples, please refer to [reference needed]. Figure 1 The diagram illustrates a method for generating video descriptions, used to train an event location prediction model and a first event description model. The method includes the following steps:

[0068] Step 1: Predict the locations of S events in the sample video 102 using the event location prediction model 101.

[0069] Among them, sample video 102 corresponds to S reference event descriptions, which represent the real event descriptions corresponding to the S event positions in sample video 102.

[0070] The sample video 102 includes at least two video frames. Frame embeddings and frame position embeddings corresponding to the at least two video frames are generated. S event embeddings are randomly initialized. The S event embeddings, the frame embeddings corresponding to the at least two video frames, and the frame position embeddings are input into the event position prediction model 101. The event position prediction model includes an attention layer, a first linear layer, and a second linear layer. The attention layer decodes the S event embeddings and the frame embeddings corresponding to the at least two video frames to obtain S decoded embeddings. The S decoded embeddings are input into the first linear layer to obtain the time period centers corresponding to the S events. The S decoded embeddings are input into the second linear layer to obtain the time period widths corresponding to the S events. Here, the time period center refers to the time period center corresponding to the occurrence time of the event in the sample video data, and the time period width refers to the time period width corresponding to the occurrence time of the event in the sample video data. The time period center and the time period width are used as the event position.

[0071] Step 2: For each event position 103, generate positive mask information and negative mask information.

[0072] For each of the S event locations, execute steps 2 through 5.

[0073] Taking the nth event position as an example, the positive mask information and the negative mask information are determined based on the time span width and time span center of the nth event position.

[0074] The forward mask information is used to perform forward masking processing on the frame embeddings corresponding to at least two video frames, wherein the forward mask information includes the forward mask values ​​corresponding to at least two frame embeddings. The reverse mask information is used to perform reverse masking processing on the frame embeddings corresponding to at least two video frames, wherein the reverse mask information includes the reverse mask values ​​corresponding to at least two frame embeddings. For a given frame embedding, its corresponding forward mask value and reverse mask value are complementary, if their sum is 1.

[0075] Step 3: Perform forward masking processing on the video frames in sample video 102 using the forward masking information to obtain the forward masking feature representation; perform reverse masking processing on the video frames in sample video 102 using the reverse masking information to obtain the reverse masking feature representation.

[0076] Optionally, forward masking is performed on at least two frame embeddings using the forward masking information to obtain at least two frame embeddings after forward masking as forward masking feature representations. For example, if the frame embeddings include frame embedding a, and the forward masking value of frame embedding a in the forward masking information is 0.8, then forward masking refers to performing weighted processing on frame embedding a using 0.8 as a weight.

[0077] Optionally, inverse masking is performed on at least two frame embeddings using the inverse masking information, resulting in at least two frame embeddings after inverse masking, which are then used as inverse masking feature representations. For example, if a frame embedding includes frame embedding a, and the inverse masking value of frame embedding a in the inverse masking information is 0.2, then the inverse masking process refers to performing weighted processing on frame embedding a using 0.2 as a weight.

[0078] Step 4: Analyze the positive mask feature representation using the first event description model 104 to obtain the first event description, and analyze the negative mask feature representation to obtain the second event description.

[0079] Input the positive mask feature representation into the first event description model 104 and decode to obtain the first event description corresponding to the nth event; input the negative mask feature representation into the first event description model 104 and decode to obtain the second event description corresponding to other events, where other events refer to events other than the nth event among the S events.

[0080] Among them, the first event description and the second event description complement each other to form the event description of S events.

[0081] Step 5: Train event location prediction model 101 and first event description model 104 based on S reference event descriptions, the first event description and the second event description.

[0082] The S reference event descriptions include the reference event description corresponding to the nth event and the reference event descriptions corresponding to the other events. Optionally, a positive masking loss is determined based on the difference between the reference event description corresponding to the nth event and the first event description; an inverse masking loss is determined based on the difference between the reference event descriptions corresponding to the other events and the second event description; and an event location prediction model and a first event description model are trained based on the positive masking loss and the inverse masking loss.

[0083] In summary, the video description generation method provided in this application indirectly evaluates the accuracy of location prediction by comparing the differences between S reference event descriptions and the first and second event descriptions. This enables the event location prediction model to automatically adjust the event location prediction, so that the event description generated by the first event description model can accurately correspond to the corresponding event location. In the absence of location labels, it achieves effective implicit alignment between event location and event description, thereby improving the training efficiency of the model.

[0084] Secondly, the computer system used to implement the video description generation method provided in this application will be introduced.

[0085] Figure 2This is a structural block diagram of a computer system provided in an exemplary embodiment of this application. The computer system can implement a system architecture for a video description generation method. The computer system includes a terminal 210 and a server 220. The terminal 210 is connected to the server 220 via a wireless network or a wired network.

[0086] Terminal 210 can be an electronic device such as a mobile phone, tablet computer, vehicle terminal (vehicle system), wearable device, or PC (Personal Computer). A client application for a first application can be installed and run on terminal 210. This first application can be an application used for training video description models (such as event location prediction models and first event description models), or it can be an application that provides training functions for video description models; this application does not limit the specific form of the first application. Furthermore, this application does not limit the form of the first application, including but not limited to apps, mini-programs, etc., installed on terminal 210, and it can also be in web page form.

[0087] Server 220 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, a cloud server providing basic cloud computing services, or a node in a blockchain system. Server 220 can be the backend server of the aforementioned first application, used to provide backend services to the clients of the first application.

[0088] The video description generation method (for training a video description model) provided in this application embodiment can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. For example... Figure 2 As shown, the video description generation method can be executed by terminal 210 (e.g., by a client of the first application installed and running in terminal 210), by server 220, or by a combination of terminal 210 and server 220. This application does not limit the specific method.

[0089] Those skilled in the art will understand that the number of terminals 210 described above can be more or less. For example, there may be only one terminal 210, or there may be dozens or hundreds of terminals 210, or even more. This application does not limit the number or type of terminals 210 in its embodiments.

[0090] In some embodiments, the trained video description model is used to provide video description generation functionality. A client application for a second application can be installed and running on the terminal 210. This second application can be an application that provides video description generation functionality. Furthermore, this application does not limit the form of the second application, including but not limited to apps, mini-programs, etc., installed on the terminal 210, and can also be in web page form.

[0091] Next, the training scheme for the video description model provided in this application will be introduced.

[0092] Figure 3 This is a flowchart illustrating a method for generating a video description according to an embodiment of this application. The method is executed by a computer device, which may be... Figure 2 The terminal 210 and / or server 220 are shown. The method includes steps 310 to 350.

[0093] Step 310: Obtain the reference event descriptions corresponding to the S events in the sample video data.

[0094] Sample video data refers to the data used to train the first event description model and the event location prediction model. Optionally, a sample dataset is obtained, which includes at least two sample video data sets. Optionally, the sample video data includes video frames, typically at least two.

[0095] The reference event description is used to describe the content of the event, where S is a positive integer. Illustratively, the reference event description refers to a true description of the event occurring in the sample video data, used to describe the specific content of the event. In this application, "description" refers to descriptive text.

[0096] Step 320: Predict the S event locations of the sample video data using the event location prediction model.

[0097] The S event positions refer to the individual event locations corresponding to S events, where the nth event position represents the time period of the nth event in the sample video data, 0 < n ≤ S, and n is an integer. Illustratively, an event position refers to the time period during which an event occurs in the sample video data.

[0098] In some embodiments, the event feature representations corresponding to S events are obtained; the event feature representations corresponding to the S events, the frame feature representations corresponding to the video frames in the sample video data, and the frame position feature representations are input into the event position prediction model, and the S event positions of the sample video data are output.

[0099] Optionally, the event feature representation is a randomly initialized feature representation. For example, S event feature representations are randomly initialized to represent S events. That is, the preset feature representations corresponding to the S events are obtained; the S preset feature representations, the frame feature representations corresponding to the video frames in the sample video data, and the frame position feature representations are input into the event position prediction model, and the model outputs the S event positions of the sample video data.

[0100] In the above scheme, the randomly initialized S event feature representations provide an initial framework for multi-feature fusion. Although these feature representations are random at the beginning, they provide a starting point for the event location prediction model. During the model training process, the randomly initialized event feature representations will be adjusted according to the prediction loss and other feedback information. The model will automatically learn how to effectively combine the event feature representations with the frame feature representations and frame location feature representations to improve the accuracy of event location prediction.

[0101] Frame feature representation refers to the representation extracted from each video frame of the sample video data that describes the features of that frame's content, such as vectors or matrices. Frame feature representation is used to characterize information such as objects, scenes, textures, shapes, and colors in the video frame. Optionally, the frame feature representation of each video frame of the sample video data is generated by a frame-level spatial encoder in the first event description model. The frame-level spatial encoder can be a pre-trained model, meaning its parameters remain unchanged during the current training process; or, the frame-level spatial encoder can be a pre-trained model, meaning it is the model to be trained. Optionally, the frame-level spatial encoder can be implemented as a convolutional neural network, etc. In some embodiments, the frame-level spatial encoder can employ multi-scale feature fusion technology to extract features from video frames at different scales; then, the features corresponding to different scales are fused to obtain the frame feature representation. Multi-scale feature fusion helps to simultaneously capture local detail information and global scene information in the video frame, thereby generating a richer and more comprehensive frame feature representation, providing more valuable feature input for subsequent tasks such as event location prediction.

[0102] Frame position feature representation is used to describe the positional information of video frames in sample video data. This positional information can be implemented as the order of the specified video frame within the video frame sequence. Illustratively, the sequential number of each video frame in the video can be used as the frame position feature representation; alternatively, the proportion of the frame within the entire video duration can be calculated (e.g., if the frame is the 50th frame in a video with a total of 100 frames, its temporal position can be represented as 0.5), and this can be used as the frame position feature representation. No specific limitations are imposed here.

[0103] Optionally, the frame feature representation and the frame position feature representation are concatenated to obtain the frame concatenated feature representation corresponding to the video frame in the sample video data. Illustratively, assume the frame feature representation is a vector F of length D, and the frame position feature representation is a vector P of length T. The concatenation operation involves connecting these two vectors in terms of dimensions to generate a new vector FP of length D+T. Then, the frame concatenated feature representation corresponding to the video frame in the sample video data and the event feature representations corresponding to S events are input into the event position prediction model. In the event position prediction model, the frame concatenated feature representation corresponding to the video frame in the sample video data is fused into the event feature representations corresponding to the S events to obtain the decoded feature representations corresponding to the S events. Illustratively, the frame concatenated feature representation corresponding to the video frame in the sample video data and the event feature representations corresponding to the S events are fused using attention mechanisms or other methods to obtain the S decoded feature representations. Finally, by analyzing the S decoded feature representations, the S event positions are obtained.

[0104] In some embodiments, the event location includes the time period center and time period width corresponding to the time period of the event in the sample video data. The example given is the nth event location including the time period center and time period width corresponding to the time period of the nth event in the sample video data.

[0105] Here, the time segment center refers to a central point in time that represents the period during which the nth event occurs in the sample video data, and the time segment width refers to the duration of the period during which the nth event occurs in the sample video data. For example, if the nth event's time segment in the sample video data is from second 10 to second 30, then the time segment center could be second 20, and the time segment width would be 30 - 10 = 20 seconds.

[0106] In some embodiments, the event location prediction model includes a decoder, a first linear layer, and a second linear layer.

[0107] Optionally, the event feature representations corresponding to the S events, the frame feature representations corresponding to the video frames in the sample video data, and the frame position feature representations are input into the decoder to output the decoded feature representations corresponding to the S events; the decoded feature representations corresponding to the S events are input into the first linear layer to output the time period centers corresponding to the S events; and the decoded feature representations corresponding to the S events are input into the second linear layer to output the time period widths corresponding to the S events.

[0108] Optionally, the decoder can be implemented as an attention-based decoder, for example, as an attention layer. Illustratively, the frame feature representation and frame position feature representation are concatenated to obtain the frame concatenated feature representation corresponding to the video frames in the sample video data. The frame concatenated feature representations corresponding to the video frames in the sample video data and the event feature representations corresponding to S events are input into the attention layer. In the attention layer, for each event feature representation: First, the similarity between each frame concatenated feature representation and the event feature representation is calculated to obtain the attention weight. The attention weight reflects the relevance between each video frame and the current event (e.g., the nth event). The attention weight can be calculated using a dot product attention mechanism, such as calculating the dot product between the event feature representation (as the query vector) and the frame concatenated feature representation (as the key vector), and then processing it through a scaling and normalization function (e.g., the softmax function) to obtain the attention weight. Then, the attention weight is used to perform a weighted summation of the frame concatenated feature representations corresponding to all video frames in the sample video data to obtain the decoded feature representation. The decoded feature representation is then input into the first linear layer, outputting the time segment center corresponding to the event; the decoded feature representation is then input into the second linear layer, outputting the time segment width corresponding to the event.

[0109] In the above scheme, the decoder is responsible for fusing and decoding the multiple input feature representations to generate decoded feature representations. This process can fully explore and utilize the relevant information between event feature representations, frame feature representations, and frame position feature representations. Through comprehensive processing of these features, the decoder can output more representative and discriminative decoded feature representations, thereby improving the accuracy of subsequent event position prediction.

[0110] The first linear layer focuses on mapping the decoded feature representation to the center of the time segment, while the second linear layer is responsible for mapping the decoded feature representation to the width of the time segment. This clear division of labor allows the model to more accurately predict the specific location of events on the timeline. The center of the time segment determines the approximate time of the event, while the width of the time segment reflects the duration of the event. Together, they constitute a complete description of the event's location, helping to more accurately determine the location of each event in the video.

[0111] Step 330: For the nth event location, obtain the positive mask feature representation and the negative mask feature representation of the video frame in the sample video data.

[0112] The positive mask feature representation is used to characterize the frame feature representation of the video frame associated with the nth event position in the sample video data. Illustratively, in the positive mask feature representation, the frame feature representation of the video frame associated with the nth event position is retained, while the frame feature representations of other unrelated video frames are ignored or have their influence reduced. The negative mask feature representation is used to characterize the frame feature representation of the video frames associated with other event positions in the sample video data. Other event positions refer to event positions other than the nth event position out of S event positions. Illustratively, in the negative mask feature representation, the frame feature representation of the video frame associated with the nth event position is ignored or has its influence reduced, while the frame feature representations of other unrelated video frames are retained.

[0113] In some embodiments, for the nth event location, based on the frame feature representation of the video frame in the sample video data and the nth event location, the positive mask information and the negative mask information corresponding to the nth event location are determined; positive mask processing is performed on the frame feature representation of the video frame in the sample video data according to the positive mask information to obtain the positive mask feature representation, and negative mask processing is performed on the frame feature representation of the video frame in the sample video data according to the negative mask information to obtain the negative mask feature representation.

[0114] The positive mask information is used to mask the video frames corresponding to all event positions except the nth event position in the sample video data, while the negative mask information is used to mask the video frames corresponding to the nth event position in the sample video data.

[0115] Optionally, the sample video data includes at least two video frames, the positive mask information includes the positive mask values ​​corresponding to at least two video frames, and the negative mask information includes the negative mask values ​​corresponding to at least two video frames. Illustratively, for positive mask processing, at least two positive mask values ​​in the positive mask information are multiplied element-wise with the frame feature representation. For negative mask processing, at least two negative mask values ​​in the negative mask information are multiplied element-wise with the frame feature representation. Specifically, if the positive or negative mask value of a video frame is 1, its corresponding frame feature representation retains all information; if the positive or negative mask value of a video frame is 0.5, its corresponding frame feature representation retains half information; if the positive or negative mask value of a video frame is 0, its corresponding frame feature representation retains no information.

[0116] In some embodiments, the method for obtaining the positive mask feature representation and the negative mask feature representation further includes the following steps:

[0117] Step 1: Generate frame feature representations of video frames in the sample video data.

[0118] Optionally, a frame feature representation of each video frame of the sample video data is generated by a frame-level spatial encoder in the first event description model.

[0119] Step 2: Based on the position of the nth event and the frame position of the video frame in the sample video data, determine the positive mask value and the negative mask value corresponding to the video frame in the sample video data.

[0120] Indicatively, the sample video data includes at least two video frames. Based on the position of the nth event and the frame positions corresponding to the at least two video frames in the sample video data, the positive mask value and the negative mask value corresponding to the at least two video frames are determined respectively.

[0121] In some embodiments, the location of the nth event includes the time period center and the time period width corresponding to the time period of the nth event in the sample video data.

[0122] Optionally, a first distance between the frame position of the video frame and the center of the time period is determined; a positive mask value of the video frame is determined based on the ratio between the first distance and the width of the time period; and the difference between a preset value and the positive mask value is obtained as an inverse mask value.

[0123] Optionally, regarding the first distance: obtain the number of video frames in the sample video data; determine the first distance between the frame position of the video frame and the center of the time period based on the ratio between the frame position of the video frame and the number of video frames. Here, the frame position of the video frame refers to the order of the video frame in the frame sequence corresponding to the sample video data. Illustratively, the frame position can be represented by the frame number; for example, in a video of 30 frames per second, the frame position of the 10th frame in the 1st second can be represented as 10. Therefore, if the number of video frames in the sample video data is 100, the first distance is 10 / 100 = 0.1.

[0124] In the above scheme, the number of video frames in the sample video data is obtained, and the first distance between the frame position and the center of the time period is determined based on the ratio between the frame position and the number of video frames, thus achieving accurate distance calculation. Using the ratio of the frame position to the number of video frames as the distance measurement standard is essentially a normalization process. Normalization eliminates the impact of differences in video length on distance calculation, enabling the model to stably calculate the relative distance between video frames and the center of the event time period even when faced with video data of different lengths, thereby ensuring the stability and consistency of the model across different video data.

[0125] Optionally, the time period center and time period width are values ​​obtained after normalization. Normalization is used to map the time period center and time period width to a specific range, such as (0,1) (or [0,1], which is not limited here). If the time period of an event is from the 10th second to the 30th second, then the time period center is the 20th second and the time period width is 20 seconds. If the total duration of the sample video data is 80 seconds, then the normalized time period center value is 0.25 and the normalized time period width value is 0.25.

[0126] After obtaining the first distance, the positive mask value of the video frame is determined based on the ratio between the first distance and the time segment width. Illustratively, after calculating the first distance, the square or absolute value of the first distance is obtained to ensure that the first distance is non-negative. Taking the absolute value of the first distance as an example, the positive mask value is determined based on the ratio between the absolute value of the first distance and the time segment width. After obtaining the positive mask value, the difference between a preset value and the positive mask value is calculated as the inverse mask value, where the preset value can be set to 1. The positive mask value represents the degree of correlation between the video frame and the event position; a higher value indicates that the video frame position is closer to the event position. The inverse mask value represents the degree of discordation between the video frame and the event position; a higher value indicates that the video frame position is farther from the event position.

[0127] In the above scheme, representing the position of the nth event as the center and width of the time segment allows for a more accurate description of the event's temporal location within the video. By calculating the first distance between the video frame's position and the center of the time segment, and determining the positive mask value based on the ratio of this distance to the time segment width, the correlation between each video frame and the event's time segment can be measured more precisely. The inverse mask value is determined by subtracting the positive mask value from a preset value. This setting ensures that the sum of the positive and inverse mask values ​​is a constant (i.e., a preset value), thus maintaining numerical stability during the weighted processing of feature representation.

[0128] Step 3: Weight the feature representation of the t-th frame with the positive mask value corresponding to the t-th video frame to obtain the positive mask feature representation of the t-th frame, where t is a positive integer.

[0129] The positive mask values ​​corresponding to at least two video frames include the positive mask value corresponding to the t-th video frame. The positive mask value corresponding to the t-th video frame is used as a weight and multiplied element-wise with the feature representation of the t-th frame to obtain the positive mask feature representation of the t-th frame.

[0130] Step 4: Weight the feature representation of the t-th frame with the inverse mask value corresponding to the t-th video frame to obtain the t-th inverse mask feature representation.

[0131] The inverse mask values ​​corresponding to at least two video frames include the inverse mask value corresponding to the t-th video frame. The inverse mask value corresponding to the t-th video frame is used as a weight and multiplied element-wise with the feature representation of the t-th frame to obtain the t-th inverse mask feature representation.

[0132] In the above scheme, the positive and negative mask values ​​are determined based on the position of the nth event and the frame position of the video frame. The positive mask value is used to weight the frame feature representation to obtain the positive mask feature representation, enabling the model to more accurately extract video frame features related to the nth event. The positive mask value highlights video frame features within the event's occurrence time period, allowing the model to more accurately capture detailed information related to the event when generating the first event description corresponding to the nth event, thereby improving the accuracy and quality of the event description. The negative mask value, used to weight the frame feature representation, helps the model effectively distinguish features of events other than the nth event. The negative mask value suppresses video frame features within the nth event's occurrence time period, highlighting features of video frames in other time periods, enabling the model to more accurately identify and describe other events when generating the second event description corresponding to other events.

[0133] This is illustrative; please refer to it. Figure 4 It shows a schematic diagram of the construction of mask information.

[0134] (1) Event location prediction

[0135] like Figure 4 As shown, taking S events as events 1, 2, and 3 as an example, after inputting at least two video frames 401 from the sample video data into the frame-level spatial encoder 402 in the first event description model, the output is the frame embedding 403 corresponding to at least two video frames respectively. Indicates all frame embeddings, N v This indicates the number of video frames in the sample video data. For each video frame, its corresponding frame embedding 403 and frame position embedding are concatenated to obtain the frame concatenation embedding. This indicates that all frames are concatenated and embedded. The event embeddings corresponding to events 1, 2, and 3 are randomly initialized. (Embedded learning is possible), where n is a positive integer representing the nth event. Indicates all event embeddings, N S This represents the number of events in the sample video data. The event location prediction model 404 can be implemented as a Transformer-based decoder, and includes an attention layer, a first linear layer, and a second linear layer.

[0136] Will Input attention layer, output in This refers to all decoding embeddings, and the decoding formula corresponding to the attention layer is shown in Formula 1 below:

[0137]

[0138] d represents the embedding dimension, and TransformerDecoder represents the attention layer. Optionally, the attention layer refers to the cross-attention layer, which is applied to each... e n As a query (Q), As keys (K) and values ​​(V). In the attention layer, for each query e n Calculate its relationship with key v a +θ p The dot product similarity; to prevent the dot product from becoming too large, the calculated dot product similarity can be divided by... That is to say Then along the dimensions of the key (e.g., each row) Normalized to attention weights; the value v is then adjusted using these attention weights. a +θ p We obtain the decoded feature representation h corresponding to the h-th event by weighted summation. n .

[0139] In obtaining Then, for each decoding embedding h n Inputting into the first linear layer yields the time period center μ corresponding to the nth event. n h n Input the second linear layer to obtain the time span width σ corresponding to the nth event. n The transformation formulas corresponding to the first linear layer and the second linear layer are shown in Formula 2 below:

[0140] Formula 2: μ n =Sigmoid(FC1(h n )),σ n =Sigmoid(FC2(h n ))

[0141] Where FC1 represents the first linear layer, FC2 represents the second linear layer, and Sigmoid is the normalization function used to normalize μ. n and σ n The range of values ​​is mapped to (0, 1).

[0142] (2) Differentiable mask construction

[0143] For each event, based on the center μ of the time period corresponding to the event. n and time period width σn Generate Gaussian mask (Positive mask information), where the positive mask information corresponds to the nth event. The construction formula is shown in Formula 3 below:

[0144]

[0145] Here, τ is a hyperparameter that controls the steepness of the Gaussian curve. This represents the positive mask value of the t-th video frame. If t is 3, then N represents the number of video frames in the sample video data. v If it is 10, then t / N v The value is 3 / 10. It should be noted that other functions for generating soft masks can also be used as alternatives to Gaussian masks, such as the Logistic function.

[0146] Among them, the inverse mask information corresponding to the nth event The construction formula is shown in Formula 4 below:

[0147]

[0148] like Figure 4 As shown, M n Indicates positive mask information or inverse mask information When M n Indicates positive mask information Through positive mask information Perform positive masking on frame embedding 403 to obtain positive mask embedding, that is: When M n Indicates inverse mask information Through positive mask information Perform inverse masking on frame embedding 403 to obtain the inverse masked embedding, that is: in, That is, the mask information M is applied. n The frame embedding is then input into the first event description model for further processing.

[0149] Step 340: Analyze the positive mask feature representation using the first event description model to obtain the first event description corresponding to the nth event; and analyze the negative mask feature representation using the first event description model to obtain the second event description corresponding to events other than the nth event.

[0150] In some embodiments, the first event description model includes a video-level temporal encoder and a description decoder. The video-level temporal encoder can be implemented as an attention-based encoder, such as a Transformer-based encoder; the description decoder can be implemented as an attention-based decoder, such as a Transformer-based decoder or a recurrent neural network-based decoder.

[0151] Optionally, the positive mask feature representation is input into the video-level time encoder, and the output is a positive mask video feature representation; the positive mask video feature representation is input into the description decoder, and the decoding yields the first event description corresponding to the nth event. The inverse mask feature representation is input into the video-level time encoder, and the output is an inverse mask video feature representation; the inverse mask video feature representation is input into the description decoder, and the decoding yields the second event description corresponding to events other than the nth event.

[0152] Indicatively, a video-level temporal encoder includes a self-attention layer. After the positive mask feature representation (or inverse mask feature representation) is input into the self-attention layer of the video-level temporal encoder, the positive mask feature representation (or inverse mask feature representation) is first linearly transformed into three vectors: query, key, and value. For each query vector: First, the similarity between each key vector and the query vector is calculated to obtain the attention weight. The attention weight can be calculated through a dot product attention mechanism, such as calculating the dot product between the query vector and the key vector, and then processed by a scaling and normalization function (such as the softmax function) to obtain the attention weight. Then, the attention weight is used to perform a weighted summation of all value vectors to obtain the positive mask video feature representation (or inverse mask video feature representation).

[0153] After obtaining the positive mask video feature representation (or inverse mask video feature representation), the positive mask video feature representation (or inverse mask video feature representation) is input into the description decoder. The description decoder can be implemented as an autoregressive language model based on the Transformer decoder. In the description decoder, the positive mask video feature representation (or inverse mask video feature representation) is the prefix of the sequence. The description decoder continuously generates the probability distribution of the next word through a self-attention mechanism, thereby realizing text generation.

[0154] Step 350: Train the event location prediction model and the first event description model based on S reference event descriptions, the first event description, and the second event description.

[0155] Among them, the S reference event descriptions include the reference event description corresponding to the nth event and the reference event descriptions corresponding to the events other than the nth event.

[0156] Optionally, a positive masking loss is determined based on the difference between the reference event description and the first event description corresponding to the nth event; an inverse masking loss is determined based on the difference between the reference event description and the second event description corresponding to events other than the nth event; and an event location prediction model and a first event description model are trained based on the positive masking loss and the inverse masking loss.

[0157] The loss function for calculating the positive masking loss can be implemented as at least one of the following: cross-entropy loss function, negative log-likelihood loss function, etc. Similarly, the loss function for calculating the inverse masking loss can be implemented as at least one of the following: cross-entropy loss function, negative log-likelihood loss function, etc.

[0158] To illustrate, after calculating the positive masking loss and the inverse masking loss, the sum or weighted sum of the positive masking loss and the inverse masking loss is calculated.

[0159] In the above scheme, the positive masking loss is determined based on the difference between the reference event description and the first event description corresponding to the nth event, and the inverse masking loss is determined based on the difference between the reference event description and the second event description corresponding to events other than the nth event. This allows for a precise measurement of the model's prediction error in event descriptions. This accurate loss calculation method effectively guides the model to adjust its parameters, enabling the model to more accurately align event positions and event descriptions during training, achieving effective implicit alignment.

[0160] In some embodiments, in order to enable the positive masking information corresponding to the S events to mask different video parts of the sample video data, the training loss of the model may also include diversity loss.

[0161] Optionally, a diversity loss is determined based on the similarity between the positive mask information corresponding to the nth event location and the positive mask information corresponding to events other than the nth event. The positive mask information includes the positive mask value corresponding to the video frame in the sample video data. The event location prediction model and the first event description model are trained based on the positive mask loss, the inverse mask loss and the diversity loss.

[0162] The similarity calculation method can use cosine similarity, which is based on the cosine similarity between the positive mask information corresponding to the nth event position and the positive mask information corresponding to all other events, to determine the diversity loss. The total loss is obtained by calculating the sum or weighted sum of the positive mask loss, inverse mask loss, and diversity loss.

[0163] In the above scheme, the diversity loss can measure the difference between the positive mask information of different events, prompting the model to increase the difference between the positive mask information of different events, thereby improving the accuracy of mask processing.

[0164] Optionally, after obtaining the total loss, the event location prediction model and the first event description model are trained based on the total loss. Optionally, the model parameters of the video-level event encoder, description decoder, and the first event description model are updated based on the total loss, thereby completing one training process corresponding to one sample video data. The above training process is performed using multiple sample video data until the calculated total loss is less than the preset loss, or the training times reach the preset number, at which point training stops. The location prediction model and event description model obtained at this point are the models used subsequently, namely the target event location prediction model and the target event description model.

[0165] In summary, the video description generation method provided in this application automatically generates event locations in sample video data through an event location prediction model, eliminating reliance on manually labeled locations and reducing labeling costs. For each event location, a positive mask feature representation and an inverse mask feature representation are obtained. The first event description model uses the positive mask feature representation to focus on video frame features related to a specified event, thereby generating a first event description corresponding to the specified event. It uses the inverse mask feature representation to focus on video frame features related to other events, thereby generating a second event description corresponding to other events. By comparing the differences between S reference event descriptions and the first and second event descriptions, the accuracy of location prediction is indirectly evaluated. This allows the event location prediction model to automatically adjust event location predictions, ensuring that the event descriptions generated by the first event description model accurately correspond to the corresponding event locations. Even without provided location labels, this achieves effective implicit alignment between event locations and event descriptions, improving model training efficiency.

[0166] In some embodiments, the first event description model is obtained by training a second event description model. Please refer to [reference needed]. Figure 5 ,exist Figure 3 In the illustrated embodiment, steps 501 and 503 are included before training the first event description model.

[0167] Step 501: Obtain the video feature representation of the video frame in the target sample video data.

[0168] Optionally, the target sample video data refers to the data used to train the second event description model. Optionally, a sample dataset is obtained, which includes at least two target sample video data sets. Wherein, the target sample video data includes at least two video frames, then the video feature representations corresponding to the at least two video frames in the target sample video data are obtained.

[0169] The target sample video data corresponds to K events, each with a corresponding reference event description. K is a positive integer.

[0170] Optionally, frame feature representations and frame position feature representations of video frames in the target sample video data are generated; based on the frame feature representations and frame position feature representations, the video feature representations of the video frames are determined.

[0171] Optionally, the frame feature representation of each video frame of the sample video data is generated by the frame-level spatial encoder in the second event description model, wherein the frame-level spatial encoder can be a pre-trained model, that is, the model parameters of the frame-level spatial encoder remain unchanged during this training process; or, the frame-level spatial encoder can be a pre-trained model, that is, the frame-level spatial encoder is the model to be trained.

[0172] Frame position feature representation is used to describe the positional information of video frames in the target sample video data. This positional information can be implemented as the order of the specified video frames within the video frame sequence. Illustratively, the sequential number of each video frame in the video is used as the frame position feature representation.

[0173] Optionally, the frame feature representation and the frame position feature representation are spliced ​​together to obtain the frame splicing feature representation corresponding to the video frame in the target video data; the frame splicing feature representation is input into the video-level time encoder in the second event description model to obtain the video feature representation of the video frame in the target sample video data.

[0174] To illustrate, a video-level temporal encoder includes a self-attention layer. After the frame-stitched feature representation is input into the self-attention layer of the video-level temporal encoder, the frame-stitched feature representation is first linearly transformed into three vectors: query, key, and value. For each query vector: first, the similarity between each key vector and the query vector is calculated to obtain the attention weight; then, all value vectors are weighted and summed using the attention weight to obtain the video feature representation.

[0175] In the above scheme, the video feature representation integrates frame feature representation and frame position feature representation, which can more effectively express the features of video frames. This fusion method allows the video feature representation to include not only the static content information of video frames but also their dynamic position information in the video sequence, thereby enhancing the expressive power of the feature representation and providing more accurate training data for model training. When training the second event description model, based on these more accurate video feature representations, the model can more accurately learn the mapping relationship between video features and event descriptions.

[0176] Step 502: Analyze the video feature representation using the second event description model to obtain the third event descriptions corresponding to the K events.

[0177] Optionally, the video feature representation is analyzed by the description decoder in the second event description model to obtain the third event descriptions corresponding to the K events respectively.

[0178] In a schematic way, the video feature representation is input into the description decoder, and the K events are decoded to obtain the third event descriptions corresponding to them. The description decoder can be implemented as an autoregressive language model based on the Transformer decoder.

[0179] Step 503: Train the second event description model based on the differences between the reference event descriptions corresponding to the K events and the third event descriptions corresponding to the K events, and obtain the first event description model.

[0180] Optionally, the full description loss is determined based on the difference between the reference event descriptions corresponding to the K events and the third event descriptions corresponding to the K events, and the second event description model is trained based on the full description loss to obtain the first event description model.

[0181] Optionally, the loss function used to calculate the fully descriptive loss can be implemented as at least one of the negative log-likelihood loss function, cross-entropy loss function, etc.

[0182] In illustrative terms, the model parameters of the video-level temporal encoder in the second event description model are updated based on the full description loss, thus completing one training cycle for a target sample video data. This training process is repeated using multiple target sample video data until the calculated full description loss is less than the preset loss, or the preset number of training iterations is reached. Training then stops, and the resulting event description model is the model used subsequently, i.e., the first event description model.

[0183] In summary, the video description generation method provided in this application training method is based on target sample video data to train a second event description model. During the training process, the second event description model can learn the mapping relationship between rich video features and event descriptions, thereby enabling the trained first event description model to generate event descriptions more accurately.

[0184] The video description model provided in this application integrates two main components: a dense video description generation module for event description generation and a complementary mask generation module for event location prediction. The dense video description generation module has two operating modes: a full description generation mode for capturing the global narrative, and a mask description generation mode for enhancing location awareness through differentiable masks. The complementary mask generation module, in the mask description generation mode, utilizes positive and negative masks to predict temporal location. The positive mask focuses on the description of a specific event, while the negative mask handles the remaining context, thus ensuring alignment under weak supervision. Through these two components, the video description model can effectively align event descriptions with event locations. The dense video description generation module includes an event description model, and the complementary mask generation module includes an event location prediction model, etc., which are described below. Figure 6The training process of the video description model provided in this application is described, which includes two stages.

[0185] (I) First Training Phase

[0186] Step 601: Obtain the first training dataset.

[0187] The first training dataset is used to train the second event description model. The first training dataset includes first sample video data (i.e., the aforementioned target sample video data), and each of the first sample video data has a corresponding reference event description for K events. The reference event descriptions for the K events are ordered from front to back according to the order in which the K events occur in the first sample video data.

[0188] Step 602: Input the first sample video data into the frame-level spatial encoder in the second event description model to obtain the frame feature representation of the video frame in the first sample video data.

[0189] In some embodiments, the first sample video data includes at least two video frames.

[0190] Optionally, at least two video frames from the first sample video data are input into the frame-level spatial encoder in the second event description model to generate frame feature representations corresponding to the at least two video frames respectively.

[0191] Step 603: Input the frame feature representation and frame position feature representation of the video frames in the first sample video data into the video-level time encoder in the second event description model to obtain the video feature representation of the video frames in the first sample video data.

[0192] Optionally, obtain the frame position feature representations corresponding to at least two video frames respectively; concatenate the frame position feature representations and the frame feature representations to obtain the frame concatenation feature representations corresponding to at least two video frames respectively; input the frame concatenation feature representations corresponding to at least two video frames respectively into the frame-level spatial encoder in the second event description model, and output the video feature representations corresponding to at least two video frames respectively.

[0193] Step 604: Input the video feature representations of the video frames in the first sample video data and the first description sequence into the description decoder in the second event description model to obtain the third event descriptions corresponding to the K events respectively.

[0194] Optionally, the first description sequence includes multiple text embeddings, which are used to indicate the first prompt text and the reference event descriptions corresponding to the K events respectively. The first prompt text is used to guide the description decoder to decode the event descriptions corresponding to the K events, i.e., the full video description. Illustratively, the first prompt text and the reference event descriptions corresponding to the K events are concatenated and then segmented into words. Then, text embeddings corresponding to each word are generated to obtain the first description sequence.

[0195] Indicatively, the first prompt text includes a first prompt symbol and event count information. The first prompt symbol can be implemented as [FULL]. The event count information is used to guide the description decoder on the number of event descriptions it needs to generate. For example, if the event count information is implemented as "3events", it means that the description decoder is guided to generate 3 event descriptions.

[0196] Optionally, the video feature representations corresponding to at least two video frames and the first description sequence are input into the description decoder in the second event description model to obtain the third event descriptions corresponding to K events.

[0197] To illustrate, taking an autoregressive language model based on a Transformer decoder as an example, the description decoder generates words in the third event description one by one based on the input video feature representation (as a prefix sequence) and the text embeddings in the first description sequence. During the generation process, the description decoder considers the video feature representation and existing text information. For example, when generating the i-th text embedding (or word) in the third event description, it focuses on the first i-1 text embeddings in the first description sequence, where i is an integer greater than 1.

[0198] Step 605: Determine the full description loss based on the differences between the reference event descriptions corresponding to the K events and the third event descriptions corresponding to the K events; train the second event description model based on the full description loss to obtain the first event description model.

[0199] To illustrate, the model parameters of the video-level temporal encoder and descriptive decoder in the second event description model are updated according to the full description loss, thus completing one training process corresponding to one first sample video data. The above training process is performed with multiple first sample video data until the calculated full description loss is less than the preset loss, or the training times reach the preset number, at which point training stops. The event description model obtained at this point is the model used subsequently, i.e., the first event description model.

[0200] Please refer to Figure 7 It illustrates a schematic diagram of full video description generation.

[0201] (1) Spatial-temporal video coding

[0202] like Figure 7 As shown, the second event description model includes a frame-level spatial encoder 701 and a video-level temporal encoder 702. The frame-level spatial encoder 701 uses E... a It indicates that the video-grade time encoder 702 uses E... b Indicates. Given sample video data (That is, video frame 703), where H and W represent the height and width of each video frame, respectively, and N v This represents the number of video frames in the sample video data.

[0203] Video frame 703 is input into frame-level spatial encoder 701, which extracts visual embeddings from each frame to generate frame embeddings. That is: Among them, v t Let d represent the t-th video frame, and d be the embedding dimension. In some embodiments, the parameters of the frame-level spatial encoder 701 remain unchanged during training.

[0204] Next, the video-level time encoder 702 processes the frame embedding v a and frame position embedding θ p To capture time information and generate video embeddings v b , that is: v a +θ p This indicates the embedding of spliced ​​frames and the embedding of frame positions. Optionally, the video-level time encoder 702 can be implemented as a randomly initialized Transformer encoder, outputting v b The characteristics of the frames and their temporal relationships are encoded.

[0205] (2) Prompt-based full video description decoding

[0206] like Figure 7 As shown, the second event description model also includes a description decoder 704, which processes the video embedding v. b This is a sequential description of all events in the generated sample video data v. For each sample video data v, there is a corresponding given description. (i.e., refer to the event description), where N S C represents the number of events in the sample video data. n This indicates the event description corresponding to the nth event. The prompt P "[FULL]N" will be displayed. S events:” is concatenated with all event descriptions into a paragraph, which is then tokenized and embedded to form the first description sequence r, i.e.:

[0207]

[0208] In the prompt P, [FULL] indicates that the model is required to generate all descriptions, N S events specifies the number of event descriptions. N S N represents the number of events in the sample video data, that is, the number of event descriptions. r This represents the number of word segments (or text embeddings) in the first description sequence r.

[0209] In the description decoder 704, the video is embedded in v b As a prefix visual marker, and associated with the first descriptive sequence Connection (r) i The input sequence is obtained by representing the i-th text embedding, that is: The input sequence Z is input to the description decoder 704. Illustratively, the description decoder 704 starts decoding from the position corresponding to r2. For the i-th text embedding, the description decoder 704 is based on v... b The i-th predicted text embedding is obtained by predicting the first i-1 reference text embeddings (i.e., the first i-1 text embeddings in the first description sequence) and the first i-1 reference text embeddings. Alternatively, the i-th probability distribution can be obtained, where the i-th probability distribution includes the probability that a word in the dictionary is the i-th text embedding. If i is greater than 2, the description decoder 704 will predict the i-th text embedding based on v. b Predict the i-th predicted text embedding or the i-th probability distribution by combining it with the first i-1 predicted text embeddings (or the first i-1 probability distributions).

[0210] It should be noted that, in In the diagram, r1 to r3 correspond to cue P, where r1 represents the text embedding corresponding to [FULL], and r2 represents N. S The corresponding text embedding and r3 represent the text embedding corresponding to events. In other words, when decoding, the description decoder 704 will first predict the number of events or the number of event descriptions in the sample video data.

[0211] (3) Full video description generation training loss

[0212] For full video description generation, the training objective is to minimize the negative log-likelihood, and the corresponding loss function is... As shown in Formula 5 below:

[0213]

[0214] Among them, v b For video embedding, N1 represents the number of text embeddings predicted by decoder 704, N1 = N r -1. θ E,GThis represents the model parameters of the frame-level spatial encoder 701, the video-level temporal encoder 702, and the description decoder 704. logp(r i |v b ,r <i ;θ E,G ) indicates that given v b r i In the case of the previous text embedding (text embedding in the first description sequence), the description decoder 704 decodes to obtain r. i The logarithmic probability.

[0215] After obtaining the full descriptive loss using the L1 loss function, the model parameters of the video-level temporal encoder 702 and the descriptive decoder 704 are updated based on the full descriptive loss. The above training process is then performed on the second event description model using multiple sample video data until the calculated full descriptive loss is less than the preset loss or the preset number of training iterations is reached. Training stops at this point, and the resulting event description model is the first event description model. Illustratively, 10 batches (epochs) of training data (i.e., 10 training datasets) are obtained. After training the second event description model using these 10 batches of training data, the first event description model is obtained. Then, the first event description model and the event location prediction model are trained using these 10 batches of training data.

[0216] (II) Second Training Phase

[0217] Step 606: Obtain the second training dataset.

[0218] The second training dataset includes second sample video data, each with S corresponding reference event descriptions. These reference event descriptions are ordered sequentially from first to last according to the order in which the S events occurred in the second sample video data.

[0219] Optionally, the first training dataset and the second training dataset can be the same dataset or different datasets; that is, the first sample video data and the second sample video data can be the same data or different data.

[0220] Step 607: Input the second sample video data into the frame-level spatial encoder in the second event description model to obtain the frame feature representation of the video frames in the second sample video data.

[0221] In some embodiments, the second sample video data includes at least two video frames.

[0222] Optionally, at least two video frames from the second sample video data are input into the frame-level spatial encoder in the second event description model, and the frame feature representations corresponding to the at least two video frames are output.

[0223] Step 608: Input the event feature representations corresponding to the S events, the frame feature representations of the video frames in the second sample video data, and the frame position feature representations into the attention layer of the event position prediction model to obtain the decoded feature representations corresponding to the S events.

[0224] Optionally, obtain the frame position feature representations corresponding to at least two video frames respectively; concatenate the frame feature representations and the frame position feature representations to obtain the frame concatenation feature representations corresponding to at least two video frames respectively; input the frame concatenation feature representations corresponding to at least two video frames and the event feature representations corresponding to S events respectively into the attention layer in the event position prediction model to obtain the decoding feature representations corresponding to S events respectively.

[0225] Step 609: Input the decoded feature representations corresponding to the S events into the first linear layer to obtain the time period centers corresponding to the S events; input the decoded feature representations corresponding to the S events into the second linear layer to obtain the time period widths corresponding to the S events.

[0226] The first and second linear layers are used to perform linear transformations on the input features.

[0227] To illustrate, let's take the nth event as an example. The decoded feature representation of the nth event is input into the first linear layer. The first linear layer performs a linear transformation on the decoded feature representation based on its internal weight and bias parameters. The correlation between each feature element in the decoded feature representation and the time segment center is learned by the first linear layer. After the linear transformation, the time segment center of the nth event is output. The decoded feature representation of the nth event is then input into the second linear layer. The second linear layer performs a linear transformation on the decoded feature representation based on its internal weight and bias parameters. The correlation between each feature element in the decoded feature representation and the time segment width is learned by the second linear layer. After the linear transformation, the time segment width of the nth event is output.

[0228] Optionally, the first linear layer and the second linear layer include a normalization layer. The normalization layer is used to map the linearly transformed input value to a preset normalization range. That is, the value range of the time period center output by the first linear layer and the value range of the time period width output by the second linear layer are the preset normalization range, wherein the preset normalization range can be implemented as (0,1).

[0229] Step 610: For the time period center and time period width corresponding to the nth event, obtain the positive mask information and the negative mask information.

[0230] The positive mask information includes positive mask values ​​corresponding to at least two video frames, and the negative mask information includes negative mask values ​​corresponding to at least two video frames.

[0231] In some embodiments, based on the time period center and time period width corresponding to the nth event and the frame positions corresponding to at least two video frames in the second video data, the positive mask value and the negative mask value corresponding to at least two video frames are determined respectively.

[0232] Optionally, using the time period center and time period width as distribution parameters, positive mask value distribution curves and negative mask value distribution curves are constructed through preset distribution expressions. The positive mask value distribution curve indicates the mapping relationship between the frame position of a video frame and the positive mask value of the video frame, and the negative mask value distribution curve indicates the mapping relationship between the frame position of a video frame and the negative mask value of the video frame. Based on the frame positions corresponding to at least two video frames, the positive mask values ​​corresponding to at least two video frames are determined in the positive mask value distribution curve; based on the frame positions corresponding to at least two video frames, the negative mask values ​​corresponding to at least two video frames are determined in the negative mask value distribution curve. The sum of the negative mask value and the positive mask value corresponding to a specified video frame is a preset value, such as 1.

[0233] The distribution parameters can be implemented as Gaussian distribution parameters, and the preset distribution expression can be implemented as Gaussian distribution expression.

[0234] Step 611: Perform forward masking processing on the video frames in the second sample video data using the forward masking information to obtain the forward masking feature representation of the video frames.

[0235] Optionally, the frame feature representations of at least two video frames are weighted by the positive mask values ​​corresponding to at least two video frames respectively, wherein the frame feature representation of the t-th video frame is weighted by the positive mask value corresponding to the t-th video frame to obtain the positive mask feature representations corresponding to at least two video frames respectively, where t is a positive integer.

[0236] Step 612: Input the positive mask feature representation into the video-level temporal encoder in the second event description model to obtain the positive mask video feature representation of the video frame in the second sample video data.

[0237] Optionally, the positive mask feature representations corresponding to at least two video frames are input into the video-level temporal encoder in the second event description model, and the positive mask video feature representations corresponding to at least two video frames are output.

[0238] Step 613: Input the positive mask video feature representation and the second description sequence into the description decoder in the second event description model to obtain the first event description corresponding to the nth event.

[0239] Optionally, the second description sequence includes multiple text embeddings, which are used to indicate the second prompt text and the reference event description corresponding to the nth event. The second prompt text guides the description decoder to decode the event description corresponding to the nth event, i.e., the positive mask description. Illustratively, the second prompt text and the reference event description corresponding to the nth event are concatenated and then segmented into words. Then, text embeddings corresponding to each word are generated, thus obtaining the second description sequence.

[0240] Indicatively, the second prompt text includes a second prompt symbol and event count information. The second prompt symbol can be implemented as [MASK], and the event count information is used to guide the description decoder on the number of event descriptions it needs to generate. For example, if the event count information is implemented as "1event", it means that the description decoder is guided to generate 1 event description.

[0241] Optionally, the positive mask video feature representations and the second description sequence corresponding to at least two video frames are input into the description decoder in the first event description model to obtain the first event description corresponding to the nth event.

[0242] To illustrate, taking an autoregressive language model based on a Transformer decoder as an example, the description decoder generates words in the first event description one by one based on the input positive mask video feature representation (as a prefix sequence) and the text embeddings in the second description sequence. During the generation process, the description decoder considers the positive mask video feature representation and existing text information. For example, when generating the i-th text embedding (or word) in the first event description, it focuses on the first i-1 text embeddings in the first description sequence, where i is an integer greater than 1.

[0243] Step 614: Perform inverse masking processing on the video frames in the second sample video data using inverse masking information to obtain the inverse masking feature representation of the video frames.

[0244] Optionally, the frame feature representations of at least two video frames are weighted by the inverse mask values ​​corresponding to at least two video frames respectively, wherein the frame feature representation of the t-th video frame is weighted by the inverse mask value corresponding to the t-th video frame to obtain the inverse mask feature representations corresponding to at least two video frames respectively, where t is an inverse integer.

[0245] Step 615: Input the inverse mask feature representation into the video-level temporal encoder in the second event description model to obtain the inverse mask video feature representation of the video frame in the second sample video data.

[0246] Optionally, the inverse mask feature representations corresponding to at least two video frames are input into the video-level temporal encoder in the second event description model, and the inverse mask video feature representations corresponding to at least two video frames are output.

[0247] Step 616: Input the inverse mask video feature representation and the third description sequence into the description decoder in the second event description model to obtain the second event description corresponding to the events other than the nth event.

[0248] Optionally, the third description sequence includes multiple text embeddings. These text embeddings indicate the third prompt text and the reference event descriptions corresponding to events other than the nth event. The third prompt text guides the description decoder to decode the event descriptions corresponding to events other than the nth event, i.e., the inverse mask description. Illustratively, the third prompt text and the reference event descriptions corresponding to events other than the nth event are concatenated, then segmented into words, and finally the text embeddings corresponding to each word are generated, thus obtaining the third description sequence.

[0249] Indicatively, the third prompt text includes a second prompt and event count information. The second prompt can be implemented as [MASK], and the event count information is used to guide the description decoder on the number of event descriptions it needs to generate. For example, if the event count information is implemented as "S-1events", it means that the description decoder is guided to generate S-1 event descriptions.

[0250] Optionally, the inverse mask video feature representations corresponding to at least two video frames and the third description sequence are input into the description decoder in the first event description model to obtain the second event description corresponding to the events other than the nth event.

[0251] To illustrate, taking an autoregressive language model based on a Transformer decoder as an example, the description decoder generates words in the second event description one by one based on the input inverse masked video feature representation (as a prefix sequence) and the text embeddings in the third description sequence. During the generation process, the description decoder considers the inverse masked video feature representation and existing text information. For example, when generating the i-th text embedding (or word) in the second event description, it focuses on the first i-1 text embeddings in the third description sequence, where i is an integer greater than 1.

[0252] Step 617: Determine the diversity loss based on the similarity between the positive mask information corresponding to the nth event position and the positive mask information corresponding to events other than the nth event; determine the positive mask loss based on the difference between the reference event description and the first event description corresponding to the nth event; determine the inverse mask loss based on the difference between the reference event description and the second event description corresponding to events other than the nth event; train the event position prediction model and the first event description model based on the positive mask loss, the inverse mask loss, and the diversity loss.

[0253] Indicatively, the sum or weighted sum of the positive masking loss, inverse masking loss, and diversity loss is calculated as the total loss. The model parameters of the video-level temporal encoder and descriptive decoder in the event location prediction model and the first event description model are updated based on this total loss, thus completing one training cycle corresponding to one second sample video data. This training process is repeated using multiple second sample video data until the calculated sum of losses is less than a preset loss, or the preset number of training iterations is reached. Training then stops, and the resulting location prediction model and event description model are used subsequently; that is, the target event location prediction model and the target event description model.

[0254] This is illustrative; please refer to it. Figure 8 It shows a schematic diagram of a positive mask description generation.

[0255] (1) Spatial-temporal encoded video mask coding

[0256] like Figure 8 As shown, the first event description model includes a frame-level spatial encoder 801 and a video-level temporal encoder 802. The frame-level spatial encoder 801 uses E... a This indicates that the video-grade time encoder 802 uses E... b Indicates. Given sample video data (That is, video frame 803), where H and W represent the height and width of each video frame, respectively, and N v This represents the number of video frames in the sample video data.

[0257] The video frame 803 is input into the frame-level spatial encoder 801, which extracts the visual embedding from each frame to generate the frame embedding. That is: Among them, v t Let d represent the t-th video frame, and d be the embedding dimension. In some embodiments, the parameters of the frame-level spatial encoder 801 remain unchanged during training.

[0258] Next, for the nth event, the positive mask information corresponding to the nth event is obtained through the differentiable mask construction process. and inverse mask information The process of constructing a differentiable mask can be referred to in step 330. Figure 4 The details of that will not be elaborated here.

[0259] After obtaining the positive mask information corresponding to the nth event Then, based on the positive mask information Perform positive masking on frame embedding 804 to obtain positive mask embedding 805 corresponding to the nth event, that is: Then, the positive mask embedding 805 corresponding to the nth event is input into the video-level time encoder 802 to obtain the positive mask video embedding corresponding to the nth event. That is:

[0260] like Figure 9 As shown, after obtaining the inverse mask information corresponding to the nth event... Then, based on the inverse mask information Perform inverse masking on frame embedding 804 to obtain inverse masking embedding 807 corresponding to the nth event, that is: Then, the inverse mask embedding 807 corresponding to the nth event is input into the video-level time encoder 802 to obtain the inverse mask video embedding corresponding to the nth event. That is:

[0261] (2) Decoding based on prompt-based positive mask description

[0262] like Figure 8 As shown, the second event description model also includes a description decoder 806, which processes the orthogonal mask video embedding. This generates an event description for the nth event in the sample video data v. For each sample video data v, there is a given description C for the nth event. n (Refer to the event description). A prompt will be provided. “[MASK]1event:” and C n Connect them into a paragraph, then segment and embed them to form the nth second description sequence. That is:

[0263]

[0264] In the prompt In this context, [MASK] indicates the model generation part description. 1event specifies that the number of event descriptions is 1. Indicates the second description sequence The number of word segments (or the number of text embeddings) in the text.

[0265] In the description of decoder 806, the positive mask video is embedded. As a prefix visual marker, and together with the second descriptive sequence The concatenation yields the nth positive mask input sequence, which is: Input sequence with positive mask Input description decoder 806, description decoder 806 starts decoding from the position corresponding to r2, for the i-th text embedding Description decoder 806 based and the embedding of the first i-1 reference texts (i.e., the second description sequence) The first i-1 text embeddings in The dictionary is used to predict the i-th text embedding or the i-th probability distribution, where the i-th probability distribution includes the probability that a word in the dictionary is the i-th text embedding. If i is greater than 2, the description decoder 806 will base its prediction on... Predict the i-th predicted text embedding or the i-th probability distribution by combining it with the first i-1 predicted text embeddings (or the first i-1 probability distributions).

[0266] It should be noted that, in In the middle, r1 to r3 correspond to the prompts. This indicates the text embedding corresponding to [MASK]. This indicates the text embedding corresponding to 1. This indicates the text embedding corresponding to the event.

[0267] (3) Decoding based on prompt-based inverse mask description

[0268] like Figure 9 As shown, the decoder 806 processes inverse masked video embeddings. This generates event descriptions for all events in the sample video data v except for the nth event. For each sample video data v, there are given descriptions for all events except the nth event. (Refer to the event description). A prompt will be provided. "[MASK]N" S -1events:”and Connect them into a paragraph, then segment and embed them to form the nth third description sequence. That is:

[0269]

[0270] In the prompt In the text, [MASK] indicates the model generation part of the description. S -1events specifies the number of event descriptions, which is N. S -1. Indicates the third description sequence The number of word segments (or the number of text embeddings) in the text.

[0271] In the description decoder 806, the inverse masked video is embedded. As a prefix visual marker, and together with the second descriptive sequence The concatenation yields the nth inverse mask input sequence, which is: Input the inverse mask sequence Input description decoder 806, description decoder 806 starts decoding from the position corresponding to r2, for the i-th text embedding Description decoder 806 based and the embedding of the first i-1 reference texts (i.e., the third description sequence) The first i-1 text embeddings in The dictionary is used to predict the i-th text embedding or the i-th probability distribution, where the i-th probability distribution includes the probability that a word in the dictionary is the i-th text embedding. If i is greater than 2, the description decoder 806 will base its prediction on... Predict the i-th predicted text embedding or the i-th probability distribution by combining it with the first i-1 predicted text embeddings (or the first i-1 probability distributions).

[0272] It should be noted that, in In the middle, r1 to r3 correspond to the prompts. This indicates the text embedding corresponding to [MASK]. N represents S -1 corresponds to the text embedding. This indicates the text embedding corresponding to the event.

[0273] (4) Mask description generates training loss

[0274] For positive mask description generation, the training objective is to minimize the negative log-likelihood, and the corresponding loss function is... As shown in Formula Six below:

[0275]

[0276] in, For embedding the positive mask video corresponding to the nth event, N S This indicates the number of events in the sample video data. For the nth second description sequence The number of text embeddings in N2. N2 represents the description of decoder 806 for N. S The total number of text embeddings predicted for each event. θ E,G This represents the frame-level spatial encoder 801, the video-level temporal encoder 802, the description decoder 806, and the event location prediction model (see reference). Figure 4 The model parameters. Indicates that in a given In the case of the previous text embedding (text embedding in the second description sequence), the description decoder 806 decodes to obtain The logarithmic probability.

[0277] For inverse mask description generation, the training objective is to minimize the negative log-likelihood, and the corresponding loss function is... As shown in Formula 7 below:

[0278]

[0279] in, For the inverse masked video embedding corresponding to the nth event, N S This indicates the number of events in the sample video data. For the nth third description sequence The number of text embeddings in N2. N2 represents the description of decoder 806 for N. S The total number of text embeddings predicted for each event. θ E,G This represents the frame-level spatial encoder 801, the video-level temporal encoder 802, the description decoder 806, and the event location prediction model (see reference). Figure 4 The model parameters. Indicates that in a given In the case of the previous text embedding (text embedding in the third description sequence), the description decoder 806 decodes to obtain The logarithmic probability.

[0280] To ensure that the positive mask information covers different parts of the video, a diversity loss based on cosine similarity is introduced, and the loss function is shown in Equation 8 below:

[0281]

[0282] in, γ is a hyperparameter that controls the overlap between the positive mask information corresponding to the event.

[0283] The positive mask loss is obtained through loss function L2, the negative mask loss is obtained through loss function L3, and the loss function is obtained through... After obtaining the diversity loss, the three losses are added together to obtain the loss sum. The model parameters of the video-level temporal encoder 802, the description decoder 806, and the event location prediction model are updated based on the loss sum. The above training process is performed on the first event description model and the event location prediction model using multiple sample video data until the calculated loss sum is less than the preset loss or the number of training iterations reaches the preset number. Training is then stopped. The event description model obtained at this time is the target event description model, and the event location prediction model obtained is the target event location prediction model.

[0284] The video description generation method provided in this application simplifies the complex event location annotation process by introducing a complementary masking mechanism, reduces the reliance on cumbersome event location determination steps, improves training efficiency, reduces computational costs, and makes the training process more efficient under weak supervision.

[0285] Next, we will introduce the application scheme of the video description model provided in this application.

[0286] Figure 10 This is a flowchart illustrating a method for generating a video description according to an embodiment of this application. The method is executed by a computer device, which may be... Figure 2 The terminal 210 and / or server 220 are shown. The method includes steps 1010 to 1030.

[0287] Step 1010: Obtain the first video data.

[0288] The first video data includes at least two video frames.

[0289] Step 1020: Predict the event descriptions corresponding to the R events in the first video data using the target event description model.

[0290] R is a positive integer.

[0291] In some embodiments, the target event description model includes a frame-level spatial encoder, a video-level temporal encoder, and a description decoder.

[0292] Optionally, at least two video frames from the first video data are input into a frame-level spatial encoder in the target event description model to obtain frame feature representations corresponding to the at least two video frames respectively; frame position feature representations corresponding to the two video frames are obtained respectively; the frame position feature representations and frame feature representations are concatenated to obtain frame concatenated feature representations corresponding to the at least two video frames respectively; and the frame concatenated feature representations corresponding to the at least two video frames are input into a video-level temporal encoder in the target event description model to obtain video feature representations corresponding to the at least two video frames respectively.

[0293] Then, the video feature representations corresponding to at least two video frames and the first cue are input into the description decoder in the target event description model to obtain the event descriptions corresponding to R events. Illustratively, the first cue can be implemented as [FULL]. Under the [FULL] cue, the description decoder first predicts the number of event descriptions (or the number of events) in the first video data. After predicting the number R of event descriptions, the description decoder begins to predict the event descriptions corresponding to the R events in sequence.

[0294] Step 1030: Predict the event locations of R events in the first video data using the target event location prediction model.

[0295] Event location is used to express the time period of the event within the first video data.

[0296] The target event description model and the target event location prediction model are trained using sample video data and the corresponding reference event descriptions.

[0297] In some embodiments, the event location includes the time period center and time period width corresponding to the time period of the event in the first video data.

[0298] Optionally, the event feature representations corresponding to the R events are obtained; the target event location prediction model is used to analyze the event feature representations corresponding to the R events, the frame feature representations corresponding to the video frames in the first video data, and the frame location feature representations to generate the decoding feature representations corresponding to the R events; the target event location prediction model is used to analyze the decoding feature representations corresponding to the R events to determine the time period center and time period width corresponding to the R events.

[0299] Schematic, the target event description model includes a decoder, a first linear layer, and a second linear layer; it concatenates frame feature representations and frame position feature representations to obtain frame concatenation feature representations corresponding to at least two video frames; it inputs the event feature representations corresponding to R events and the frame concatenation feature representations corresponding to at least two video frames into the decoder in the target event description model to obtain decoded feature representations corresponding to R events; it inputs the decoded feature representations corresponding to R events into the first linear layer to obtain the time period centers corresponding to R events; and it inputs the decoded feature representations corresponding to R events into the second linear layer to obtain the time period widths corresponding to R events.

[0300] In the above scheme, by representing the event location as a time segment center and a time segment width, the temporal location of the event in the video can be described more accurately. The time segment center determines the approximate time point of the event, while the time segment width reflects the duration of the event. This representation provides a more precise reference for determining video events, helping the model to more accurately understand and describe the time of event occurrence.

[0301] Optionally, after obtaining the time period center and time period width corresponding to each of the R events, the time period corresponding to each of the R events is determined, and the time period is used to indicate the start and end time of the time.

[0302] Indicatively, let μ be the center of the time interval for the nth event. n The time period width is σ n The starting time is s n The termination time is e nThen the normalized start time of the nth event is... The termination time is It should be noted that if the calculation result s n If s < 0, then let s n =0; if e n >1, then let e n =1.

[0303] Finally, s n e n By multiplying each timeline by the duration of the first video data, we can obtain the start time and end time, which are also the timestamps of the nth event.

[0304] In some embodiments, after predicting the event positions of R events in the first video data using the target event position prediction model, the method further includes: for the j-th event position, obtaining the positive mask feature representation of the video frame in the first video data, where 0 < j ≤ R; and analyzing the positive mask feature representation using the target event description model to obtain the target event description corresponding to the j-th event.

[0305] Optionally, based on the j-th event position and the frame position of the video frame in the first video data, the positive mask value corresponding to the video frame in the first video data is determined; the feature representation of the t-th frame is weighted by the positive mask value corresponding to the t-th video frame to obtain the t-th positive mask feature representation, where t-th is a positive integer.

[0306] Schematic, the positive mask information corresponding to the j-th event is determined based on the location of the j-th event and the frame locations corresponding to at least two video frames in the first video data. The positive mask information includes the positive mask values ​​corresponding to at least two video frames. Taking the positive mask value corresponding to the t'-th video frame as an example, the feature representation of the t'-th frame is weighted and processed using the positive mask value corresponding to the t'-th video frame to obtain the t'-th positive mask feature representation. The location of the j-th event includes the time segment center and time segment width corresponding to the j-th event in the first video data.

[0307] Optionally, a first distance is determined between the frame position of the video frame and the center of the time segment; based on the ratio between the first distance and the width of the time segment, the positive mask value of the video frame is determined. Details regarding the specific generation of the positive mask feature representation can be found in the training-side scheme described above (e.g., step 330), and are not limited here.

[0308] In the above scheme, the positive mask value determination method based on distance and width can reasonably set the positive mask value, so that the model can more accurately focus on the video frame features related to the event when generating event descriptions.

[0309] Taking the j-th event as an example, after obtaining the positive mask feature representation of the j-th event, this positive mask feature representation is input into the video-level temporal encoder in the target event description model to obtain the positive mask video feature representation. The positive mask video feature representation and the second prompt text are then input into the description decoder in the target event description model to obtain the target event description corresponding to the j-th event. Illustratively, the second prompt text can be implemented as [MASK]1event. Under the [MASK]1event prompt, the description decoder directly predicts the target event description corresponding to the j-th event. The target event description is an event description generated based on the positive mask feature representation of the j-th event. The positive mask feature representation of the j-th event can highlight the video frame features related to the j-th event, enabling the model to more accurately capture the detailed information related to the event when generating the event description, thereby improving the accuracy and quality of the event description.

[0310] Taking R=3 as an example, the positive mask information corresponding to the three events is applied to the first video data to obtain three masked video embeddings. These three masked video embeddings and "[MASK]1event" are then input into the description decoder again to generate positive descriptions, thereby regenerating the description of each event and improving the description quality. The inputs to the description decoder are: masked video embedding 1 + "[MASK]1event", masked video embedding 2 + "[MASK]1event", and masked video embedding 3 + "[MASK]1event". The corresponding outputs of the description decoder are: "[MASK]1event: event description 1", "[MASK]1event: event description 2", and "[MASK]1event: event description 3".

[0311] The video description generation method provided in this application, by using a trained target event description model and a target event location prediction model, can more accurately predict event descriptions and event locations in the first video data.

[0312] The video description generation method provided in this application can be widely applied to various scenarios such as video understanding, video summarization, and video search. For example, online short video platforms can use this application to generate a summary title for each segment of a video to help users quickly understand the video content or jump to segments of interest; online education platforms can use this application to generate temporally localized descriptions for video courses to help students better understand the course content; and video monitoring systems can use this application to automatically generate event descriptions, improving the readability and retrieval efficiency of monitored videos.

[0313] To illustrate, taking an online short video platform as an example, the method for generating video descriptions in this application further includes the following steps:

[0314] Step 1: Receive the description generation operation for the target video on the video playback interface.

[0315] As an illustration, the terminal runs a client of an online short video platform, which displays a video playback interface used to play the target video.

[0316] Optionally, the video playback interface includes a description generation control; receiving a trigger operation on the description generation control is used as a description generation operation for the target video; or, the video playback interface includes an interactive interface, receiving a sending operation for a specified message on the interactive interface is used as a description generation operation for the target video, wherein the specified message is used to trigger the generation of description text for the target video, and illustratively, the specified message can be implemented as "@Account1 Please generate video description".

[0317] Step 2: In response to the description generation operation, predict the event descriptions corresponding to the P events of the target video using the target event description model, and predict the time period center and time period width corresponding to the P events using the target event location prediction model.

[0318] P is a positive integer. Illustratively, in response to the description generation operation, the target event description model and target event location prediction model deployed in the client are invoked; or the target event description model and target event location prediction model deployed in a remote server are invoked. The target event description model is used to predict the event description of each event in the target video, and for each event, to determine the time segment center and time segment width corresponding to its occurrence time in the target video.

[0319] Optionally, for each of the P events, the positive mask feature representation of the video frame in the target video is obtained, where 0 < j ≤ R; the target event description is obtained by analyzing the positive mask feature representation through the target event description model. That is, a more refined event description is generated through the positive mask feature representation, improving the accuracy of the event description.

[0320] Step 3: Determine the timestamps corresponding to each of the P events based on the time period center and time period width.

[0321] In a schematic way, for each event, the start and end times of the event are determined based on the center and width of the time period, which are also the timestamps of the time.

[0322] Step 4: Determine the video description text based on the event description and timestamp corresponding to each of the P events.

[0323] After determining the event descriptions and timestamps corresponding to each of the P events, the P events are sorted from front to back according to their occurrence order in the target video, and a corresponding event stamp is added to each event to obtain the video description text. An example of a video description text is as follows:

[0324] "0:00-0:15: The chef is chopping onions and carrots;

[0325] 0:15-0:30: The water in the pot begins to heat up, and steam is emitted;

[0326] 0:30-0:45: The chef puts the chopped vegetables into the pot and stir-fries them, with flames churning.

[0327] 0:45-1:00: The stir-fried dishes are served in bowls, looking and smelling delicious;

[0328] 1:00-1:10: The dishes are served and ready to be enjoyed.

[0329] Step 5: Display the video description text on the video playback interface.

[0330] After obtaining the video description text, display the video description text on the video playback interface, such as displaying the video description text in the introduction area of ​​the video playback interface or displaying the video description text in the interactive area of ​​the video playback interface, such as sending the video description text to the interactive area in the form of replying to a specified message.

[0331] This is illustrative; please refer to it. Figure 11 It shows a structural block diagram of a video description generation device, such as Figure 11 As shown, the device includes:

[0332] The first acquisition module 1110 is used to acquire reference event descriptions corresponding to S events in the sample video data, wherein the reference event descriptions are used to describe the event content, and S is a positive integer.

[0333] Model training module 1120 is used to predict S event locations of the sample video data through an event location prediction model, wherein the nth event location is used to express the time period of the nth event in the sample video data, 0 < n ≤ S;

[0334] The model training module 1120 is used to obtain the positive mask feature representation and the negative mask feature representation of the video frame in the sample video data for the nth event position;

[0335] The model training module 1120 is used to analyze the positive mask feature representation through the first event description model to obtain a first event description corresponding to the nth event; and to analyze the inverse mask feature representation through the first event description model to obtain a second event description corresponding to events other than the nth event.

[0336] The model training module 1120 is used to train the event location prediction model and the first event description model based on S reference event descriptions, the first event description, and the second event description.

[0337] In some embodiments, the model training module 1120 is configured to: generate frame feature representations of video frames in the sample video data; determine the positive mask value and inverse mask value corresponding to the video frame in the sample video data based on the nth event position and the frame position of the video frame in the sample video data; weight the tth frame feature representation with the positive mask value corresponding to the tth video frame to obtain the tth positive mask feature representation, where t is a positive integer; and weight the tth frame feature representation with the inverse mask value corresponding to the tth video frame to obtain the tth inverse mask feature representation.

[0338] In some embodiments, the location of the nth event includes the time period center and the time period width corresponding to the time period of the nth event in the sample video data; the model training module 1120 is configured to: determine a first distance between the frame position of the video frame and the time period center; determine the positive mask value of the video frame based on the ratio between the first distance and the time period width; and obtain the difference between a preset value and the positive mask value as the inverse mask value.

[0339] In some embodiments, the model training module 1120 is configured to: obtain the number of video frames in the sample video data; and determine a first distance between the frame position of the video frame and the center of the time period based on the ratio between the frame position of the video frame and the number of video frames.

[0340] In some embodiments, the model training module 1120 is configured to: obtain the event feature representations corresponding to the S events respectively; input the event feature representations corresponding to the S events respectively, the frame feature representations and the frame position feature representations corresponding to the video frames in the sample video data into the event position prediction model, and output the S event positions of the sample video data.

[0341] In some embodiments, the event location prediction model includes a decoder, a first linear layer, and a second linear layer; the model training module 1120 is configured to: input the event feature representations corresponding to the S events, the frame feature representations corresponding to the video frames in the sample video data, and the frame location feature representations into the decoder, and output the decoded feature representations corresponding to the S events; input the decoded feature representations corresponding to the S events into the first linear layer, and output the time period centers corresponding to the S events; input the decoded feature representations corresponding to the S events into the second linear layer, and output the time period widths corresponding to the S events.

[0342] In some embodiments, the S reference event descriptions include a reference event description corresponding to the nth event and reference event descriptions corresponding to events other than the nth event; the model training module 1120 is configured to: determine a positive masking loss based on the difference between the reference event description corresponding to the nth event and the first event description; determine an inverse masking loss based on the difference between the reference event descriptions corresponding to events other than the nth event and the second event description; and train the event location prediction model and the first event description model based on the positive masking loss and the inverse masking loss.

[0343] In some embodiments, the model training module 1120 is configured to: determine a diversity loss based on the similarity between the positive mask information corresponding to the nth event location and the positive mask information corresponding to events other than the nth event, wherein the positive mask information includes the positive mask value corresponding to the video frame in the sample video data; and train the event location prediction model and the first event description model based on the positive mask loss, the inverse mask loss and the diversity loss.

[0344] In some embodiments, the model training module 1120 is configured to: acquire video feature representations of video frames in target sample video data, wherein the target sample video data corresponds to K reference event descriptions for each of the events, where K is a positive integer; analyze the video feature representations using a second event description model to obtain third event descriptions for each of the K events; and train the second event description model based on the differences between the reference event descriptions for each of the K events and the third event descriptions for each of the K events to obtain the first event description model.

[0345] In some embodiments, the model training module 1120 is configured to: generate frame feature representations and frame position feature representations of video frames in the target sample video data; and determine the video feature representation of the video frame based on the frame feature representations and the frame position feature representations.

[0346] In summary, the video description generation apparatus provided in this application automatically generates event locations in sample video data through an event location prediction model, eliminating reliance on manually labeled locations and reducing labeling costs. For each event location, a positive mask feature representation and an inverse mask feature representation are obtained. The first event description model uses the positive mask feature representation to focus on video frame features related to a specified event, thereby generating a first event description corresponding to the specified event. It uses the inverse mask feature representation to focus on video frame features related to other events, thereby generating a second event description corresponding to other events. By comparing the differences between S reference event descriptions and the first and second event descriptions, the accuracy of location prediction is indirectly evaluated. This allows the event location prediction model to automatically adjust event location predictions, ensuring that the event descriptions generated by the first event description model accurately correspond to the corresponding event locations. Even without provided location labels, this achieves effective implicit alignment between event locations and event descriptions, improving model training efficiency.

[0347] This is illustrative; please refer to it. Figure 12 It shows a structural block diagram of a video description generation device, such as Figure 12 As shown, the device includes:

[0348] The second acquisition module 1210 is used to acquire the first video data;

[0349] Model prediction module 1220 is used to predict the event descriptions corresponding to R events of the first video data through the target event description model, where R is a positive integer;

[0350] The model prediction module 1220 is used to predict the event positions of the R events in the first video data through a target event position prediction model, wherein the event position is used to express the time period of the event in the first video data;

[0351] The target event description model and the target event location prediction model are obtained by training sample video data and the reference event description corresponding to the sample video data.

[0352] In some embodiments, the event location includes the time period center and time period width corresponding to the time period of the event in the first video data; the model prediction module 1220 is configured to: obtain the event feature representations corresponding to the R events respectively; analyze the event feature representations corresponding to the R events respectively, the frame feature representations corresponding to the video frames in the first video data, and the frame position feature representations through the target event location prediction model to generate the decoding feature representations corresponding to the R events respectively; analyze the decoding feature representations corresponding to the R events respectively through the target event location prediction model to determine the time period center and time period width corresponding to the R events respectively.

[0353] In some embodiments, the model prediction module 1220 is configured to: for the j-th event location, obtain the positive mask feature representation of the video frame in the first video data, 0 < j ≤ R; and analyze the positive mask feature representation through the target event description model to obtain the target event description corresponding to the j-th event.

[0354] In some embodiments, the model prediction module 1220 is configured to: generate frame feature representations of video frames in the first video data; determine the positive mask value corresponding to the video frame in the first video data based on the j-th event position and the frame position of the video frame in the first video data; and weight the t`-th frame feature representation with the positive mask value corresponding to the t`-th video frame to obtain the t`-th positive mask feature representation, where t` is a positive integer.

[0355] In some embodiments, the location of the j-th event includes the time period center and time period width corresponding to the time period of the j-th event in the first video data; the model prediction module 1220 is configured to: determine a first distance between the frame position of the video frame and the time period center; and determine the positive mask value of the video frame based on the ratio between the first distance and the time period width.

[0356] In summary, the video description generation apparatus provided in this application simplifies the complex event location annotation process by introducing a complementary masking mechanism, reduces the reliance on cumbersome event location determination steps, improves training efficiency, reduces computational costs, and makes the training process more efficient under weak supervision.

[0357] It should be noted that the specific limitations of the embodiments of the one or more video description generation devices provided above can be found in the limitations of the video description generation method above, and will not be repeated here. Each module of the above device can be implemented entirely or partially by software, hardware, or a combination thereof. Each module can be embedded in the processor of the computer device in hardware form or independent of the processor, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0358] This application also provides a computer device, which includes: a processor and a memory, wherein the memory stores a computer program; the processor is used to execute the computer program in the memory to implement the video description generation method provided in the above method embodiments.

[0359] For example, Figure 13 This is a structural block diagram of a computer device 1300 provided in an exemplary embodiment of this application. Typically, the computer device 1300 includes a processor 1301 and a memory 1302.

[0360] Processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1301 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 1301 may also include a main processor and a coprocessor. The main processor, also known as a Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1301 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1301 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.

[0361] The memory 1302 may include one or more computer-readable storage media, which may be non-transitory. The memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 are used to store at least one instruction, which is executed by the processor 1301 to implement the video description generation method provided in the various method embodiments of this application.

[0362] In some embodiments, the computer device 1300 may optionally include an input interface 1303 and an output interface 1304. The processor 1301, memory 1302, and input interfaces 1303 and 1304 can be connected via a bus or signal lines. Various peripheral devices can be connected to the input interfaces 1303 and 1304 via a bus, signal lines, or a circuit board. The input interfaces 1303 and 1304 can be used to connect at least one input / output (I / O) related peripheral device to the processor 1301 and memory 1302. In some embodiments, the processor 1301, memory 1302, and input interfaces 1303 and 1304 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1301, memory 1302, and input interfaces 1303 and 1304 can be implemented on separate chips or circuit boards, and this application does not limit this.

[0363] Those skilled in the art will understand that Figure 13 The structure shown does not constitute a limitation on the computer device 1300, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0364] In an exemplary embodiment, this application provides a chip including programmable logic circuits and / or program instructions, which, when run on a computer device, are used to implement the video description generation method provided in the above-described method embodiments.

[0365] In an exemplary embodiment, this application provides a computer-readable storage medium storing a computer program that is loaded and executed by a processor to implement the video description generation method provided in the above-described method embodiments.

[0366] In an exemplary embodiment, this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the processor of the computer device to load and execute the video description generation method provided in the above-described method embodiments.

[0367] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0368] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0369] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0370] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for generating video descriptions, characterized in that, The method includes: Obtain reference event descriptions corresponding to S events in the sample video data. The reference event descriptions are used to describe the event content, and S is a positive integer. The event location prediction model predicts S event locations in the sample video data, where the nth event location is used to express the time period of the nth event in the sample video data, and 0 < n ≤ S; For the nth event location, obtain the positive mask feature representation and the negative mask feature representation of the video frame in the sample video data; The first event description corresponding to the nth event is obtained by analyzing the positive mask feature representation using the first event description model; and the second event description corresponding to events other than the nth event is obtained by analyzing the negative mask feature representation using the first event description model. The event location prediction model and the first event description model are trained based on S reference event descriptions, the first event description, and the second event description.

2. The method according to claim 1, characterized in that, For the nth event location, obtaining the positive and negative mask feature representations of the video frames in the sample video data includes: Generate frame feature representations of video frames in the sample video data; Based on the position of the nth event and the frame position of the video frame in the sample video data, determine the positive mask value and the negative mask value corresponding to the video frame in the sample video data; The feature representation of the t-th frame is weighted by the positive mask value corresponding to the t-th video frame to obtain the positive mask feature representation of the t-th frame, where t is a positive integer; The feature representation of the t-th frame is weighted by the inverse mask value corresponding to the t-th video frame to obtain the inverse mask feature representation of the t-th frame.

3. The method according to claim 2, characterized in that, The location of the nth event includes the time period center and time period width corresponding to the time period of the nth event in the sample video data; The step of determining the positive and negative mask values ​​corresponding to the video frames in the sample video data based on the position of the nth event and the frame position of the video frames in the sample video data includes: Determine the first distance between the frame position of the video frame and the center of the time period; Based on the ratio between the first distance and the time period width, the positive mask value of the video frame is determined; The difference between the preset value and the positive mask value is obtained as the inverse mask value.

4. The method according to claim 3, characterized in that, Determining the first distance between the frame position of the video frame and the center of the time period includes: Obtain the number of video frames in the sample video data; Based on the ratio between the frame position of the video frame and the number of video frames, a first distance between the frame position of the video frame and the center of the time period is determined.

5. The method according to any one of claims 1 to 4, characterized in that, The prediction of S event locations in the sample video data using the event location prediction model includes: Obtain the event feature representations corresponding to the S events respectively; The event feature representations corresponding to the S events, the frame feature representations corresponding to the video frames in the sample video data, and the frame position feature representations are input into the event position prediction model, and the S event positions of the sample video data are output.

6. The method according to claim 5, characterized in that, The event location prediction model includes a decoder, a first linear layer, and a second linear layer. The step of inputting the event feature representations corresponding to the S events, the frame feature representations corresponding to the video frames in the sample video data, and the frame position feature representations into the event position prediction model, and outputting the S event positions of the sample video data, includes: The event feature representations corresponding to the S events, the frame feature representations corresponding to the video frames in the sample video data, and the frame position feature representations are input into the decoder, and the decoded feature representations corresponding to the S events are output. The decoded feature representations corresponding to the S events are input into the first linear layer, and the time period centers corresponding to the S events are output. The decoded feature representations corresponding to the S events are input into the second linear layer, and the time period widths corresponding to the S events are output.

7. The method according to any one of claims 1 to 4, characterized in that, The S reference event descriptions include the reference event description corresponding to the nth event and the reference event descriptions corresponding to events other than the nth event; The process of training the event location prediction model and the first event description model based on S reference event descriptions, the first event description, and the second event description includes: The positive mask loss is determined based on the difference between the reference event description corresponding to the nth event and the first event description; The inverse mask loss is determined based on the difference between the reference event description and the second event description corresponding to events other than the nth event. The event location prediction model and the first event description model are trained based on the positive mask loss and the inverse mask loss.

8. The method according to claim 7, characterized in that, The method further includes: Based on the similarity between the positive mask information corresponding to the nth event location and the positive mask information corresponding to events other than the nth event, the diversity loss is determined, wherein the positive mask information includes the positive mask value corresponding to the video frame in the sample video data; The step of training the event location prediction model and the first event description model based on the positive mask loss and the inverse mask loss includes: The event location prediction model and the first event description model are trained based on the positive masking loss, the inverse masking loss, and the diversity loss.

9. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain video feature representations of video frames in target sample video data, wherein the target sample video data corresponds to K reference event descriptions for each of the events, where K is a positive integer; By analyzing the video feature representation using the second event description model, third event descriptions corresponding to the K events are obtained respectively; The second event description model is trained based on the differences between the reference event descriptions corresponding to the K events and the third event descriptions corresponding to the K events, thereby obtaining the first event description model.

10. The method according to claim 9, characterized in that, The acquisition of video feature representations of video frames in the target sample video data includes: Generate frame feature representations and frame position feature representations of video frames in the target sample video data; Based on the frame feature representation and the frame position feature representation, the video feature representation of the video frame is determined.

11. A method for generating video descriptions, characterized in that, The method includes: Obtain the first video data; Predict the event descriptions corresponding to R events in the first video data using the target event description model, where R is a positive integer; The target event location prediction model predicts the event locations of the R events in the first video data, where the event location is used to express the time period of the event in the first video data; The target event description model and the target event location prediction model are obtained by training sample video data and the reference event description corresponding to the sample video data.

12. The method according to claim 11, characterized in that, The event location includes the time period center and time period width corresponding to the time period of the event in the first video data; The step of predicting the event locations of the R events in the first video data using a target event location prediction model includes: Obtain the event feature representations corresponding to the R events respectively; The target event location prediction model is used to analyze the event feature representations corresponding to the R events, the frame feature representations corresponding to the video frames in the first video data, and the frame location feature representations to generate the decoding feature representations corresponding to the R events. By analyzing the decoded feature representations corresponding to the R events through the target event location prediction model, the time period center and time period width corresponding to the R events are determined.

13. The method according to claim 11 or 12, characterized in that, After predicting the event locations of the R events in the first video data using the target event location prediction model, the method further includes: For the j-th event position, obtain the positive mask feature representation of the video frame in the first video data, where 0 < j ≤ R; The target event description corresponding to the j-th event is obtained by analyzing the positive mask feature representation through the target event description model.

14. The method according to claim 13, characterized in that, The step of obtaining the positive mask feature representation of the video frame in the first video data for the j-th event position includes: Generate frame feature representations of the video frames in the first video data; Based on the position of the j-th event and the frame position of the video frame in the first video data, determine the positive mask value corresponding to the video frame in the first video data; The feature representation of the t-th frame is weighted by the positive mask value corresponding to the t-th video frame to obtain the positive mask feature representation of the t-th frame, where t is a positive integer.

15. The method according to claim 14, characterized in that, The location of the j-th event includes the time period center and time period width corresponding to the time period of the j-th event in the first video data; The step of determining the positive mask value corresponding to the video frame in the first video data based on the j-th event position and the frame position of the video frame in the first video data includes: Determine the first distance between the frame position of the video frame and the center of the time period; The positive mask value of the video frame is determined based on the ratio between the first distance and the time period width.

16. A video description generation apparatus, characterized in that, The device includes: The first acquisition module is used to acquire reference event descriptions corresponding to S events in the sample video data, wherein the reference event descriptions are used to describe the event content, and S is a positive integer. The model training module is used to predict S event locations in the sample video data through the event location prediction model, wherein the nth event location is used to express the time period of the nth event in the sample video data, and 0 < n ≤ S; The model training module is used to obtain the positive mask feature representation and the negative mask feature representation of the video frame in the sample video data for the nth event position; The model training module is used to analyze the positive mask feature representation through the first event description model to obtain a first event description corresponding to the nth event; and to analyze the inverse mask feature representation through the first event description model to obtain a second event description corresponding to events other than the nth event. The model training module is used to train the event location prediction model and the first event description model based on S reference event descriptions, the first event description, and the second event description.

17. A video description generation apparatus, characterized in that, The device includes: The second acquisition module is used to acquire the first video data; The model prediction module is used to predict the event descriptions corresponding to R events of the first video data through the target event description model, where R is a positive integer; The model prediction module is used to predict the event positions of the R events in the first video data through a target event position prediction model, wherein the event position is used to express the time period of the event in the first video data; The target event description model and the target event location prediction model are obtained by training sample video data and the reference event description corresponding to the sample video data.

18. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the video description generation method as described in any one of claims 1 to 15.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one program, which is loaded and executed by a processor to implement the video description generation method as described in any one of claims 1 to 15.

20. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method for generating a video description as described in any one of claims 1 to 15.