Video description text generation method and device, computer equipment and storage medium
By acquiring target video feature data, a multi-path feature extraction model and cross-attention fusion technology are used to generate coherent video description text, which solves the problem of fragmented temporal information and improves the coherence and quality of video description text.
Patent Information
- Application Number
- CN202510950480.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-11-07
AI Technical Summary
Existing methods for generating video description text struggle to achieve coherence in the face of fragmented temporal information, resulting in poor quality of the generated video description text.
By acquiring feature data of the target video, the first path feature extraction model outputs the current frame features, the second path feature extraction model outputs the future frame features, and the database is updated based on the future frame features. The current frame features and database features are then fused together using cross-attention, and finally, a coherent video description text is generated through a text decoder.
It achieves the generation of temporally coherent video description text, solves the problem of fragmented temporal information, and improves the quality of video description text.
Smart Images

Figure CN120912894A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, and is suitable for the fields of finance and medicine, and in particular relates to a video description text generation method and device, computer equipment and a storage medium. BACKGROUND
[0002] In recent years, short video content has shown explosive growth, and is widely used in the fields of finance and medicine. For example, financial short videos are often used to analyze market trends and interpret stock trends, and medical short videos focus on the dissemination of knowledge such as health popularization and first aid protection. The main problem of existing video description text methods is that the fragmentation of time sequence information makes it difficult to present the time development context coherently, resulting in poor quality of the generated video description text.
[0003] Therefore, how to generate a time sequence coherent video description text has become a problem to be solved. SUMMARY
[0004] The present application provides a video description text generation method, device, computer equipment and storage medium, aiming to generate a time sequence coherent video description text.
[0005] In a first aspect, the present application provides a video description text generation method, which comprises:
[0006] obtaining feature data of a target video;
[0007] inputting the feature data into a first path feature extraction model, and outputting current frame features through the first path feature extraction model;
[0008] inputting the feature data into a second path feature extraction model, and outputting future frame features through the second path feature extraction model;
[0009] updating a database according to the future frame features to obtain updated database features;
[0010] fusing the current frame features and the database features to obtain fused features;
[0011] decoding the fused features to generate a video description text corresponding to the target video.
[0012] In a second aspect, the present application further provides a video description text generation device, which comprises:
[0013] a data preprocessing module configured to obtain feature data of a target video;
[0014] The feature extraction module is configured to input the feature data into a first path feature extraction model, output current frame features through the first path feature extraction model, and input the feature data into a second path feature extraction model, and output future frame features through the second path feature extraction model.
[0015] The database updating module is configured to update a database according to the future frame features, and obtain updated database features.
[0016] The data fusion module is configured to fuse the current frame features and the database features, and obtain fused features.
[0017] The text generation module is configured to decode the fused features, and generate a video description text corresponding to the target video.
[0018] In a third aspect, the present application further provides a computer device, which comprises a memory and a processor.
[0019] The memory is configured to store a computer program.
[0020] The processor is configured to execute the computer program and implement the video description text generation method as described above when executing the computer program.
[0021] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to make the processor implement the video description text generation method as described above.
[0022] The present application discloses a video description text generation method and device, a computer device and a storage medium. The feature data of a target video is obtained, the feature data is input into a first path feature extraction model and a second path feature extraction model, current frame features are output through the first path feature extraction model, future frame features are output through the second path feature extraction model, a database is updated according to the future frame features, updated database features are obtained, the current frame features are fused with the database features, fused features are obtained, and the fused features are decoded to generate a video description text corresponding to the target video. The method integrates the time sequence context information, solves the time sequence information fragmentation problem, and realizes the generation of a time sequence coherent video description text. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0024] Figure 1 is a step schematic flow chart of a video description text generation method provided by an embodiment of the present application;
[0025] Figure 2 is an architecture schematic diagram of a video description text generation method provided by an embodiment of the present application;
[0026] Figure 3 is an architecture schematic diagram of a text generation module provided by an embodiment of the present application;
[0027] Figure 4 is a schematic block diagram of a video description text generation apparatus provided by an embodiment of the present application;
[0028] Figure 5 is a structural schematic block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without any creative work fall within the scope of protection of the present application.
[0030] The flow charts shown in the drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be further decomposed, combined or partially merged, so the actual execution order can be changed according to the actual situation.
[0031] It should be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clearly indicated by the context, the singular forms “a”, “an” and “the” are intended to include the plural forms.
[0032] It should also be understood that the term “and / or” used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0033] Embodiments of the present application provide a video description text generation method, device, computer equipment and storage medium, wherein the video description text generation method can be applied to a video description text generation device, which can be a server or a terminal or the like, and the method comprises the following steps: obtaining feature data of a target video; inputting the feature data into a first path feature extraction model and a second path feature extraction model; outputting current frame features through the first path feature extraction model; outputting future frame features through the second path feature extraction model; updating a database according to the future frame features; obtaining updated database features; fusing the current frame features and the database features to obtain fused features; and decoding the fused features to generate a video description text corresponding to the target video. The method integrates temporal context information, solves the problem of temporal information fragmentation, and thus realizes the generation of a temporally coherent video description text.
[0034] The server can be a stand-alone server or a server cluster. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer or the like.
[0035] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.
[0036] Please refer to Figure 1 , Figure 1 is a step schematic flowchart of a video description text generation method provided by an embodiment of the present application.
[0037] As shown in Figure 1 , the video description text generation method comprises steps S10 to S60.
[0038] Step S10: obtaining feature data of a target video.
[0039] Exemplarily, the target video can include short videos in different fields such as finance, health, propaganda, introduction, music, product, education and film and television.
[0040] In some embodiments, the feature data of the target video is obtained by: frame extraction on the target video to obtain video frames; standardization processing on the video frames to obtain standard video frames; segmentation processing on the video frames to obtain segments with a fixed number of frames; and inputting the segments with the fixed number of frames into a 3D convolutional neural network for processing to obtain the feature data, which is represented as wherein V t is the feature data, N is the number of frames, and d is the feature dimension.
[0041] The feature data is obtained by sequentially performing frame extraction, standardization processing and segmentation processing on the target video, and then using a 3D convolutional neural network for feature extraction, thereby providing directly inputtable data for subsequent processing.
[0042] For example, the target video is a financial short video, such as Figure 2 As shown, the video description text generation method architecture includes three parts, a data preprocessing module, a text generation module, and a post-processing module. First, input the financial video into the data preprocessing module for frame extraction, standardization processing, segmentation processing, and feature extraction in sequence to obtain corresponding feature data; second, input the feature data into the text generation module to process the feature data through the trained neural network model to obtain the video description text corresponding to the financial video, wherein the trained neural network model includes a first path feature extraction model, a second path feature extraction model, and a text decoder model, and the specific processing process is as follows: input the obtained feature data into the first path feature extraction model for processing to obtain current frame features, process the current frame features through the second path feature extraction model to obtain future frame features, store the obtained future frame features in a database to update the database features, obtain updated database features, cross-attention fuse the current frame features and the database features to obtain fused features, and input the fused features into the text decoder model for decoding to obtain the video description text corresponding to the financial video; finally, input the video description text into the post-processing module for text optimization to obtain the optimized video description text.
[0043] As shown in Figure 3 The text generation module architecture includes a first path feature extraction model, a second path feature extraction model, cross-attention fusion, and a text decoder model. After obtaining the corresponding feature data of the financial short video, input the feature data into the first path feature extraction model for multi-head attention mechanism calculation to obtain current frame features, input the feature data into the second path feature extraction model for processing to obtain future frame features, store the obtained future frame features in a database to update the database features, obtain updated database features, cross-attention fuse the current frame features and the database features based on an attention weight matrix and a layer normalization function to obtain fused features, and input the fused features into the text decoder model for decoding to obtain the video description text corresponding to the financial video.
[0044] In some embodiments, before the video description text generation method processes the feature data through the trained neural network model to obtain the video description text, the method further includes: determining a total loss function corresponding to the neural network model, wherein the neural network model includes the first path feature extraction model, the second path feature extraction model, and the text decoder model; and performing model training on the neural network model until the total loss function converges to obtain the trained neural network model.
[0045] Exemplarily, a total loss function corresponding to the neural network model is determined, the neural network model comprising a first path feature extraction model, a second path feature extraction model and a text decoder model, after obtaining the feature data of the sample video, the feature data is input into the neural network model for model training until the total loss function converges, and a trained neural network model is obtained.
[0046] In order to improve the accuracy of neural network model training, a total loss function corresponding to the neural network model is determined, comprising: determining a first loss function, the first loss function representing standard cross-entropy loss calculation on generated text and real text; determining a second loss function, the second loss function representing loss calculation on predicted future frame features and real future frame features; determining a third loss function, the third loss function representing memory consistency loss calculation between adjacent features in the database; and determining the total loss function according to the first loss function, the second loss function and the third loss function.
[0047] Exemplarily, first, a first loss function is determined, the first loss function representing standard cross-entropy loss calculation on generated text and real text; second, a formula is used to determine a second loss function, the second loss function representing loss calculation on predicted future frame features and real future frame features, wherein, is the second loss function, t is the time, Δ is the prediction time span, V t+Δ is the real future frame feature, is the predicted future frame feature; third, a formula is used to determine a third loss function, the third loss function representing memory consistency loss calculation between adjacent features in the database, wherein, is the third loss function, K is the database capacity, m l is the kth feature stored in the database, m k+1 is the k+1th feature stored in the database; and finally, a formula is used to determine a total loss function, wherein, is the total loss function, is the first loss function, λ1 is a preset hyperparameter, is the second loss function, λ2 is a preset hyperparameter, is the third loss function.
[0048] Step S20, inputting the feature data into the first path feature extraction model, and outputting the current frame feature through the first path feature extraction model.
[0049] The first path feature extraction model includes but is not limited to a Swin Transformer model. Taking the Swin Transformer model as an example, the feature data is input into the Swin Transformer model for processing to obtain the current frame feature output by the Swin Transformer model.
[0050] Specifically, the feature data is input into the Swin Transformer model, and the feature data is processed through the multi-head attention mechanism of the local window and the moving window to obtain the current frame feature output by the Swin Transformer model.
[0051] Step S30, input the feature data into the second path feature extraction model, and output the future frame feature through the second path feature extraction model.
[0052] The second path feature extraction model includes but is not limited to a time convolution layer network model. Taking the time convolution layer network model as an example, the feature data is input into the time convolution layer network model for processing to obtain the future frame feature output by the time convolution layer network model.
[0053] In some embodiments, inputting the feature data into the second path feature extraction model and outputting the future frame feature through the second path feature extraction model includes: performing time convolution layer processing on the feature data based on the second path feature extraction model to obtain a one-dimensional convolution result; and adjusting the one-dimensional convolution result according to offset data corresponding to the time position encoding to obtain the future frame feature.
[0054] For example, inputting the feature data into the second path feature extraction model includes first performing time convolution layer processing on the feature data to obtain a one-dimensional convolution result, and then adjusting the one-dimensional convolution result according to offset data corresponding to the time position encoding to obtain the future frame feature. The mathematical representation is: wherein, is the future frame feature, V t is the feature data, E Δ is the offset data, and Conv1D is a one-dimensional convolution function.
[0055] Step S40, updating the database according to the future frame feature to obtain updated database features.
[0056] After obtaining the future frame feature, it is determined whether to store in the database according to the priority score of the future frame. When the future frame feature is stored in the database, the database is updated to obtain updated database features.
[0057] In some embodiments, the database features are updated according to future frame features, and the updated database features are obtained by: when the database capacity reaches an upper threshold, performing priority scoring on the future frame features to obtain a first scoring result, and performing priority scoring on each feature stored in the database to obtain a second scoring result; and if the smallest second scoring result is smaller than the first scoring result, replacing the feature corresponding to the smallest second scoring result with the future frame feature to update the database and obtain the updated database features.
[0058] For example, when the database capacity reaches the upper threshold, the priority scoring is first performed on the future frame features, and the first scoring result is obtained by the formula , wherein, is the first scoring result, W a is a learnable weight matrix, H t is the current frame feature, is the future frame feature, and σ is a Sigmoid function; then the priority scoring is performed on each feature stored in the database, and the second scoring result is obtained by the formula k = σ(W a [H t ⊙m k ]), wherein, k is the second scoring result, W a is a learnable weight matrix, H t is the current frame feature, m k is the kth feature stored in the database, and σ is a Sigmoid function; finally, the first scoring result and the smallest second scoring result are compared, and if the smallest second scoring result is smaller than the first scoring result, the feature corresponding to the smallest second scoring result is replaced with the future frame feature to update the database and obtain the updated database features.
[0059] In step S50, the current frame feature is fused with the database features to obtain a fused feature.
[0060] After obtaining the current frame feature and the database features, the two are fused by a cross-attention mechanism to obtain a fused feature.
[0061] In some embodiments, the current frame feature is fused with the database features to obtain a fused feature, including: performing cross-attention calculation on the current frame feature and the database features to obtain an attention weight matrix; performing activation function processing on the attention weight matrix to obtain a gating vector; and performing fusion calculation based on the attention weight matrix and the gating vector to obtain the fused feature.
[0062] For example, the cross-attention calculation is first performed on the current frame feature and the database features, and the first scoring result is obtained by the formula Calculate the attention weight matrix, where C t H is the attention weight matrix. t For the features of the current frame, W Q To query the weight matrix, W K W is the key weight matrix. V M is the value weight matrix. t For database features, The dimension is , and Softmax is the normalized exponential function; subsequently, the attention weight matrix is processed by an activation function, using formula G. t =σ(W g [H t C t The gating vector is calculated, where G t W is the gate vector. g H is the gated weight matrix. t C is the feature of the current frame. t Let F be the attention weight matrix, and σ be the sigmoid function. Finally, a fusion calculation is performed based on the attention weight matrix and the gating vector, using formula F. t =LayerNorm(H t +G t ⊙C t ), calculate to obtain fusion features, where F t For fusion features, H t C is the feature of the current frame. t G is the attention weight matrix. t is the gate vector, LayerNorm is the layer normalization function, and ⊙ represents element-wise multiplication.
[0063] Step S60: Decode the fused features to generate video description text corresponding to the target video.
[0064] Text decoder models include, but are not limited to, Transformer decoder models. Taking the Transformer decoder model as an example, the fused features are input into the Transformer decoder model for decoding to generate the video description text corresponding to the target video.
[0065] Specifically, the fused features are input into the Transformer decoder model, and the model is decoded through an autoregressive generation mechanism to generate video description text corresponding to the target video word by word.
[0066] The video description text generation method provided by the embodiment has the advantages that the feature data of a target video is acquired; the feature data is input into a first-path feature extraction model, and current frame features are output by the first-path feature extraction model; the feature data is input into a second-path feature extraction model, and future frame features are output by the second-path feature extraction model; the database is updated according to the future frame features, and updated database features are acquired; the current frame features are fused with the database features, and fused features are acquired; and the fused features are decoded to generate a video description text corresponding to the target video, so that the problem of time sequence information fragmentation is solved, and the generation of a time sequence coherent video description text is implemented.
[0067] Please refer to Figure 4 , Figure 4 The embodiment of the application further provides a schematic block diagram of a video description text generation device 1000, which is used for executing the foregoing video description text generation method. The video description text generation device can be configured in a server or a terminal.
[0068] As shown in Figure 4 , the video description text generation device 1000 comprises a data preprocessing module 1001, a feature extraction module 1002, a database updating module 1003, a data fusion module 1004 and a text generation module 1005.
[0069] The data preprocessing module 1001 is configured to acquire feature data of a target video.
[0070] The feature extraction module 1002 is configured to input the feature data into a first-path feature extraction model, and output current frame features by the first-path feature extraction model; and input the feature data into a second-path feature extraction model, and output future frame features by the second-path feature extraction model.
[0071] The database updating module 1003 is configured to update a database according to the future frame features, and acquire updated database features.
[0072] The data fusion module 1004 is configured to fuse the current frame features with the database features, and acquire fused features.
[0073] The text generation module 1005 is configured to decode the fused features, and generate a video description text corresponding to the target video.
[0074] It should be noted that, for the convenience and brevity of description, the specific working process of the device and each module described above can refer to the corresponding process in the foregoing method embodiments, which will not be described herein.
[0075] The device described above can be implemented in the form of a computer program, which can run on a computer device such as a server or a terminal.Figure 5 The computer device shown runs on.
[0076] Please refer to Figure 5 , Figure 5 is a structural schematic block diagram of a computer device provided by the embodiment of the present application.
[0077] Please refer to Figure 5 , the computer device comprises a processor, a memory and a network interface connected through a system bus, wherein the memory can comprise a storage medium and an internal memory. The storage medium can be a non-volatile storage medium or a volatile storage medium.
[0078] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0079] The internal memory provides an environment for the running of a computer program in the storage medium, and the computer program, when executed by the processor, can make the processor execute any one of the video description text generation methods.
[0080] The network interface is used for network communication, such as sending assigned tasks, etc., which can be understood by those skilled in the art, Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can comprise more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0081] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0082] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:
[0083] Obtaining feature data of a target video; inputting the feature data into a first path feature extraction model, outputting current frame features through the first path feature extraction model; inputting the feature data into a second path feature extraction model, outputting future frame features through the second path feature extraction model; updating a database according to the future frame features to obtain updated database features; fusing the current frame features and the database features to obtain fused features; and decoding the fused features to generate a video description text corresponding to the target video.
[0084] In one embodiment, the processor, when implementing inputting the feature data into the second path feature extraction model and outputting the future frame features through the second path feature extraction model, is configured to implement:
[0085] performing time convolution layer processing on the feature data based on the second path feature extraction model to obtain one-dimensional convolution results; and adjusting the one-dimensional convolution results according to offset data corresponding to the time position encoding to obtain the future frame features.
[0086] In one embodiment, the processor, when implementing updating the database according to the future frame features to obtain the updated database features, is configured to implement:
[0087] When the database capacity reaches an upper threshold, performing priority scoring on the future frame features to obtain a first scoring result, and performing priority scoring on each feature stored in the database to obtain a second scoring result; if the smallest second scoring result is smaller than the first scoring result, replacing the feature corresponding to the smallest second scoring result with the future frame feature to update the database to obtain the updated database features.
[0088] In one embodiment, the processor, when implementing performing priority scoring on each feature stored in the database to obtain the second scoring result, is configured to implement:
[0089] obtaining the second scoring result based on a formula α k = σ(W a [H t ⊙m k ]);
[0090] wherein, α k is the second scoring result, W a is a learnable weight matrix, H t is the current frame features, m k is the kth feature stored in the database, σ is a Sigmoid function, and ⊙ is an element-wise multiplication.
[0091] In one embodiment, the processor, when implementing fusing the current frame features and the database features to obtain the fused features, is configured to implement:
[0092] The cross-attention calculation is performed on the current frame features and the database features to obtain an attention weight matrix; an activation function is performed on the attention weight matrix to obtain a gating vector; and fusion calculation is performed based on the attention weight matrix and the gating vector to obtain fused features.
[0093] In one embodiment, the processor, when implementing the video description text generation method, is further configured to implement:
[0094] determining a total loss function corresponding to the neural network model, the neural network model comprising the first path feature extraction model, the second path feature extraction model, and the text decoder model; and performing model training on the neural network model until the total loss function converges, to obtain the trained neural network model.
[0095] In one embodiment, the processor, when determining the total loss function corresponding to the neural network model, is configured to implement:
[0096] determining a first loss function, the first loss function representing standard cross-entropy loss calculation on the generated text and the real text; determining a second loss function, the second loss function representing loss calculation on the predicted future frame features and the real future frame features; determining a third loss function, the third loss function representing memory consistency loss calculation between the database adjacent features; and determining the total loss function according to the first loss function, the second loss function, and the third loss function.
[0097] In an embodiment of the present application, a computer readable storage medium is also provided, which stores a computer program, the computer program comprising program instructions, and a processor executes the program instructions to implement any one of the video description text generation methods provided in the embodiments of the present application.
[0098] The computer readable storage medium can be an internal storage unit of the computer device, such as a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital card (SD Card), a flash card, etc.
[0099] Further, the computer readable storage medium can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required by a function, etc.; and the data storage area can store data created according to the use of the blockchain node, etc.
[0100] The blockchain referred to in the present application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism and encryption algorithm. The blockchain is essentially a decentralized database, and is a series of data blocks associated using cryptographic methods, each data block containing information of a batch of network transactions, for verifying the validity (anti-fake) of the information and generating the next block. The blockchain can include a blockchain underlying platform, a platform product service layer and an application service layer.
[0101] The above merely provides a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements shall be encompassed within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method of generating a video description text, characterized by, The method comprises: obtaining feature data of a target video; inputting the feature data into a first path feature extraction model, and outputting current frame features through the first path feature extraction model; inputting the feature data into a second path feature extraction model, and outputting future frame features through the second path feature extraction model; updating a database according to the future frame features to obtain updated database features; fusing the current frame features and the database features to obtain fused features; decoding the fused features to generate a video description text corresponding to the target video.
2. The video description text generating method of claim 1, wherein, The method comprises: performing time convolution layer processing on the feature data based on the second path feature extraction model to obtain one-dimensional convolution results; adjusting the one-dimensional convolution results according to offset data corresponding to a time position encoding to obtain the future frame features.
3. The method of claim 1, wherein, The method comprises: when the database capacity reaches an upper threshold, performing priority scoring on the future frame features to obtain a first scoring result, and performing priority scoring on each feature stored in the database to obtain a second scoring result; if the smallest second scoring result is smaller than the first scoring result, replacing the feature corresponding to the smallest second scoring result with the future frame feature to update the database and obtain the updated database features.
4. The method of claim 3, wherein, The method comprises: Based on the formula α k = σ(W a [H t ⊙m k ]) to calculate the second score result; wherein α k is the second score result, W a is a learnable weight matrix, H t is the current frame feature, m k is the kth feature stored in the database, σ is a Sigmoid function, and is an element-wise multiplication.
5. The method of claim 1, wherein, The method comprises: performing cross-attention calculation on the current frame features and the database features to obtain an attention weight matrix; performing activation function processing on the attention weight matrix to obtain a gating vector; performing fusion calculation based on the attention weight matrix and the gating vector to obtain the fused features.
6. The method of claim 1, wherein, The method further comprises: determining a total loss function corresponding to a neural network model, wherein the neural network model comprises the first path feature extraction model, the second path feature extraction model, and a text decoder model; performing model training on the neural network model until the total loss function converges, to obtain a trained neural network model.
7. The video description text generating method of claim 6, wherein, The method comprises: determining a first loss function, wherein the first loss function represents standard cross-entropy loss calculation on generated text and real text; determining a second loss function, wherein the second loss function represents loss calculation on predicted future frame features and real future frame features; determining a third loss function, wherein the third loss function represents memory consistency loss calculation between adjacent features in the database; determining the total loss function according to the first loss function, the second loss function, and the third loss function.
8. An apparatus for video description text generation, the apparatus comprising: The apparatus comprises: a data preprocessing module configured to obtain feature data of a target video; The feature extraction module is configured to input the feature data into a first path feature extraction model, output a current frame feature through the first path feature extraction model, input the feature data into a second path feature extraction model, and output a future frame feature through the second path feature extraction model. The database updating module is configured to update a database according to the future frame feature, and obtain an updated database feature. The data fusion module is configured to fuse the current frame feature and the database feature, and obtain a fused feature. The text generation module is configured to decode the fused feature, and generate a video description text corresponding to the target video.
9. A computer device, comprising: The computer device comprises a memory and a processor. The memory is configured to store a computer program. The processor is configured to execute the computer program and implement the video description text generation method according to any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program enables the processor to implement the video description text generation method according to any one of claims 1-7 when executed by the processor.
Citation Information
Patent Citations
Video description generation method and device based on deep learning model
CN117292293A