A method, device, and storage medium for generating video descriptions

By constructing a video description generation model including an encoder, a semantic detector and a decoder, feature analysis and semantic analysis of the video are solved, and the problem of ignoring the correlation of visual content in the prior art is achieved, and a semantic-rich and accurate video description is generated.

CN114386260BActive Publication Date: 2025-06-13GUILIN UNIV OF ELECTRONIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111640894.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-06-13
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

Existing video description generation methods ignore the correlation between sentence semantics and visual content, and fail to fully consider prominent features in the video.

Method used

By building a training model including an encoder, a semantic detector and a decoder, feature analysis, semantic analysis and decoding of the trained video, a video description generation model is generated, and the video description results are generated.

Benefits of technology

This method can explore the correlation between the generated description and visual content, generate semantic rich sentences, fully consider outstanding features, and improve the accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114386260B_ABST
    Figure CN114386260B_ABST
Patent Text Reader

Abstract

The present invention provides a method, an apparatus, and a storage medium for generating video descriptions, belonging to the technical field of video processing. The method includes: S1: Import the video to be trained and construct an encoder, a semantic detector, and a decoder; S2: Perform feature analysis on the video to be trained through the encoder to obtain the features to be processed and visual features; S3: Perform semantic analysis on the features to be processed through the semantic detector to obtain semantic attributes; S4: Decode the visual features through the decoder to obtain a predicted label vector; S5: Perform loss analysis on the semantic attributes and the predicted label vector to obtain a video description generation model; S6: Perform video description on the video to be described through the video description generation model to generate a video description result. The present invention can explore the correlation between the generated description and the visual content, generate sentences with rich semantics, fully consider the prominent features, and improve the accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of video processing, and particularly relates to a video description generation method, apparatus, and storage medium. Background Art

[0002] The purpose of video description is to automatically generate a concise and accurate video description, which requires technologies of computer vision (CV) and natural language processing (NLP). The sequence-to-sequence learning method of deep learning can learn dense vectors from discrete color arrays and generate natural language sequences without human interference. However, most of the existing methods compress the entire video shot or frame into a static representation without considering prominent features. In addition, most of the existing translation methods model translation errors but ignore the correlation between sentence semantics and visual content. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a video description generation method, apparatus, and storage medium for the deficiencies of the prior art.

[0004] The technical solution of the present invention to solve the above technical problem is as follows: A video description generation method includes the following steps:

[0005] S1: Import the video to be trained and construct a training model, where the training model includes an encoder, a semantic detector, and a decoder;

[0006] S2: Perform feature analysis on the video to be trained through the encoder to obtain the feature to be processed and visual features;

[0007] S3: Perform semantic analysis on the feature to be processed through the semantic detector to obtain semantic attributes;

[0008] S4: Decode the visual features through the decoder to obtain a predicted label vector;

[0009] S5: Perform loss analysis on the semantic attributes and the predicted label vector to obtain a video description generation model;

[0010] S6: Import the video to be described, and perform video description on the video to be described through the video description generation model to generate a video description result.

[0011] Another technical solution of the present invention to solve the above technical problem is as follows: A video description generation apparatus includes:

[0012] A model construction module, configured to import the video to be trained and construct a training model, where the training model includes an encoder, a semantic detector, and a decoder;

[0013] A feature analysis module, configured to perform feature analysis on the video to be trained through the encoder to obtain features to be processed and visual features;

[0014] A semantic analysis module, configured to perform semantic analysis on the features to be processed through the semantic detector to obtain semantic attributes;

[0015] A feature decoding module, configured to decode the visual features through the decoder to obtain a predicted label vector;

[0016] A loss analysis module, configured to perform loss analysis on the semantic attributes and the predicted label vector to obtain a video description generation model;

[0017] A video description result generation module, configured to import a video to be described, and perform video description on the video to be described through the video description generation model to generate a video description result.

[0018] Another technical solution for the present invention to solve the above technical problems is as follows: A video description generation device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the video description generation method as described above is implemented.

[0019] Another technical solution for the present invention to solve the above technical problems is as follows: A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the video description generation method as described above is implemented.

[0020] The beneficial effects of the present invention are as follows: Through the feature analysis of the video to be trained by the encoder, features to be processed and visual features are obtained; through the semantic analysis of the features to be processed by the semantic detector, semantic attributes are obtained; through the decoding of the visual features by the decoder, a predicted label vector is obtained; through the loss analysis of the semantic attributes and the predicted label vector, a video description generation model is obtained; through the video description generation model, video description of the video to be described generates a video description result, which can explore the correlation between the generated description and the visual content, generate sentences with rich semantics, fully consider prominent features, and improve the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic flowchart of a video description generation method provided by an embodiment of the present invention;

[0022] Figure 2 It is a module block diagram of a video description generation device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The principles and features of the present invention will be described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0024] Figure 1 It is a schematic flowchart of a video description generation method provided by an embodiment of the present invention.

[0025] As Figure 1 shown, a video description generation method includes the following steps:

[0026] S1: Import the video to be trained and construct a training model, where the training model includes an encoder, a semantic detector, and a decoder;

[0027] S2: Perform feature analysis on the video to be trained through the encoder to obtain the feature to be processed and visual features;

[0028] S3: Perform semantic analysis on the feature to be processed through the semantic detector to obtain semantic attributes;

[0029] S4: Decode the visual features through the decoder to obtain a predicted label vector;

[0030] S5: Perform loss analysis on the semantic attributes and the predicted label vector to obtain a video description generation model;

[0031] S6: Import the video to be described, and perform video description on the video to be described through the video description generation model to generate a video description result.

[0032] In the above embodiment, the feature analysis of the video to be trained by the encoder obtains the feature to be processed and visual features, the semantic analysis of the feature to be processed by the semantic detector obtains semantic attributes, the decoding of the visual features by the decoder obtains a predicted label vector, the loss analysis of the semantic attributes and the predicted label vector obtains a video description generation model, and the video description of the video to be described by the video description generation model generates a video description result, which can explore the correlation between the generated description and visual content, generate sentences with rich semantics, fully consider prominent features, and improve the accuracy of the model.

[0033] Optionally, as an embodiment of the present invention, the encoder includes a 2D-CNN convolutional neural network and a 3D-CNN convolutional neural network, and the process of step S2 includes:

[0034] Perform global feature extraction on the video to be trained through the 2D-CNN convolutional neural network to obtain global features;

[0035] Extract motion features from the video to be trained through the 3D-CNN convolutional neural network, obtain motion features, and use the global features and the motion features together as features to be processed;

[0036] Concatenate the global features and the motion features to obtain visual features.

[0037] It should be understood that the 2D-CNN convolutional neural network refers to the convolutional kernel performing a sliding window operation in the two-dimensional space of the input image. 2D convolution only considers spatial features and does not consider temporal features. The input and output data of 2D-CNN are three-dimensional. It is mainly used for image data.

[0038] It should be understood that the 3D-CNN convolutional neural network refers to the convolutional kernel performing a sliding window operation in the three-dimensional space of the input image. 3D convolution has an additional depth channel, and this depth is generally consecutive frames in a video or different slices of a stereoscopic image. The input and output data of 3D-CNN are four-dimensional. It is usually used in the field of video processing (detecting actions and human behaviors).

[0039] Specifically, given a video V to be described (i.e., the video to be trained), first extract the global feature Va, the motion feature Vm, and other visual feature representations of the video V (i.e., the video to be trained). Then, by passing these two feature vectors through an encoder, a visual spatio-temporal representation of the video (i.e., the visual feature) can be obtained.

[0040] In the above embodiment, the features to be processed and the visual features are obtained through the encoder's analysis of the features of the video to be trained, providing a basis for exploring the correlation between the generated description and the visual content, and improving the accuracy of the model.

[0041] Optionally, as an embodiment of the present invention, the process of step S3 includes:

[0042] Perform semantic analysis on the global features to obtain multiple global feature semantic attributes;

[0043] Perform semantic analysis on the motion features to obtain multiple motion feature semantic attributes;

[0044] Use all the global feature semantic attributes and all the motion feature semantic attributes as semantic attributes.

[0045] It should be understood that the obtained global features and motion features are input into a semantic detector, and a multi-label classification method is used to learn the semantic attributes of the video. Vi represents the feature vector of the i-th video. Through the training instances {vi, yi}, an MLP is used to learn a function f: R m →RK , where m is the number of input dimensions, K is the number of output dimensions, and K is equal to the number of semantic attributes.

[0046] In the above embodiments, through the global feature semantic analysis of the global features, multiple global feature semantic attributes are obtained, and through the motion feature semantic analysis of the motion features, multiple motion feature semantic attributes are obtained. Taking all the global feature semantic attributes and all the motion feature semantic attributes as semantic attributes can filter out the most appropriate descriptions and improve the accuracy of the model.

[0047] Optionally, as an embodiment of the present invention, the global features include multiple global feature vectors, and the process of performing global feature semantic analysis on the global features to obtain multiple global feature semantic attributes includes:

[0048] Calculate the global feature similarity between each of the global feature vectors and each of the word vectors in the preset feature library respectively, and obtain multiple global feature similarities corresponding to each of the global feature vectors;

[0049] Sort the multiple global feature similarities corresponding to each of the global feature vectors according to the size of the global feature similarity respectively, and obtain multiple sorted global feature similarities corresponding to each of the global feature vectors;

[0050] Use the Spacy Tagging Tool to perform global feature screening on each of the sorted global feature similarities respectively, and after screening, obtain multiple screened global feature similarities corresponding to each of the global feature vectors;

[0051] Take the word vectors corresponding to the first K screened global feature similarities corresponding to each of the global feature vectors as the global feature semantic attributes.

[0052] It should be understood that K is a positive integer and can be 1, 2, 3,....

[0053] It should be understood that the Spacy Tagging Tool, namely spaCy, is the fastest industrial-grade natural language processing tool in the world. It supports multiple basic natural language processing functions; the main functions of spaCy include tokenization, part-of-speech tagging, stemming, named entity recognition, noun phrase extraction, and so on.

[0054] It should be understood that first, all the words extracted from the training set (i.e., the global features) are sorted according to their frequencies, then some function words (such as "a", "the") are deleted, and finally, the first K words including verbs, nouns, and adjectives are selected as the semantic attributes (i.e., the global feature semantic attributes).

[0055] In the above embodiments, semantic analysis of the global features of the global features obtains multiple global feature semantic attributes, which can screen out the closest descriptions and improve the accuracy of the model.

[0056] Optionally, as an embodiment of the present invention, the motion features include multiple motion feature vectors, and the process of performing semantic analysis of the motion features to obtain multiple motion feature semantic attributes includes:

[0057] Calculate the motion feature similarity between each of the motion feature vectors and each of the word vectors in the preset feature library respectively, and obtain multiple motion feature similarities corresponding to each of the motion feature vectors;

[0058] Sort the multiple motion feature similarities corresponding to each of the motion feature vectors according to the magnitude of the motion feature similarity respectively, and obtain multiple sorted motion feature similarities corresponding to each of the motion feature vectors;

[0059] Use the Spacy Tagging Tool to screen the motion features of each of the sorted motion feature similarities respectively, and after screening, obtain multiple screened motion feature similarities corresponding to each of the motion feature vectors;

[0060] Take the word vectors corresponding to the top K of the screened motion feature similarities corresponding to each of the motion feature vectors as the motion feature semantic attributes.

[0061] It should be understood that K is a positive integer and can be 1, 2, 3,....

[0062] It should be understood that first, all the words extracted from the training set (i.e., the motion features) are sorted according to their frequencies, then some function words (such as "a", "the") are deleted, and finally the top K words including verbs, nouns, and adjectives are selected as the semantic attributes (i.e., the motion feature semantic attributes).

[0063] In the above embodiments, semantic analysis of the motion features of the motion features obtains multiple motion feature semantic attributes, which can screen out the closest descriptions and improve the accuracy of the model.

[0064] Optionally, as an embodiment of the present invention, the process of step S4 includes:

[0065] Decode the visual features based on the LSTM long short-term memory network to obtain a predicted label vector.

[0066] It should be understood that the extracted visual features are input into a decoder composed of LSTM to obtain a predicted label vector of the video.

[0067] In the above embodiments, the predicted label vector is obtained by decoding the visual features based on the LSTM long short-term memory network, so as to obtain accurate visual content and improve the accuracy of the model.

[0068] Optionally, as an embodiment of the present invention, the process of step S5 includes:

[0069] Calculating the loss value between the semantic attribute and the predicted label vector by using the cross-entropy loss algorithm to obtain the loss value;

[0070] Updating the parameters of the decoder according to the loss value, and returning to step S2 until a preset number of iterations is reached, and using the updated training model as the video description generation model.

[0071] It should be understood that the decoder decodes the predicted label vector, calculates the cross-entropy loss with the semantic attribute vector (i.e., the semantic attribute) of the video, the loss layer calculates the reverse gradient and feeds back, and finally completes the model training.

[0072] In the above embodiments, the video description generation model is obtained by analyzing the loss between the semantic attribute and the predicted label vector, which can explore the correlation between the generated description and the visual content, generate sentences with rich semantics, fully consider the prominent features, and improve the accuracy of the model.

[0073] Optionally, as an embodiment of the present invention, after step S4, it further includes:

[0074] Retrieving the predicted label vector according to a preset corpus to obtain video description information.

[0075] In the above embodiments, the video description information is obtained by retrieving the predicted label vector through a preset corpus, which can intuitively obtain the description result and generate sentences with rich semantics.

[0076] Figure 2 It is a block diagram of a video description generation device provided by an embodiment of the present invention.

[0077] Optionally, as another embodiment of the present invention, as Figure 2 shown, a video description generation device includes:

[0078] A model construction module, configured to import a video to be trained and construct a training model, where the training model includes an encoder, a semantic detector, and a decoder;

[0079] A feature analysis module, configured to perform feature analysis on the video to be trained through the encoder to obtain to-be-processed features and visual features;

[0080] A semantic analysis module, configured to perform semantic analysis on the to-be-processed feature through the semantic detector to obtain semantic attributes;

[0081] A feature decoding module, configured to decode the visual feature through the decoder to obtain a predicted label vector;

[0082] A loss analysis module, configured to perform loss analysis on the semantic attributes and the predicted label vector to obtain a video description generation model;

[0083] A video description result generation module, configured to import the video to be described, and perform video description on the video to be described through the video description generation model to generate a video description result.

[0084] Optionally, another embodiment of the present invention provides a video description generation device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the video description generation method as described above is implemented. The device may be a computer or the like.

[0085] Optionally, another embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the video description generation method as described above is implemented.

[0086] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0087] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0088] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0089] The unit described as a separate component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0090] In addition, each functional unit in the embodiments of the present invention may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0091] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0092] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for generating video descriptions, characterized in that, it includes the following steps: S1: Import the video to be trained and construct a training model, where the training model includes an encoder, a semantic detector, and a decoder; S2: Perform feature analysis on the video to be trained through the encoder to obtain features to be processed and visual features; S3: Perform semantic analysis on the features to be processed through the semantic detector to obtain semantic attributes; S4: Decode the visual features through the decoder to obtain a predicted label vector; S5: Perform loss analysis on the semantic attributes and the predicted label vector to obtain a video description generation model; S6: Import the video to be described, and generate a video description result by performing video description on the video to be described through the video description generation model.

2. The method for generating video descriptions according to claim 1, characterized in that, the encoder includes a 2D-CNN convolutional neural network and a 3D-CNN convolutional neural network, and the process of step S2 includes: Perform global feature extraction on the video to be trained through the 2D-CNN convolutional neural network to obtain global features; Perform motion feature extraction on the video to be trained through the 3D-CNN convolutional neural network to obtain motion features, and use the global features and the motion features together as features to be processed; Concatenate the global features and the motion features to obtain visual features.

3. The method for generating video descriptions according to claim 2, characterized in that, the process of step S3 includes: Perform semantic analysis of the global features to obtain multiple global feature semantic attributes; Perform semantic analysis of the motion features to obtain multiple motion feature semantic attributes; Use all the global feature semantic attributes and all the motion feature semantic attributes as semantic attributes.

4. The method for generating video descriptions according to claim 3, characterized in that, the global features include multiple global feature vectors, and the process of performing semantic analysis of the global features to obtain multiple global feature semantic attributes includes: Calculate the global feature similarity between each global feature vector and each word vector in the preset feature library respectively to obtain multiple global feature similarities corresponding to each global feature vector; Sort the multiple global feature similarities corresponding to each global feature vector according to the size of the global feature similarity respectively to obtain multiple sorted global feature similarities corresponding to each global feature vector; Use the Spacy Tagging Tool to screen the global features of each sorted global feature similarity respectively, and after screening, obtain multiple screened global feature similarities corresponding to each global feature vector; Use the word vectors corresponding to the first K screened global feature similarities corresponding to each global feature vector as global feature semantic attributes.

5. The method for generating video descriptions according to claim 4, characterized in that, The motion features include a plurality of motion feature vectors. The process of performing semantic analysis of the motion features to obtain a plurality of motion feature semantic attributes includes: Calculating the motion feature similarity between each of the motion feature vectors and each of the word vectors in the preset feature library respectively, to obtain a plurality of motion feature similarities corresponding to each of the motion feature vectors; Sorting the plurality of motion feature similarities corresponding to each of the motion feature vectors respectively according to the magnitude of the motion feature similarity, to obtain a plurality of sorted motion feature similarities corresponding to each of the motion feature vectors; Using the Spacy Tagging Tool to perform screening of the motion features on each of the sorted motion feature similarities respectively, and after screening, obtaining a plurality of screened motion feature similarities corresponding to each of the motion feature vectors; Taking the word vectors corresponding to the top K of the screened motion feature similarities corresponding to each of the motion feature vectors as the motion feature semantic attributes.

6. The video description generation method according to claim 1, wherein, the process of step S4 includes: Decoding the visual features based on the LSTM long short-term memory network to obtain a predicted label vector.

7. The video description generation method according to claim 1, wherein, the process of step S5 includes: Calculating the loss value of the semantic attribute and the predicted label vector by using the cross-entropy loss algorithm to obtain a loss value; Updating the parameters of the decoder according to the loss value, and returning to step S2 until a preset number of iterations is reached, and taking the updated training model as the video description generation model.

8. The video description generation method according to claim 1, wherein, after step S4, it further includes: Retrieving the predicted label vector according to a preset corpus to obtain video description information.

9. A video description generation device, wherein, it includes: A model construction module, configured to import a video to be trained and construct a training model, where the training model includes an encoder, a semantic detector, and a decoder; A feature analysis module, configured to perform feature analysis on the video to be trained through the encoder to obtain features to be processed and visual features; A semantic analysis module, configured to perform semantic analysis on the features to be processed through the semantic detector to obtain semantic attributes; A feature decoding module, configured to decode the visual features through the decoder to obtain a predicted label vector; A loss analysis module, configured to perform loss analysis on the semantic attribute and the predicted label vector to obtain a video description generation model; A video description result generation module, configured to import a video to be described, and perform video description on the video to be described through the video description generation model to generate a video description result.

10. A computer-readable storage medium, the computer-readable storage medium stores a computer program, wherein, when the computer program is executed by a processor, it implements the video description generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video content description method, system and device based on multi-modal attention mechanism

    CN111079601A

  • Video description method and device based on convolutional neural network

    CN111325068A