Training method, application method, device and medium for video content description model
By extracting verbs and templates in the video content description for encoding processing and feature sampling, combining visual features to approximate the encoder and decoder data, generating an open video description, solving the problem of single description in the prior art, and achieving diverse and efficient description generation.
Patent Information
- Application Number
- CN202111499121.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-09
AI Technical Summary
The prior art cannot generate open-ended video descriptions, cannot generate diverse descriptions using different expressions from multiple angles, cannot meet the ambiguity of video description needs, and the search space is low.
By obtaining the target video and content description, extracting verbs and templates for encoding processing and feature sampling, combining the target visual features to perform data approximation processing on the prior hidden variable encoder and language decoder, and generating a video content description model.
It realizes the generation of open video descriptions, which can generate diverse descriptions using different expressions from multiple angles, meets the needs of diverse video descriptions, and improves the efficiency of description generation.
Smart Images

Figure CN114386480B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, specifically to the field of video content description technology, and relates to a training method, an application method, a device and a medium of a video content description model. Background Art
[0002] In recent years, with the continuous development of information technology and the iterative upgrade of intelligent devices, people are more inclined to use videos to convey information, making the scale of various types of video data increasingly large, and at the same time bringing huge challenges. For example, hundreds of video data are uploaded to the server every minute on video content sharing websites. If it is manually reviewed whether these videos comply with the rules, it is very time-consuming and laborious. By means of video description methods, the efficiency of the review work can be significantly improved, saving a large amount of time and labor costs. Video content description technology can be mainly widely applied in actual scenarios such as video title generation, video retrieval, and helping visually impaired people understand videos. The methods in the related technologies only focus on generating a single and as accurate as possible statement, and cannot generate open-ended descriptions like humans. Summary of the Invention
[0003] The present invention aims to at least solve one of the technical problems existing in the prior art. For this purpose, the present invention provides a training method, an application method, a device and a medium of a video content description model, which can generate open-ended descriptions, that is, generate various descriptions from many angles and using different expressions.
[0004] The training method of the video content description model according to the first aspect embodiment of the present invention includes:
[0005] Obtain a target video and the content description of the target video;
[0006] Extract a verb and a template from the content description, where the template is the remaining part after separating the verb from the target content description corresponding to the target video;
[0007] Perform encoding processing and feature sampling on the verb and the template to obtain a first latent variable feature;
[0008] Extract features from the target video to obtain a target visual feature;
[0009] Perform data approximation processing on a prior latent variable encoder and a language decoder according to the first latent variable feature and the target visual feature to obtain a video content description model.
[0010] The training method according to the first aspect embodiment of the present invention has at least the following beneficial effects: obtaining a target video and a content description of the target video, extracting verbs and templates from the content description, encoding and feature sampling the verbs and templates to obtain a first latent variable feature, extracting features from the target video to obtain a target visual feature, and performing data approximation processing on the prior latent variable encoder and the language decoder according to the first latent variable feature and the target visual feature to obtain a video content description model. According to the video description model of the present invention, an open description can be generated, that is, descriptions can be generated in various ways from many perspectives using different expressions.
[0011] According to some embodiments of the present invention, the encoding and feature sampling of the verbs and the templates to obtain a first latent variable feature includes:
[0012] Inputting the verbs and the templates into a posterior latent variable encoder for encoding to obtain an encoded action latent variable and an encoded template latent variable;
[0013] Inputting the encoded template latent variable into a preset first neural network model to obtain a posterior distribution of the template latent variable;
[0014] Inputting the encoded action latent variable and the posterior distribution of the template latent variable into the first neural network model for processing to obtain a posterior distribution of the action latent variable;
[0015] Performing feature sampling on the posterior distribution of the action latent variable and the posterior distribution of the template latent variable to obtain a first action latent variable feature and a first template latent variable feature;
[0016] Obtaining a first latent variable feature according to the first action latent variable feature and the first template latent variable feature.
[0017] According to some embodiments of the present invention, the data approximation processing of the prior latent variable encoder and the language decoder according to the first latent variable feature and the target visual feature to obtain a video content description model includes:
[0018] Substituting the first latent variable feature and the visual feature into the language decoder to obtain an initial decoding result;
[0019] Inputting the initial decoding result and the first latent variable feature into a preset loss function to train the prior latent variable encoder and the language decoder to obtain a video content description model.
[0020] According to some embodiments of the present invention, the inputting the initial decoding result and the first latent variable feature into a preset loss function to train the prior latent variable encoder and the language decoder to obtain a video content description model includes:
[0021] Input the initial decoding result and the latent variable feature into a preset loss function to obtain a loss value;
[0022] Adjust the parameters of the language decoder according to the loss value to obtain a trained language decoder;
[0023] Adjust the parameters of the prior latent variable encoder according to the loss value to make the target prior distribution close to the first latent variable feature, and obtain a trained prior latent variable encoder;
[0024] Obtain a video content description model based on the trained language decoder and the trained prior latent variable encoder.
[0025] According to some embodiments of the present invention, the feature extraction of the target video to obtain the target visual feature includes:
[0026] Use a pre-trained deep convolutional neural network to extract features from the video to obtain the target visual feature.
[0027] According to some embodiments of the present invention, before the feature extraction of the target video to obtain the target visual feature, the training method further includes:
[0028] Preprocess the target video to obtain a preprocessed target video.
[0029] According to an application method of video content description according to the second aspect embodiment of the present invention, it includes:
[0030] Obtain a video to be analyzed, and extract the video visual feature of the video to be analyzed;
[0031] Obtain a second latent variable feature according to the video content description model;
[0032] Substitute the second latent variable feature and the video visual feature into the video content description model to obtain a video content description, and the video content description model is trained according to any training method of the first aspect embodiment of the present invention.
[0033] According to the application method of the second aspect embodiment of the present invention, it has at least the following beneficial effects: Obtain the second latent variable feature from the trained video content description model, generate a new video content description according to the second latent variable feature and the extracted video visual feature. When the obtained second latent variable features are different, the obtained video content descriptions are also different, that is, an open video description can be generated, which can be used in a variety of different application scenarios to meet the actual demand for diverse video descriptions.
[0034] According to some embodiments of the present invention, the obtaining of the second latent variable feature according to the video content description model includes:
[0035] Obtain the prior distribution of the action latent variable and the prior distribution of the template latent variable according to the video content description model;
[0036] Sample the prior distribution of the action latent variable and the prior distribution of the template latent variable respectively to obtain the second action latent variable feature and the second template latent variable feature;
[0037] Obtain the second latent variable feature according to the second action latent variable feature and the second template latent variable feature.
[0038] A computer device according to an embodiment of the third aspect of the present invention, the computer device includes a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor is used to execute:
[0039] The training method according to any one of the embodiments of the first aspect of the present invention; or
[0040] The application method according to any one of the embodiments of the second aspect of the present invention.
[0041] A storage medium according to an embodiment of the fourth aspect of the present invention, the storage medium is a computer-readable storage medium, and the computer-readable medium stores a computer program, and when the computer program is executed by a computer, the computer is used to execute:
[0042] The training method according to any one of the embodiments of the first aspect of the present invention; or
[0043] The application method according to any one of the embodiments of the second aspect of the present invention.
[0044] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. Description of the Drawings
[0045] The following further describes the present invention with reference to the drawings and embodiments, wherein:
[0046] Figure 1 Is a flowchart of a training method for a video content description model provided by an embodiment of the present invention;
[0047] Figure 2 Is a flowchart of a training method for a video content description model provided by another embodiment of the present invention;
[0048] Figure 3 Is a flowchart of a training method for a video content description model provided by another embodiment of the present invention;
[0049] Figure 4Flowchart of the training method for the video content description model provided by another embodiment of the present invention;
[0050] Figure 5 Flowchart of the training method for the video content description model provided by another embodiment of the present invention;
[0051] Figure 6 Flowchart of the training method for the video content description model provided by another embodiment of the present invention;
[0052] Figure 7 Flowchart of the application method for the video content description model provided by an embodiment of the present invention;
[0053] Figure 8 Flowchart of the application method for the video content description model provided by another embodiment of the present invention. Detailed implementation manners
[0054] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0055] In the description of the present invention, it should be understood that the orientation descriptions such as up, down, front, back, left, right, etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. It is only for convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.
[0056] In the description of the present invention, the meaning of several is more than one, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0057] In the description of the present invention, unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0058] In the description of the present invention, the descriptions with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0059] Video Captioning is a technology that uses a single sentence to assertively describe given video content. The maturity of this technology can assist future artificial intelligence in many scenarios: massive short-video platforms such as Douyin, Kuaishou, or news platforms need automatic natural language descriptions (e.g., titles); the blind can use voice assistants to describe the surrounding scenes to help themselves navigate; other video-related tasks such as retrieval and visual question answering also require video descriptions for auxiliary training. On the other hand, the video scene itself is complex and diverse, and combined with the ambiguity of natural language, humans tend to use different and open descriptions to depict the surrounding scenes. This uncertain description is particularly important in some video captioning applications that emphasize openness. For example, in educational applications such as "describe what you see in the picture", intelligent assistants often need to give diverse descriptions to guide students; the blind may not be satisfied with or understand the description of the current scene by the voice assistant and need it to "speak" in another way; different users may need different languages to describe video media segments to suit the understanding of different groups. However, most current video captioning applications do not take this openness into account, cannot generate diverse video descriptions, cannot meet ambiguous scenarios, and the search space of related technologies is large and the efficiency is low.
[0060] Based on this, the present invention provides a training method, application method, device, and medium for a video content description model. The training method includes obtaining a target video and the content description of the target video, extracting verbs and templates from the content description, performing encoding processing and feature sampling on the verbs and templates to obtain a first latent variable feature, extracting target visual features from the target video, and performing data approximation processing on the prior latent variable encoder and the language decoder according to the first latent variable feature and the target visual features to obtain a video content description model. The present invention can generate open descriptions according to the video description model, that is, generate various descriptions from many angles and using different expressions.
[0061] The following further elaborates on the embodiments of the present invention with reference to the accompanying drawings.
[0062] Reference Figure 1, an embodiment of the present invention provides a training method for a video content description model, and the training method includes but is not limited to:
[0063] Step S110, obtaining a target video and a content description of the target video;
[0064] Step S120, extracting a verb and a template from the content description, where the template is the remaining part after separating the verb from the target content description corresponding to the target video;
[0065] Step S130, performing encoding processing and feature sampling on the verb and the template to obtain a first latent variable feature;
[0066] Step S140, extracting features from the target video to obtain a target visual feature;
[0067] Step S150, performing data approximation processing on the prior latent variable encoder and the language decoder according to the first latent variable feature and the target visual feature to obtain a video content description model.
[0068] Specifically, let the verb be v, and the verb v in the content description of the target video is identified and separated by using a natural language part-of-speech tagging tool; the remaining part m is called a template. For example, in "The band is playing the guitar", the verb "playing" is extracted, and the template is "The band is [empty] the guitar". It should be noted that [empty] represents a special character.
[0069] Specifically, under the variational autoencoder, encoding processing and feature sampling are performed on the verb and the template to obtain a first latent variable feature. Since the first latent variable feature is obtained by performing encoding processing and feature sampling on the verb and the template, the first latent variable has the function of reflecting the video content description. In this embodiment, combining the first latent variable feature with the video visual feature and inputting them into the decoder will obtain a new video content description.
[0070] In one embodiment, data approximation processing is performed on the prior latent variable encoder and the language decoder according to the first latent variable feature and the target visual feature to obtain a video content description model. First, combining the first latent variable feature with the video visual feature and inputting them into the decoder can obtain a new video content description. Then, data approximation processing is performed on the prior latent variable encoder and the language decoder according to the first latent variable and the obtained new video content description, so that a prior latent variable can be obtained from the prior latent variable encoder, and a decoder that can sufficiently recover the original description according to the latent variable can be obtained, and then a video description model that can generate diverse video descriptions can be obtained.
[0071] Another embodiment of the present invention also provides a training method, as Figure 2 shown, Figure 2 is Figure 1Schematic diagram of another embodiment of the refined process of step S130, which includes but is not limited to:
[0072] Step S210: Input the verb and the template into the posterior latent variable encoder for encoding to obtain the encoded action latent variable and the encoded template latent variable.
[0073] Step S220: Input the encoded template latent variable into a preset first neural network model to obtain the posterior distribution of the template latent variable.
[0074] Step S230: Input the encoded action latent variable and the posterior distribution of the template latent variable into the first neural network model for processing to obtain the posterior distribution of the action latent variable.
[0075] Step S240: Perform feature sampling on the posterior distribution of the action latent variable and the posterior distribution of the template latent variable to obtain the first action latent variable feature and the first template latent variable feature.
[0076] Step S250: Obtain the first latent variable feature according to the first action latent variable feature and the first template latent variable feature.
[0077] Specifically, the posterior latent variable encoder first encodes each time point t(m t ) of the template m to obtain the encoded template latent variable. Among them, the preset first neural network model is a long short-term memory model (LSTM) and a fully connected network (FCNs). The encoded template latent variable infers the distribution of the template latent variable through a long short-term memory model (LSTM) and a fully connected network (FCNs). Under the architecture of variational auto-encoders (VAEs), this distribution is an approximate posterior distribution of the template latent variable. This posterior distribution is often an independent multivariate Gaussian distribution, and the output of the FCNs is the mean and variance of the Gaussian distribution of the template information.
[0078] Specifically, the posterior latent variable encoder encodes the verb to obtain the encoded action latent variable, and inputs the encoded action latent variable and the posterior distribution of the template latent variable into a long short-term memory model (LSTM) and a fully connected network (FCNs) to infer the distribution of the action latent variable. Similarly, this distribution is an approximate posterior distribution of the action latent variable. The output of the FCNs is the mean and variance of the Gaussian distribution of the action information.
[0079] In one embodiment, feature sampling is performed on the posterior distribution of the action latent variable and the posterior distribution of the template latent variable to obtain the action latent variable feature and the template latent variable feature. The first latent variable feature is obtained based on the action latent variable feature and the template latent variable feature. The first latent variable feature can be used as a parameter to generate a new video content description.
[0080] Another embodiment of the present invention also provides a training method, as Figure 3 shown Figure 3 is Figure 1 a schematic diagram of another embodiment of the refined process of step S150 in
[0081] Step S310: Substitute the first latent variable feature and the visual feature into the language decoder to obtain an initial decoding result;
[0082] Step S320: Input the initial decoding result and the first latent variable feature into a preset loss function to train the prior latent variable encoder and the language decoder, and obtain a video content description model.
[0083] Specifically, combining the first latent variable feature with the video visual feature and inputting them into the decoder can obtain a new video content description, and then obtain an initial decoding result.
[0084] In one embodiment, inputting the initial decoding result and the first latent variable feature into a preset loss function to train the prior latent variable encoder and the language decoder aims to obtain a video content description model according to the initial decoding result and the first latent variable feature through the preset loss function.
[0085] Another embodiment of the present invention also provides a training method, as Figure 4 shown Figure 4 is Figure 3 a schematic diagram of another embodiment of the refined process of step S320 in
[0086] Step S410: Input the initial decoding result and the first latent variable feature into a preset loss function to obtain a loss value;
[0087] Step S420: Adjust the parameters of the language decoder according to the loss value to obtain a trained language decoder;
[0088] Step S430: Adjust the parameters of the latent variable encoder according to the loss value so that the target prior distribution is close to the first latent variable feature, and obtain a trained prior latent variable encoder;
[0089] Step S440: Obtain a video content description model based on the trained language decoder and the trained prior latent variable encoder.
[0090] Specifically, input the initial decoding result and the first latent variable feature into a preset loss function to train the prior latent variable encoder and the language decoder, and obtain a video content description model. The preset loss function is as follows:
[0091]
[0092] Among them, q and p are the posterior distribution and the prior distribution of the action latent variable and the template latent variable respectively. Since they are both Gaussian distributions, their KL distance can be explicitly obtained. x t-1 is the single word predicted by the decoder at the previous time point, and z t is the latent variable decoded by the current latent variable decoder. x t is the single word predicted at the current moment, and the value predicted by the neural network is used as the probability of obtaining x t by the model. Adjust the parameters of the language decoder according to the loss value to make the predicted decoding result close to the actual decoding result, and obtain the trained language decoder. Adjust the parameters of the latent variable encoder according to the loss value to make the target prior distribution close to the first latent variable feature, and obtain the trained prior latent variable encoder. This objective function aims to obtain a decoder that can recover the original description sufficiently based on the latent variable and a prior latent variable encoder that approximates the posterior distribution.
[0093] It should be noted that this loss function comes from the VAE framework. By approximating the actual posterior and the approximate posterior of the data, the lower bound of the data likelihood function (Evidence Lower Bound, ELBO) is finally deduced. The former term is responsible for the reconstruction loss, aiming to increase the faithfulness of the latent variable to recover the original description, manifested as the conditional likelihood function (conditional on the latent variable). The latter term is the distance between the prior and the posterior (using KL to represent the distance between two distributions). By shortening this distance, the prior distribution can be made closer to the posterior distribution. The posterior distribution is learned from the data and is informative, so the prior gradually acquires this ability. In the application stage, the posterior is unknown, while the prior distribution close to the posterior distribution, which is known in the prior distribution training stage, has the potential to "recover" the data.
[0094] Another embodiment of the present invention also provides a training method, as Figure 5 shown, Figure 5 is Figure 1 a schematic diagram of another embodiment of the refined process of step S140 in
[0095] Step S510: Use a pre-trained deep convolutional neural network to extract features from the video to obtain target visual features.
[0096] In one embodiment, the video encoder uses a pre-trained deep convolutional neural network (Convolutional Neural Networks, CNNs) to extract features from the video to obtain target visual features, aiming to input the target visual features and the first latent variable features into the language decoder for training the video content description model.
[0097] Another embodiment of the present invention also provides a training method. As Figure 6 shown, before Figure 5 step S510, it further includes but is not limited to:
[0098] Step S610: Preprocess the target video to obtain a preprocessed target video.
[0099] In one embodiment, before extracting features from the target video, it is necessary to preprocess the video, such as sampling, to meet the input requirements of the deep model.
[0100] Refer to Figure 7 , one embodiment of the present invention provides an application method of a video content description model. The application method includes but is not limited to:
[0101] Step S710: Obtain the video to be analyzed and extract the video visual features of the video to be analyzed;
[0102] Step S720: Obtain the second latent variable features according to the video content description model;
[0103] Step S730: Substitute the second latent variable features and the video visual features into the video content description model to obtain a video content description.
[0104] In one embodiment, the second latent variable features are obtained from the trained video content description model, and new video content descriptions are generated according to the second latent variable features and the extracted video visual features. When the obtained second latent variable features are different, the resulting video content descriptions are also different, that is, open-ended video descriptions can be generated, which can be used in a variety of different application scenarios to meet the actual demand for diverse video descriptions.
[0105] It should be noted that the video content description in the embodiments of the present invention is generated by obtaining the second latent variable features and video vision, effectively improving the generation efficiency of the video content description.
[0106] Another embodiment of the present invention also provides an application method of a video content description model. As Figure 8 shown, Figure 8 isFigure 7 Schematic diagram of another embodiment of the refined process of step S720, which includes but is not limited to:
[0107] Step S810, obtaining the prior distribution of the action latent variable and the prior distribution of the template latent variable according to the video content description model;
[0108] Step S820, respectively sampling the prior distribution of the action latent variable and the prior distribution of the template latent variable to obtain the second action latent variable feature and the second template latent variable feature;
[0109] Step S830, obtaining the second latent variable feature according to the second action latent variable feature and the second template latent variable feature.
[0110] In one embodiment, the prior distribution of the action latent variable and the prior distribution of the template latent variable are obtained from the video content description model. The prior distribution learned during the training phase already has the decoding function. Sampling from the learned prior distribution of the action latent variable and the prior distribution of the template latent variable decodes the corresponding description. Since the latent variable features sampled each time are different, the generated descriptions are also different, but they can all be used to describe the corresponding video.
[0111] In one embodiment, the prior distributions of the action latent variable and the template latent variable are obtained from the trained prior latent variable encoder. The action and template latent variable features can be sampled from the prior distribution, and then they are concatenated and used as the input of the decoder together with the video visual features to predict a new description. In the application stage, the present invention does not need to adopt the traditional beam search method, but samples the latent variable multiple times from the prior distribution. For each sampled latent variable, it can be input into the neural network to obtain a final sentence. At the same time, this sampling method belongs to the parallel method, that is, the generation processes between sentences can be independent of each other, which can effectively avoid the problems of the very wide search space and low efficiency caused by the serial search in the beam search method.
[0112] An embodiment of the present invention provides a training device for a video content description model, which includes but is not limited to:
[0113] The first acquisition module: used to acquire the target video and the content description of the target video;
[0114] The first processing module: used to extract the verb and the template from the content description, and the template is the remaining part after separating the verb from the target content description corresponding to the target video;
[0115] The feature processing module: used to perform encoding processing and feature sampling on the verb and the template to obtain the first latent variable feature;
[0116] Second acquisition module: used to extract features from the target video to obtain target visual features;
[0117] Model generation module: performs data approximation processing on the prior latent variable encoder and the language decoder using the first latent variable feature and the target visual feature pair to obtain a video content description model.
[0118] In one embodiment, the training device of the present invention obtains a target video and a content description of the target video, extracts verbs and templates from the content description, encodes and performs feature sampling on the verbs and templates to obtain a first latent variable feature, extracts features from the target video to obtain target visual features, and performs data approximation processing on the prior latent variable encoder and the language decoder according to the first latent variable feature and the target visual feature to obtain a video content description model. The present invention can generate open-ended descriptions according to the video description model, that is, generate various descriptions from many angles using different expressions.
[0119] An embodiment of the present invention provides an application device for a video content description model, and the application device includes but is not limited to:
[0120] Third acquisition module: used to acquire the video to be analyzed and extract the video visual features of the video to be analyzed;
[0121] Fourth acquisition module: used to obtain a second latent variable feature according to the video content description model;
[0122] Video content description generation module: used to substitute the second latent variable feature and the video visual feature into the video content description model to obtain a video content description, and the video content description model is trained according to the training method of any of the above embodiments of the present invention.
[0123] In one embodiment, the application device provided by the present invention obtains a second latent variable feature from a trained video content description model, generates a new video content description according to the second latent variable feature and the extracted video visual features. When the obtained second latent variable features are different, the obtained video content descriptions are also different, that is, open-ended video descriptions can be generated, and various different application scenarios are used to meet the actual needs for diverse video descriptions.
[0124] Among them, the specific execution steps of an application device for a video content description model refer to the above application method for a video content description model, and will not be elaborated here.
[0125] An embodiment of the present invention further provides a computer device, including: at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method and the application method in any one of the foregoing method embodiments.
[0126] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by one or more control processors. The one or more control processors execute the training method in the foregoing method embodiments. For example, the one or more control processors execute the Figure 1 method steps S110 to S150 in Figure 2 the method steps S210 to S250 in Figure 3 the method steps S310 to S320 in Figure 4 the method steps S410 to S440 in Figure 5 the method step S510 in Figure 6 the method step S610 in. The one or more control processors execute the application method in the foregoing method embodiments. For example, the one or more control processors execute the Figure 7 method steps S710 to S730 in Figure 8 the method steps S810 to S830 in.
[0127] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0128] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention. In addition, the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
Claims
1. A training method for a video content description model, characterized in that Including: Obtain a target video and a content description of the target video; Extract a verb and a template from the content description, where the template is the remaining part after separating the verb from the target content description corresponding to the target video; Perform encoding processing and feature sampling on the verb and the template to obtain a first latent variable feature; Extract features from the target video to obtain a target visual feature; Substitute the first latent variable feature and the target visual feature into a language decoder to obtain an initial decoding result; Input the initial decoding result and the first latent variable feature into a preset loss function to obtain a loss value; Adjust the parameters of the language decoder according to the loss value to obtain a trained language decoder; Adjust the parameters of the prior latent variable encoder according to the loss value so that the target prior distribution is close to the first latent variable feature to obtain a trained prior latent variable encoder; Obtain a video content description model based on the trained language decoder and the trained prior latent variable encoder.
2. The training method of the video content description model according to claim 1, wherein The performing encoding processing and feature sampling on the verb and the template to obtain a first latent variable feature includes: Input the verb and the template into a posterior latent variable encoder for encoding processing to obtain an encoded action latent variable and an encoded template latent variable; Input the encoded template latent variable into a preset first neural network model to obtain a posterior distribution of the template latent variable; Input the encoded action latent variable and the posterior distribution of the template latent variable into the first neural network model for processing to obtain a posterior distribution of the action latent variable; Perform feature sampling on the posterior distribution of the action latent variable and the posterior distribution of the template latent variable to obtain a first action latent variable feature and a first template latent variable feature; Obtain a first latent variable feature according to the first action latent variable feature and the first template latent variable feature.
3. The training method of the video content description model according to claim 1, wherein The extracting features from the target video to obtain a target visual feature includes: Use a pre-trained deep convolutional neural network to extract features from the video to obtain the target visual feature.
4. The training method of the video content description model according to claim 3, wherein Before extracting features from the target video to obtain a target visual feature, the training method further includes: Perform preprocessing on the target video to obtain a preprocessed target video.
5. A method for applying a video content description model, characterized in that, Including: Obtain a video to be analyzed and extract video visual features of the video to be analyzed; Obtain a second latent variable feature according to the video content description model; Substitute the second latent variable feature and the video visual feature into the video content description model to obtain a video content description, where the video content description model is trained according to the method described in any one of claims 1 to 4.
6. The application method of the video content description model according to claim 5, wherein, The obtaining a second latent variable feature according to the video content description model includes: Obtain a prior distribution of the action latent variable and a prior distribution of the template latent variable according to the video content description model; Respectively sample the prior distribution of the action latent variable and the prior distribution of the template latent variable to obtain a second action latent variable feature and a second template latent variable feature; Obtain a second latent variable feature according to the second action latent variable feature and the second template latent variable feature.
7. A computer device, characterized in that, The computer device includes a memory and a processor. Among them, a program is stored in the memory, and when the program is executed by the processor, the processor is used to execute: The training method according to any one of claims 1 to 4; or The application method according to any one of claims 5 to 6.
8. A storage medium, the storage medium being a computer-readable storage medium, characterized in that, The computer-readable storage stores a computer program, and when the computer program is executed by a computer, the computer is used to execute: The training method according to any one of claims 1 to 4; or The application method according to any one of claims 5 to 6.
Citation Information
Patent Citations
Video generation method combining variational auto-encoder and generative adversarial network
CN110572696A
Video content extraction method and device, computer equipment and storage medium
CN112966150A