Video generation method and device, electronic equipment, storage medium and computer product

By performing visual semantic feature extraction and content semantic feature extraction on the reference image, combined with semantic segmentation and directional modulation technology, the problem that the video generated in the prior art appears stiff and unnatural, and the generation of videos with variability and semantic coherence is achieved.

CN120198522AInactive Publication Date: 2025-06-24CHINA MOBILE M2M +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510176662.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to generate coherent, natural dynamic video sequences based on single still images, resulting in low video availability.

Method used

By performing visual semantic feature extraction and content semantic feature extraction on the reference image, the semantic feature map and content semantic encoding vector are obtained, semantic segmentation and directional modulation are performed, and the target video is generated.

Benefits of technology

The usability of videos generated when video generation is generated based on a single image is improved, so that the video content has variability and semantic coherence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198522A_ABST
    Figure CN120198522A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides a video generation method and device, electronic equipment, a storage medium and a computer product, and the method comprises the steps: carrying out the visual semantic feature extraction and content semantic feature extraction of a reference image, and obtaining a semantic feature map and a content semantic coding vector; performing semantic segmentation on the content semantic coding vector to obtain a first sequence containing at least one frame granularity coding vector; on the basis of each frame granularity coding vector in the first sequence, directional modulation of features is carried out on the semantic feature map, and a second sequence containing at least one directional modulation feature map is obtained; and performing video generation based on the second sequence to obtain a target video. According to the method and the device, the target video with variability and semantic coherence in content can be generated, and the availability of the video generated during video generation based on the single image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a video generation method, apparatus, electronic device, storage medium, and computer product. Background Art

[0002] With the rapid development of digital media technologies, the creation and generation of video content have become an important part of information dissemination and entertainment consumption. Traditional video production methods usually rely on manual shooting, editing, and post-processing, which is not only time-consuming and laborious but also difficult to quickly adapt to diverse content creation requirements. In recent years, breakthroughs in computer vision and deep learning technologies have brought revolutionary changes to the field of video generation. In particular, the image-to-video conversion technology has provided new possibilities for realizing automated and intelligent video generation.

[0003] In the research of video generation, how to generate a coherent and meaningful dynamic video sequence based on a single static image has always been a hot and difficult point in this field. Early research attempts mainly focused on using simple image deformation, color gradient, or motion models to simulate video effects. However, these methods often lack an understanding of the deep semantics of images, and the generated video content often appears rigid and unnatural. As a result, the usability of the generated videos when generating videos based on a single image is low. Summary of the Invention

[0004] This application aims to at least solve one of the technical problems existing in the related technologies. For this purpose, this application provides a video generation method, apparatus, electronic device, storage medium, and computer product to solve the problem that the currently generated video content appears rigid and unnatural, and to improve the usability of the generated videos when generating videos based on a single image.

[0005] According to the video generation method of the first aspect embodiment of this application, it includes: Performing visual semantic feature extraction and content semantic feature extraction on a reference image respectively to obtain a semantic feature map and a content semantic coding vector; Performing semantic segmentation on the content semantic coding vector to obtain a first sequence including at least one frame-level coding vector; Based on each frame-level coding vector in the first sequence, respectively performing directional modulation on the features of the semantic feature map to obtain a second sequence including at least one directionally modulated feature map; Generating a target video based on the second sequence.

[0006] According to an embodiment of the present application, the method of respectively performing feature directional modulation on the semantic feature map based on each frame-level encoding vector in the first sequence to obtain a second sequence including at least one directionally modulated feature map includes: Respectively extract the implicit common features between each frame-level encoding vector in the first sequence and the semantic feature map to obtain corresponding content-visual implicit common semantic feature vectors; Based on each of the content-visual implicit common semantic feature vectors, respectively perform dual feature modulation on the semantic feature map to obtain a second sequence including at least one directionally modulated feature map.

[0007] According to an embodiment of the present application, the method of respectively performing dual feature modulation on the semantic feature map based on each of the content-visual implicit common semantic feature vectors to obtain a second sequence including at least one directionally modulated feature map includes: Respectively input each of the content-visual implicit common semantic feature vectors into a feature preliminary screening module to obtain first-level semantic feature preliminary screening weight vectors output corresponding to the feature preliminary screening module; wherein, the feature preliminary screening module is determined based on a gated unit; Based on each of the first-level semantic feature preliminary screening weight vectors, respectively perform preliminary feature screening on the semantic feature map to obtain a preliminary screening content semantic guidance reference image semantic feature map; Respectively input each of the content-visual implicit common semantic feature vectors into a feature secondary screening module to obtain second-level semantic feature screening weight vectors output corresponding to the feature secondary screening module; wherein, the feature secondary screening module is determined based on a multi-level gated unit; For each of the content-visual implicit common semantic feature vectors, based on the second-level semantic feature screening weight vectors, perform secondary feature screening on the preliminary screening content semantic guidance reference image semantic feature map to obtain a directionally modulated feature map; Based on each of the directionally modulated feature maps, construct a second sequence.

[0008] According to an embodiment of the present application, performing preliminary feature screening on the semantic feature map based on the first-level semantic feature preliminary screening weight vectors to obtain a preliminary screening content semantic guidance reference image semantic feature map includes: Using the respective eigenvalues of the first-level semantic feature preliminary screening weight vectors as weights, perform position-wise weighting on each feature matrix of the semantic feature map along the channel dimension and then perform activation processing to obtain a preliminary screening content semantic guidance reference image semantic feature map.

[0009] According to an embodiment of the present application, the feature preliminary screening module is used for: After multiplying the preset weight parameter matrix with the content-visual implicit common semantic feature vector, perform a dot addition of the product with a preset bias term to obtain a first-level content-visual implicit common semantic modulation vector; Perform a non-linear activation process on the first-level content-visual implicit common semantic modulation vector to obtain a first-level content-visual implicit common semantic modulation activation vector; Input the first-level content-visual implicit common semantic modulation activation vector into a mask unit to obtain a first-level semantic feature preliminary screening weight vector output by the mask unit.

[0010] According to an embodiment of the present application, extracting the implicit common features between the frame-level encoding vector and the semantic feature map to obtain a content-visual implicit common semantic feature vector includes: Perform feature distillation on the semantic feature map to obtain a semantic distillation feature vector; Perform implicit correlation feature capture based on the frame-level encoding vector and the semantic distillation feature vector to obtain a content-visual implicit common semantic feature vector.

[0011] According to an embodiment of the present application, the performing feature distillation on the semantic feature map to obtain a semantic distillation feature vector includes: Perform point convolution encoding on the semantic feature map to obtain a channel modulation reference image semantic feature map; Perform global average pooling processing along the channel dimension on the channel modulation reference image semantic feature map and then perform an activation process to obtain the semantic distillation feature vector.

[0012] According to the video generation device of the second aspect embodiment of the present application, it includes: An extraction module, configured to perform visual semantic feature extraction and content semantic feature extraction on a reference image respectively to obtain a semantic feature map and a content semantic encoding vector; A segmentation module, configured to perform semantic segmentation on the content semantic encoding vector to obtain a first sequence including at least one frame-level encoding vector; A modulation module, configured to respectively perform directional modulation of features on the semantic feature map based on each frame-level encoding vector in the first sequence to obtain a second sequence including at least one directionally modulated feature map; A generation module, configured to perform video generation based on the second sequence to obtain a target video.

[0013] According to the electronic device of the third aspect embodiment of the present application, it includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the video generation method as described in any one of the above.

[0014] A storage medium according to an embodiment of the fourth aspect of the present application, the storage medium is a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the video generation method as described in any one of the above.

[0015] A computer program product according to an embodiment of the fifth aspect of the present application, including a computer program, and when the computer program is executed by a processor, it implements the video generation method as described in any one of the above.

[0016] One or more of the above technical solutions in the embodiments of the present application have at least the following technical effects: After respectively performing visual semantic feature extraction and content semantic feature extraction on a reference image to obtain a semantic feature map and a content semantic coding vector, by performing semantic segmentation on the content semantic coding vector, a first sequence including at least one frame-level coding vector is obtained. Furthermore, based on each frame-level coding vector in the first sequence, the semantic feature map can be respectively subjected to directional modulation of features to obtain a second sequence including at least one directionally modulated feature map, so that video generation can be performed based on the second sequence to obtain a target video. Since semantic segmentation is performed on the content semantic coding vector, the overall semantic information of the reference image is refined to a more specific frame-level, and then according to each frame-level coding vector, directional modulation of features is performed on the semantic feature map respectively, the semantic information of the frame-level coding vector can be incorporated into the visual features. Furthermore, according to the rich semantic information and visual features contained in each directionally modulated feature map, a target video with variable content and coherent semantics can be generated, improving the usability of the video generated when generating a video based on a single image.

[0017] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a schematic flowchart of the video generation method provided by the embodiments of the present application.

[0020] Figure 2 It is a schematic structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following further describes the implementation manners of the present application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.

[0022] In the description of the embodiments of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the embodiments of the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation on the embodiments of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0023] In the description of the embodiments of the present application, it should be noted that unless otherwise clearly specified and defined, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific situations.

[0024] In the embodiments of the present application, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on" the second feature can be that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is at a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature can be that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is at a lower horizontal height than the second feature.

[0025] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0026] This application provides a video generation method, apparatus, electronic device, storage medium, and computer product.

[0027] Figure 1 It is a schematic flowchart of the video generation method provided by an embodiment of this application. As Figure 1 shown, the video generation method includes: Step 110: Extract visual semantic features and content semantic features from a reference image respectively to obtain a semantic feature map and a content semantic coding vector.

[0028] Step 120: Perform semantic segmentation on the content semantic coding vector to obtain a first sequence including at least one frame-level coding vector.

[0029] Step 130: Based on each frame-level coding vector in the first sequence, respectively perform directional modulation on the features of the semantic feature map to obtain a second sequence including at least one directionally modulated feature map.

[0030] Step 140: Generate a target video based on the second sequence.

[0031] It should be noted that the execution subject of the video generation method provided by the embodiment of this application can be a computer device, etc. The computer device can be, for example, a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an Ultra-mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc. It should be noted that all the data that needs to be obtained in this application is obtained through formal channels after being authorized by relevant users.

[0032] A video generation apparatus can be set or connected in the computer device of this application, whereby the video generation apparatus can be controlled to execute the video generation method of this application.

[0033] It should be noted that currently, diffusion models are mostly used to generate videos based on a single image. However, although diffusion models have made remarkable progress in video generation, there is still a risk that the image content of each video frame in the generated frame sequence is exactly the same, resulting in the generated video content not having a content change effect.

[0034] Based on this, the present application uses image processing technology based on deep learning to extract multi-level features of visual semantics and content semantics from the reference image. After semantic segmentation of the content semantic encoding features of the reference image, each local content semantic feature after segmentation is used as the semantic guidance for each image frame of the target-generated video, and the visual semantic features of the reference image are directionally modulated respectively, so as to intelligently generate a sequence of target video image frames. In this way, the problems existing in the traditional video generation method can be effectively solved, and the generated video content has variability and semantic coherence.

[0035] Specifically, the present application can obtain a reference image. It can be understood that the reference image is the object of algorithm analysis in the present application, and it contains the initial visual information and semantic information required for video generation. By analyzing this image, the visual elements, objects and their relationships in the image can be extracted, and these information will be used to guide the generation of subsequent video frames.

[0036] More specifically, the user can select a photo stored on the device or directly take a photo through the camera in a mobile application or a desktop application as the reference image. In addition, there is also a situation in a web application where the user may upload an image as the basis for generating a video. Of course, the above are just examples. The specific process of obtaining the reference image mainly depends on the user's needs and application environment, and can be completed by selection, shooting or other means.

[0037] It should be understood that considering that video generation is not simply copying or transforming a single image, but requires understanding the visual elements (such as color, texture, shape) and content semantics (such as objects, scenes) in the image, and using this information to create a series of coherent image frames to form a video with a dynamic effect. Visual semantic features help capture the appearance attributes of the image, while content semantic features focus on the specific objects in the image and their relationships. The combination of the two makes the generated video have both visual beauty and logical coherence.

[0038] Therefore, in order to fully extract the visual information in the reference image, the present application uses a dilated convolutional neural network model with excellent performance in image processing tasks as the semantic feature extractor of the reference image to extract features from the reference image, capture visual semantic features such as color distribution and object shape in the reference image, so as to obtain the semantic feature map of the reference image. It is worth mentioning that the dilated convolutional neural network optimizes the convolutional operation through dilated convolution, which can expand the receptive field while maintaining the resolution of the feature map, and further improve the efficiency and effect of feature extraction.

[0039] Specifically, the reference image is input into the semantic feature extractor of the reference image based on the dilated convolutional neural network model, and the semantic feature map of the reference image output by the semantic feature extractor is obtained.

[0040] Among them, the dilated convolutional neural network model is a special convolutional neural network architecture. It expands the receptive field by changing the sampling step between the convolutional kernel and the input, that is, introducing the so-called "dilation rate" or "expansion rate", so as to obtain a larger field of view without increasing the number of parameters. This method allows the network to effectively capture a larger range of context information without changing the input size.

[0041] Next, considering that it is usually difficult to generate variable video content relying solely on the visual semantic information of the reference image. Therefore, in order to enhance the coherence and variability of the generated video content, the present application also extracts the image content semantic information from the reference image, converts the visual information in the image into a text feature description, so as to generate the content semantic coding vector of the reference image, thereby using the semantic changes in the text feature description to guide the generation of video content, so as to generate video content with content variability and semantic coherence.

[0042] Specifically, the present application first processes the reference image using an image semantic recognizer based on an image encoding and decoding structure. Among them, the image encoding and decoding structure can include an image encoder and an image semantic decoder. The image encoder can be a depthwise separable convolutional neural network model. By separating the spatial convolution and the channel convolution, the computational amount is reduced and the processing speed is increased, and the image features of the reference image are fully extracted. The image semantic decoder uses an RNN model to utilize the serialization information processing ability of the RNN model to process the image features and decode the encoded image features into a text feature description, that is, the content semantic coding vector of the reference image, so as to realize the in-depth understanding and text-based feature representation of the image content.

[0043] Among them, the depthwise separable convolutional neural network model is a convolutional neural network architecture. It reduces the computational amount and improves the network efficiency by decomposing the standard convolution operation into two independent steps. In the traditional convolutional layer, each output channel is formed by adding the results of convolving all channels of the input through a set of weights.

[0044] The RNN model (Recurrent Neural Network Model) is an artificial neural network model specifically designed to process sequential data. Different from traditional feedforward neural networks, the RNN has a memory function and can process input sequences with time dependence, making it very suitable for tasks such as time series prediction, natural language processing, and speech recognition. In the RNN, the state of the hidden layer depends not only on the input at the current moment but also on the hidden state at the previous moment, which enables the network to capture the time dynamic characteristics in the input data.

[0045] In the video generation method based on the semantics of a single image, the RNN model is used as part of the image semantic decoder to utilize its ability to process sequential data to decode image features and convert them into text feature descriptions, and then generate a semantic encoding vector for the content of the reference image. This helps to generate video content with coherence and variability, ensuring that the generated video is not only visually beautiful but also logically coherent.

[0046] Furthermore, the semantic encoding vector of the content of the reference image can be semantically segmented to refine the overall semantic information of the image to a more specific frame granularity level, generating a sequence of semantic frame granularity encoding vectors for the content of the reference image (which can be abbreviated as frame granularity encoding vectors hereafter) to achieve content guidance and feature differentiation control for each image frame in the target video.

[0047] In an embodiment of the present application, the semantic encoding vector of the content of the reference image can be evenly segmented based on the number of generated image frames to obtain a sequence of semantic frame granularity encoding vectors for the content of the reference image.

[0048] In another embodiment of the present application, the semantic encoding vector of the content of the reference image can also be textually segmented based on a large language model to obtain a sequence of semantic frame granularity encoding vectors for the content of the reference image and use it as the first sequence.

[0049] Furthermore, each frame granularity encoding vector in the sequence of semantic frame granularity encoding vectors for the content of the reference image is used to respectively perform feature-oriented modulation on the underlying visual features of the reference image, that is, the semantic feature map of the reference image, so as to integrate the semantic information of the frame granularity encoding vector into the visual features to facilitate the generation of video frames with different image contents.

[0050] It is worth mentioning that, in order to improve the effect of feature fusion, the present application proposes a fusion method based on multi-modal implicit association, which can effectively identify and utilize the potential connection between the frame-level encoded vector and the semantic feature map of the reference image, and through the feature selection and recombination mechanism, to achieve deeper semantic fusion and obtain a second sequence containing at least one content semantic-guided reference image semantic directional modulation feature map (which can be abbreviated as directional modulation feature map hereinafter).

[0051] Finally, each directional modulation feature map in the sequence of directional modulation feature maps is input into the image frame generator of the target video based on the diffusion model for processing, so as to utilize the powerful image generation ability of the diffusion model and generate a sequence of image frames of the target video based on the rich semantic information and visual features contained in each directional modulation feature map.

[0052] Among them, the diffusion model is a framework based on a probabilistic generative model. In the field of video generation, the diffusion model generates new data points, such as images or video frames, through a process of gradually adding and removing noise. Its core idea is to gradually restore the random noise into an approximation of the original data through a reverse process, thereby realizing the generation of new data.

[0053] In the application of video generation, the diffusion model is used to generate a sequence of image frames of the target video. Specifically, when generating a video based on a single image, the sequence of content semantic-guided reference image semantic directional modulation feature maps obtained through the previous steps will be input into the image frame generator of the target video based on the diffusion model. The diffusion model utilizes its powerful generation ability and generates a series of coherent and diverse image frames based on the semantic information and visual features contained in these feature maps.

[0054] Thus, the target video can be generated according to each image frame in the sequence.

[0055] According to the video generation method of the embodiments of the present application, after visually semantically extracting and content semantically extracting the reference image respectively to obtain a semantic feature map and a content semantic encoding vector, by semantically segmenting the content semantic encoding vector, a first sequence containing at least one frame-level encoding vector is obtained. Furthermore, based on each frame-level encoding vector in the first sequence, the semantic feature map can be respectively subjected to feature directional modulation to obtain a second sequence containing at least one directionally modulated feature map, so that video generation can be performed based on the second sequence to obtain a target video. Since the content semantic encoding vector is semantically segmented, the overall semantic information of the reference image is refined to a more specific frame-level, and then according to each frame-level encoding vector, the semantic feature map is respectively subjected to feature directional modulation, the semantic information of the frame-level encoding vector can be incorporated into the visual features. Furthermore, according to the rich semantic information and visual features contained in each directionally modulated feature map, a target video with variable content and coherent semantics can be generated, improving the usability of the video generated when generating a video based on a single image.

[0056] Based on the above embodiments, based on each frame-level encoding vector in the first sequence, the semantic feature map is respectively subjected to feature directional modulation to obtain a second sequence containing at least one directionally modulated feature map, including: Respectively extract the implicit common features between each frame-level encoding vector in the first sequence and the semantic feature map to obtain corresponding content-visual implicit common semantic feature vectors; Based on each content-visual implicit common semantic feature vector, the semantic feature map is respectively subjected to dual feature modulation to obtain a second sequence containing at least one directionally modulated feature map.

[0057] Specifically, the present application can perform feature distillation on the semantic feature map of the reference image to obtain a semantic distillation feature vector of the reference image.

[0058] Furthermore, each frame-level encoding vector in the first sequence can be respectively input into a multi-modal implicit association feature capture network together with the semantic distillation feature vector of the reference image to obtain a content-visual implicit common semantic feature vector corresponding to the corresponding frame-level encoding vector of the reference image.

[0059] Furthermore, for each content-visual implicit common semantic feature vector, it can be respectively input into a feature preliminary screening module based on a gated unit to obtain a first-level reference image semantic feature preliminary screening weight vector (subsequently can be simply referred to as a first-level semantic feature preliminary screening weight vector).

[0060] Furthermore, based on the first-level semantic feature preliminary screening weight vector, the semantic feature map of the reference image can be preliminarily screened for features to obtain a preliminarily screened content semantic guided reference image semantic feature map.

[0061] For each content-visual implicit common semantic feature vector, it can also be input into a feature secondary screening module based on a multi-level gated unit respectively to obtain a secondary reference image semantic feature screening weight vector (which can be abbreviated as a secondary semantic feature screening weight vector hereinafter).

[0062] Furthermore, based on the secondary semantic feature screening weight vector, feature secondary screening can be performed on the pre-screened content semantic guidance reference image semantic feature map to obtain an orientation modulation feature map.

[0063] Based on the above process, a sequence of orientation modulation feature maps can be obtained and used as the second sequence.

[0064] In this application, each frame-level encoding vector is used to perform feature orientation modulation on the underlying visual features of the reference image, that is, the semantic feature map, so as to integrate the semantic information of the frame-level encoding vector into the visual features, facilitating the generation of video frames with different image contents. Therefore, a target video with variable content and coherent semantics can be generated, improving the usability of the video generated when generating a video based on a single image.

[0065] Based on the above embodiments, implicit common features between the frame-level encoding vector and the semantic feature map are extracted to obtain a content-visual implicit common semantic feature vector, including: Performing feature distillation on the semantic feature map to obtain a semantic distillation feature vector; Performing implicit associated feature capture based on the frame-level encoding vector and the semantic distillation feature vector to obtain a content-visual implicit common semantic feature vector.

[0066] Specifically, this application first performs feature distillation on the semantic feature map of the reference image to remove redundant information therein, obtain a more compact and representative semantic distillation feature vector of the reference image, and unify the feature dimensions to promote effective feature interaction in subsequent stages.

[0067] Then, the frame-level encoding vector and the semantic distillation feature vector of the above reference image are input into an implicit associated feature capture network to identify and capture potential common patterns or structures between the two, thereby constructing a content-visual implicit common semantic feature vector of the reference image.

[0068] Specifically, the frame-level encoding vector and the semantic distillation feature vector can be cascaded and fused and then input into a neural network layer based on the tanh function, whereby a content-visual implicit common semantic feature vector of the reference image can be obtained. The tanh function, that is, the Hyperbolic Tangent Function, is an activation function commonly used in machine learning and deep learning.

[0069] The present application can effectively identify and utilize the potential connection between the frame-level encoded vector and the semantic feature map, and through the feature selection and recombination mechanism, achieve deeper semantic fusion, facilitating the subsequent generation of video frames with different image contents based on the content-visual implicit common semantic feature vector. Therefore, it is possible to generate a target video with variable content and coherent semantics, improving the usability of the video generated when generating a video based on a single image.

[0070] Based on the above embodiments, feature distillation is performed on the semantic feature map to obtain a semantic distilled feature vector, including: Perform point convolution encoding on the semantic feature map to obtain a channel-modulated reference image semantic feature map; Perform global average pooling processing along the channel dimension on the channel-modulated reference image semantic feature map and then perform activation processing to obtain a semantic distilled feature vector.

[0071] Specifically, when performing feature distillation on the semantic feature map, point convolution encoding can be specifically performed on the semantic feature map to obtain a channel-modulated reference image semantic feature map.

[0072] Furthermore, global average pooling processing along the channel dimension can be performed on the channel-modulated reference image semantic feature map and then input into the Sigmoid function for activation processing, thereby obtaining the semantic distilled feature vector of the reference image. The Sigmoid function is a common activation function widely used in the fields of machine learning and deep learning.

[0073] The present application performs feature distillation on the semantic feature map to remove redundant information therein, obtain a more compact and representative semantic distilled feature vector of the reference image, unify the feature dimensions, promote effective feature interaction in the subsequent stage, contribute to generating a target video with variable content and coherent semantics, and improve the usability of the video generated when generating a video based on a single image.

[0074] Based on the above embodiments, based on each content-visual implicit common semantic feature vector, double feature modulation is respectively performed on the semantic feature map to obtain a second sequence including at least one directional modulation feature map, including: Input each content-visual implicit common semantic feature vector into the feature preliminary screening module respectively to obtain a first-level semantic feature preliminary screening weight vector output corresponding to the feature preliminary screening module; wherein, the feature preliminary screening module is determined based on a gated unit; Based on each first-level semantic feature preliminary screening weight vector, perform preliminary feature screening on the semantic feature map respectively to obtain a preliminary screening content semantic guidance reference image semantic feature map; Input each content-visual implicit common semantic feature vector into the feature secondary screening module to obtain the secondary semantic feature screening weight vector output by the feature secondary screening module; wherein, the feature secondary screening module is determined based on a multi-level gated unit. For each content-visual implicit common semantic feature vector, based on the secondary semantic feature screening weight vector, perform feature secondary screening on the initially screened content semantic guidance reference image semantic feature map to obtain an orientation modulation feature map. Based on each orientation modulation feature map, construct a second sequence.

[0075] Specifically, the present application can perform gated masking processing on each content-visual implicit common semantic feature vector based on the gated masking mechanism to achieve importance evaluation and selection of features, and generate a corresponding weight vector as the primary semantic feature initial screening weight vector.

[0076] Specifically, each content-visual implicit common semantic feature vector can be input into the feature initial screening module.

[0077] The feature initial screening module can multiply the reference image content-visual implicit common semantic feature vector by a preset weight parameter matrix (to distinguish it from the subsequent weight parameter matrix, the weight parameter matrix used by the feature initial screening module can be defined as the first weight parameter matrix), and then perform dot addition with a preset bias term (to distinguish it from the subsequent bias term, the bias term used by the feature initial screening module can be defined as the first bias term), thereby obtaining a primary reference image content-visual implicit common semantic modulation vector (which can be abbreviated as the primary content-visual implicit common semantic modulation vector).

[0078] Furthermore, input the primary content-visual implicit common semantic modulation vector into the Sigmoid function for non-linear activation processing, thereby obtaining a primary reference image content-visual implicit common semantic modulation activation vector (which can be abbreviated as the primary content-visual implicit common semantic modulation activation vector).

[0079] Furthermore, the primary content-visual implicit common semantic modulation activation vector can be input into the masking unit, thereby obtaining the primary reference image semantic feature initial screening weight vector output by the masking unit (which can be abbreviated as the primary semantic feature initial screening weight vector).

[0080] Specifically, the first masking threshold and the second masking threshold can be set according to actual needs, where the second masking threshold is twice the first masking threshold.

[0081] If the eigenvalue in the primary content-visual implicit common semantic modulation activation vector is greater than the second masking threshold, it remains unchanged.

[0082] If the eigenvalue in the first-level content - visually implicit common semantic modulation activation vector is greater than the first mask threshold and less than or equal to the second mask threshold, it is processed by halving.

[0083] If the eigenvalue in the first-level content - visually implicit common semantic modulation activation vector is less than or equal to the first mask threshold, it is set to 0.

[0084] Thus, the first-level semantic feature pre-screening weight vector can be obtained.

[0085] Furthermore, the semantic feature map can be preliminarily screened according to the first-level semantic feature pre-screening weight vector, so as to suppress the irrelevant or secondary parts therein, and obtain the pre-screened content semantic guidance reference image semantic feature map focusing on the content - visual associated features of the reference image.

[0086] Next, to further improve the model performance, the content - visually implicit common semantic feature vector of the reference image is subjected to multi-level gating screening again. The process of information transmission is more finely managed through a multi-level gating architecture, and the corresponding weight vector is generated as the second-level semantic feature screening weight vector to perform a secondary refinement process on the previously formed pre-screened content semantic guidance reference image semantic feature map.

[0087] Specifically, each content - visually implicit common semantic feature vector can be input into the feature secondary screening module.

[0088] The feature secondary screening module can multiply the first-level content - visually implicit common semantic modulation vector by a preset weight parameter matrix (which can be defined as the second weight parameter matrix), and then perform a dot addition with a preset bias term (which can be defined as the second bias term), thereby obtaining the second-level reference image content - visually implicit common semantic modulation vector (which can be abbreviated as the second-level content - visually implicit common semantic modulation vector).

[0089] Furthermore, the second-level content - visually implicit common semantic modulation vector is input into the Sigmoid function for non-linear activation processing, thereby obtaining the second-level reference image content - visually implicit common semantic modulation activation vector (which can be abbreviated as the second-level content - visually implicit common semantic modulation activation vector).

[0090] The second-level content - visually implicit common semantic modulation activation vector is input into the mask unit to obtain the second-level reference image semantic feature screening weight vector output by the mask unit (which can be abbreviated as the second-level semantic feature screening weight vector). Among them, the process of obtaining the weight vector through the mask unit can refer to the above process of determining the first-level reference image semantic feature pre-screening weight vector.

[0091] Furthermore, the weight vector can be screened based on the secondary semantic features, and the semantic features map of the reference image guided by the pre-screened content is secondarily screened for features, thereby obtaining the orientation modulation feature map.

[0092] Specifically, using each eigenvalue of the weight vector for screening the secondary semantic features, after weighted processing of each feature matrix along the channel dimension of the semantic features map of the reference image guided by the pre-screened content according to positions, it is then input into the Sigmoid function for activation processing to obtain the orientation modulation feature map.

[0093] Thus, through two targeted feature selections and reconstructions, a content semantic guided reference image semantic orientation modulation feature map (which can be abbreviated as the orientation modulation feature map) rich in significant correlation information between the content semantic information and the visual semantic information of the reference image is finally obtained.

[0094] Furthermore, a sequence can be constructed based on each orientation modulation feature map and determined as the second sequence.

[0095] In this application, the frame-level encoding vector and the semantic features map are processed with the orientation modulation formula of the features to obtain the content semantic guided reference image semantic orientation modulation feature map. Among them, the formula for feature orientation modulation is: ; ; ; ; ; ; ; ; ; Among them, represents the frame-level encoding vector of the content semantics of the reference image, represents the semantic features map of the reference image, represents point convolution, represents the average pooling operation, represents the Sigmoid activation function, represents the semantic distillation feature vector of the reference image, represents concatenation, 、 and respectively represent different weight parameter matrices, 、 and respectively represent different bias terms, represents the hyperbolic tangent function, represents the reference image content - visual implicit common semantic feature vector, represents the first - level reference image content - visual implicit common semantic modulation activation vector, represents the masking function, represents the first - level reference image semantic feature screening weight vector, represents the second - level reference image content - visual implicit common semantic modulation activation vector, represents the second - level reference image semantic feature screening weight vector, represents the first masking threshold, represents the matrix multiplication operation, represents the dot product, represents the preliminary screening content semantic - guided reference image semantic feature map, represents the content semantic - guided reference image semantic directional modulation feature map.

[0096] In this way, each frame - granularity encoding vector in the sequence of frame - granularity encoding vectors can be interactively processed with the semantic feature map respectively, so that a sequence of content semantic - guided reference image semantic directional modulation feature maps with different content characteristics can be generated. Therefore, a target video with variable content and coherent semantics can be generated, improving the usability of the video generated when generating a video based on a single image.

[0097] Based on the above - mentioned embodiments, based on the first - level semantic feature preliminary screening weight vector, the semantic feature map is preliminarily screened for features to obtain the preliminary screening content semantic - guided reference image semantic feature map, including: Using each eigenvalue of the first - level semantic feature preliminary screening weight vector as a weight, each feature matrix of the semantic feature map along the channel dimension is weighted by position and then activated to obtain the preliminary screening content semantic - guided reference image semantic feature map.

[0098] Specifically, using each eigenvalue of the first - level semantic feature preliminary screening weight vector as a weight, each feature matrix of the semantic feature map along the channel dimension is weighted by position and then input into the tanh function for activation processing, thereby obtaining the preliminary screening content semantic - guided reference image semantic feature map.

[0099] This application uses the preliminary screening weight vector to weight the feature map by position and then processes it through an activation function, which can highlight important features related to semantics and suppress unimportant features at the same time, enabling the model to focus more on features meaningful for the task, thus enhancing the model's understanding and expression ability of semantic information and helping to improve the usability of the video generated when generating a video based on a single image.

[0100] The video generation device provided by the present application will be described below. The video generation device described below can be correspondingly referred to the video generation method described above.

[0101] Furthermore, the present application also provides a video generation device.

[0102] The video generation device includes: An extraction module, configured to perform visual semantic feature extraction and content semantic feature extraction on a reference image respectively, to obtain a semantic feature map and a content semantic encoding vector; A segmentation module, configured to perform semantic segmentation on the content semantic encoding vector, to obtain a first sequence including at least one frame-level encoding vector; A modulation module, configured to respectively perform directional modulation of features on the semantic feature map based on each frame-level encoding vector in the first sequence, to obtain a second sequence including at least one directionally modulated feature map; A generation module, configured to perform video generation based on the second sequence, to obtain a target video.

[0103] For the video generation device of the present application, after performing visual semantic feature extraction and content semantic feature extraction on a reference image respectively to obtain a semantic feature map and a content semantic encoding vector, by performing semantic segmentation on the content semantic encoding vector, a first sequence including at least one frame-level encoding vector is obtained. Furthermore, based on each frame-level encoding vector in the first sequence, directional modulation of features can be respectively performed on the semantic feature map, to obtain a second sequence including at least one directionally modulated feature map, so that video generation can be performed based on the second sequence to obtain a target video. Since semantic segmentation is performed on the content semantic encoding vector, the overall semantic information of the reference image is refined to a more specific frame-level, and then based on each frame-level encoding vector, directional modulation of features is performed on the semantic feature map respectively, so that the semantic information of the frame-level encoding vector can be incorporated into the visual features. Furthermore, based on the rich semantic information and visual features contained in each directionally modulated feature map, a target video with variable content and coherent semantics can be generated, improving the usability of the video generated when generating a video based on a single image.

[0104] In one embodiment, the modulation module is specifically configured to: Extract implicit common features between each frame-level encoding vector in the first sequence and the semantic feature map respectively, to obtain corresponding content-visual implicit common semantic feature vectors; Based on each of the content-visual implicit common semantic feature vectors, perform dual feature modulation on the semantic feature map respectively, to obtain a second sequence including at least one directionally modulated feature map.

[0105] In one embodiment, the modulation module is further configured to: Input each of the content-visual implicit common semantic feature vectors into the feature preliminary screening module to obtain the first-level semantic feature preliminary screening weight vectors output corresponding to the feature preliminary screening module; wherein, the feature preliminary screening module is determined based on a gated unit; Based on each of the first-level semantic feature preliminary screening weight vectors, perform preliminary feature screening on the semantic feature map to obtain a preliminary screening content semantic guidance reference image semantic feature map; Input each of the content-visual implicit common semantic feature vectors into the feature secondary screening module to obtain the second-level semantic feature screening weight vectors output corresponding to the feature secondary screening module; wherein, the feature secondary screening module is determined based on a multi-level gated unit; For each of the content-visual implicit common semantic feature vectors, based on the second-level semantic feature screening weight vectors, perform secondary feature screening on the preliminary screening content semantic guidance reference image semantic feature map to obtain an orientation modulation feature map; Construct a second sequence based on each of the orientation modulation feature maps.

[0106] In one embodiment, the modulation module is further configured to: Using each eigenvalue of the first-level semantic feature preliminary screening weight vectors as weights, perform position-wise weighting on each feature matrix along the channel dimension of the semantic feature map and then perform activation processing to obtain a preliminary screening content semantic guidance reference image semantic feature map.

[0107] In one embodiment, the modulation module is further configured to: Perform feature distillation on the semantic feature map to obtain a semantic distillation feature vector; Perform implicit associated feature capture based on the frame granularity encoding vector and the semantic distillation feature vector to obtain a content-visual implicit common semantic feature vector.

[0108] In one embodiment, the modulation module is further configured to: Perform point convolution encoding on the semantic feature map to obtain a channel modulation reference image semantic feature map; Perform global average pooling processing along the channel dimension on the channel modulation reference image semantic feature map and then perform activation processing to obtain the semantic distillation feature vector.

[0109] Figure 2 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 2As shown, the electronic device may include: a processor 210, a communications interface 220, a memory 230, and a communication bus 240. Among them, the processor 210, the communications interface 220, and the memory 230 communicate with each other through the communication bus 240. The processor 210 may call the logical instructions in the memory 230 to execute the following method: respectively perform visual semantic feature extraction and content semantic feature extraction on a reference image to obtain a semantic feature map and a content semantic coding vector; Perform semantic segmentation on the content semantic coding vector to obtain a first sequence including at least one frame-granularity coding vector; Based on each frame-granularity coding vector in the first sequence, respectively perform directional modulation of features on the semantic feature map to obtain a second sequence including at least one directionally modulated feature map; Generate a target video based on the second sequence.

[0110] In addition, when the logical instructions in the above-mentioned memory 230 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0111] On the other hand, an embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the methods provided in the above-mentioned various embodiments. For example, it includes: respectively performing visual semantic feature extraction and content semantic feature extraction on a reference image to obtain a semantic feature map and a content semantic coding vector; Perform semantic segmentation on the content semantic coding vector to obtain a first sequence including at least one frame-granularity coding vector; Based on each frame-granularity coding vector in the first sequence, respectively perform directional modulation of features on the semantic feature map to obtain a second sequence including at least one directionally modulated feature map; Video generation is performed based on the second sequence to obtain a target video.

[0112] In another aspect, an embodiment of the present application also provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above embodiments. For example, it includes: respectively performing visual semantic feature extraction and content semantic feature extraction on a reference image to obtain a semantic feature map and a content semantic coding vector; Performing semantic segmentation on the content semantic coding vector to obtain a first sequence including at least one frame-level coding vector; Based on each frame-level coding vector in the first sequence, respectively performing directional modulation of features on the semantic feature map to obtain a second sequence including at least one directionally modulated feature map; Video generation is performed based on the second sequence to obtain a target video.

[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications, or equivalent replacements of the technical solutions of the present application do not depart from the spirit and scope of the technical solutions of the present application.

Claims

1. A video generation method, characterized in that: include: Perform visual semantic feature extraction and content semantic feature extraction on the reference image to obtain a semantic feature map and a content semantic encoding vector; Performing semantic segmentation on the content semantic coding vector to obtain a first sequence including at least one frame granularity coding vector; Based on each frame granularity coding vector in the first sequence, directional modulation of features is performed on the semantic feature map respectively to obtain a second sequence including at least one directional modulation feature map; Video generation is performed based on the second sequence to obtain a target video.

2. The video generation method according to claim 1, characterized in that: The step of performing directional modulation of features on the semantic feature maps based on each frame granularity coding vector in the first sequence to obtain a second sequence including at least one directional modulation feature map comprises: Respectively extracting implicit common features between each frame granularity coding vector in the first sequence and the semantic feature map to obtain corresponding content-visual implicit common semantic feature vectors; Based on each of the content-visual implicit common semantic feature vectors, dual feature modulation is performed on the semantic feature map to obtain a second sequence including at least one directional modulation feature map.

3. The video generation method according to claim 2, characterized in that: The method of performing dual feature modulation on the semantic feature map based on each of the content-visual implicit common semantic feature vectors to obtain a second sequence including at least one directional modulation feature map comprises: Inputting each of the content-visual implicit common semantic feature vectors into a feature preliminary screening module respectively, and obtaining a primary semantic feature preliminary screening weight vector outputted by the feature preliminary screening module; wherein the feature preliminary screening module is determined based on a gating unit; Based on the primary semantic feature preliminary screening weight vectors, the semantic feature maps are respectively subjected to preliminary feature screening to obtain a primary screening content semantics guided reference image semantic feature map; Input each of the content-visual implicit common semantic feature vectors into a feature secondary screening module to obtain a secondary semantic feature screening weight vector corresponding to the output of the feature secondary screening module; wherein the feature secondary screening module is determined based on a multi-level gating unit; For each of the content-visual implicit common semantic feature vectors, based on the secondary semantic feature screening weight vector, the primary screening content semantics guided reference image semantic feature map is subjected to secondary feature screening to obtain a directional modulation feature map; Based on each of the directional modulation feature maps, a second sequence is constructed.

4. The video generation method according to claim 3, characterized in that: Based on the primary semantic feature preliminary screening weight vector, the semantic feature map is subjected to preliminary feature screening to obtain a preliminary screening content semantics guided reference image semantic feature map, including: Taking each eigenvalue of the primary semantic feature preliminary screening weight vector as a weight, each feature matrix of the semantic feature map along the channel dimension is weighted by position and then activated to obtain a semantic feature map of the reference image guided by the preliminary screening content semantics.

5. The video generation method according to claim 3, characterized in that: The feature initial screening module is used for: After multiplying the content-visual implicit common semantic feature vector by a preset weight parameter matrix, a point addition is performed on the product and a preset bias term to obtain a first-level content-visual implicit common semantic modulation vector; Performing nonlinear activation processing on the primary content-visual implicit common semantic modulation vector to obtain the primary content-visual implicit common semantic modulation activation vector; The primary content-visual implicit common semantic modulation activation vector is input into the mask unit to obtain the primary semantic feature preliminary screening weight vector output by the mask unit.

6. The video generation method according to claim 2, characterized in that: Extracting implicit common features between the frame granularity coding vector and the semantic feature map to obtain a content-visual implicit common semantic feature vector, including: Performing feature distillation on the semantic feature map to obtain a semantic distillation feature vector; Based on the frame granularity encoding vector and the semantic distillation feature vector, implicit correlation features are captured to obtain a content-visual implicit common semantic feature vector.

7. The video generation method according to claim 6, characterized in that: The performing feature distillation on the semantic feature map to obtain a semantic distillation feature vector includes: Performing point convolution coding on the semantic feature map to obtain a channel modulated reference image semantic feature map; The semantic feature map of the channel-modulated reference image is subjected to global mean pooling processing along the channel dimension and then to activation processing to obtain the semantic distillation feature vector.

8. A video generating device, characterized in that: include: An extraction module is used to extract visual semantic features and content semantic features of the reference image to obtain a semantic feature map and a content semantic coding vector; A segmentation module, used for performing semantic segmentation on the content semantic coding vector to obtain a first sequence including at least one frame granularity coding vector; A modulation module, configured to perform directional modulation of features on the semantic feature map based on each frame granularity coding vector in the first sequence, to obtain a second sequence including at least one directional modulation feature map; A generation module is used to generate a video based on the second sequence to obtain a target video.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the video generating method according to any one of claims 1 to 7 is implemented.

10. A storage medium, the storage medium being a non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by a processor, the video generating method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the video generating method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Video generation method and device, electronic equipment and storage medium

    CN120475234A