Video generation method and device, computer equipment, readable storage medium and program product
The two-stage image generation model enhances video quality and efficiency by iteratively refining frames to form a coherent target video set, addressing the inefficiencies in existing image-to-video generation technologies.
Patent Information
- Application Number
- CN202510544582.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, the quality of image-generated videos is poor, resulting in low efficiency of generating videos.
By using the pre-trained first and second image generation models, the initial image set is gradually constructed, and the target video is generated after the preset number threshold is met, and the target video is generated using the target image set.
The accuracy of each frame of the image and the consistency of multiple images are improved, thereby improving the quality and efficiency of the generated video.
Smart Images

Figure CN120321470A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and in particular, to a video generation method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] Generating a video from an image refers to a technology of generating a dynamic video based on a given static image. With the development of artificial intelligence, although the technology of generating a video from an image has been widely applied, there are still problems with the poor quality of the generated video, which in turn leads to low efficiency in generating a video from an image.
[0003] Therefore, there is an urgent need for a method that can effectively improve the quality of the generated video to improve the efficiency of generating a video from an image. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a video generation method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the quality of the generated video to improve the efficiency of generating a video from an image.
[0005] In a first aspect, the present application provides a video generation method, including:
[0006] Obtain a first image, and input the first image into a pre-trained first image generation model to obtain a second image output by the image generation model;
[0007] Generate an initial image set based on the first image and the second image, and input the initial image set into a pre-trained second image generation model to obtain a new image output by the second image generation model;
[0008] Add the new image to the initial image set, and after adding, input the initial image set into the second image generation model again to obtain a new image output by the second image generation model again;
[0009] Continue to execute the step of adding the new image to the initial image set, and after adding, input the initial image set into the second image generation model again until the number of images in the initial image set is greater than a preset number threshold, determine the initial image set as a target image set, and generate a target video using the target image set.
[0010] In one embodiment, the first image generation model includes a key point detector, a latent behavior encoder, and an image frame encoder. Inputting the first image into the pre-trained first image generation model to obtain a second image output by the image generation model includes: inputting the first image into the key point detector to obtain a first key point mask feature vector of the first image output by the key point detector; inputting the first image and the first key point mask feature into the latent behavior encoder to obtain a first latent behavior hidden vector of the first image output by the latent behavior encoder; inputting the first image into the image frame encoder to obtain a first image hidden vector of the first image output by the image frame encoder, and determining the second image based on the first image hidden vector and the first latent behavior hidden vector.
[0011] In one embodiment, the first image generation model further includes a latent feature encoder and an image frame decoder. Determining the second image based on the first image hidden vector and the first latent behavior hidden vector includes: inputting the first image hidden vector and the first latent behavior hidden vector into the latent feature encoder to obtain a second image hidden vector of the second image output by the latent feature encoder; inputting the second image hidden vector into the image frame decoder to obtain the second image output by the image frame decoder.
[0012] In one embodiment, the second image generation model includes a key point detector and a latent behavior encoder. Inputting the initial image set into the pre-trained second image generation model to obtain a new image output by the second image generation model includes inputting the initial image set into the key point detector to obtain a set of key point mask feature vectors of the initial image set output by the key point detector; inputting the set of key point mask feature vectors and the initial image set into the latent behavior encoder to obtain a set of latent behavior hidden vectors of the initial image set output by the latent behavior encoder; obtaining a set of image hidden vectors of the initial image set, and determining the new image according to the set of image hidden vectors and the set of latent behavior hidden vectors.
[0013] In one embodiment, the second image generation model further includes a latent feature encoder and an image frame decoder. Determining the new image according to the set of image hidden vectors and the set of latent behavior hidden vectors includes: inputting the set of image hidden vectors and the set of latent behavior hidden vectors into the latent feature encoder to obtain an image hidden vector of the new image output by the latent feature encoder; inputting the image hidden vector of the new image into the image frame decoder to obtain the new image output by the image frame decoder.
[0014] In one embodiment, generating a target video by using the target image set includes: determining the timing information of each image according to the timestamps at which the images in the target image set are added to the initial image set; and encoding each image in the target image set by using the timing information to generate the target video.
[0015] In a second aspect, the present application further provides a video generation device, including:
[0016] An acquisition module, configured to acquire a first image and input the first image into a pre-trained first image generation model to obtain a second image output by the image generation model;
[0017] A first execution module, configured to generate an initial image set based on the first image and the second image and input the initial image set into a pre-trained second image generation model to obtain a new image output by the second image generation model;
[0018] A second execution module, configured to add the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again to obtain a new image output by the second image generation model again;
[0019] A generation module, configured to continue to execute the step of adding the new image to the initial image set and, after the addition, inputting the initial image set into the second image generation model until the number of images in the initial image set is greater than a preset number threshold, determining the initial image set as the target image set, and generating a target video by using the target image set.
[0020] In a third aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the embodiments in the first aspect are implemented.
[0021] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the embodiments in the first aspect are implemented.
[0022] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in any one of the embodiments in the first aspect are implemented.
[0023] The above video generation method, device, computer device, computer-readable storage medium, and computer program product first obtain a first image and input the first image into a pre-trained first image generation model to obtain a second image output by the image generation model. Then, an initial image set is generated based on the first image and the second image, and the initial image set is input into a pre-trained second image generation model to obtain a new image output by the second image generation model. The new image is added to the initial image set, and after the addition, the initial image set is input into the second image generation model again to obtain a new image output by the second image generation model again. Then, continue to execute the steps of adding the new image to the initial image set and, after the addition, inputting the initial image set into the second image generation model until the number of images in the initial image set is greater than a preset number threshold. The initial image set is determined as the target image set, and the target video is generated using the target image set. In the video generation method provided in this application, each frame of the image used to construct the target video is generated based on all the generated images used to construct the target video. On the one hand, the accuracy of each generated image can be improved. On the other hand, multiple images used to form the target video can have coherence, that is, the quality of the target image set is higher, so that the quality of the target video generated based on the target image set is higher, thereby effectively improving the efficiency of generating video from images. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0025] Figure 1 It is a flowchart of the video generation method in an embodiment;
[0026] Figure 2 It is a flowchart of the method for obtaining the second image output by the image generation model in an embodiment;
[0027] Figure 3 It is a flowchart of the method for determining the second image based on the first image latent vector and the first latent behavior latent vector in an embodiment;
[0028] Figure 4 It is a flowchart of the method for obtaining the new image output by the second image generation model in an embodiment;
[0029] Figure 5Flow diagram of a method for determining a new image according to a set of image latent vectors and a set of latent behavior vectors in an embodiment;
[0030] Figure 6 Flow diagram of a method for generating a target video using a target image set in an embodiment;
[0031] Figure 7 Schematic diagram of the training process of a latent behavior encoder in an embodiment;
[0032] Figure 8 Schematic diagram of the training process of an image frame encoder and an image frame decoder in an embodiment;
[0033] Figure 9 Schematic diagram of the training process of latent feature encoding in an embodiment;
[0034] Figure 10 Flow diagram of a video generation method in another embodiment;
[0035] Figure 11 Flow diagram of inputting a first image into a pre-trained first image generation model to obtain a second image output by the image generation model in an embodiment;
[0036] Figure 12 Flow diagram of inputting an initial image set into a second image generation model again to obtain a new image output again by the second image generation model in an embodiment;
[0037] Figure 13 Structural block diagram of a video generation device in an embodiment;
[0038] Figure 14 Internal structure diagram of a computer device in an embodiment;
[0039] Figure 15 Internal structure diagram of a computer device in another embodiment. Detailed implementation manners
[0040] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0041] Image-to-video generation refers to a technology for generating a dynamic video based on a given static image. With the development of artificial intelligence, although image-to-video generation technology has been widely used, there are still problems such as poor video generation quality, which in turn leads to low efficiency of image-to-video generation.
[0042] Therefore, there is an urgent need for a method that can effectively improve the quality of generated videos to improve the efficiency of generating videos from images.
[0043] In view of this, the present application provides a video generation method. Each image used to construct the target video is generated based on all the previously generated images used to construct the target video. This can not only improve the accuracy of each generated image but also make multiple images used to form the target video coherent. That is to say, the quality of the target image set is higher, so that the quality of the target video generated based on the target image set is higher, thereby effectively improving the efficiency of generating videos from images.
[0044] The video generation method provided by the present application can be executed by a computer device, which can be a terminal or a server.
[0045] In an exemplary embodiment, as Figure 1 shown, a video generation method is provided. The method includes the following steps:
[0046] Step 101, obtain a first image and input the first image into a pre-trained first image generation model to obtain a second image output by the image generation model.
[0047] Optionally, the first image is provided by the user and is an image used to generate the target video. The first image generation model can be a model pre-trained by those skilled in the art according to actual needs. The second image is an image generated by the first image generation model based on the first image and is used to generate the target video.
[0048] In an alternative embodiment, the first image can be an image containing human face information.
[0049] Exemplarily, the target video to be generated can be composed of multiple images. The first image and the second image are images among the multiple images. Specifically, the first image is the first frame image of the target video to be generated. The second image is the second frame image of the target video to be generated.
[0050] In some exemplary embodiments, the computer device can first obtain a first image provided by the user and used to generate the target video. After obtaining the first image, the computer device can input the first image into a pre-trained first image generation model, so that the first image generation model uses the first image as the first frame image of the target video to be generated and generates the second frame image of the target video to be generated based on the first frame image, that is, the second image.
[0051] Step 102: Generate an initial image set based on the first image and the second image, and input the initial image set into a pre-trained second image generation model to obtain a new image output by the second image generation model.
[0052] An image set refers to a set of images used to generate a target video. An initial image set refers to an image set in an initial state, and the images in the initial image set do not yet meet the conditions for generating the target video. For example, the image quality does not meet the requirements or the number of images does not meet the requirements.
[0053] The second image generation model can be a model pre-trained by a technician according to actual needs. The second image generation model can be used to generate new images for constructing the target video based on the initial image set.
[0054] Exemplarily, as described above, the target video can be composed of multiple images. The first image among the multiple images is provided by the user, the second image among the multiple images is generated by the first image generation model based on the first image, and the other images among the multiple images can be generated by the second image generation model.
[0055] In some exemplary embodiments, after the computer device acquires the first image, it can input the first image into the first image generation model to obtain the second image output by the first image generation model.
[0056] Further, the computer device can first construct an initial image set based on the first image and the second image, and then input the initial image set into the second image generation model to obtain a new image output by the second image generation model.
[0057] Specifically, as described above, the first image is the first frame image of the target video to be generated, the second image is the second frame image of the target video to be generated, the new image output by the second image generation model is the third frame image of the target video to be generated, and the images in the initial image set are arranged in chronological order.
[0058] Step 103: Add the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again to obtain a new image output again by the second image generation model.
[0059] In some exemplary embodiments, after the computer device generates an initial image set based on the first image and the second image, it can input the initial image set into the second image generation model to obtain a new image output by the second image generation model.
[0060] Further, after obtaining a new image, the computer device may add the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again to obtain a new image output by the second image generation model again.
[0061] Step 104: Continue to execute the step of adding the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again until the number of images in the initial image set is greater than a preset number threshold. Then, determine the initial image set as the target image set, and generate a target video using the target image set.
[0062] The preset number threshold can be set in advance by those skilled in the art according to actual needs. It can also be determined according to the duration of the target video to be generated.
[0063] In some exemplary embodiments, after the computer device obtains a new image output by the second image generation model again, it may add the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again until the number of images in the initial image set is greater than a preset number threshold.
[0064] Further, when the number of images in the initial image set is greater than the preset number threshold, the initial image set may be determined as the target image set, and a target video may be generated using the target image set.
[0065] Specifically, taking the preset number threshold as 4 as an example, the whole process is described as follows. The computer device may first obtain a first image, input the first image into the first image generation model to obtain a second image, then construct an initial image set based on the first image and the second image, input the initial image set into the second image generation model to obtain a new image, that is, the third image, then add the third image to the initial image set, and then input the initial image set into the second image generation model to obtain a new image, that is, the fourth image, then add the fourth image to the initial image set, and input the initial image set into the second image generation model again to obtain a new image, that is, the fifth image, and add the fifth image to the initial image set. At this time, there are five images in the initial image set, which is greater than the preset number threshold, so the initial image set can be determined as the target image set.
[0066] Further, the computer device may input the target image set into a pre-trained target video generation model to obtain a target video output by the target video generation model.
[0067] The above video generation method first obtains a first image and inputs the first image into a pre-trained first image generation model to obtain a second image output by the image generation model. Then, an initial image set is generated based on the first image and the second image, and the initial image set is input into a pre-trained second image generation model to obtain a new image output by the second image generation model. The new image is added to the initial image set, and after the addition, the initial image set is input into the second image generation model again to obtain a new image output by the second image generation model again. Then, continue to execute the steps of adding the new image to the initial image set and, after the addition, inputting the initial image set into the second image generation model until the number of images in the initial image set is greater than a preset number threshold. The initial image set is determined as the target image set, and the target video is generated using the target image set. In the video generation method provided by this application, each frame of the image used to construct the target video is generated based on all the generated images used to construct the target video. On the one hand, the accuracy of each generated image can be improved. On the other hand, multiple images used to form the target video can be made coherent, that is, the quality of the target image set is higher, so that the quality of the target video generated based on the target image set is higher, thereby effectively improving the efficiency of generating a video from images.
[0068] In an exemplary embodiment, as Figure 2 shown, the first image generation model includes a key point detector, a latent behavior encoder, and an image frame encoder. The step of inputting the first image into a pre-trained first image generation model to obtain a second image output by the image generation model includes the following steps:
[0069] Step 201: Input the first image into the key point detector to obtain a first key point mask feature vector of the first image output by the key point detector.
[0070] Optionally, the key point detector can be used to detect key points in the image. Specifically, as described above, the first image can be an image containing human face information, and the key point detector can be used to detect the key points for characterizing the face features in the first image.
[0071] For example, these key points for characterizing face features can include the endpoints and midpoints of the eyebrows, the corners and pupil centers of the eyes, the tip of the nose and both sides of the nostrils, the corners of the mouth, the midpoints of the upper and lower lips, etc. By accurately detecting these key points, each part of the human face can be accurately located, and expression analysis can also be achieved.
[0072] Exemplarily, after the computer device inputs the first image containing face information into the key point detector, the key point detector will detect and mark the key points in the first image. The first key point mask can be regarded as a representation of the positions and distributions of these key points on the first image. The feature vector refers to a set of numerical feature representations extracted from the first key point mask.
[0073] In some exemplary embodiments, after the computer device obtains the first image, it can input the first image into the key point detector in the first image generation model to obtain the first key point mask feature vector of the first image output by the key point detector.
[0074] Step 202: Input the first image and the first key point mask feature into the latent behavior encoder to obtain the first latent behavior hidden vector of the first image output by the latent behavior encoder.
[0075] Optionally, the latent behavior encoder refers to an encoder that can be used to analyze latent behaviors. Specifically, as described above, the first image can be an image containing face information, and the latent behavior encoder can analyze possible behavioral intentions based on the first image and the first key point mask feature, such as whether nodding, shaking the head, smiling, etc. are being performed, or the person is in a certain emotional state.
[0076] The first latent behavior hidden vector can contain the key information related to the latent behavior in the first image and is presented in the form of a vector. Each dimension in the vector corresponds to a latent behavioral feature or pattern and can reflect the behavioral state corresponding to the first image.
[0077] In some exemplary embodiments, after the computer device obtains the first key mask feature, it can input the first image and the first key point mask feature into the latent behavior encoder to obtain the first latent behavior hidden vector of the first image output by the latent behavior encoder.
[0078] Step 203: Input the first image into the image frame encoder to obtain the first image hidden vector of the first image output by the image frame encoder, and determine the second image based on the first image hidden vector and the first latent behavior hidden vector.
[0079] Optionally, the image frame encoder refers to an encoder that can be used to analyze images. Specifically, as described above, the first image can be an image containing face information, and the image frame encoder can analyze features such as the outline of the face and the shapes of facial features based on the first image.
[0080] The first image hidden vector can contain the key information of the first image, such as color, shape, texture, etc., and is presented in the form of a vector.
[0081] In some exemplary embodiments, the computer device may further input the first image into an image frame encoder to obtain a first image latent vector of the first image output by the image frame encoder, and determine a second image based on the first image latent vector and the first latent behavior latent vector.
[0082] Specifically, the first image latent vector and the first latent behavior latent vector may be first fused, and then the second image is determined based on the fused vector.
[0083] In the method of inputting the first image into the key point detector to obtain the first key point mask feature vector of the first image output by the key point detector, inputting the first image and the first key point mask feature into the latent behavior encoder to obtain the first latent behavior latent vector of the first image output by the latent behavior encoder, inputting the first image into the image frame encoder to obtain the first image latent vector of the first image output by the image frame encoder, and determining the second image based on the first image latent vector and the first latent behavior latent vector, on the one hand, the key point detector can determine the prior probability distribution of the key points, and the rich local information can assist the subsequent encoders to achieve more accurate performance, and is robust to image viewpoint changes and illumination changes. On the other hand, adding the fusion of the key mask feature vectors can reduce the sensitivity to subtle differences and maintain the visual consistency of the target video to be generated. Moreover, fusing the key point feature vector with the image latent vector brings prior knowledge that enhances the interpretability of the encoder and the semantic representation of the face features, which is beneficial to exploring local information to capture fine-grained features, including the positions and shapes of organs such as eyes and noses, and enhancing the understanding of global information such as expressions and postures, and can better depict the geometric structure and attributes of the face, including the generation of detailed information such as expressions, to obtain better realism, and at the same time improve the generalization ability in complex scenarios such as occlusion and illumination changes.
[0084] In an exemplary embodiment, as Figure 3 shown, the first image generation model further includes a latent feature encoder and an image frame decoder. Determining the second image based on the first image latent vector and the first latent behavior latent vector includes the following steps:
[0085] Step 301: Input the first image latent vector and the first latent behavior latent vector into the latent feature encoder to obtain a second image latent vector of the second image output by the latent feature encoder.
[0086] Exemplarily, a latent feature encoder can be used to perform a fusion process on the first image latent vector and the first latent behavior latent vector. Specifically, as described above, the first image can be an image containing facial information. The first image latent vector may represent the basic contour and facial feature of the face, and the first latent behavior latent vector represents the facial expression and movement. Then the latent feature encoder will synthesize these two aspects of information and extract higher-level feature combinations, such as "round face with a smile" and "long face with a frown".
[0087] The second image latent vector can contain key information of the second image, such as color, shape, texture, etc., and is presented in the form of a vector.
[0088] In some exemplary embodiments, after obtaining the first image latent vector and the first latent behavior latent vector, the computer device can input the first image latent vector and the first latent behavior latent vector into the latent feature encoder to obtain the second image latent vector of the second image output by the latent feature encoder.
[0089] Step 302: Input the second image latent vector into the image frame decoder to obtain the second image output by the image frame decoder.
[0090] Optionally, the image frame decoder refers to a decoder used to convert an image latent vector into an image.
[0091] In some exemplary embodiments, after obtaining the second image latent vector of the second image output by the latent feature encoder, the computer device can input the second image latent vector of the second image into the image frame decoder to obtain the second image output by the image frame decoder.
[0092] In an exemplary embodiment, as Figure 4 shown, the second image generation model includes a key point detector and a latent behavior encoder. Inputting the initial image set into the pre-trained second image generation model to obtain a new image output by the second image generation model includes the following steps:
[0093] Step 401: Input the initial image set into the key point detector to obtain the key point mask feature vector set of the initial image set output by the key point detector.
[0094] In some exemplary embodiments, after obtaining the initial image set, the computer device can input the initial image set into the key point detector to obtain the key point mask feature vector set of the initial image set output by the key point detector.
[0095] For example, if the initial image set includes Image A, Image B, and Image C, the set of key-point mask feature vectors of the initial image set output by the key-point detector includes the key-point mask feature vector of Image A, the key-point mask feature vector of Image B, and the key-point mask feature vector of Image C.
[0096] Step 402: Input the set of key-point mask feature vectors and the initial image set into the latent behavior encoder to obtain the set of latent behavior hidden vectors of the initial image set output by the latent behavior encoder.
[0097] In some exemplary embodiments, after the computer device obtains the set of key-point mask feature vectors of the initial image set, it may input the set of key-point mask feature vectors and the initial image set into the latent behavior encoder to obtain the set of latent behavior hidden vectors of the initial image set output by the latent behavior encoder.
[0098] For example, if the initial image set includes Image A, Image B, and Image C, and the set of key-point mask feature vectors of the initial image set output by the key-point detector includes the key-point mask feature vector of Image A, the key-point mask feature vector of Image B, and the key-point mask feature vector of Image C, then the set of latent behavior hidden vectors of the initial image set includes the latent behavior hidden vector of Image A, the latent behavior hidden vector of Image B, and the latent behavior hidden vector of Image C.
[0099] Step 403: Obtain the set of image hidden vectors of the initial image set, and determine the new image according to the set of image hidden vectors and the set of latent behavior hidden vectors.
[0100] Exemplarily, the set of image hidden vectors of the initial image set can be directly obtained. The initial image set includes the first image, the second image, and the new images generated in each loop process. The set of image hidden vectors of the initial image set includes the image hidden vectors of each image in the initial image set. Among them, the first image hidden vector and the second image hidden vector of the first image and the second image are obtained from the first image generation model, and the others are obtained from the second image generation model.
[0101] For example, after generating the initial image set based on the first image and the second image, the initial image set can be input into a pre-trained second image generation model. During the process of the second image generation model generating new images, it will also generate the image hidden vectors of the new images. Therefore, after adding the new images to the initial image set and inputting the initial image set into the second image generation model again, during the process of the second image generation model generating new images, it can directly obtain the image hidden vectors of the new images generated in the previous loop process.
[0102] In some exemplary embodiments, after the computer device obtains the set of image latent vectors of the initial image set, it may determine the new image according to the set of image latent vectors and the set of potential behavior latent vectors.
[0103] Specifically, the computer device may perform a fusion process on the set of image latent vectors and the set of potential behavior latent vectors to obtain a new image.
[0104] In one exemplary embodiment, as Figure 5 shown, the second image generation model further includes a latent feature encoder and an image frame decoder. Determining the new image according to the set of image latent vectors and the set of potential behavior latent vectors includes the following steps:
[0105] Step 501: Input the set of image latent vectors and the set of potential behavior latent vectors into the latent feature encoder to obtain the image latent vector of the new image output by the latent feature encoder.
[0106] In some exemplary embodiments, after the computer device obtains the set of image latent vectors and the set of potential behavior latent vectors, it may input the set of image latent vectors and the set of potential behavior latent vectors into the latent feature encoder to obtain the image latent vector of the new image output by the latent feature encoder.
[0107] Step 502: Input the image latent vector of the new image into the image frame decoder to obtain the new image output by the image frame decoder.
[0108] In some exemplary embodiments, after the computer device obtains the image latent vector of the new image output by the latent feature encoder, it may input the image latent vector of the new image into the image frame decoder to obtain the new image output by the image frame decoder.
[0109] In one exemplary embodiment, as Figure 6 shown, generating the target video using the target image set includes the following steps:
[0110] Step 601: Determine the timing information of each image according to the timestamps at which the images in the target image set are added to the initial image set.
[0111] Exemplarily, since the initial image set is generated based on the first image and the second image, the timing information of the first image and the second image is not based on the timestamps added to the initial image set, but on their generation timestamps.
[0112] In some exemplary embodiments, after determining that the initial image set is the target image set, the computer device may first obtain the timestamps when each image in the target image set is added to the initial image set, and determine the timing information of each image based on the timestamps of each image.
[0113] For example, the target image set includes Image A, Image B, Image C, Image D, and Image E, where Image A and Image B are the first image and the second image, the timestamp of Image C is t, the timestamp of Image D is t + 1, and the timestamp of Image E is t + 2. Then, the timing information of each image can be determined as the first image < the second image < the third image < the fourth image < the fifth image.
[0114] Step 602: Use this timing information to encode each image in the target image set to generate the target video.
[0115] In some exemplary embodiments, after the computer device obtains the timing information of each image, it can use the timing information to encode each image in the target image set to generate the target video.
[0116] Specifically, the computer device can adopt a video coding algorithm to sequentially convert the images into video frames according to the timing information of each image. For example, take Image A as the starting frame of the target video, and perform appropriate preprocessing on it according to information such as the resolution and color mode of the image, such as resizing and color space conversion, to make it meet the requirements of video coding. Read the next image according to the timing information, that is, Image B, calculate the difference between the current image and the previous frame image. Frame interpolation prediction technology can be used to reduce data redundancy. By analyzing the movement of pixels in the two frame images, predict which parts of the current frame are similar to the previous frame, and only encode the different parts and record the motion vectors. For subsequent images, repeat the above process, continuously convert the images into video frames, and use the timing relationship to optimize the inter-frame coding to obtain the target video.
[0117] In an alternative embodiment of the present application, both the above-mentioned first image generation model and the second image generation model are trained using unsupervised learning methods.
[0118] First, the training process for the latent behavior encoder can be as Figure 7 shown. With VQ-VAE as the basic architecture, the latent behavior encoder is a spatio-temporal self-attention network, and the key point detector is the SiLK model. Obtain training data, where the training data is video data, and perform frame extraction on the video data. For the images from the 1st frame to the tth frame In the input key-point detector, each image corresponds to an output key-point mask feature vector. The key-point mask feature vector is a binary sparse vector, whose size is the same as the length and width of the image, and has values only at the key-point coordinate positions. The t key-point mask feature vectors are denoted as . The spatial representation of the input and output of the key-point detector can be , , where t is the time span, H is the height of the video frame, W is the width of the video frame, and C is the number of channels of the video frame.
[0119] Furthermore, the t-frame images and their corresponding key-point mask feature vectors are input into the initial latent behavior encoder, and the latent behavior hidden vector at the next moment is predicted through the original images from the 1st frame to the t-th moment , and the spatial representation is , is the spatial size of the behavior hidden vector, is the discrete vector after the nearest neighbor reconstruction by the codebook, and the codebook size of the latent behavior is set to 32× , and 32 is the number limit of the latent actions.
[0120] The image set corresponding to the video data and the latent behavior hidden vector are input into the latent behavior decoder, and the image at the (t + 1)-th frame is decoded and obtained. The reconstructed image has the same size as the original image, that is . The entire training process follows the training paradigm of VQ-VAE, and the latent behavior encoder and the codebook parameters are iteratively trained through the prediction results and the real images.
[0121] Second, the training process for the image frame encoder and the image frame decoder can be as Figure 8 shown. With VQ-VAE as the basic architecture, the image frame encoder and the image frame decoder are spatio-temporal self-attention networks. The t-th frame image in the video data is input into the image frame encoder, and the image frame encoder compresses the image into a discrete vector , which is the hidden feature vector of the t-th frame image, and its spatial representation is , where is the spatial size of the behavior hidden feature vector , is the discrete vector after the nearest neighbor reconstruction by the codebook. Here, the codebook size of the hidden features can be set to , and is the number limit of the latent features.
[0122] Furthermore, the image hidden vector is then An input image frame decoder to obtain the reconstructed t-th frame image , the reconstructed image has the same size as the input image . The entire training process follows the training paradigm of VQ-VAE, and through the prediction result and the real image , iteratively train the image frame encoder, the image frame decoder, and the corresponding codebook parameters.
[0123] Thirdly, the training process for latent feature encoding can be as shown Figure 9 . The latent feature encoder is based on the Transformer architecture. The images from the 1st frame to the t-th frame can be obtained, and the image frame latent vectors are concatenated with the latent behavior latent vectors , and its feature space representation is .
[0124] Furthermore, the concatenated vector is input into the latent feature encoder to obtain the next frame image latent vector . The prediction result and the real image latent vector use cross-entropy to calculate the reconstruction loss to update the feature encoder parameters.
[0125] The above model training method based on unsupervised learning develops a more refined generative model by deeply studying the laws of human face movement, improving the naturalness and smoothness of the video, and enhancing the video generation quality. Using the unsupervised VQ-VAE architecture to compress and learn the latent representation from the data, including extracting latent behavior features and original image features, capturing the hidden patterns and structures in the data. On the one hand, it compresses the high-dimensional images and video frames into a low-dimensional space to achieve data dimensionality reduction; on the other hand, by controlling the codebook and the latent representation dimension, it can further analyze the semantics therein, map the vector parameters to specific attribute information, and realize the generation of the target video by manipulating the image generation process. Using Transformer can enhance the fusion of features in the time and space dimensions, avoid the problem of unsmooth visual experience caused by inconsistent changes between frames, and improve the coherence.
[0126] Furthermore, using the training method of unsupervised learning can utilize more existing training data without spending a large amount of manpower and time on annotation. The encoder and decoder structures of VQ-VAE allow the model to learn the compressed representation of the data. In the process of learning the structure and pattern of the data itself, it can learn the hidden relationships and diversities in the data, retain the key and general feature information therein, thereby generating richer and more diverse new samples and having strong generalization ability.
[0127] The VQ-VAE uses vector quantization and entropy constraints to discretize the latent encoding. The discrete representation makes the output of the model interpretable and enables better learning of the meaning and structure of the data generated by the model. The discrete latent encodings stored in the codebook contribute to learning more decoupled feature representations, which makes it easier to manipulate specific attributes during the generation process. By understanding the meaning of each vector in the codebook, semantic editing of the latent vectors can be performed during the inference stage, allowing for more precise control of the features and attributes of the generated images. For example, adjusting the weights of specific vectors or adding new vectors to generate new features. This fine-grained control ability can improve the flexibility of generating new images and enable more refined image editing and modification.
[0128] In an exemplary embodiment, as Figure 10 shown, another video generation method is provided, which includes the following steps:
[0129] Step 1001: Obtain a first image, input the first image into the key point detector in the first image generation model to obtain the first key point mask feature vector of the first image output by the key point detector; input the first image and the first key point mask feature into the latent behavior encoder in the first image generation model to obtain the first latent behavior hidden vector of the first image output by the latent behavior encoder; input the first image into the image frame encoder in the first image generation model to obtain the first image hidden vector of the first image output by the image frame encoder.
[0130] Step 1002: Input the first image hidden vector and the first latent behavior hidden vector into the latent feature encoder in the first image generation model to obtain the second image hidden vector of the second image output by the latent feature encoder; input the second image hidden vector into the image frame decoder in the first image generation model to obtain the second image output by the image frame decoder.
[0131] Step 1003: Generate an initial image set based on the first image and the second image, and input the initial image set into the key point detector in the second image generation model to obtain a set of key point mask feature vectors of the initial image set output by the key point detector; input the set of key point mask feature vectors and the initial image set into the latent behavior encoder in the second image generation model to obtain a set of latent behavior hidden vectors of the initial image set output by the latent behavior encoder.
[0132] Step 1004. Obtain the set of image latent vectors of the initial image set, and input the set of image latent vectors and the set of latent behavior latent vectors into the latent feature encoder in the second image generation model to obtain the image latent vector of the new image output by the latent feature encoder; input the image latent vector of the new image into the image frame decoder in the second image generation model to obtain the new image output by the image frame decoder.
[0133] Step 1005. Add the new image to the initial image set, and after adding, input the initial image set into the second image generation model again to obtain the new image output by the second image generation model again; continue to execute the step of adding the new image to the initial image set and, after adding, inputting the initial image set into the second image generation model until the number of images in the initial image set is greater than the preset number threshold, and determine the initial image set as the target image set.
[0134] Step 1006. Determine the timing information of each image according to the timestamp when each image in the target image set is added to the initial image set; use the timing information to encode each image in the target image set to generate the target video.
[0135] In one embodiment, the complete process of obtaining the first image and inputting the first image into the pre-trained first image generation model to obtain the second image output by the first image generation model can be as Figure 11 shown. And the complete process of inputting the initial image set into the second image generation model to obtain the new image output by the second image generation model can be as Figure 12 shown. That is, when there is only one image, the process shown in Figure 11 needs to be executed. When generating a new image based on one image and constructing the initial image set, the process shown in Figure 12 needs to be executed.
[0136] Among them, Figure 11 I1 in it is the first image, k1 is the first key point mask feature vector of the first image, a1 is the first latent behavior latent vector of the first image, z1 is the first image latent vector of the first image, z2 is the second image latent vector of the second image, and I2 is the second image.
[0137] Figure 12 I in 1:t is the initial image set, k 1:t is the set of key point mask feature vectors of the initial image set, a 1:t is the set of latent behavior latent vectors of the initial image set, z 1:t is the set of image latent vectors of the initial image set, z t+1 is the image latent vector of the new image, and I t+1 is the new image.
[0138] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0139] Based on the same inventive concept, the embodiments of the present application also provide a video generation device for implementing the above-mentioned video generation method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the video generation device provided below can refer to the limitations on the video generation method in the above text, and will not be repeated here.
[0140] In an exemplary embodiment, as Figure 13 shown, a video generation device 1300 is provided, including: an acquisition module 1301, a first execution module 1302, a second execution module 1303, and a generation module 1304, where:
[0141] The acquisition module 1301 is configured to acquire a first image and input the first image into a pre-trained first image generation model to obtain a second image output by the image generation model;
[0142] The first execution module 1302 is configured to generate an initial image set based on the first image and the second image, and input the initial image set into a pre-trained second image generation model to obtain a new image output by the second image generation model;
[0143] The second execution module 1303 is configured to add the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again to obtain a new image output again by the second image generation model;
[0144] The generation module 1304 is configured to continue to execute the step of adding the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again until the number of images in the initial image set is greater than a preset number threshold, determine the initial image set as a target image set, and generate a target video using the target image set.
[0145] In one embodiment, the first image generation model includes a key point detector, a latent behavior encoder, and an image frame encoder. The obtaining module 1301 is specifically configured to input the first image into the key point detector to obtain a first key point mask feature vector of the first image output by the key point detector; input the first image and the first key point mask feature into the latent behavior encoder to obtain a first latent behavior hidden vector of the first image output by the latent behavior encoder; input the first image into the image frame encoder to obtain a first image hidden vector of the first image output by the image frame encoder, and determine the second image based on the first image hidden vector and the first latent behavior hidden vector.
[0146] In one embodiment, the first image generation model further includes a latent feature encoder and an image frame decoder. The obtaining module 1301 is specifically configured to input the first image hidden vector and the first latent behavior hidden vector into the latent feature encoder to obtain a second image hidden vector of the second image output by the latent feature encoder; input the second image hidden vector into the image frame decoder to obtain the second image output by the image frame decoder.
[0147] In one embodiment, the second image generation model includes a key point detector and a latent behavior encoder. The first execution module 1302 is specifically configured to input the initial image set into the key point detector to obtain a set of key point mask feature vectors of the initial image set output by the key point detector; input the set of key point mask feature vectors and the initial image set into the latent behavior encoder to obtain a set of latent behavior hidden vectors of the initial image set output by the latent behavior encoder; obtain a set of image hidden vectors of the initial image set, and determine the new image according to the set of image hidden vectors and the set of latent behavior hidden vectors.
[0148] In one embodiment, the second image generation model further includes a latent feature encoder and an image frame decoder. The first execution module 1302 is specifically configured to input the set of image hidden vectors and the set of latent behavior hidden vectors into the latent feature encoder to obtain an image hidden vector of the new image output by the latent feature encoder; input the image hidden vector of the new image into the image frame decoder to obtain the new image output by the image frame decoder.
[0149] In one embodiment, the generating module 1304 is specifically configured to determine the timing information of each image according to the timestamps at which the images in the target image set are added to the initial image set; use the timing information to encode each image in the target image set to generate the target video.
[0150] Each module in the above video generation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0151] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 14 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a video generation method.
[0152] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 15As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a video generation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0153] Those skilled in the art can understand that Figure 14 and Figure 15 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0154] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps described in any of the above embodiments.
[0155] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, it implements the steps described in any of the above embodiments.
[0156] In an embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, it implements the steps described in any of the above embodiments.
[0157] It should be noted that the user information involved in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to images containing face information for analysis, stored data, displayed data, etc.) are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0158] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0159] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0160] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. A video generation method, characterized in that, The method includes: Obtain a first image, and input the first image into a pre-trained first image generation model to obtain a second image output by the image generation model; Generate an initial image set based on the first image and the second image, and input the initial image set into a pre-trained second image generation model to obtain a new image output by the second image generation model; Add the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again to obtain a new image output by the second image generation model again; Continue to execute the step of adding the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again until the number of images in the initial image set is greater than a preset number threshold. Determine the initial image set as the target image set, and generate a target video using the target image set.
2. The method according to claim 1, characterized in that The first image generation model includes a key point detector, a latent behavior encoder, and an image frame encoder. The step of inputting the first image into the pre-trained first image generation model to obtain a second image output by the image generation model includes: Input the first image into the key point detector to obtain a first key point mask feature vector of the first image output by the key point detector; Input the first image and the first key point mask feature into the latent behavior encoder to obtain a first latent behavior hidden vector of the first image output by the latent behavior encoder; Input the first image into the image frame encoder to obtain a first image hidden vector of the first image output by the image frame encoder, and determine the second image based on the first image hidden vector and the first latent behavior hidden vector.
3. The method according to claim 2, wherein The first image generation model further includes a latent feature encoder and an image frame decoder. The step of determining the second image based on the first image hidden vector and the first latent behavior hidden vector includes: Input the first image hidden vector and the first latent behavior hidden vector into the latent feature encoder to obtain a second image hidden vector of the second image output by the latent feature encoder; Input the second image hidden vector into the image frame decoder to obtain the second image output by the image frame decoder.
4. The method according to claim 1, wherein The second image generation model includes a key point detector and a latent behavior encoder. The step of inputting the initial image set into the pre-trained second image generation model to obtain a new image output by the second image generation model includes: Input the initial image set into the key point detector to obtain a set of key point mask feature vectors of the initial image set output by the key point detector; Input the set of key point mask feature vectors and the initial image set into the latent behavior encoder to obtain a set of latent behavior hidden vectors of the initial image set output by the latent behavior encoder; Obtain the set of image latent vectors of the initial image set, and determine the new image according to the set of image latent vectors and the set of latent behavior latent vectors.
5. The method according to claim 4, wherein The second image generation model further includes a latent feature encoder and an image frame decoder. The determining the new image according to the set of image latent vectors and the set of latent behavior latent vectors includes: Input the set of image latent vectors and the set of latent behavior latent vectors into the latent feature encoder to obtain the image latent vector of the new image output by the latent feature encoder; Input the image latent vector of the new image into the image frame decoder to obtain the new image output by the image frame decoder.
6. The method according to claim 1, characterized in that, The generating the target video by using the target image set includes: Determine the timing information of each image according to the timestamps at which the images in the target image set are added to the initial image set; Encode each image in the target image set by using the timing information to generate the target video.
7. A video generation device, characterized in that, The apparatus includes: An acquisition module, configured to acquire a first image and input the first image into a pre-trained first image generation model to obtain a second image output by the image generation model; A first execution module, configured to generate an initial image set based on the first image and the second image, and input the initial image set into a pre-trained second image generation model to obtain a new image output by the second image generation model; A second execution module, configured to add the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again to obtain a new image output again by the second image generation model; A generation module, configured to continue to execute the step of adding the new image to the initial image set, and after the addition, input the initial image set into the second image generation model again until the number of images in the initial image set is greater than a preset number threshold, determine the initial image set as the target image set, and generate a target video by using the target image set.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Image processing method and related device
CN121078270A