Video synthesis method and device, equipment and storage medium
By extracting and encoding compensation features in the video stream synthesis model, the problem of accurately modeling action and facial expression details in dynamic digital human video synthesis is solved, generating realistic and high-quality dynamic digital human videos.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies face the challenge of accurately modeling action and facial expression details when generating realistic and dynamic digital humans, resulting in synthesized videos that are not realistic enough, have stiff movements, and have low image quality.
The feature extraction component extracts identity appearance and dynamic change features, which are then input into a video stream synthesis model trained by joint learning. The encoding compensation component is used to perform encoding compensation and fusion on the feature encoding results to generate realistic videos.
It achieves more realistic dynamic digital human video synthesis, improves image quality, and can replace real people in the fintech field.
Smart Images

Figure CN121888052A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and is applied to the scenario of dynamic digital human video synthesis using video stream synthesis models. It relates to a video synthesis method, apparatus, device and storage medium. Background Technology
[0002] Facial or head motion video generation technology aims to create a realistic video that retains the identity of the source image while mimicking the movements of the driving video, based on a source image and a driving video. This technology has broad application prospects in fields such as video conferencing, film production, and virtual reality. For example, in financial management scenarios, it can composite an image of a bank broker in professional attire and wearing a bank broker identification badge with a financial management introductory video.
[0003] Current technologies still face significant challenges in achieving accurate capture of pose and facial expression details. Facial expression details, especially subtle local expressions, are extremely complex and difficult to model accurately. Both unsupervised keypoint methods and methods relying on predefined models suffer from limited motion representation capabilities, often failing to capture certain dynamic facial motion information, resulting in a lack of static and dynamic feature information. This lack of static and dynamic feature information ultimately leads to synthesized dynamic digital humans that are not realistic enough, have overly stiff movements, and exhibit low image quality. Summary of the Invention
[0004] The purpose of this application is to provide a video synthesis method, apparatus, device, and storage medium to synthesize more realistic and high-quality dynamic digital humans.
[0005] In a first aspect, embodiments of this application provide a video synthesis method, which employs the following technical solution: A video synthesis method includes the following steps: Acquire a source image and a driving video, wherein the source image contains the identity and appearance information of the target synthetic object, and the driving video contains reference action posture and reference facial expression; The source image and the driving video are used as grouped data and input into a preset feature extraction component to extract identity appearance features and dynamic change features; The identity appearance features and the dynamic change features are input together into the video stream synthesis model trained through joint learning; The feature encoding component in the video stream synthesis model is used to encode the identity appearance features and the dynamic change features respectively, to obtain preliminary feature encoding results; The initial feature encoding result is processed by the encoding compensation component learned during joint learning training to obtain the feature encoding result after encoding compensation. The feature encoding results after encoding compensation are subjected to feature fusion processing to obtain the feature fusion result; The feature decoding component in the video stream synthesis model is used to decode the feature fusion result to obtain the target synthesized video.
[0006] Secondly, embodiments of this application also provide a video synthesis apparatus, which adopts the technical solution described below: A video compositing apparatus, comprising: The synthesis resource acquisition module is used to acquire source images and driving videos, wherein the source images contain the identity and appearance information of the target synthesis object, and the driving videos contain reference action postures and reference facial expressions; The feature extraction module is used to input the source image and the driving video as grouped data into a preset feature extraction component to extract identity appearance features and dynamic change features; The synthetic feature input module is used to input the identity appearance features and the dynamic change features into the video stream synthesis model trained through joint learning. The feature encoding processing module is used to perform feature encoding on the identity appearance features and the dynamic change features respectively using the feature encoding component in the video stream synthesis model to obtain preliminary feature encoding results; The encoding compensation processing module is used to perform encoding compensation processing on the preliminary feature encoding result using the encoding compensation component learned during joint learning training, so as to obtain the feature encoding result after encoding compensation processing. The feature fusion processing module is used to perform feature fusion processing on the feature encoding results after encoding compensation processing to obtain the feature fusion result; The video decoding module is used to decode the feature fusion result using the feature decoding component in the video stream synthesis model to obtain the target synthesized video.
[0007] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below: A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the video synthesis method described above.
[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below: A computer-readable storage medium storing computer-readable instructions that, when executed by a processor, implement the steps of the video synthesis method described above.
[0009] Compared with the prior art, the embodiments of this application have the following main advantages: The video synthesis method described in this application involves acquiring source images and driving videos; inputting them in groups into a preset feature extraction component to extract identity appearance features and dynamic change features; inputting these features into a jointly trained video stream synthesis model; encoding the identity appearance features and dynamic change features separately to obtain preliminary feature encoding results; using an encoding compensation component learned during joint learning training to perform encoding compensation processing on the preliminary feature encoding results to obtain encoded compensation-processed feature encoding results; performing feature fusion processing on the encoded compensation-processed feature encoding results to obtain feature fusion results; and using a feature decoding component to decode the feature fusion results to obtain the target synthesized video. By introducing an encoding compensation component, the video stream synthesis model obtained through final training is ensured to utilize the encoding compensation component in the trained video stream synthesis model to perform encoding compensation on the static image features of the input source image and encoding compensation on the dynamic change features of the input driving video during actual video synthesis, thereby generating synthesized videos that are more realistic and closer to real videos. Applying the video synthesis method described above to the field of financial technology can enable more realistic and high-quality dynamic digital humans to replace existing human workers in scenarios such as financial courses, financial marketing, and product launches. Attached Figure Description
[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of a video synthesis method according to this application; Figure 3 This is a flowchart of a specific embodiment of setting the encoding compensation component of the video stream synthesis model in the video synthesis method described in this application; Figure 4 This is a flowchart of a specific embodiment of the video synthesis method described in this application, which involves calling and processing a video stream synthesis model; Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 205 shown; Figure 6 This is a schematic diagram of one embodiment of a video synthesis apparatus according to this application; Figure 7 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0013] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0014] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0015] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0016] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0017] It should be noted that the video synthesis method provided in this application embodiment is generally executed by a server, and correspondingly, a video synthesis device is generally set in the server.
[0018] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0019] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of a video synthesis method according to this application. The video synthesis method includes the following steps: Step 201: Obtain the source image and driving video.
[0020] The source image contains the identity and appearance information of the target synthesized object, and the driving video contains reference action postures and reference facial expressions. The identity and appearance information includes the person's head image information and clothing appearance information. In this embodiment, the source image is labeled with at least the identity and appearance information of the target synthesized object, while the driving video is labeled with at least the reference action postures and reference facial expressions used during video synthesis.
[0021] In this embodiment, the source image refers to an image containing the clothing and identity information of the target synthesized object, such as an image of a bank broker wearing a professional suit and a bank broker identity badge in a financial management business scenario; while the driving video refers to the source video used to synthesize the source image into a dynamic digital human. By fitting the source image and the driving video, a dynamic digital human corresponding to the bank broker is generated. The purpose is to generate a dynamic digital human corresponding to the static figure in the source image by using the dynamic features in the driving video as a reference.
[0022] Specifically, by applying the aforementioned video synthesis method to the field of financial technology, dynamic digital humans can replace existing human workers in scenarios such as financial courses, financial marketing, and product launches. For example, dynamic digital humans can replace existing financial analysts in explaining financial business; dynamic digital humans can replace existing marketers or product launchers in launching marketing products.
[0023] Step 202: The source image and the driving video are input as grouped data into a preset feature extraction component to extract identity appearance features and dynamic change features.
[0024] In this embodiment, the dynamic change features include action posture change features and facial expression change features. Before inputting the source image and the driving video into the preset feature extraction component, for the convenience of feature extraction, a grouping identification method is used to mark the source image and the driving video as a set of data for video synthesis. Then, the source image and the driving video are input into the preset feature extraction component in combination with the grouping identification, and features are extracted from the source image and the driving video respectively according to the preset feature extraction component.
[0025] Specifically, the preset feature extraction component includes a first extraction sub-component and a second extraction sub-component. The first extraction sub-component is an image feature extraction sub-component, such as an image feature extraction sub-component based on the OpenCV architecture, which can extract static image features from static images. The second extraction sub-component is a video feature extraction sub-component, such as a video feature extraction sub-component based on the SlowFast model, which can extract dynamic image change features from video streams.
[0026] Step 203: Input the identity appearance features and the dynamic change features into the video stream synthesis model trained by joint learning.
[0027] In this embodiment, the video stream synthesis model trained by joint learning includes a video stream synthesis model in the form of encoding and decoding based on the Transformer neural network architecture.
[0028] Step 204: Use the feature encoding component in the video stream synthesis model to perform feature encoding on the identity appearance features and the dynamic change features respectively, and obtain preliminary feature encoding results.
[0029] Step 205: The initial feature encoding result is processed by encoding compensation using the encoding compensation component learned during joint learning training to obtain the feature encoding result after encoding compensation.
[0030] In this embodiment, after obtaining the preliminary feature encoding results corresponding to the identity appearance features and the dynamic change features respectively, the encoding compensation component learned during joint learning training is used to perform encoding compensation processing on the feature encoding results corresponding to the identity appearance features and the dynamic change features respectively. The purpose is to correct the preliminary feature encoding results through encoding compensation, so that the subsequent video synthesis results are more realistic and closer to the level of real-life videos.
[0031] In this embodiment, when using the encoding compensation component learned during joint learning training to perform encoding compensation processing on the feature encoding results corresponding to the identity appearance features and the dynamic change features, the corresponding encoding compensation scale can be selected for encoding compensation based on the encoding feature scale when encoding the identity appearance features and the dynamic change features respectively. For example, when encoding the source image, if the feature encoding size is a matrix size of N multiplied by M, then when performing encoding compensation on the feature encoding corresponding to the source image, a matrix size of N multiplied by M is also used for complementation processing, ensuring that the encoding compensation component can perform multi-scale encoding compensation processing based on the encoding scale of the initial feature encoding results.
[0032] Step 206: Perform feature fusion processing on the feature encoding results after encoding compensation processing to obtain the feature fusion results.
[0033] In this embodiment, the feature fusion processing of the feature encoding result after the encoding compensation process to obtain the feature fusion result is essentially a cross-modal feature fusion method that combines static image features (identity and appearance features) with dynamic video stream features (dynamic change features such as posture, action or facial expression), thereby improving the quality and realism of the synthesized video.
[0034] Step 207: Using the feature decoding component in the video stream synthesis model, the feature fusion result is decoded to obtain the target synthesized video.
[0035] Specifically, the cross-modal features obtained by cross-modal feature fusion in step 206 are decoded, and the corresponding decoded video is output as the target synthesized video.
[0036] In this embodiment, source images and driving videos are acquired and input into a preset feature extraction component to extract identity appearance features and dynamic change features. These are then input into a video stream synthesis model trained through joint learning. The identity appearance features and dynamic change features are respectively encoded to obtain preliminary feature encoding results. An encoding compensation component learned during joint learning is used to perform encoding compensation processing on the preliminary feature encoding results, resulting in encoded compensation-processed feature encoding results. The encoded compensation-processed feature encoding results are then subjected to feature fusion processing to obtain feature fusion results. Finally, a feature decoding component is used to decode the feature fusion results to obtain the target synthesized video. By introducing an encoding compensation component, it is ensured that the video stream synthesis model obtained through final training can, during actual video synthesis, utilize the encoding compensation component in the trained video stream synthesis model to perform encoding compensation on the static image features of the input source image and encoding compensation on the dynamic change features of the input driving video, thereby generating a more realistic synthesized video that closely resembles a real video. Applying the video synthesis method described above to the field of financial technology can enable more realistic and high-quality dynamic digital humans to replace existing human workers in scenarios such as financial courses, financial marketing, and product launches.
[0037] In this embodiment, the step of inputting the source image and the driving video as grouped data into a preset feature extraction component to extract identity appearance features and dynamic change features specifically includes: using the first extraction sub-component to extract identity appearance features from the source image; and using the second extraction sub-component to extract dynamic change features from the driving video, wherein the dynamic change features include action posture change features and facial expression change features.
[0038] Specifically, static features, such as identity and appearance features, are extracted from the source image, and dynamic features, such as movement and posture changes and facial expression changes, are extracted from the driving video. This facilitates the subsequent synthesis of the target video using the identity and appearance features, movement and posture changes, and facial expression changes.
[0039] In some specific embodiments, before step 203, a step of preparing training materials before joint learning training is included. A flowchart of a specific embodiment of the video synthesis method described in this application, involving the preparation of training materials before joint learning training, includes: Obtain several source images for training and construct a source image training set, where different source images contain different identity and appearance information; The first extraction component is used to extract the identity appearance features from all source images in the source image training set, and the identity appearance features are distinguished and labeled according to the different source images. Obtain several training driving videos and construct a driving video training set, wherein different driving videos contain different motion posture changes and / or facial expression changes; The second extraction sub-component is used to extract the motion posture change features and facial expression change features from all driving videos in the driving video training set. The extraction results of the second extraction sub-component are then distinguished and labeled according to the different driving videos.
[0040] Specifically, the source image training set can contain different identity and appearance information from different source images. For example, the source image training set may contain 100 images, including 10 images of teachers, 20 images of chefs, and 10 images of delivery workers in specific colored clothing. Several training source images can be obtained from multiple acquisition channels to construct the source image training set, which can then be used for joint learning and training of the video stream synthesis model.
[0041] Specifically, since the source image training set is used in the joint learning training of the video stream synthesis model, after obtaining the source image training set, the first extraction component is first used to extract the identity and appearance features of all images in the source image training set and perform differentiation labeling to facilitate the recognition of image feature information during the training process.
[0042] Similarly, since the driving video training set is used in the joint learning training of the video stream synthesis model, after obtaining the driving video training set, the second extraction component is first used to distinguish and label the action posture change features and facial expression change features in all driving videos in the driving video training set, so as to facilitate the recognition of dynamic change features and information during the training process.
[0043] In this embodiment, the video stream synthesis model includes a feature encoding component, an encoding compensation component, an encoding fusion component, and a decoding output component.
[0044] Continue to refer to Figure 3 In some specific implementations, before step 203, a step of setting the encoding compensation component of the video stream synthesis model is also included. Figure 3 This is a flowchart of a specific embodiment of setting the encoding compensation component of the video stream synthesis model in the video synthesis method described in this application, including: Step 301: Input the identity appearance features in all source images in the source image training set and the dynamic change features in all driving videos in the driving video training set into the video stream synthesis model to be jointly learned and trained. Step 302: Use the feature encoding component to perform feature encoding on the identity appearance features in all source images in the source image training set and the dynamic change features in all driving videos in the driving video training set to obtain preliminary feature encoding results; Step 303: Using a dual-loop method, different source images and driving videos are selected sequentially as fusion combinations, and feature encoding fusion processing is performed using the encoding fusion component to obtain the feature fusion encoding result; Specifically, assuming the source image training set contains i source images and the driving video training set contains j driving videos, a dual-loop filtering mechanism with an outer loop of i times and an inner loop of j times can be constructed to filter out a total of i multiplied by j fusion combinations.
[0045] Step 304: The decoding output component is used to decode and output all feature fusion encoding results of the dual loop output one by one to obtain the fused video corresponding to different fusion combinations. Step 305: Input different fused videos into the preset feature extraction component to extract the identity appearance features and dynamic change features corresponding to different fused videos; Step 306: Compare the identity appearance features corresponding to different fused videos with the identity appearance features of the corresponding source images to obtain the differences in identity appearance features. Based on the differences in identity appearance features, construct an identity appearance feature coding compensation component. Specifically, the identity appearance feature encoding compensation component can generate a first encoding compensation matrix based on the differences in identity appearance features. After encoding the identity appearance features of a new source image, the first encoding compensation matrix is used to compensate the corresponding identity appearance feature encoding matrix according to the row and column correspondence.
[0046] Step 307: Compare the dynamic change features corresponding to different fused videos with the dynamic change features of the corresponding driving videos to obtain motion feature differences. Based on the motion feature differences, construct a motion feature coding compensation component, wherein the motion feature differences include action posture feature differences and facial expression feature differences. Similarly, the motion feature coding compensation component can also generate a second coding compensation matrix based on the motion feature differences. When the new driving video undergoes dynamic change feature coding, the second coding compensation matrix is applied to the corresponding dynamic change feature coding matrix according to the row and column correspondence.
[0047] Step 308: Set the identity appearance feature encoding compensation component and the motion feature encoding compensation component as two compensation sub-components of the encoding compensation component to obtain the video stream synthesis model to be optimized.
[0048] Specifically, in the initial training stage of the video stream synthesis model, the feature encoding component, encoding fusion component, and decoding output component are first activated to obtain a preliminarily generated fused video. Then, the preset feature extraction component is used to extract the identity appearance features and dynamic change features contained in the preliminarily generated fused video. The identity appearance features contained in the preliminarily generated fused video are compared with the identity appearance features corresponding to the source images, and the dynamic change features contained in the preliminarily generated fused video are compared with the dynamic change features corresponding to the driving video. Based on the two comparison results, an encoding compensation component is constructed to obtain the video stream synthesis model to be optimized.
[0049] Figure 3 The provided video synthesis and encoding / decoding steps utilize the feature differences in the encoding / decoding process to construct an encoding compensation component. Therefore, in Figure 3 The encoding compensation component was not used in the code.
[0050] Continue to refer to Figure 4 In some specific implementations, after step 308, a step of calling the video stream synthesis model is also included. Figure 4 This is a flowchart of a specific embodiment of the video synthesis method described in this application, which involves calling and processing a video stream synthesis model, including: Step 401: Using a dual-loop method, different source images and driving videos are selected sequentially as fusion combinations to identify the preliminary feature encoding results corresponding to different fusion combinations. Step 402: Input the preliminary feature encoding results corresponding to different fusion combinations into the encoding compensation component to obtain the feature encoding results after encoding compensation; Step 403: Input the feature encoding results after encoding compensation corresponding to different fusion combinations into the feature fusion component to obtain fused features; Step 404: Use the decoding output component to decode and output all fusion features one by one to obtain the fusion video corresponding to different fusion combinations; Step 405: Input different fused videos into the preset feature extraction component to extract the identity appearance features and dynamic change features corresponding to different fused videos; Step 406: Compare the identity appearance features corresponding to different fused videos with the identity appearance features of the corresponding source images for feature compensation loss. After summarizing and calculating, obtain the identity appearance feature compensation loss corresponding to all fused videos. Step 407: Compare the dynamic change features corresponding to different fused videos with the dynamic change features of the corresponding driving videos for feature compensation loss. After summarizing and calculating, obtain the motion feature compensation loss corresponding to all fused videos. Step 408: Identify whether the identity appearance feature compensation loss and the motion feature compensation loss meet the preset loss conditions; Step 409: If the preset loss condition is not met, the hyperparameters of the video stream synthesis model are readjusted and model tuning is performed until the preset loss condition is met, and then the model tuning is stopped. Step 410: If the preset loss condition is met, the optimized video stream synthesis model is obtained as the video stream synthesis model trained by the joint learning.
[0051] Specifically, during the optimization and training phase of the video stream synthesis model, a feature encoding component is enabled. After obtaining the preliminary feature encoding results corresponding to different fusion combinations, the encoding compensation component set in step 308 is used to compensate the preliminary feature encoding results to obtain the compensated feature encoding results. Then, the encoded fusion component and the decoding output component are used to obtain the fused video generated after encoding compensation. Afterward, the preset feature extraction component is used to extract the identity appearance features and dynamic change features contained in the fused video generated after encoding compensation. The identity appearance features are compared with the identity appearance features corresponding to the source images, and the dynamic change features are compared with the dynamic change features corresponding to the driving videos. Based on the two comparison results, the identity appearance feature compensation loss and motion feature compensation loss are calculated. The video stream synthesis model is optimized using the identity appearance feature compensation loss and the motion feature compensation loss until the final video stream synthesis model trained by joint learning is obtained.
[0052] Figure 4 In the provided video synthesis encoding and decoding process, an encoding compensation component is first used after encoding. Finally, by combining the difference between the feature extraction results of the decoded video and the feature extraction results of the training set, the identity appearance feature compensation loss and motion feature compensation loss are calculated. This process is used to fine-tune the video stream synthesis model repeatedly until the final video stream synthesis model trained by joint learning is obtained.
[0053] The video synthesis method provided in this embodiment, after initial training of the video stream synthesis model, introduces an encoding compensation component based on encoding-decoding differences. During the model optimization and training phase, the model is optimized based on the compensation loss of the encoding compensation component. This ensures that the final trained video stream synthesis model can, in actual video synthesis, utilize the encoding compensation component in the trained model to perform encoding compensation for the static image features of the input source image and the dynamic change features of the input driving video, respectively, thereby generating more realistic and near-realistic synthesized videos. Applying this video synthesis method to the field of financial technology, it can replace existing human workers with more realistic and high-quality dynamic digital humans in scenarios such as financial courses, financial marketing, and product launches. For example, dynamic digital humans can replace existing financial analysts in explaining financial business; dynamic digital humans can replace existing marketers or product launchers in launching marketing products.
[0054] Continue to refer to Figure 5 , Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 205 shown includes: Step 501: Using the identity appearance feature encoding compensation component, the identity appearance features in the preliminary feature encoding result are compensated to obtain the identity appearance features after encoding compensation. Step 502: Using the motion feature encoding compensation component, the dynamic change features in the preliminary feature encoding result are compensated to obtain the dynamic change features after encoding compensation.
[0055] In this embodiment, the video stream synthesis model further includes a feature enhancement processing component. Before performing feature fusion processing on the feature encoding result after encoding compensation processing to obtain the feature fusion result, i.e., before step 206, the method further includes: inputting the identity appearance features after encoding compensation processing into a preset feature enhancement processing component, wherein the preset feature enhancement processing component includes an image super-resolution decoder; using the image super-resolution decoder, performing upsampling and super-resolution fusion processing on the identity appearance features after encoding compensation processing a preset number of times to obtain the feature-enhanced identity appearance features.
[0056] In order to obtain better identity appearance features based on the source image feature encoding compensation, an image super-resolution decoder can be introduced after the identity appearance feature encoding compensation component to decode and obtain the enhanced identity appearance features, so as to obtain a composite video with better image quality of the target composite object in subsequent video synthesis.
[0057] In this embodiment, source images and driving videos are acquired and input into a preset feature extraction component to extract identity appearance features and dynamic change features. These are then input into a video stream synthesis model trained through joint learning. The identity appearance features and dynamic change features are respectively encoded to obtain preliminary feature encoding results. An encoding compensation component learned during joint learning is used to perform encoding compensation processing on the preliminary feature encoding results, resulting in encoded compensation-processed feature encoding results. The encoded compensation-processed feature encoding results are then subjected to feature fusion processing to obtain feature fusion results. Finally, a feature decoding component is used to decode the feature fusion results to obtain the target synthesized video. By introducing an encoding compensation component, it is ensured that the video stream synthesis model obtained through final training can, during actual video synthesis, utilize the encoding compensation component in the trained video stream synthesis model to perform encoding compensation on the static image features of the input source image and encoding compensation on the dynamic change features of the input driving video, thereby generating a more realistic synthesized video that closely resembles a real video. Applying the video synthesis method described above to the field of financial technology can enable more realistic and high-quality dynamic digital humans to replace existing human workers in scenarios such as financial courses, financial marketing, and product launches.
[0058] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0059] Foundational technologies in artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0060] Further reference Figure 6 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a video synthesis apparatus, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0061] like Figure 6As shown, the video synthesis device 600 described in this embodiment includes: a synthesis resource acquisition module 601, a feature extraction module 602, a synthesis feature input module 603, a feature encoding processing module 604, an encoding compensation processing module 605, a feature fusion processing module 606, and a video decoding acquisition module 607. Wherein: The compositing resource acquisition module 601 is used to acquire source images and driving videos, wherein the source images contain the identity and appearance information of the target compositing object, and the driving videos contain reference action postures and reference facial expressions; The feature extraction module 602 is used to input the source image and the driving video as grouped data into a preset feature extraction component to extract identity appearance features and dynamic change features; The synthetic feature input module 603 is used to input the identity appearance features and the dynamic change features into the video stream synthesis model trained by joint learning; The feature encoding processing module 604 is used to perform feature encoding on the identity appearance feature and the dynamic change feature respectively using the feature encoding component in the video stream synthesis model to obtain preliminary feature encoding results; The encoding compensation processing module 605 is used to perform encoding compensation processing on the preliminary feature encoding result using the encoding compensation component learned during joint learning training, so as to obtain the feature encoding result after encoding compensation processing. The feature fusion processing module 606 is used to perform feature fusion processing on the feature encoding result after encoding compensation processing to obtain the feature fusion result; The video decoding acquisition module 607 is used to decode the feature fusion result using the feature decoding component in the video stream synthesis model to obtain the target synthesized video.
[0062] This application acquires source images and driving videos; inputs them in groups into a preset feature extraction component to extract identity appearance features and dynamic change features; inputs these features into a video stream synthesis model trained through joint learning; encodes the identity appearance features and dynamic change features separately to obtain preliminary feature encoding results; uses an encoding compensation component learned during joint learning to perform encoding compensation processing on the preliminary feature encoding results to obtain encoded compensation-processed feature encoding results; performs feature fusion processing on the encoded compensation-processed feature encoding results to obtain feature fusion results; and uses a feature decoding component to decode the feature fusion results to obtain the target synthesized video. By introducing an encoding compensation component, this application ensures that the video stream synthesis model obtained through final training can, during actual video synthesis, use the encoding compensation component in the trained video stream synthesis model to perform encoding compensation on the static image features of the input source image and encoding compensation on the dynamic change features of the input driving video, thereby generating more realistic synthesized videos that are closer to real videos. Applying the video synthesis method described above to the field of financial technology can enable more realistic and high-quality dynamic digital humans to replace existing human workers in scenarios such as financial courses, financial marketing, and product launches.
[0063] In this embodiment, the feature extraction module 602 includes a first feature extraction unit and a second feature extraction unit. Wherein: The first feature extraction unit is used to extract the identity appearance features in the source image using the first extraction component; The second feature extraction unit is used to extract dynamic change features in the driving video using the second extraction component, wherein the dynamic change features include action posture change features and facial expression change features.
[0064] In this embodiment, the video synthesis device 600 further includes a source image training set construction module, an identity appearance feature extraction and labeling module, a driving video training set construction module, and a dynamic change feature extraction and labeling module. Wherein: The source image training set construction module is used to obtain several source images for training and construct a source image training set, wherein different source images contain different identity and appearance information; The identity appearance feature extraction and labeling module is used to extract the identity appearance features from all source images in the source image training set using the first extraction sub-component, and to perform distinguishing labeling processing on the identity appearance features according to the different source images; The driving video training set construction module is used to acquire several driving videos for training and construct a driving video training set, wherein different driving videos contain different motion posture changes and / or facial expression changes. The dynamic change feature extraction and labeling module is used to extract the action posture change features and facial expression change features in all driving videos in the driving video training set using the second extraction sub-component, and to perform differential labeling processing on the extraction results of the second extraction sub-component according to the different driving videos.
[0065] In this embodiment, the video synthesis device 600 further includes a training input module, a preliminary feature encoding module, a feature encoding fusion module, a fused video decoding and acquisition module, a feature extraction module, an identity and appearance feature encoding compensation component construction module, a motion feature encoding compensation component construction module, and an encoding compensation component setting module. Wherein: The training input module is used to input the identity appearance features from all source images in the source image training set and the dynamic change features from all driving videos in the driving video training set into the video stream synthesis model to be jointly learned and trained. The preliminary feature encoding module is used to perform feature encoding processing on the identity appearance features in all source images in the source image training set and the dynamic change features in all driving videos in the driving video training set using the feature encoding component, so as to obtain the preliminary feature encoding result; The feature encoding fusion module is used to select different source images and driving videos as fusion combinations in a dual loop manner, and to perform feature encoding fusion processing using the encoding fusion component to obtain the feature fusion encoding result; The fusion video decoding and acquisition module is used to decode and output all feature fusion encoding results of the dual loop output one by one using the decoding output component to obtain fusion videos corresponding to different fusion combinations. The feature extraction module is used to input different fused videos into the preset feature extraction component to extract the identity appearance features and dynamic change features corresponding to different fused videos; The identity appearance feature encoding compensation component construction module is used to compare the identity appearance features corresponding to different fused videos with the identity appearance features of the corresponding source images to obtain the identity appearance feature differences, and construct the identity appearance feature encoding compensation component based on the identity appearance feature differences. The motion feature coding compensation component construction module is used to compare the dynamic change features corresponding to different fused videos with the dynamic change features of the corresponding driving videos to obtain motion feature differences. Based on the motion feature differences, a motion feature coding compensation component is constructed, wherein the motion feature differences include action posture feature differences and facial expression feature differences. The encoding compensation component setting module is used to set the identity appearance feature encoding compensation component and the motion feature encoding compensation component as two compensation sub-components of the encoding compensation component, so as to obtain the video stream synthesis model to be optimized.
[0066] In this embodiment, the video synthesis device 600 further includes a preliminary feature encoding and recognition module, an encoding compensation processing module, a feature fusion module, a fused video decoding and acquisition module, a feature extraction module, an identity and appearance feature compensation loss calculation module, a motion feature compensation loss calculation module, a loss condition recognition module, a model optimization processing module, and a video stream synthesis model training completion module. Wherein: The preliminary feature encoding and recognition module is used to select different source images and driving videos as fusion combinations in a dual loop manner, and identify the preliminary feature encoding results corresponding to different fusion combinations. The encoding compensation processing module is used to input the preliminary feature encoding results corresponding to different fusion combinations into the encoding compensation component to obtain the feature encoding results after encoding compensation. The feature fusion module is used to input the feature encoding results after encoding compensation corresponding to different fusion combinations into the feature fusion component to obtain fused features; The fused video decoding and acquisition module is used to decode and output all fused features one by one using the decoding output component to obtain fused videos corresponding to different fused combinations. The feature extraction module is used to input different fused videos into the preset feature extraction component to extract the identity appearance features and dynamic change features corresponding to different fused videos; The identity appearance feature compensation loss calculation module is used to compare the identity appearance features corresponding to different fused videos with the identity appearance features of the corresponding source images for feature compensation loss calculation. After summarizing and calculating, the identity appearance feature compensation loss corresponding to all fused videos is obtained. The motion feature compensation loss calculation module is used to compare the dynamic change features corresponding to different fused videos with the dynamic change features of the corresponding driving videos for feature compensation loss calculation. After summarizing and calculating, the motion feature compensation loss corresponding to all fused videos is obtained. The loss condition identification module is used to identify whether the identity appearance feature compensation loss and the motion feature compensation loss meet preset loss conditions. The model tuning module is used to readjust the hyperparameters of the video stream synthesis model and perform model tuning if the preset loss conditions are not met, until the preset loss conditions are met, and then stop the model tuning process. The video stream synthesis model training completion module is used to obtain the optimized video stream synthesis model as the jointly trained video stream synthesis model if the preset loss condition is met.
[0067] In this embodiment, the video synthesis device 600 further includes a feature enhancement input module and a feature enhancement processing module. Wherein: A feature enhancement input module is used to input the identity appearance features after encoding compensation processing into a preset feature enhancement processing component, wherein the preset feature enhancement processing component includes an image super-resolution decoder; The feature enhancement processing module is used to perform upsampling and super-resolution fusion processing on the encoded and compensated identity appearance features a preset number of times using the image super-resolution decoder to obtain the feature-enhanced identity appearance features.
[0068] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0069] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0070] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 7 , Figure 7 This is a basic structural block diagram of the computer device in this embodiment.
[0071] The computer device 7 includes a memory 7a, a processor 7b, and a network interface 7c that are interconnected via a system bus. It should be noted that... Figure 7Only a computer device 7 with component memory 7a, processor 7b, and network interface 7c is shown. However, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0072] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0073] The memory 7a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 7a may be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 7a may also be an external storage device of the computer device 7, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 7. Of course, the memory 7a may also include both the internal storage unit and its external storage device of the computer device 7. In this embodiment, the memory 7a is typically used to store the operating system and various application software installed on the computer device 7, such as computer-readable instructions for a video synthesis method. In addition, the memory 7a can also be used to temporarily store various types of data that have been output or will be output.
[0074] In some embodiments, the processor 7b may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 7b is typically used to control the overall operation of the computer device 7. In this embodiment, the processor 7b is used to execute computer-readable instructions stored in the memory 7a or to process data, for example, to execute computer-readable instructions for the video compositing method described above.
[0075] The network interface 7c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 7 and other electronic devices.
[0076] The computer device proposed in this embodiment belongs to the field of artificial intelligence technology and is applied to a dynamic digital human video synthesis scenario using a video stream synthesis model. This application acquires source images and driving videos; inputs them in groups into a preset feature extraction component to extract identity appearance features and dynamic change features; inputs these features into a jointly trained video stream synthesis model; encodes the identity appearance features and dynamic change features respectively to obtain preliminary feature encoding results; uses an encoding compensation component learned during joint learning to perform encoding compensation processing on the preliminary feature encoding results to obtain encoded compensation-processed feature encoding results; performs feature fusion processing on the encoded compensation-processed feature encoding results to obtain feature fusion results; and uses a feature decoding component to decode the feature fusion results to obtain the target synthesized video. By introducing an encoding compensation component, it ensures that the video stream synthesis model obtained through final training can, during actual video synthesis, use the encoding compensation component in the trained video stream synthesis model to perform encoding compensation on the static image features of the input source image and encoding compensation on the dynamic change features of the input driving video, thereby generating a more realistic and lifelike synthesized video. Applying the video synthesis method described above to the field of financial technology can enable more realistic and high-quality dynamic digital humans to replace existing human workers in scenarios such as financial courses, financial marketing, and product launches.
[0077] This application also provides another embodiment, namely, a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to perform the steps of the video synthesis method described above.
[0078] The computer-readable storage medium proposed in this embodiment belongs to the field of artificial intelligence technology and is applied to the dynamic digital human video synthesis scenario using a video stream synthesis model. This application acquires source images and driving videos; inputs them in groups into a preset feature extraction component to extract identity appearance features and dynamic change features; inputs these features into a jointly trained video stream synthesis model; encodes the identity appearance features and dynamic change features separately to obtain preliminary feature encoding results; uses an encoding compensation component learned during joint learning to perform encoding compensation processing on the preliminary feature encoding results to obtain encoded compensation-processed feature encoding results; performs feature fusion processing on the encoded compensation-processed feature encoding results to obtain feature fusion results; and uses a feature decoding component to decode the feature fusion results to obtain the target synthesized video. By introducing an encoding compensation component, it ensures that the video stream synthesis model obtained through final training can, during actual video synthesis, use the encoding compensation component in the trained video stream synthesis model to perform encoding compensation on the static image features of the input source image and encoding compensation on the dynamic change features of the input driving video, thereby generating a more realistic and lifelike synthesized video. Applying the video synthesis method described above to the field of financial technology can enable more realistic and high-quality dynamic digital humans to replace existing human workers in scenarios such as financial courses, financial marketing, and product launches.
[0079] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0080] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to make the disclosure of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
Claims
1. A video synthesis method, characterized in that, Includes the following steps: Acquire a source image and a driving video, wherein the source image contains the identity and appearance information of the target synthetic object, and the driving video contains reference action posture and reference facial expression; The source image and the driving video are used as grouped data and input into a preset feature extraction component to extract identity appearance features and dynamic change features; The identity appearance features and the dynamic change features are input together into the video stream synthesis model trained through joint learning; The feature encoding component in the video stream synthesis model is used to encode the identity appearance features and the dynamic change features respectively, to obtain preliminary feature encoding results; The initial feature encoding result is processed by the encoding compensation component learned during joint learning training to obtain the feature encoding result after encoding compensation. The feature encoding results after encoding compensation are subjected to feature fusion processing to obtain the feature fusion result; The feature decoding component in the video stream synthesis model is used to decode the feature fusion result to obtain the target synthesized video.
2. The video synthesis method according to claim 1, characterized in that, The preset feature extraction component includes a first extraction sub-component and a second extraction sub-component. The step of inputting the source image and the driving video as grouped data into the preset feature extraction component to extract identity appearance features and dynamic change features specifically includes: The first extraction component is used to extract the identity appearance features from the source image; The second extraction component is used to extract dynamic change features from the driving video, wherein the dynamic change features include action posture change features and facial expression change features.
3. The video synthesis method according to claim 2, characterized in that, Before performing the step of inputting the identity appearance features and the dynamic change features into the jointly trained video stream synthesis model, the method further includes: Obtain several source images for training and construct a source image training set, where different source images contain different identity and appearance information; The first extraction component is used to extract the identity appearance features from all source images in the source image training set, and the identity appearance features are distinguished and labeled according to the different source images. Obtain several training driving videos and construct a driving video training set, wherein different driving videos contain different motion posture changes and / or facial expression changes; The second extraction component is used to extract the motion posture change features and facial expression change features from all driving videos in the driving video training set. The extraction results of the second extraction component are then distinguished and labeled according to the different driving videos.
4. The video synthesis method according to claim 3, characterized in that, The video stream synthesis model includes a feature encoding component, an encoding compensation component, an encoding fusion component, and a decoding output component. Before performing the step of inputting the identity appearance features and the dynamic change features into the jointly learned and trained video stream synthesis model, the method further includes: The identity and appearance features of all source images in the source image training set, and the dynamic change features of all driving videos in the driving video training set, are input into the video stream synthesis model to be jointly learned and trained. The feature encoding component is used to perform feature encoding on the identity and appearance features in all source images in the source image training set and the dynamic change features in all driving videos in the driving video training set to obtain preliminary feature encoding results. A dual-loop approach is adopted, in which different source images and driving videos are selected sequentially as fusion combinations, and feature encoding fusion processing is performed using the encoding fusion component to obtain the feature fusion encoding result; The decoding output component is used to decode and output each feature fusion encoding result of the dual loop output one by one to obtain the fused video corresponding to different fusion combinations; Different fused videos are input into the preset feature extraction component to extract the identity appearance features and dynamic change features corresponding to the different fused videos; The identity appearance features corresponding to different fused videos are compared with the identity appearance features of the corresponding source images to obtain the differences in identity appearance features. Based on the differences in identity appearance features, an identity appearance feature coding compensation component is constructed. The dynamic change features corresponding to different fused videos are compared with the dynamic change features of the corresponding driving videos to obtain motion feature differences. Based on the motion feature differences, a motion feature coding compensation component is constructed, wherein the motion feature differences include action posture feature differences and facial expression feature differences. The identity appearance feature encoding compensation component and the motion feature encoding compensation component are set as two compensation sub-components of the encoding compensation component to obtain the video stream synthesis model to be optimized.
5. The video synthesis method according to claim 4, characterized in that, After performing the step of setting the identity appearance feature encoding compensation component and the motion feature encoding compensation component as two compensation sub-components of the encoding compensation component to obtain the video stream synthesis model to be optimized, the method further includes: A dual-loop approach is adopted, sequentially selecting different source images and driving videos as fusion combinations to identify the preliminary feature encoding results corresponding to different fusion combinations; The preliminary feature encoding results corresponding to different fusion combinations are input into the encoding compensation component to obtain the feature encoding results after encoding compensation. The feature encoding results after encoding compensation corresponding to different fusion combinations are input into the feature fusion component to obtain fused features; The decoding output component is used to decode and output all fusion features one by one to obtain fusion videos corresponding to different fusion combinations. Different fused videos are input into the preset feature extraction component to extract the identity appearance features and dynamic change features corresponding to the different fused videos; The feature compensation loss is compared between the identity appearance features corresponding to different fused videos and the identity appearance features corresponding to the source images. After summarizing and calculating, the identity appearance feature compensation loss for all fused videos is obtained. The motion feature compensation loss is compared with the motion feature compensation loss of the corresponding driving video and the dynamic change features of the different fused videos. After summarizing and calculating, the motion feature compensation loss of all fused videos is obtained. Identify whether the loss for compensation of the identity appearance features and the loss for compensation of the motion features meet the preset loss conditions; If the preset loss condition is not met, the hyperparameters of the video stream synthesis model are readjusted and model tuning is performed until the preset loss condition is met, at which point the model tuning is stopped. If the preset loss condition is met, the optimized video stream synthesis model is obtained as the video stream synthesis model trained by the joint learning.
6. The video synthesis method according to claim 4 or 5, characterized in that, The step of using the encoding compensation component learned during joint learning training to perform encoding compensation processing on the preliminary feature encoding result to obtain the encoding compensation processed feature encoding result specifically includes: The identity appearance feature encoding compensation component is used to compensate the identity appearance features in the preliminary feature encoding result to obtain the identity appearance features after encoding compensation. The motion feature encoding compensation component is used to compensate for the dynamic change features in the preliminary feature encoding result, resulting in dynamic change features after encoding compensation.
7. The video synthesis method according to claim 1 or 4, characterized in that, The video stream synthesis model further includes a feature enhancement processing component. Before performing the step of performing feature fusion processing on the feature encoding result after encoding compensation processing to obtain the feature fusion result, the method further includes: The identity appearance features after encoding compensation are input into a preset feature enhancement processing component, wherein the preset feature enhancement processing component includes an image super-resolution decoder; Using the image super-resolution decoder, the identity appearance features after encoding compensation are upsampled and super-resolution fusion processed a preset number of times to obtain the enhanced identity appearance features.
8. A video synthesis apparatus, characterized in that, include: The synthesis resource acquisition module is used to acquire source images and driving videos, wherein the source images contain the identity and appearance information of the target synthesis object, and the driving videos contain reference action postures and reference facial expressions; The feature extraction module is used to input the source image and the driving video as grouped data into a preset feature extraction component to extract identity appearance features and dynamic change features; The synthetic feature input module is used to input the identity appearance features and the dynamic change features into the video stream synthesis model trained through joint learning. The feature encoding processing module is used to perform feature encoding on the identity appearance features and the dynamic change features respectively using the feature encoding component in the video stream synthesis model to obtain preliminary feature encoding results; The encoding compensation processing module is used to perform encoding compensation processing on the preliminary feature encoding result using the encoding compensation component learned during joint learning training, so as to obtain the feature encoding result after encoding compensation processing. The feature fusion processing module is used to perform feature fusion processing on the feature encoding results after encoding compensation processing to obtain the feature fusion result; The video decoding module is used to decode the feature fusion result using the feature decoding component in the video stream synthesis model to obtain the target synthesized video.
9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the video synthesis method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the video synthesis method as described in any one of claims 1 to 7.