Video generation method and device, electronic equipment and storage medium

By adjusting the rotation angle of the video generation model, a second video generation model adapted to the video to be generated is generated, which solves the problem of image quality degradation when the video generation model exceeds the length of the sample video, realizes the generation of high-quality videos, and reduces training and deployment costs.

CN122069412APending Publication Date: 2026-05-19BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-02-13
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing video generation models experience rapid degradation in image quality when the length of the generated video exceeds that of the sample video, resulting in poor spatiotemporal coherence of the video.

Method used

By adjusting the reference rotation angle, the rotation angle of the video to be generated is dynamically determined and replaced in the video generation model to generate a second video generation model. Input generation prompts and noisy video to generate a high-quality video.

Benefits of technology

When the video length exceeds the length of the training sample videos, the model improves the color and structural stability of the videos, generates high-quality videos, and reduces the cost of training and deploying new video generation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069412A_ABST
    Figure CN122069412A_ABST
Patent Text Reader

Abstract

The invention relates to a video generation method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a generation prompt message, a noise video, a duration of a first video and a first video generation model; the first video is a video to be generated; based on the duration of the first video, adjusting the reference rotation angle to obtain a first rotation angle; the reference rotation angle is a rotation angle used in rotation position embedding in the first video generation model; replacing a reference rotation angle in the first video generation model with the first rotation angle to obtain a second video generation model; and inputting the generation prompt information and the noise video into a second video generation model to obtain a first video. The purpose of generating a high-quality video under the condition that the duration of a to-be-generated video is greater than the duration of a sample video used for training a first video generation model can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video generation technology, and in particular to a video generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of video generation technology, video generation has demonstrated significant value in fields such as film and television production, advertising creativity, and virtual reality, significantly reducing the threshold for content creation and improving production efficiency and expressiveness.

[0003] In related technologies, video generation models are used to generate videos. These models are trained based on sample videos of a preset length. In practical applications, if the length of the video to be generated exceeds the length of the sample videos, the quality of the generated video deteriorates rapidly over time. Summary of the Invention

[0004] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this application provides a video generation method, apparatus, electronic device and storage medium.

[0005] Firstly, this application provides a video generation method, including: Obtain the generation prompt information, noisy video, duration of the first video, and the first video generation model; the first video is the video to be generated; Based on the duration of the first video, the reference rotation angle is adjusted to obtain the first rotation angle; the reference rotation angle is the rotation angle used in the rotation position embedding in the first video generation model. The first rotation angle is used to replace the reference rotation angle in the first video generation model to obtain the second video generation model; The generated prompt information and the noisy video are input into the second video generation model to obtain the first video.

[0006] Secondly, this application also provides a video generation apparatus, comprising: The acquisition module is used to acquire the generation prompt information, the noisy video, the duration of the first video, and the first video generation model; the first video is the video to be generated. An adjustment module is used to adjust a reference rotation angle based on the duration of the first video to obtain a first rotation angle; the reference rotation angle is the rotation angle used in the rotation position embedding in the first video generation model. The replacement module is used to replace the reference rotation angle in the first video generation model with the first rotation angle to obtain the second video generation model. The generation module is used to input the generation prompt information and the noisy video into the second video generation model to obtain the first video.

[0007] Thirdly, this application also provides an electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described above.

[0008] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video generation method described above.

[0009] The technical solution provided in this application has the following advantages compared with the prior art: The technical solution provided in this application involves setting and acquiring generation prompt information, noisy video, the duration of a first video, and a first video generation model. The first video is the video to be generated. Based on the duration of the first video, a reference rotation angle is adjusted to obtain a first rotation angle. The reference rotation angle is the rotation angle used in the first video at the rotation position. The reference rotation angle in the first video generation model is replaced with the first rotation angle to obtain a second video generation model. The generation prompt information and noisy video are input into the second video generation model to obtain the first video. Essentially, it dynamically determines a first rotation angle that matches the duration of the first video to be generated, and then replaces the reference rotation angle in the first video generation model with this first rotation angle. This reduces the spectral deviation caused by the length of the first video exceeding the duration of the sample videos used to train the first video generation model, thereby improving the color and structural stability of the first video and enhancing its visual effect. Furthermore, by employing the technical solution provided in this application, when the duration of the video to be generated is longer than the duration of the sample videos used to train the first video generation model, it is not necessary to train a new video generation model specifically for the duration of the video to be generated, which can reduce the training and deployment costs of the new video generation model. In summary, it can achieve the goal of generating high-quality videos when the duration of the video to be generated is longer than the duration of the sample videos used to train the first video generation model. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A schematic diagram illustrating a use case of a video generation method provided in an embodiment of this application; Figure 2 A flowchart illustrating a video generation method provided in this application embodiment; Figures 3-6 A schematic diagram illustrating a video generation method provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a video generation device according to an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0013] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0014] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.

[0015] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0016] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this application's technical solution, based on the prompt message.

[0017] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0018] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this application. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.

[0019] The video generation method provided in this application embodiment can be applied to... Figure 1 The video generation system shown. For example... Figure 1 As shown, the video generation system may include terminal devices, servers, and database servers. A medium (e.g., a network) may be provided to connect the terminal devices to the servers and database servers. This network may include various connection types, such as wired or wireless communication links.

[0020] Various applications (APPs) or software can be installed on terminal devices, such as video playback applications or software, live streaming applications or software, collaborative office applications or software, image generation or processing applications or software, video generation or processing applications or software, media resource management applications or software, video conferencing applications or software, reading applications or software, social networking applications or software, payment applications or software, web browsers, and instant messaging tools. In some embodiments, these applications or software can be iterated and changed based on a continuous integration / continuous delivery (CI / CD) model.

[0021] Terminal devices include, but are not limited to, smartphones, tablets, e-book readers, MP3 players, laptops, and desktop computers (PCs).

[0022] A server can be a backend device that provides background support for applications or software running on terminal devices, performing functions such as business logic, model inference, and task scheduling. A database server can be a dedicated server that provides data-related services, such as storing, managing, and querying data involved in the video generation process, including but not limited to user-input generation prompts, generation parameter configurations, intermediate feature caching, historical generation results, and their metadata. It is understandable that if the server can perform the functions of a database server, a database server may not be necessary in the video generation system.

[0023] The servers and database servers here can be implemented as a distributed server cluster consisting of multiple servers, or as a single server.

[0024] When a user uses a terminal device, the device displays a user interface. The user can issue commands to the application or software through input operations, including but not limited to text input, voice commands, touch interaction, or gesture operations. After receiving the input operations, the application or software processes them based on its functional logic and outputs corresponding feedback information to the user. This feedback information can be presented in the form of a visual user interface, voice broadcast, notification prompts, etc., thus completing a round of human-computer interaction.

[0025] It should be noted that the video generation method provided in this application embodiment can be executed by a terminal device, by a server, or by interactive execution of various devices in the video generation system. It should be understood that... Figure 1 The number of terminal devices, servers, and database servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, servers, and database servers can be included.

[0026] Figure 2 This is a flowchart illustrating a video generation method provided in an embodiment of this application. Figure 2 As shown, the method may specifically include: S110. Obtain the generation prompt information, noisy video, duration of the first video, and first video generation model; the first video is the video to be generated.

[0027] The first video generation model can be, for example, a pre-trained video generation model. The first video generation model corresponds to a first preset duration. For example, the correspondence between the first video generation model and the first preset duration can mean that, during the training process of the first video generation model, the longest duration of the sample videos used is equal to the first preset duration. If the video generated using the first video generation model is less than or equal to the first preset duration, the generated video has good spatiotemporal coherence, meaning the generated video meets the criteria for high-quality spatiotemporal coherence. If the video generated using the first video generation model is longer than the first preset duration, the generated video has poor spatiotemporal coherence, meaning the generated video does not meet the criteria for high-quality spatiotemporal coherence and has poor image quality. This application does not limit the specific criteria for high-quality spatiotemporal coherence. For example, the criteria for high-quality spatiotemporal coherence are related to at least one of the following features: video image quality, video content coherence, structural consistency of the same object in different image frames, smoothness of object motion, and temporal logic of events.

[0028] The purpose of adopting the technical solution provided in this application is to generate a first video. In other words, the first video is the video to be generated. The duration of the first video can be user-specified or a default duration can be used. The duration of the first video is longer than the first preset duration corresponding to the first video generation model. That is, all sample videos used when training the first video generation model are shorter than the duration of the first video.

[0029] See Figure 3 The method for generating the first video involves modifying the first video generation model, which is then referred to as the second video generation model. Generated prompts and noisy video are input into the second video generation model to generate the first video. The generated first video exhibits better spatiotemporal coherence.

[0030] The generated prompts can be external input signals used to guide the generation of video content, and can take the form of text descriptions, reference images, audio, or existing video clips. In some scenarios, the generated prompts are specified, entered, or confirmed by the user, reflecting the user's needs, i.e., the content or visual effects the user expects the first video to include.

[0031] The noisy video can be, for example, a video to be denoised. Under the control of generating prompt information, the second video generation model generates the first video by denoising the noisy video. Optionally, the noisy video includes multiple noisy video blocks.

[0032] In some scenarios, the process of generating the first video includes: determining the number of first video blocks included in the first video based on a first value and the duration of the first video; the first value is preset to limit the number of image frames included in a video block; determining noisy video based on the number of first video blocks included in the first video; the number of noisy video blocks included in the noisy video is equal to the number of first video blocks included in the first video, and the number of noisy image frames included in any noisy video block is equal to the first value; performing noise reduction processing on each noisy video block in sequence to obtain multiple first video blocks; and stitching the first video blocks together to obtain the first video.

[0033] A video block is the basic processing unit for temporal-spatial joint features in a video generation model (such as the first video generation model). The first video block is a video block in the first video, and the noisy video block is a video block in the noisy video.

[0034] Optionally, in practice, the user specifies the duration of the first video to be generated, calculates the number of image frames included in the first video based on the duration, and then determines the number of first video blocks included in the first video based on the number of image frames included in the first video and a first value. Optionally, the number of first video blocks included in the first video is equal to the ratio of the number of image frames included in the first video to the first value.

[0035] See Figure 4 For example, if the first video includes 126 image frames and the first value is 3 frames, the number of first video blocks in the first video is determined to be 42, and each first video block includes 3 image frames. To generate this first video, the noisy video is set to include 42 noisy video blocks, and each noisy video block includes 3 noisy image frames. Each noisy video block corresponds one-to-one with a first video block. First video block 1 is obtained by denoising noisy video block 1, first video block 2 is obtained by denoising noisy video block 2, and so on.

[0036] S120. Based on the duration of the first video, the reference rotation angle is adjusted to obtain the first rotation angle; the reference rotation angle is the rotation angle used in the rotation position embedding in the first video generation model.

[0037] Rotational position embedding can be achieved by injecting positional information by rotating the feature vector of a video block by a specific angle, enabling the attention mechanism to perceive the relative distance between different video blocks.

[0038] The rotation angle can be, for example, a hyperparameter used in the rotation position embedding mechanism to limit the basic rotation speed of each feature dimension. The reference rotation angle is the rotation angle used in the rotation position embedding of the first video generation model; the reference rotation angle can be regarded as an inherent attribute value in the first video generation model. During the training of the first video generation model, the reference rotation angle remains fixed and is not updated as a learnable parameter.

[0039] It should be noted that in practice, during video generation, noisy video blocks are encoded as latent space vectors, which are high-dimensional vectors. The first video generation model includes multiple reference rotation angles, each corresponding to a different latent space dimension (also known as a feature dimension). When embedding each noisy video block at its rotation position, the rotation angle used for a specific latent space dimension is determined based on the position information of the noisy video block within the noisy video and the rotation angle corresponding to that feature dimension. In this application, when adjustments to the reference rotation angles are required, only some reference rotation angles may be adjusted, or all reference rotation angles may be adjusted.

[0040] The first rotation angle is the result after adjusting the reference rotation angle.

[0041] S130. Replace the reference rotation angle in the first video generation model with the first rotation angle to obtain the second video generation model.

[0042] It should be noted that when adjusting the reference rotation angle for a specific latent space dimension, the resulting first rotation angle still corresponds to that specific latent space dimension. When it is necessary to replace a reference rotation angle, it must be ensured that the reference rotation angle before replacement and the first rotation angle after replacement correspond to the same latent space dimension. For example, in Figure 3 In this context, the first rotation angle θ1' is the result of adjusting the reference rotation angle θ1, and the first rotation angle θ1' and the reference rotation angle θ1 correspond to the same latent space dimension. The reference rotation angle θ1 is replaced with the first rotation angle θ1'. The relationships between other first rotation angles and reference rotation angles follow the same pattern.

[0043] In the first video generation model, when denoising a specific noisy video block, a first actual rotation angle corresponding to the noisy video block is determined based on a reference rotation angle and the position information of the noisy video block within the noisy video. Based on the first actual rotation angle, the rotation position of the noisy video block is embedded.

[0044] In the second video generation model, since the reference rotation angle in the first video generation model is replaced by the first rotation angle, when denoising a specific noisy video block, a second actual rotation angle corresponding to the noisy video block is determined based on the first rotation angle and the position information of the noisy video block in the noisy video. Based on the second actual rotation angle, the rotation position of the noisy video block is embedded.

[0045] S140. Input the generated prompt information and noisy video into the second video generation model to obtain the first video.

[0046] The second video generation model is used to denoise the noisy video based on the generated prompt information to obtain the first video.

[0047] In practice, for the first video generation model, if the length of the video to be generated exceeds the length of the sample videos used to train the first video generation model, the position index in the time dimension continues to increase and far exceeds the training range, resulting in severe spectral bias, which disrupts the attention mechanism's interpretation of spatiotemporal context and causes rapid degradation in image quality. The above technical solution obtains the generated prompt information, the noisy video, the duration of the first video, and the first video generation model by setting the parameters. The first video is the video to be generated. Based on the duration of the first video, the reference rotation angle is adjusted to obtain the first rotation angle. The reference rotation angle is the rotation angle used in the rotation position embedding within the first video. The first rotation angle is used to replace the reference rotation angle in the first video generation model to obtain the second video generation model. The generated prompt information and the noisy video are input into the second video generation model to obtain the first video. Essentially, this method dynamically determines a first rotation angle adapted to the duration of the first video to be generated. This first rotation angle is then used to replace the reference rotation angle in the first video generation model. This mitigates the spectral bias caused by the first video's length exceeding the duration of the sample videos used to train the model, thereby improving the color and structural stability of the first video and enhancing its visual appeal. Furthermore, this technique eliminates the need to train a new video generation model specifically for the duration of the first video when the video to be generated exceeds the duration of the sample videos used to train the model, reducing the training and deployment costs of a new model. In summary, this method can generate high-quality videos even when the duration of the video to be generated exceeds the duration of the sample videos used to train the model.

[0048] Based on the above technical solution, optionally, the noisy video includes multiple noisy video blocks; S120 may include: determining a reference rotation angle and the number of training rotation cycles; the number of training rotation cycles is the number of rotation cycles completed during the training of the first video generation model by embedding the rotation position with the reference rotation angle; and adjusting the reference rotation angle based on the number of training rotation cycles and the duration of the first video to obtain the first rotation angle.

[0049] Here, the number of training rotation cycles is used as a parameter to measure whether the training of the first video generation model is sufficient. If the first video generation model is not sufficiently trained, it means that the reference rotation angle needs to be adjusted to stabilize the encoding effect of out-of-distribution images (such as image frames in the video to be generated whose frame numbers exceed the maximum number of frames in the sample videos used during the training phase). If the first video generation model is sufficiently trained, it means that no adjustment of the reference rotation angle is required.

[0050] Optionally, "determining the number of training rotation cycles" may include: determining the maximum number of frames in the sample video; the sample video being the video used during the training of the first video generation model; and determining the number of training rotation cycles based on the number of frames in the sample video and the reference rotation angle.

[0051] For example, the number of training rotation cycles is directly proportional to the maximum number of frames in the sample video, and inversely proportional to the first wavelength. The first wavelength is the number of frames required to complete one rotation at a reference rotation angle. Optionally, the first wavelength is inversely proportional to the reference rotation angle.

[0052] There are multiple ways to implement "adjusting the reference rotation angle based on the number of training rotation cycles and the duration of the first video to obtain the first rotation angle," and this application does not limit this. "Adjusting the reference rotation angle based on the number of training rotation cycles and the duration of the first video to obtain the first rotation angle" includes: if the number of training rotation cycles is less than a first threshold, reducing the reference rotation angle based on the duration of the first video to obtain the first rotation angle; if the number of training rotation cycles is greater than a second threshold, using the reference rotation angle as the first rotation angle; the second threshold is greater than the first threshold; if the number of training rotation cycles is greater than or equal to the first threshold and less than or equal to the second threshold, reducing the reference rotation angle based on the duration of the first video to obtain the second rotation angle; and determining the first rotation angle based on the second rotation angle and the reference rotation angle.

[0053] The first and second thresholds can be pre-set values, used to measure whether the training of the first video generation model is sufficient. This application does not limit the specific values ​​of the first and second thresholds. The first threshold is less than the second threshold.

[0054] If, for any latent space dimension, the number of training rotation cycles for the reference rotation angle corresponding to that latent space dimension is less than a first threshold, it means that the positional coverage experienced by that latent space dimension during training is insufficient. The first video generation model has failed to fully learn the positional dependencies within that latent space dimension, leading to a higher probability of spectral deviations in out-of-distribution scenes. This easily disrupts the attention mechanism's ability to interpret spatiotemporal context, necessitating adjustment of the reference rotation angle corresponding to that latent space dimension. The adjustment method is to reduce the reference rotation angle. Reducing the reference rotation angle slows down the phase growth rate of video blocks in that latent space dimension, making phase growth smoother, thereby reducing positional confusion, lowering the probability of spectral deviations, improving the attention mechanism's ability to interpret spatiotemporal context, and enhancing the color and structural stability of the generated image.

[0055] There are various specific implementation methods for "reducing the reference rotation angle based on the duration of the first video", and this application does not limit this one. For example, "reducing the reference rotation angle based on the duration of the first video" includes: determining the number of frames in the first video based on the duration of the first video; determining the reduction ratio based on the maximum number of frames in the sample video and the number of frames in the first video; the sample video is the video used in the training process of the first video generation model; and reducing the reference rotation angle based on the reduction ratio.

[0056] For example, the reduction ratio is directly proportional to the number of frames in the first video and inversely proportional to the maximum number of frames in the sample video. When the maximum number of frames in the sample video is less than the number of frames in the first video, the reduction ratio is greater than 1.

[0057] By setting the maximum number of frames based on the sample video and the number of frames in the first video, the first rotation angle can be determined. This ensures that the subsequently determined first rotation angle is suitable for generating the first video, which helps to ensure that the first video has better color and structural stability.

[0058] For any latent space dimension, if the number of training rotation cycles corresponding to the reference rotation angle of that latent space dimension is greater than the first threshold, it means that the position coverage experienced by that latent space dimension during training is sufficient, and the first video generation model has fully learned the positional dependencies under that latent space dimension. The probability of it having spectral deviation in out-of-distribution scenes is small, which will not destroy the attention mechanism's interpretation of the spatiotemporal context, and there is no need to adjust the reference rotation angle corresponding to that latent space dimension.

[0059] For any latent space dimension, if the number of training rotation cycles for the reference rotation angle corresponding to that latent space dimension is greater than or equal to a first threshold and less than or equal to a second threshold, it means that the latent space dimension is in a transitional state during training and the reference rotation angle corresponding to that latent space dimension needs to be adjusted. Specifically, in this scenario, the method for adjusting the reference rotation angle can be as follows: First, determine the reduction ratio based on the maximum number of frames in the sample video and the number of frames in the first video; the sample video is the video used during the training of the first video generation model; based on the reduction ratio, reduce the reference rotation angle to obtain a second rotation angle; based on the second rotation angle and the reference rotation angle, determine the first rotation angle. Further, determining the first rotation angle based on the second rotation angle and the reference rotation angle can be set by: performing a weighted calculation on the second rotation angle and the reference rotation angle to obtain the first rotation angle.

[0060] By setting a strategy for determining the first rotation angle based on the relationship between the number of training rotation cycles and the magnitudes of the first and second thresholds, the essence is to mitigate the long-term video generation collapse caused by spectral bias by mapping back to the training phase domain through interpolation in the low-frequency latent space dimension. In the high-frequency latent space dimension, the fine-grained resolution of the first video generation model is preserved, reducing the probability of high-frequency loss during inference; and in the mid-frequency latent space dimension, global and local performance are balanced through weighted combination.

[0061] Furthermore, in the noisy video, multiple noisy video blocks are arranged sequentially to form a first sequence; the second video generation model is used to perform noise reduction processing on the noisy video based on the generated prompt information to obtain the first video, including: for any noisy video block, the second video generation model performs T rounds of noise reduction on the noisy video block based on the generated prompt information to obtain the first video block corresponding to the noisy video block; determine the sequence number of the first video block, the sequence number of the corresponding noisy video block and the first video block are the same, and the sequence number of the noisy video block is the number of the noisy video block in the first sequence; based on the sequence number of the first video block, the first video blocks are spliced ​​to obtain the first video. T is a positive integer, and this application does not limit the specific value of T.

[0062] For implementation details, see [link / reference]. Figure 4 The noisy video consists of multiple noisy video blocks, arranged sequentially to form a first sequence. In this first sequence, the sequence number of noisy video block 1 is 1, the sequence number of noisy video block 2 is 2, and so on. A second video generation model performs T rounds of denoising on noisy video block 1 based on generated prompts, resulting in first video block 1. The sequence number of first video block 1 is 1. The second video generation model then performs T rounds of denoising on noisy video block 2 based on generated prompts, resulting in first video block 2. The sequence number of first video block 2 is 2. This process is repeated until denoising has been performed on all noisy video blocks. The first video blocks are then concatenated according to their sequence numbers to obtain the first video. It should be noted that in practice, before denoising, the noisy video blocks in the noisy video are encoded as feature vectors (i.e., latent space vectors). During the denoising and concatenation processes, both the noisy video blocks and the first video block are represented using feature vectors. After concatenating the feature vectors of the first video blocks, a decoder restores them to the video in image space.

[0063] Based on the above technical solution, optionally, the noisy video includes multiple noisy video blocks; the second video generation model performs noise reduction processing on the noisy video blocks, including: determining the feature vector of the video block to be denoised; determining the feature vector of a first reference image frame; the first reference image frame is the first preset number of image frames of the first video; determining the feature vector of a second reference image frame; the second reference image frame is the second preset number of image frames adjacent to the video block to be denoised; the frame number of the second reference image frame is less than the minimum frame number of the image frames in the video block to be denoised; concatenating the feature vector of the first reference image frame, the feature vector of the second reference image frame, and the feature vector of the video block to be denoised to obtain a first feature vector; and performing noise reduction processing on the video block to be denoised based on the first feature vector to obtain a noise reduction result corresponding to the video block to be denoised.

[0064] The first preset quantity and the second preset quantity can be, for example, a pre-specified value. This application does not restrict the specific values ​​of the first preset quantity and the second preset quantity. In practice, the values ​​of the first preset quantity and the second preset quantity can be equal or unequal.

[0065] In practice, if the user specifies one or more starting image frames as the beginning of the first video, in some cases, the first reference image frame may only include all or some of the image frames specified by the user; in other cases, the first reference image frame, in addition to including all or some of the image frames specified by the user, also includes an image frame generated by the second video generation model that follows the image frame specified by the user. If the user does not specify an image frame as the beginning of the first video, the first reference image frame is the image frame generated by the second video generation model that is located at the beginning of the first video.

[0066] It should be noted that the first and second reference image frames are sharp image frames, which are image frames from the first video. Here, "sharp" means that if the first and second reference image frames are generated by the second video generation model, they are denoised image frames, or generated image frames. The feature vector of the first reference image frame is its vector representation in the latent space. The feature vector of the second reference image frame is its vector representation in the latent space.

[0067] For example, see Figure 5The first preset quantity is set to 3, and the second preset quantity is set to 12. One video block includes 3 image frames. Assuming the current video block to be denoised is noisy video block 21, this means that all other noisy video blocks with sequence numbers lower than noisy video block 21 have already been denoised, i.e., clear image frames 1-60 in the first video have been obtained. When denoising noisy video block 21, image frames 1 to 3 are used as the first reference image frames, and image frames 49 to 60 are used as the second reference image frames. The feature vectors of the first reference image frames, the second reference image frames, and the noisy video block 21 are concatenated to obtain the first feature vector. Subsequently, based on the first feature vector, noise prediction is performed, and then noisy video block 21 is denoised to obtain the denoised result of noisy video block 21.

[0068] Assuming the current video block to be denoised is noisy video block 22, this means that all other noisy video blocks with sequence numbers lower than noisy video block 22 have already been denoised, i.e., clear image frames 1-63 in the first video have been obtained. When denoising noisy video block 22, image frames 1 to 3 are used as the first reference image frames, and image frames 61 to 63 are used as the second reference image frames. The feature vectors of the first reference image frames, the second reference image frames, and the noisy video block 22 are concatenated to obtain the first feature vector. Subsequently, based on the first feature vector, noise prediction is performed, and then noisy video block 22 is denoised to obtain the denoised result of noisy video block 22.

[0069] By concatenating the feature vectors of the video block to be denoised with the feature vectors of the first reference image frame before denoising, the aim is to force attention to the first few frames of the first video during the attention calculation process, establishing long-term semantic constraints, reducing the probability of semantic shift, and improving the global spatiotemporal consistency of the first video during the expansion process. Similarly, by concatenating the feature vectors of the video block to be denoised with the feature vectors of the second reference image frame before denoising, the aim is to improve the coherence of adjacent video blocks in the first video in terms of image content.

[0070] Based on the first feature vector, denoising is performed on the noisy video block to obtain the denoising result corresponding to the video block to be denoised. This includes: projecting the first feature vector onto the query subspace, key subspace, and value subspace respectively to obtain a first query feature vector, a first key feature vector, and a first value feature vector; determining a second actual rotation angle corresponding to the noisy video block, which is based on the first rotation angle and the position information of the video block to be denoised in the noisy video; rotating the first query feature vector and the first key feature vector based on the second actual rotation angle to obtain a second query feature vector and a second key feature vector; and performing T rounds of denoising starting from the second query feature vector, the second key feature vector, and the first value feature vector to obtain the denoising result corresponding to the video block to be denoised.

[0071] Based on the above technical solution, optionally, the noisy video includes multiple noisy video blocks; acquiring the noisy video includes: determining a first value; the first value is the number of noisy image frames in the noisy video block; for any noisy video block, determining the first noisy image frame in the noisy video block; letting i=1, repeating the determination step until i+1 equals the first value; wherein, the determination step includes: based on the i-th noisy image frame, determining the (i+1)-th noisy image frame in the noisy video block; updating the value of i so that the difference between the updated value of i and the original value of i is 1; splicing the first noisy image frame to the (i+1)-th noisy image frame in the noisy video block to obtain the noisy video block.

[0072] In a noisy video block, the first noisy image frame may be, for example, a pre-specified noisy image. This application does not limit the specific method or image used as the first noisy image frame. In a noisy video block, other noisy image frames besides the first noisy image frame are obtained based on the image frame located in the preceding frame of the same noisy video block.

[0073] For example, see Figure 6 When determining noisy video block 1, a preset noisy image can be used as the first noisy image frame (i.e., noisy image frame 1); a second noisy image frame (i.e., noisy image frame 2) is obtained based on noisy image frame 1; and a third noisy image frame (i.e., noisy image frame 3) is obtained based on noisy image frame 2. When determining noisy video block 2, a preset noisy image can be used as the first noisy image frame (i.e., noisy image frame 4); a second noisy image frame (i.e., noisy image frame 5) is obtained based on noisy image frame 4; and a third noisy image frame (i.e., noisy image frame 6) is obtained based on noisy image frame 5.

[0074] Further, based on the i-th noisy image frame, determining the (i+1)-th noisy image frame in the noisy video block includes: determining a first intermediate image based on the i-th noisy image frame and a first weight; determining a second intermediate image based on a preset noisy image and a second weight; and fusing the first intermediate image and the second intermediate image to obtain the (i+1)-th noisy image frame. Optionally, the preset noisy image is a different noisy image from the first noisy image frame. The preset noisy image is a random noise image, such as a Gaussian noise image.

[0075] By setting other noise image frames in a noisy video block, excluding the first noise image frame, the results are obtained based on the image frame located in the preceding frame within the same noisy video block. Essentially, these other noise image frames, besides the first one, include not only random noise (i.e., a preset noise map) but also some noise from the previous frame's noise map. In other words, when determining the other noise images besides the first one, a temporal structural correlation between adjacent image frames is introduced, which can improve the visual effect of the generated first video.

[0076] In practice, optionally, the first weight is negative and the second weight is positive; or, the first weight is positive and the second weight is positive. When the first weight is negative, it means that in the same noisy video block, the shared noise in adjacent noisy image frames exhibits an inverse correlation. This leads to the preservation of more high-frequency temporal variation components during the denoising process, which can, to some extent, offset motion attenuation during the diffusion process, allowing the generated first video to maintain a significant and natural dynamic effect at a distance. When the first weight is positive, the noise in adjacent frames tends to be consistent, enhancing the continuity of inter-frame perturbations and helping to improve the spatiotemporal consistency of the generated video in terms of color, structural layout, visual style, and overall content.

[0077] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0078] Figure 7 This is a schematic diagram of the structure of a video generation device according to an embodiment of this application. See also... Figure 7 The video generation device specifically includes: The acquisition module 210 is used to acquire the generation prompt information, the noisy video, the duration of the first video, and the first video generation model; the first video is the video to be generated. The adjustment module 220 is used to adjust the reference rotation angle based on the duration of the first video to obtain a first rotation angle; the reference rotation angle is the rotation angle used in the rotation position embedding in the first video generation model. Replacement module 230 is used to replace the reference rotation angle in the first video generation model with the first rotation angle to obtain a second video generation model; The generation module 240 is used to input the generation prompt information and the noisy video into the second video generation model to obtain the first video.

[0079] Furthermore, module 220 is adjusted for: Determine the reference rotation angle and the number of training rotation cycles; the number of training rotation cycles is the number of rotation cycles completed during the training of the first video generation model by embedding the rotation position with the reference rotation angle. Based on the number of training rotation cycles and the duration of the first video, the reference rotation angle is adjusted to obtain the first rotation angle.

[0080] Furthermore, module 220 is adjusted for: Determine the maximum number of frames in the sample video; the sample video is the video used during the training of the first video generation model. The number of training rotation cycles is determined based on the maximum number of frames in the sample video and the reference rotation angle.

[0081] Furthermore, module 220 is adjusted for: If the number of training rotation cycles is less than a first threshold, the reference rotation angle is reduced based on the duration of the first video to obtain a first rotation angle; If the number of training rotation cycles is greater than the second threshold, the reference rotation angle is used as the first rotation angle; the second threshold is greater than the first threshold. If the number of training rotation cycles is greater than or equal to the first threshold and less than or equal to the second threshold, the reference rotation angle is reduced based on the duration of the first video to obtain a second rotation angle; based on the second rotation angle and the reference rotation angle, the first rotation angle is determined.

[0082] Furthermore, module 220 is adjusted for: The number of frames in the first video is determined based on its duration. The reduction ratio is determined based on the maximum number of frames in the sample video and the number of frames in the first video; the sample video is the video used during the training of the first video generation model. Based on the reduction ratio, the reference rotation angle is reduced in size.

[0083] Furthermore, the noisy video includes multiple noisy video blocks; Generate modules for: Determine the feature vector of the video block to be denoised; Determine the feature vector of the first reference image frame; the first reference image frame is the first preset number of image frames of the first video. The feature vector of the second reference image frame is determined; the second reference image frame is a second preset number of image frames that are adjacent to the video block to be denoised; the frame number of the second reference image frame is less than the minimum frame number of the image frames in the video block to be denoised. The feature vectors of the first reference image frame, the second reference image frame, and the video block to be denoised are concatenated to obtain the first feature vector. Based on the first feature vector, the video block to be denoised is denoised to obtain the denoising result corresponding to the video block to be denoised.

[0084] Furthermore, the noisy video includes multiple noisy video blocks; the acquisition module is used for: Determine a first value; the first value is the number of noisy image frames in the noisy video block; For any of the aforementioned noisy video blocks, determine the first noisy image frame in the noisy video block; Let i=1, and repeat the determination steps until i+1 equals the first value; The determining step includes: determining the (i+1)th noise image frame in the noise video block based on the i-th noise image frame; updating the value of i so that the difference between the updated value of i and the original value of i is 1; and splicing the first noise image frame to the (i+1)th noise image frame in the noise video block to obtain the noise video block.

[0085] Furthermore, the acquisition module is used for: Based on the i-th noisy image frame and the first weight, the first intermediate image is determined; The second intermediate image is determined based on the preset noisy image and the second weight; The first intermediate image and the second intermediate image are fused to obtain the (i+1)th noisy image frame.

[0086] Furthermore, the first weight is negative, and the second weight is positive; or, The first weight is a positive number, and the second weight is a positive number.

[0087] The video generation apparatus provided in this application embodiment can execute the steps performed by the client or server in the video generation method provided in this application method embodiment, and has the execution steps and beneficial effects, which will not be described in detail here.

[0088] The following is a detailed reference. Figure 8 The diagram illustrates a structural schematic suitable for implementing the electronic device 1000 in the embodiments of this application. The electronic device 1000 in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 8 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0089] like Figure 8 As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003 to implement the video generation method as described in the embodiments of this application. The RAM 1003 also stores various programs and information required for the operation of the electronic device 1000. The processing device 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0090] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic device 1000 to exchange information with other devices wirelessly or via wired communication. Although Figure 8 An electronic device 1000 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0091] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the video generation method as described above. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1008, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this application.

[0092] It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include information signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated information signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0093] In some implementations, clients and servers may communicate using any known or future network protocol, such as HTTP (Hypertext Transfer Protocol), and may interconnect with any form or medium of digital information communication (e.g., communication networks). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any known or future networks.

[0094] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0095] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: Obtain the generation prompt information, noisy video, duration of the first video, and the first video generation model; the first video is the video to be generated; Based on the duration of the first video, the reference rotation angle is adjusted to obtain the first rotation angle; the reference rotation angle is the rotation angle used in the rotation position embedding in the first video generation model. The first rotation angle is used to replace the reference rotation angle in the first video generation model to obtain the second video generation model; The generated prompt information and the noisy video are input into the second video generation model to obtain the first video.

[0096] Optionally, when one or more of the above-described procedures are executed by the electronic device, the electronic device may also perform other steps described in the above embodiments.

[0097] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0099] The units described in the embodiments of this application can be implemented in software or hardware. The names of the units are not, in some cases, limiting the scope of the unit itself.

[0100] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0101] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0102] According to one or more embodiments of this application, this application provides an electronic device, including: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement any of the video generation methods provided in this application.

[0103] According to one or more embodiments of this application, this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements any of the video generation methods provided in this application.

[0104] This application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the video generation method described above.

[0105] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0106] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video generation method, comprising: Obtain the generated prompt information, noisy video, duration of the first video, and the first video generation model; The first video is the video to be generated; Based on the duration of the first video, the reference rotation angle is adjusted to obtain the first rotation angle; The reference rotation angle is the rotation angle used in the rotation position embedding in the first video generation model; The first rotation angle is used to replace the reference rotation angle in the first video generation model to obtain the second video generation model; The generated prompt information and the noisy video are input into the second video generation model to obtain the first video.

2. The method according to claim 1, wherein adjusting the reference rotation angle based on the duration of the first video to obtain the first rotation angle includes: Determine the reference rotation angle and the number of training rotation cycles; The number of training rotation cycles is the number of rotation cycles completed during the training of the first video generation model, with the reference rotation angle used for rotation position embedding. Based on the number of training rotation cycles and the duration of the first video, the reference rotation angle is adjusted to obtain the first rotation angle.

3. The method according to claim 2, wherein determining the number of training rotation cycles includes: Determine the maximum number of frames in the sample video; The sample video is the video used during the training of the first video generation model; The number of training rotation cycles is determined based on the maximum number of frames in the sample video and the reference rotation angle.

4. The method according to claim 2, wherein adjusting the reference rotation angle based on the number of training rotation cycles and the duration of the first video to obtain the first rotation angle includes: If the number of training rotation cycles is less than a first threshold, the reference rotation angle is reduced based on the duration of the first video to obtain a first rotation angle; If the number of training rotation cycles is greater than the second threshold, the reference rotation angle is used as the first rotation angle; the second threshold is greater than the first threshold. If the number of training rotation cycles is greater than or equal to the first threshold and less than or equal to the second threshold, the reference rotation angle is reduced based on the duration of the first video to obtain the second rotation angle. The first rotation angle is determined based on the second rotation angle and the reference rotation angle.

5. The method according to claim 4, wherein reducing the reference rotation angle based on the duration of the first video includes: The number of frames in the first video is determined based on its duration. The reduction ratio is determined based on the maximum number of frames in the sample video and the number of frames in the first video; The sample video is the video used during the training of the first video generation model; Based on the reduction ratio, the reference rotation angle is reduced in size.

6. The method according to claim 1, wherein the noisy video comprises a plurality of noisy video blocks; the second video generation model performs noise reduction processing on the noisy video blocks, comprising: Determine the feature vector of the video block to be denoised; Determine the feature vector of the first reference image frame; The first reference image frame is the first preset number of image frames of the first video; Determine the feature vector of the second reference image frame; the second reference image frame is a second preset number of image frames that are adjacent to the video block to be denoised. The frame number of the second reference image frame is less than the minimum frame number of the image frames in the video block to be denoised; The feature vectors of the first reference image frame, the second reference image frame, and the video block to be denoised are concatenated to obtain the first feature vector. Based on the first feature vector, the video block to be denoised is denoised to obtain the denoising result corresponding to the video block to be denoised.

7. The method according to claim 1, wherein the noisy video comprises a plurality of noisy video blocks; The acquisition of the noisy video includes: Determine a first value; the first value is the number of noisy image frames in the noisy video block; For any of the aforementioned noisy video blocks, determine the first noisy image frame in the noisy video block; Let i=1, and repeat the determination steps until i+1 equals the first value; The determining step includes: determining the (i+1)th noise image frame in the noise video block based on the i-th noise image frame; updating the value of i so that the difference between the updated value of i and the original value of i is 1; and splicing the first noise image frame to the (i+1)th noise image frame in the noise video block to obtain the noise video block.

8. A video generation apparatus, comprising: The acquisition module is used to acquire the generated prompt information, noisy video, duration of the first video, and the first video generation model; The first video is the video to be generated; An adjustment module is used to adjust a reference rotation angle based on the duration of the first video to obtain a first rotation angle; The reference rotation angle is the rotation angle used in the rotation position embedding in the first video generation model; The replacement module is used to replace the reference rotation angle in the first video generation model with the first rotation angle to obtain the second video generation model. The generation module is used to input the generation prompt information and the noisy video into the second video generation model to obtain the first video.

9. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method as described in any one of claims 1-7.