A video generation method, an electronic device, a storage medium and a product
Patent Information
- Application Number
- CN202610903402.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]但是,受限于累计误差,时间一致性的交互式视频难以突破帧数限制
[0010]在一种情形下,通过获取第一信息,在已生成的视频帧中筛选与第一信息相关的第一视频帧,基于第一视频帧获取第二信息,以表征第一视频帧的三维信息,以第二信息为辅助信息,生成第一信息关联的第一视频块,通过第二信息可降低或消除生成视频过程中产生的累计误差,减少累计误差对第一视频块的影响,提高第一视频块针对虚拟场景的全局一致性,进而提高用户与虚拟场景的高质量交互时长,提高生成视频的质量。
Smart Images

Figure CN122802750A_ABST
Abstract
Description
Technical Field
[0001] The embodiments described herein relate to computer technology, and more particularly to a video generation method, electronic device, storage medium, and product. Background Technology
[0002] Interactive world modeling is widely used in fields such as virtual reality and human-computer interaction. The main requirement in these fields is to generate temporally coherent, high-fidelity video sequences based on user interactions.
[0003] However, due to accumulated errors, interactive videos with temporal consistency are difficult to overcome frame rate limitations. Summary of the Invention
[0004] This invention provides a video generation method, electronic device, storage medium, and product that reduce the impact of accumulated errors during interactive video generation and improve the frame rate of time-consistent interactive videos.
[0005] In one scenario, this paper provides a video generation method, including: Obtain first information, which is used to characterize a first viewpoint, and the first viewpoint is used to indicate the generation of a first video block; A first video frame is obtained, which is related to the first information. The first video frame is obtained from a second video frame, which is a generated video frame. Obtain second information, which is used to characterize the three-dimensional information of the first video frame; The first video block is generated based on the first information and the second information.
[0006] In one scenario, this paper also provides a video generation apparatus, comprising: The first information acquisition module is used to acquire first information, which is used to characterize a first viewpoint and the first viewpoint is used to indicate the generation of a first video block. The first video frame acquisition module is used to acquire a first video frame, which is related to the first information. The first video frame is acquired from a second video frame, which is a generated video frame. The second information acquisition module is used to acquire second information, which is used to characterize the three-dimensional information of the first video frame. The video generation module is used to generate the first video block based on the first information and the second information.
[0007] In one instance, this document also provides an electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described herein.
[0008] In one instance, this document also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform the video generation method as described herein.
[0009] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the video generation method as described herein.
[0010] In one scenario, by acquiring first information, a first video frame related to the first information is selected from the generated video frames. Second information is then acquired based on the first video frame to characterize the three-dimensional information of the first video frame. Using the second information as auxiliary information, a first video block associated with the first information is generated. The second information can reduce or eliminate the cumulative error generated during the video generation process, reduce the impact of the cumulative error on the first video block, improve the global consistency of the first video block for the virtual scene, thereby increasing the high-quality interaction time between the user and the virtual scene and improving the quality of the generated video. Attached Figure Description
[0011] The above and other features, advantages, and aspects will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0012] Figure 1 This is a scenario diagram illustrating an application scenario involving video generation using a video generation method in one particular case. Figure 2 This is a flowchart illustrating a video generation method in one specific scenario. Figure 3 This is a flowchart illustrating a video generation method in one specific scenario. Figure 4 This is a schematic diagram of the first model in one scenario; Figure 5 This is a schematic diagram of a video generation device provided in one scenario; Figure 6 This is a schematic diagram of the structure of an electronic device provided in one scenario. Detailed Implementation
[0013] The embodiments will now be described in more detail with reference to the accompanying drawings. While some embodiments are shown in the drawings, it should be understood that the technical solution can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present technical solution. It should be understood that the illustrated drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the technical solution.
[0014] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0015] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.
[0016] It should be noted that the concepts of "first" and "second" mentioned are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0019] It is understandable that before using the technical solutions disclosed in this article, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this article in an appropriate manner in accordance with relevant laws and regulations, and their authorization should be obtained.
[0020] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.
[0021] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0022] It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.
[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0024] Figure 1 This is a scenario illustration illustrating an application scenario involving video generation using a video generation method. In some cases, the provided solution can be applied to this application scenario. For example... Figure 1 As shown, the system involved in this application scenario may include client 101 and server 102. Client 101 may include, but is not limited to, web applications such as browsers, applications (Apps), Hyper Text Markup Language (HTML) applications, lightweight applications (also known as mini-programs, a type of lightweight application), or cloud applications. Client 101 may be deployed on an electronic device and relies on the operation of that device or certain applications on the device to implement its functions. The electronic device may be, for example, a device with a display screen that supports information browsing, such as a smartphone, tablet, personal computer, or other client terminal. For ease of understanding, Figure 1 The client is primarily represented by a device. Other applications can also be configured on the electronic device, such as e-shopping and entertainment applications, media content publishing applications, and conversational applications. Server 102 can be one or more servers providing various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; furthermore, it can be a server for a distributed system, a server integrating blockchain technology, a cloud server, or an intelligent cloud computing server or intelligent cloud host deployed with machine learning models, etc.
[0025] The task execution method described in this paper can be implemented through client 101 interacting with server 102, such as receiving or sending messages. For example, in this paper, client 101 can be used as an interactive terminal for video generation, which can be used to obtain interactive information. Server 102 is used to generate video, and can receive interactive information through client 101 and execute the video generation method to generate video.
[0026] In one scenario, the task execution method can be executed on client 101, and the corresponding video generation device can be deployed on client 101. During the execution of the video generation method, client 101 and server 102 can achieve data interaction and functional collaboration through network communication. It should be understood that... Figure 1 The number of clients and servers shown is for illustrative purposes only. Any number of clients and servers can be configured to meet specific implementation requirements.
[0027] Figure 2 This is a flowchart illustrating a video generation method for one scenario. One scenario is applicable to situations where auxiliary information is provided for the generation of subsequent video blocks by filtering related video frames from already generated video frames, thereby reducing accumulated errors. This method can be executed by a video generation device, which can be implemented in the form of software and / or hardware, or optionally by an electronic device, such as a mobile terminal, PC, or server.
[0028] like Figure 1 As shown, the method includes: S110. Obtain first information, the first information being used to characterize a first viewpoint, the first viewpoint being used to indicate the generation of a first video block.
[0029] S120. Obtain a first video frame, which is related to the first information. The first video frame is obtained from a second video frame, which is a generated video frame.
[0030] S130. Obtain second information, which is used to characterize the three-dimensional information of the first video frame.
[0031] S140. Generate the first video block based on the first information and the second information.
[0032] In one scenario, a video is generated based on an input image and interactive information using a first model. This first model can be a machine learning model, such as a neural network model, with video generation capabilities. The input image can be captured in real-time by a client, transmitted externally, or read from storage; this input image provides the basic information for generating the video. The interactive information is used to indicate changes in the viewpoint of the generated video.
[0033] Using the first model, multiple video blocks are generated sequentially based on the input image and interaction information. These video blocks form video data according to their temporal order. For the first video block, generation is based on the input image and interaction information. For subsequent video blocks, generation is based on the input image, interaction information, and previously generated video frames. The previously generated video frames provide auxiliary information to reduce accumulated errors. Each video block can include multiple video frames, and the interaction information used to generate different video blocks can be the same or different.
[0034] In one scenario, the first information can be understood as interactive information input by the user through a client, representing a first perspective used to generate the first video block. It is understood that the first video block may include multiple video frames, and correspondingly, the first information can be a sequence of multiple perspective information. Each perspective information is used to indicate the generation of a video frame, and multiple video frames can form the first video block. The multiple perspective information in the first information can be the same or different.
[0035] In one scenario, the first information includes trajectory information of a virtual camera, which comprises the pose information of the virtual camera at multiple moments. The virtual camera can be understood as a mathematical imaging model used to generate video, and its pose information characterizes the virtual camera's spatial position and orientation within the virtual scene. The pose information of the virtual camera at multiple moments forms the trajectory information, which can be understood as a sequence of pose information formed by the virtual camera at multiple moments. One pose information of the virtual camera corresponds to a first viewpoint, used to simulate a user's observation perspective. The trajectory information of the virtual camera includes continuously changing pose information, simulating the user's moving observation behavior within the virtual scene, used to instruct the first model to generate a first video block from the moving observation perspective.
[0036] In one scenario, the initial information can be input via a keyboard. Different keys on the keyboard represent different directions of change. The trajectory information of the virtual camera is obtained by collecting the user's key presses. A single key press generates the virtual camera's pose information, while multiple key presses generate multiple pose information sets. These multiple pose information sets form the virtual camera's trajectory information based on a temporal sequence. Specifically, the virtual camera's pose information at the initial moment is obtained. In response to pressing the first key on the keyboard, the initial pose information is transformed to obtain the pose information at the first moment. In response to pressing the second key on the keyboard, the first pose information is transformed to obtain the pose information at the second moment, and so on, obtaining the virtual camera's pose information at multiple moments to form the virtual camera's trajectory information, i.e., the initial information. It can be understood that different keys on the keyboard can correspond to different directions of change, such as forward, backward, left, and right. The keyboard can be a physical keyboard or a virtual keyboard displayed on a touchscreen.
[0037] In one scenario, the first information can be obtained through mouse input. By matching the mouse position information with the virtual camera's pose information according to the correspondence between the mouse position information and the virtual camera's pose information, the virtual camera's pose information at multiple moments can be obtained, thus yielding the virtual camera's trajectory information, i.e., the first information.
[0038] In one scenario, the first information can be input via a directional control, displayed on the client's interface. This directional control can include controls for directions such as up, down, left, and right. The system acquires the virtual camera's pose information at the initial moment. In response to a first trigger operation on the directional control, the initial pose information is transformed to obtain the pose information at the first moment. In response to a second trigger operation on the directional control, the first pose information is transformed to obtain the pose information at the second moment, and so on, obtaining the virtual camera's pose information at multiple moments to form the virtual camera's trajectory information, i.e., the first information.
[0039] In one scenario, the first information can be input via a control handle, which may include directional buttons such as up, down, left, and right. The method for generating the first information will not be elaborated here.
[0040] Understandably, the interactive information obtained through the client includes the overall trajectory information of the virtual camera from the initial moment to the current moment. Based on the number of video frames in each video block, the overall trajectory information is divided into multiple local trajectory information, and each local trajectory information is used to indicate the generation of a video block. The first information is one of the multiple local trajectory information, used to indicate the generation of the first video block.
[0041] It is understandable that the first video block is not the first video block.
[0042] To reduce accumulated errors, a first video frame is selected from the generated video frames to provide reference information for generating the first video block. The generated video frames are designated as the second video frame, which includes video frames from the generated video block.
[0043] The second video frame carries third information, which represents the second viewpoint of the second video frame. This second viewpoint can be understood as the pose information of the virtual camera.
[0044] The third information of the second video frame is compared with the first information. In response to the second viewpoint of the second video frame satisfying the first condition, the second video frame that satisfies the condition is taken as the first video frame. That is to say, the second viewpoint of the first video frame satisfies the first condition, which includes a set degree of viewpoint similarity.
[0045] In one scenario, for a second video frame, similarity data between third information and first information is obtained. If the similarity data is greater than a first similarity threshold, the second video frame is used as the first video frame. Here, the first condition includes the similarity data being greater than the first similarity threshold.
[0046] In one scenario, for the second video frame, the similarity data between the third information and the first information is obtained. The second video frames are then sorted from largest to smallest based on the similarity data. The first n second video frames, where n is a positive integer greater than 1, are used as the first video frame. Here, the first condition includes the first n similarity frames in the similarity data sorting.
[0047] In one scenario, the first information is a sequence of pose information from a virtual camera. Obtaining similarity data between the third information and the first information includes: obtaining similarity data between the third information and multiple pose information in the first information, and using the maximum similarity data as the similarity data between the third information and the first information.
[0048] In one scenario, the first information is a sequence of pose information from a virtual camera. Acquiring the first video frame includes: acquiring fourth information, where the fourth information is pose information at a specific moment in the first information; and acquiring the first video frame based on the fourth information. For example, the fourth information is pose information at a random moment in the first information; or, for instance, the fourth information is pose information at a predetermined moment in the first information, where the predetermined moment includes intermediate moments.
[0049] Obtaining the first video frame based on the fourth information includes: for the second video frame, obtaining similarity data between the third and fourth information; and in response to the similarity data being equal to a first similarity threshold, using the second video frame as the first video frame. Alternatively, the second video frames are sorted from largest to smallest based on the similarity data, and the first n second video frames are used as the first video frames according to the sorting.
[0050] In one scenario, the fourth information and the third information are the pose information of the virtual camera, which includes position and angle. Obtaining similarity data between the third information and the fourth information includes: obtaining a first distance and a first angle difference, where the first distance is the distance between the position in the third information and the position in the fourth information, and the first angle difference is the angle deviation between the angle in the third information and the angle in the fourth information; obtaining weighted data of the first distance and the first angle difference; and using the weighted data as the similarity data between the third information and the fourth information.
[0051] It is understandable that the method for obtaining the similarity data between the third information and the first information is similar to the method for obtaining the similarity data between the fourth information and the third information, and will not be repeated here.
[0052] By filtering the generated video frames to find the first video frames that are relevant to the first information, and ensuring that the number of frames in the first video frames is less than the number of frames in the generated video frames, interference from other video frames that are irrelevant to the first information is removed, thereby reducing the amount of information processed by the first model and avoiding interference from invalid information.
[0053] In one scenario, a first video frame is used as an auxiliary video frame for generating a first video block. Second information is obtained based on the first video frame and used as auxiliary information for generating the first video block. The first model generates the first video block based on the second and first information. Specifically, the first and second information can be input into the first model to obtain the first video block output by the first model.
[0054] The second information can be understood as the three-dimensional information representing the first video frame. Using the second information as auxiliary information, the first model is provided with at least one of the three-dimensional information of the objects themselves in the virtual scene, the three-dimensional information between objects, and the spatial three-dimensional information. This can reduce the problems of visual distortion, three-dimensional spatial geometric misalignment, and breakage of inter-frame temporal continuity caused by cumulative errors, such as object deformation, object drift, object disappearance, and spatial distortion.
[0055] In one scenario, the second information is obtained through a second model. Specifically, the first video frame is input into the second model to obtain the second information output by the second model. The second model can be a machine learning model, such as a neural network model, which has the function of extracting three-dimensional information.
[0056] In one technical solution, by acquiring first information, filtering out first video frames related to the first information from the generated video frames, acquiring second information based on the first video frames to characterize the three-dimensional information of the first video frames, and using the second information as auxiliary information, generating a first video block associated with the first information. The second information can reduce or eliminate the cumulative error generated during the video generation process, reduce the impact of the cumulative error on the first video block, improve the global consistency of the first video block for the virtual scene, thereby increasing the high-quality interaction time between the user and the virtual scene and improving the quality of the generated video.
[0057] In one scenario, to improve the accuracy of the second information, a secondary filtering process is performed on the second video frame. After the first video frame is obtained by filtering the second video frame using the first information in the first stage, a second stage filtering process is performed on the first video frame to obtain the third video frame. The second information is then obtained based on the third video frame.
[0058] Specifically, a third video frame is obtained from the first video frame; the second information is obtained based on the third video frame; wherein the third video frame includes at least one of the following: a fourth video frame that satisfies a second condition, the second condition being used to indicate a set time condition; a fifth video frame that satisfies a third condition, the third condition being used to indicate a set spatial relationship; and a sixth video frame that satisfies a fourth condition, the fourth condition being used to indicate a similarity relationship with the input image, the input image being used to generate the first video block.
[0059] The second condition is used to filter fourth video frames that meet a set time condition. For example, the set time condition could be a timestamp earlier than a specified timestamp. The timestamps of the first video frames are compared with the set timestamps in the second condition. If a first video frame's timestamp is earlier than the specified timestamp, then that first video frame is selected as the fourth video frame. Another example is that the set time condition could be the top m timestamps. Specifically, the first video frames are sorted from smallest to largest timestamp, and the top m first video frames are selected as the fourth video frames. It can be understood that, regarding the characteristics of cumulative error, the larger the timestamp of the first video frame, the larger the cumulative error in the first video frame; conversely, the smaller the timestamp of the first video frame, the smaller the cumulative error in the first video frame. The fourth video frames selected through the second condition have the smallest cumulative error, providing more accurate 3D information.
[0060] The third condition is used to filter the fifth video frame that satisfies the defined spatial relationship. The defined spatial relationship can include the same or similar viewpoints, that is, the same or similar poses of the virtual cameras. The process involves obtaining the third information of the first video frame, obtaining the similarity data between the third and fourth information, and selecting the fifth video frame whose similarity data is greater than the second similarity threshold. Alternatively, the first video frames can be sorted based on the similarity data, and the top p first video frames can be selected as the fifth video frames. It is understandable that for two video frames that satisfy the defined spatial relationship, the objects in the two video frames have a high degree of overlap. Filtering the fifth video frame using the third condition ensures that this fifth video frame has a high degree of overlap with the objects in the first video frame, providing more accurate 3D information.
[0061] The fourth condition is used to filter the sixth video frame that has a similarity relationship with the input image. The input image is provided by the user and is the initial image used for video generation; the input image has no cumulative error. Similarity data between the first video frame and the input image is obtained, and this similarity data can be calculated using cosine similarity. The higher the similarity data between the first video frame and the input image, the closer the virtual scene of the first video frame is to the scene of the input image, and the smaller the cumulative error.
[0062] For example, the sixth video frame whose similarity data is greater than the third similarity threshold can be selected. Alternatively, the first q first video frames can be sorted based on similarity data and selected as the sixth video frame.
[0063] The third video frame may include at least one of the fourth, fifth, and sixth video frames. Second information is obtained based on the third video frame; specifically, the second information is obtained based on the third video frame using a second model.
[0064] One technical solution involves performing a secondary filtering process on already generated video frames to obtain a third video frame. Secondary information is then obtained from this third video frame. This second information is used as auxiliary information to generate a first video block associated with the first information. The second information can reduce or eliminate accumulated errors generated during video generation, thus minimizing the impact of these accumulated errors on the first video block. Specifically, the secondary filtering of the generated video frames improves the accuracy of the second information, thereby providing more accurate three-dimensional information for generating the first video block and further enhancing its quality.
[0065] Figure 3 This is a flowchart illustrating a video generation method provided under one specific scenario. This optional technical solution can be combined with other implementation scenarios; for identical or related parts, descriptions of other implementation scenarios can be used, and will not be repeated here. For example... Figure 3 As shown, the video generation method includes the following steps: S210. Obtain first information, the first information being used to characterize a first viewpoint, the first viewpoint being used to indicate the generation of a first video block.
[0066] S220. Obtain a first video frame, which is related to the first information. The first video frame is obtained from a second video frame, which is a generated video frame.
[0067] S230. Obtain second information, the second information including first point cloud data and / or first image, the first point cloud data representing the three-dimensional information of the first video frame in the time dimension; the first image representing the three-dimensional information of the first video frame in the spatial dimension.
[0068] S240. Generate the first video block based on the first information and the second information.
[0069] In one scenario, the second information includes three-dimensional information provided by the first video frame in the temporal and / or spatial dimensions. Specifically, the first point cloud data represents the point cloud data constructed by the first video frame for the first temporal information, where the first temporal information is the start time information of the first video block, which is also the time information of the last generated video frame.
[0070] In one scenario, acquiring first point cloud data includes: acquiring second point cloud data, the second point cloud data including point cloud data constructed from the first video frame; acquiring third point cloud data, the third point cloud data including point cloud data indicated by a seventh video frame, the seventh video frame being a generated video frame, the seventh video frame being sequentially connected to the first video block; and filtering the second point cloud data to obtain the first point cloud data, the first point cloud data including at least a portion of the second point cloud data, the first point cloud data including second point cloud data located in the neighborhood of the third point cloud data.
[0071] The second point cloud data can be understood as the global point cloud data of the virtual scene constructed from the first video frame. For example, the second point cloud data can be constructed based on the first video frame using a third model. This third model can be a machine learning model, such as a neural network model. Alternatively, the second point cloud data can be constructed based on the first video frame using a point cloud data construction algorithm.
[0072] The seventh video frame is the video frame that is sequentially connected to the first video block among the generated video frames; that is, the seventh video frame is the last generated video frame at the current moment. Third point cloud data is constructed based on the seventh video frame. For example, it can be generated using a third model based on the seventh video frame, or it can be constructed using a point cloud data construction algorithm. The third point cloud data can be point cloud data of a local virtual scene, which is the local virtual scene shown in the seventh video frame.
[0073] Based on the third point cloud data, the second point cloud data is filtered to obtain the first point cloud data. Optionally, the third point cloud data is traversed, and for a point cloud data in the third point cloud data, the neighborhood range of that point cloud data is obtained, and the second point cloud data within that neighborhood range is used as the first point cloud data. The third point cloud data includes multiple point cloud data, and each of the multiple point cloud data corresponds to a second point cloud data within its neighborhood range, forming the first point cloud data.
[0074] In one scenario, the second and third point cloud data reside in the same spatial coordinate system, which is divided into voxels. Accordingly, the second and third point cloud data fall into voxels respectively. For a point cloud data point within the third point cloud data, the first voxel to which it belongs is determined, and the second voxel is obtained. The second voxel is a neighboring voxel of the first voxel. The second point cloud data within both the first and second voxels is then used as the first point cloud data. The second voxel lies within the range [-1, 0, 1] centered on the first voxel.
[0075] The seventh video frame, serving as the preceding video frame of the first video block, obtains the first point cloud data by filtering the second point cloud data within the neighborhood of the third point cloud data. This first point cloud data provides the 3D information corresponding to the start time information of the first video block. Furthermore, the first point cloud data is obtained by filtering the second point cloud data, which is global point cloud data constructed using the first video frame. This second point cloud data has a small cumulative error, and consequently, the first point cloud data also has a small cumulative error, providing high-quality 3D information for generating the first video block. Additionally, the first point cloud data is the second point cloud data within the neighborhood of the third point cloud data, expanding upon the third point cloud data to enrich the amount of 3D information provided. In summary, the first, second, and third point cloud data are all 3D point cloud data.
[0076] The first image is used to characterize the projection information of the first video frame under the first viewpoint, and serves as a priori image under the first viewpoint.
[0077] In one scenario, acquiring the first image includes: acquiring second point cloud data, the second point cloud data including point cloud data constructed from the first video frame; and acquiring the first image, the first image indicating the projection information of the second point cloud data relative to the first viewpoint. The method for acquiring the second point cloud data will not be described in detail here.
[0078] The first information includes the trajectory information of the virtual camera, that is, multiple pose information of the virtual camera within the time period displayed in the first video block. One pose information of the virtual camera corresponds to one first viewpoint. The second point cloud data is projected onto the multiple pose information of the virtual camera to obtain multiple first images. Taking the first pose information as an example, the first pose information is one pose information in the first information. The virtual camera under the first pose information is used as the observation viewpoint, and the second point cloud data is projected onto a two-dimensional plane to obtain the first image under the first pose information.
[0079] In one scenario, the first point cloud data represents the three-dimensional information of the third video frame in the time dimension; the first image is used to represent the three-dimensional information of the third video frame in the spatial dimension.
[0080] Accordingly, acquiring the first point cloud data includes: acquiring second point cloud data, the second point cloud data including the point cloud data constructed by the third video frame; acquiring third point cloud data, the third point cloud data including the point cloud data indicated by the seventh video frame, the seventh video frame being a generated video frame, the seventh video frame being sequentially connected to the first video block; and filtering the second point cloud data to obtain the first point cloud data, the first point cloud data including at least a portion of the second point cloud data, the first point cloud data including the second point cloud data located in the neighborhood of the third point cloud data.
[0081] Acquiring a first image includes: acquiring second point cloud data, the second point cloud data including point cloud data constructed from the third video frame; and acquiring a first image, the first image indicating the projection information of the second point cloud data onto the first viewpoint.
[0082] In one scenario, at least one of the first point cloud data and the first image is used as second information, and a first video block is generated based on the second information and the first information.
[0083] In one scenario, a technical solution is provided to acquire second information through a first video frame / third video frame. The second information includes first point cloud data and / or a first image. The second image provides three-dimensional information of the first video frame / third video frame in the time and / or spatial dimensions. A first video block is generated using the first and second information. The first information provides viewpoint information for generating the first video block, and the second information provides three-dimensional information for generating the first video block, thereby reducing the impact of accumulated errors on the generation of the first video block and improving the global consistency of the first video block for the virtual scene.
[0084] In one scenario, the video generation method includes: acquiring interactive information, which includes multiple trajectory information, which can be obtained by dividing the video block according to the number of frames. Each trajectory information is used to indicate the generation of a video block. The interactive information includes first information, which is used to indicate the generation of a first video block. The first information includes the trajectory information of a virtual camera, i.e., multiple pose information of the virtual camera. The method then acquires fourth information, i.e., the pose information of the virtual camera at an intermediate time point in the first information. Based on the fourth information, a secondary filtering is performed on the generated video frames: first, a first video frame satisfying a first condition, which includes a set viewpoint similarity; and then, a third video frame satisfying at least one of the second, third, and fourth conditions from the first video frame. Second point cloud data is reconstructed from the third video frame. A seventh video frame, i.e., the last video frame generated at the current time point, is acquired. Third point cloud data is reconstructed from the seventh video frame. Based on the third point cloud data, first point cloud data is filtered from the second point cloud data. The first point cloud data includes second point cloud data located within the neighborhood of the third point cloud data. The second point cloud data is projected onto multiple pose information of the first information to obtain the first image. The first image and the first point cloud data are used as the second information to provide 3D information for generating the first video block. The first information, the second information, the seventh video frame, and random noise are input into the first model, and the first video block is output through the first model. The viewpoint changes of the first video block are consistent with the first viewpoint provided by the first information.
[0085] Figure 4This is a schematic diagram of the structure of a first model provided in one scenario. The first model includes a first module and a second module. The first model is used to acquire 3D information provided by first point cloud data. The 3D information provided by the first point cloud data includes geometric constraint information. The first module includes a point encoder and a point adapter. The point encoder is used to encode the first point cloud data into a value feature space to obtain geometric constraint information, which can be in the form of a 3D geometric feature vector. The point adapter is used to convert the format of the 3D information output by the point encoder, that is, to convert it into a feature format compatible with the second model for input into the second model. The second module includes multiple serial first units. The first unit includes a self-attention unit, a point attention unit, a cross attention unit, and a feedforward network unit connected in sequence. The output of the first module is connected to the input of the point attention unit in the first unit. First information, a first image, a seventh video frame, and random noise are used as input information of the second module, and are responded to by the multiple serial first units respectively. The point attention unit of the first unit receives the geometric constraint information input by the first module to reduce the influence of accumulated error.
[0086] In one scenario, the first model further includes a third module. This third module includes a variational autoencoder region, and its output is connected to the input of the second module. It is used to respond to the input first information, the first image, the seventh video frame, and random noise to obtain image description information. The response process of the third module includes: compressing and encoding the image features after splicing random noise, mapping high-dimensional pixel information to a low-dimensional continuous latent space, and outputting a standardized latent representation adapted to the second module.
[0087] The self-attention unit in the second module is used to obtain the video's own spatiotemporal dependency information in response to the input information (such as the image description information output by the third module, or the output information of the previous first unit). This spatiotemporal dependency information contains accumulated errors. The point attention unit is used to align the geometric constraint information output by the first model with the video's own spatiotemporal dependency information, introducing geometric constraint information to suppress accumulated errors. The cross-attention unit combines the first information to generate video frames with changing viewpoints. The feedforward network unit performs nonlinear transformations to improve the expressiveness of the video frames and enrich texture details.
[0088] In one scenario, a first model is provided, which introduces geometric constraints during video generation by acquiring three-dimensional information from first point cloud data, thereby suppressing cumulative errors and improving the global consistency and video quality of the generated video blocks.
[0089] Figure 5 This is a schematic diagram of the structure of a video generation device provided in one scenario, such as... Figure 5As shown, the device includes: a first information acquisition module 310, a first video frame acquisition module 320, a second information acquisition module 330, and a video generation module 340.
[0090] The first information acquisition module 310 is used to acquire first information, which is used to characterize a first viewpoint and the first viewpoint is used to indicate the generation of a first video block. The first video frame acquisition module 320 is used to acquire a first video frame, which is related to the first information. The first video frame is acquired from a second video frame, which is a generated video frame. The second information acquisition module 330 is used to acquire second information, which is used to characterize the three-dimensional information of the first video frame. The video generation module 340 is used to generate the first video block based on the first information and the second information.
[0091] In one scenario, the provided technical solution involves acquiring first information, filtering out first video frames related to the first information from the generated video frames, acquiring second information based on the first video frames to characterize the three-dimensional information of the first video frames, and generating a first video block associated with the first information using the second information as auxiliary information. The second information can reduce or eliminate the cumulative error generated during the video generation process, reduce the impact of the cumulative error on the first video block, improve the global consistency of the first video block for the virtual scene, thereby increasing the high-quality interaction time between the user and the virtual scene and improving the quality of the generated video.
[0092] In one scenario, the second video frame carries third information, which represents a second viewpoint of the second video frame; the second viewpoint of the first video frame satisfies a first condition with the first information, the first condition including a set degree of viewpoint similarity.
[0093] In one scenario, the first information includes trajectory information of a virtual camera, the trajectory information including pose information of the virtual camera at multiple moments, and the pose information of the virtual camera at one moment representing a viewpoint; The first video frame acquisition module 320 is further configured to: acquire fourth information, wherein the fourth information is pose information at a certain moment in the first information; and acquire the first video frame based on the fourth information.
[0094] In one scenario, the second information acquisition module 330 is further configured to acquire a third video frame, which is acquired from the first video frame; and acquire the second information based on the third video frame. The third video frame includes at least one of the following: The fourth video frame that satisfies the second condition, which is used to indicate a set time condition; The fifth video frame that satisfies the third condition, which is used to indicate the established spatial relationship; The sixth video frame satisfies the fourth condition, which indicates the similarity relationship with the input image used to generate the first video block.
[0095] In one scenario, the second information includes first point cloud data and / or a first image.
[0096] In one scenario, the second information acquisition module 330 is further configured to acquire second point cloud data, the second point cloud data including point cloud data constructed from the first video frame; acquire third point cloud data, the third point cloud data including point cloud data indicated by a seventh video frame, the seventh video frame being a generated video frame, the seventh video frame being sequentially connected to the first video block; and filter the second point cloud data to obtain the first point cloud data, the first point cloud data including at least a portion of the second point cloud data, the first point cloud data including second point cloud data located in the neighborhood of the third point cloud data.
[0097] In one scenario, the second information acquisition module 330 is further configured to acquire second point cloud data, the second point cloud data including point cloud data constructed from the first video frame; and acquire a first image, the first image indicating the projection information of the second point cloud data onto the first viewpoint.
[0098] In one scenario, the provided video generation apparatus can execute the video generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0099] It is worth noting that the various units and modules included in the above-mentioned device are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection in a certain situation.
[0100] Figure 6 This is a schematic diagram of the structure of an electronic device provided in one scenario. See below for reference. Figure 6 It shows an electronic device suitable for implementing a particular situation (e.g.) Figure 6The diagram shows the structure of the terminal device or server 500. In one case, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use in any particular situation.
[0101] like Figure 6 As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0102] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0103] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 508, or installed from ROM 502. When the computer program is executed by processing device 501, it performs the functions defined in one scenario of the method.
[0104] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0105] In one scenario, the electronic device provided is based on the same inventive concept as the video generation method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0106] In one scenario, a computer storage medium is provided that stores a computer program, which, when executed by a processor, implements the video generation method provided in the above embodiments.
[0107] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0108] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0109] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0110] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire first information, the first information being used to characterize a first viewpoint, the first viewpoint being used to instruct the generation of a first video block; acquire a first video frame, the first video frame being related to the first information, the first video frame being acquired from a second video frame, the second video frame being a generated video frame; acquire second information, the second information being used to characterize the three-dimensional information of the first video frame; and generate the first video block based on the first information and the second information.
[0111] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0113] The unit described in a particular scenario can be implemented in software or hardware. The name of the unit may not, in some cases, constitute a limitation on the unit itself; for example, the first acquisition unit could also be described as "a unit that acquires at least two Internet Protocol addresses".
[0114] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0115] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0117] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0118] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video generation method, comprising: Obtain first information, which is used to characterize a first viewpoint, and the first viewpoint is used to indicate the generation of a first video block; A first video frame is obtained, which is related to the first information. The first video frame is obtained from a second video frame, which is a generated video frame. Obtain second information, which is used to characterize the three-dimensional information of the first video frame; The first video block is generated based on the first information and the second information.
2. The method according to claim 1, wherein the second video frame carries third information, the third information representing a second viewpoint of the second video frame; The second perspective of the first video frame satisfies the first condition with the first information, and the first condition includes a set degree of similarity in perspective.
3. The method according to claim 1 or 2, wherein the first information includes trajectory information of a virtual camera, the trajectory information includes pose information of the virtual camera at multiple times, and the pose information of the virtual camera at one time represents a viewpoint; The method further includes: Obtain fourth information, which is the pose information at a certain moment in the first information; The first video frame is obtained based on the fourth information.
4. The method according to claim 2, further comprising: Obtain a third video frame, which is obtained from the first video frame; The second information is obtained based on the third video frame; The third video frame includes at least one of the following: The fourth video frame that satisfies the second condition, which is used to indicate a set time condition; The fifth video frame that satisfies the third condition, which is used to indicate the established spatial relationship; The sixth video frame satisfies the fourth condition, which indicates the similarity relationship with the input image used to generate the first video block.
5. The method according to claim 1, wherein the second information includes first point cloud data and / or a first image, the first point cloud data representing the three-dimensional information of the first video frame in the time dimension; and the first image representing the three-dimensional information of the first video frame in the spatial dimension.
6. The method according to claim 5, wherein the acquisition method of the first point cloud data includes: Acquire second point cloud data, which includes point cloud data constructed from the first video frame; Acquire third point cloud data, which includes point cloud data indicated by the seventh video frame. The seventh video frame is an already generated video frame and is sequentially connected to the first video block. The first point cloud data is obtained by filtering the second point cloud data. The first point cloud data includes at least a portion of the second point cloud data and includes the second point cloud data located in the neighborhood of the third point cloud data.
7. The method according to claim 5, wherein the acquisition method of the first image includes: Acquire second point cloud data, which includes point cloud data constructed from the first video frame; Acquire a first image, which indicates the projection information of the second point cloud data onto the first viewpoint.
8. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any one of claims 1-7.
9. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video generation method as described in any one of claims 1-7.
10. A computer program product comprising a computer program that, when executed by a processor, implements the video generation method as described in any one of claims 1-7.