Three-dimensional model generation method and three-dimensional model generation device
By automatically generating and optimizing the fusion of 3D dynamic world models, the problems of low generation efficiency and high cost in existing technologies have been solved, enabling ordinary users to design efficient 3D dynamic scenes, improving generation efficiency and reducing costs.
Patent Information
- Application Number
- CN202510831801.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies require manual design and professional software modeling to create 3D dynamic scenes, resulting in low generation efficiency, high costs, and high requirements for professional knowledge, making it difficult to meet the needs of ordinary users.
By acquiring user-input prompts, an initial 3D dynamic world model is automatically generated, and then optimized and integrated based on supplementary video, ultimately generating a high-quality 3D dynamic scene, reducing professional requirements and simplifying the design process.
It improves the efficiency of generating 3D dynamic scenes and reduces costs, enabling ordinary users to design high-quality 3D dynamic scenes and enhancing the user experience.
Smart Images

Figure CN120976408A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a three-dimensional model generation method and a three-dimensional model generation device. BACKGROUND
[0002] Currently, when a three-dimensional (3-dimensional, 3D) dynamic scene is made, not only scene design needs to be performed manually, but also modeling needs to be performed by professional personnel through modeling software. In order to ensure that the dynamic effect is both beautiful and real, the entire production process needs to be constantly modified. The entire production process of the three-dimensional dynamic scene not only consumes time and effort, but also has a very long production cycle, thereby resulting in very low generation efficiency and very high generation cost, and requiring a large amount of time cost and manpower cost. SUMMARY
[0003] The present application provides a three-dimensional model generation method and a three-dimensional model generation device, which not only help to improve the generation efficiency of the three-dimensional dynamic scene and reduce the generation cost of the three-dimensional dynamic scene, but also help to ensure the presentation effect of the three-dimensional dynamic scene.
[0004] In a first aspect, a three-dimensional model generation method is provided. The method includes: obtaining prompt information input by a user, wherein the prompt information is used to indicate at least one element and an action performed by a target element in the at least one element; generating an initial video according to the prompt information, and generating an initial three-dimensional dynamic world model according to the initial video, wherein the initial three-dimensional dynamic world model includes a three-dimensional model of the at least one element and a three-dimensional model of the action performed by the target element; obtaining at least one supplementary video according to initial information, wherein the initial information includes at least one of the initial three-dimensional dynamic world model or the initial video; generating at least one supplementary three-dimensional dynamic world model according to the at least one supplementary video, wherein the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model are fused to obtain a three-dimensional dynamic world model.
[0005] In the scheme, the three-dimensional model generation apparatus can automatically generate an initial three-dimensional dynamic world model according to the prompt information input by the user, the initial three-dimensional dynamic world model including a three-dimensional model of at least one element and a three-dimensional model of the target element performing an action, so that the three-dimensional dynamic scene indicated by the prompt information can be obtained by rendering the initial three-dimensional dynamic world model, thereby simplifying the design process of the three-dimensional dynamic scene and avoiding a complex modeling process. In this way, not only the generation efficiency of the three-dimensional dynamic scene is improved, but also the generation cost of the three-dimensional dynamic scene is reduced. In addition, the three-dimensional model generation apparatus can also generate a supplementary three-dimensional dynamic world model according to the target initial information, and fuse the supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model to optimize the initial three-dimensional dynamic world model, thereby improving the quality of the target three-dimensional dynamic world model obtained by fusion, and further improving the presentation effect of the three-dimensional dynamic scene obtained by rendering the target three-dimensional dynamic world model. In addition, since the initial three-dimensional dynamic world model is optimized by the automatically generated supplementary three-dimensional dynamic world model, the speed of modifying the initial three-dimensional dynamic world model is improved, and the difficulty of modifying the three-dimensional dynamic world model is reduced. In addition, since the entire generation process reduces the professional requirements for the user, ordinary people can also design the three-dimensional dynamic scene, thereby improving the user experience.
[0006] In another possible implementation, according to the target initial information, at least one supplementary video is obtained, including: providing a first page, the first page being used to display a picture obtained by rendering the initial three-dimensional dynamic world model; and in response to a first instruction input by the user, obtaining at least one supplementary video according to the target initial information, wherein the first instruction is used to confirm any one of the initial three-dimensional dynamic world model or the supplementary initial three-dimensional dynamic world model.
[0007] In this way, by providing the first page, the automatically generated three-dimensional dynamic scene is displayed for the user, thereby helping the user to confirm whether the initial three-dimensional dynamic world model is accurate. By starting to obtain the supplementary video again in the case that the first instruction input by the user is received, the supplementary video used to optimize the three-dimensional dynamic scene is obtained again in the case that the user confirms that the content of the three-dimensional dynamic scene is correct, thereby helping to ensure the necessity of the optimization process.
[0008] In another possible implementation, the method further includes: providing a second page, wherein the second page is used to display a picture obtained by rendering the target three-dimensional dynamic world model.
[0009] In this implementation, by providing the second page, the picture obtained by rendering the target three-dimensional dynamic world model is displayed for the user, thereby helping the user to determine whether the final three-dimensional dynamic scene meets the requirements.
[0010] In a possible implementation, the initial video is generated according to the prompt information, and the initial three-dimensional dynamic world model is generated according to the initial video, including: inputting the prompt information into a video generation model to obtain an initial video output by the video generation model; providing a third page, where the third page is configured to display the initial video; and in response to a second instruction input by a user, inputting the initial video into a dynamic generation model to obtain an initial three-dimensional dynamic world model output by the dynamic generation model, where the second instruction is configured to indicate confirmation of the initial video or to indicate generation of the three-dimensional dynamic world model.
[0011] In this implementation, the initial video is displayed for the user by providing the third page, which helps the user to confirm whether the content of the initial video is correct. The initial three-dimensional dynamic world model is generated in the case where the second instruction input by the user is received, which helps to improve the accuracy of the initial three-dimensional dynamic world model.
[0012] In a possible implementation, the at least one supplementary video is obtained according to the target initial information, including: providing a fourth page, where the fourth page is configured to display a picture of at least one view obtained by rendering the initial three-dimensional dynamic world model; obtaining the picture of the at least one view, and rendering the picture of the at least one view to obtain the at least one supplementary video.
[0013] In this implementation, the supplementary video is obtained by rendering the picture of each view obtained by the initial three-dimensional dynamic world model, which helps to obtain the picture with a hole in the three-dimensional dynamic scene, thereby helping to optimize the hole in the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, and further helping to improve the presentation effect of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model.
[0014] In a possible implementation, the at least one supplementary video is obtained according to the target initial information, including: inputting a camera pose of the initial video and an initial picture rendered by the camera pose of the initial video into a neural network model to obtain at least one supplementary camera pose and at least one supplementary picture rendered by the at least one supplementary camera pose, which are output by the neural network model; and generating the at least one supplementary video according to the at least one supplementary camera pose and the at least one supplementary picture.
[0015] In this implementation, the subsequent camera pose of the camera pose of the initial video is predicted by the neural network model, which helps to obtain more supplementary videos of the camera pose, thereby not only helping to enrich the camera pose of the three-dimensional dynamic scene, but also helping to optimize the hole in the three-dimensional dynamic scene by the new camera pose.
[0016] In another possible implementation manner, the at least one supplementary three-dimensional dynamic world model is generated according to the at least one supplementary video, including: processing a hole in a first supplementary video in the at least one supplementary video to obtain a target first supplementary video; inputting the target first supplementary video into the dynamic generation model to obtain a first supplementary three-dimensional dynamic world model in at least one supplementary three-dimensional dynamic world model output by the dynamic generation model.
[0017] In this implementation manner, by processing the hole in the supplementary video, the hole in the supplementary video is reduced or eliminated, and the hole in the scene of the supplementary rendered three-dimensional dynamic world model is reduced or eliminated, so that after the supplementary three-dimensional dynamic world model is fused with the initial three-dimensional dynamic world model, the hole in the target three-dimensional dynamic world model obtained by fusing is reduced or eliminated.
[0018] In another possible implementation manner, the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model are fused to obtain a target three-dimensional dynamic world model, including: inputting the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model into a fusion model to obtain the target three-dimensional dynamic world model.
[0019] In another possible implementation manner, the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model are fused to obtain a target three-dimensional dynamic world model, including: placing the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model in a world coordinate system to obtain a first three-dimensional dynamic world model, and the first three-dimensional dynamic world model is the target three-dimensional dynamic world model.
[0020] In this implementation manner, by taking the first three-dimensional dynamic world model in the world coordinate system as the target three-dimensional dynamic world model, the efficiency of obtaining the three-dimensional dynamic world model can be improved.
[0021] In another possible implementation manner, the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model are fused to obtain a target three-dimensional dynamic world model, including: placing the target three-dimensional dynamic world model and the three-dimensional dynamic world model in a world coordinate system to obtain a first three-dimensional dynamic world model; projecting the first three-dimensional dynamic world model to obtain at least one projection video of the first three-dimensional dynamic world model; optimizing the first three-dimensional dynamic world model according to a difference between the at least one projection video and at least one target video to obtain the target three-dimensional dynamic world model, and the camera pose of the at least one projection video is the same as the camera pose of the at least one target video, and the at least one target video includes at least one of the at least one supplementary video and the initial video.
[0022] In the implementation, the first three-dimensional dynamic world model is optimized through the difference between the at least one projection video and the at least one target video, which helps to ensure the consistency of the target three-dimensional dynamic world model with the content of the target video used to generate the initial three-dimensional dynamic world model, thereby helping to improve the accuracy of the target three-dimensional dynamic world model after fusion.
[0023] In another possible implementation, the method further includes: processing a hole in the first projection video in the at least one projection video to obtain a target first projection video; inputting the target first projection video into the dynamic generation model to obtain at least one second three-dimensional dynamic world model output by the dynamic generation model; and fusing the second three-dimensional dynamic world model and the first three-dimensional dynamic world model to obtain the target three-dimensional dynamic world model.
[0024] In the implementation, the hole in the projection video is processed to obtain the target projection video, which helps to reduce or eliminate the hole in the target projection video. Then, the second three-dimensional dynamic world model is generated according to the target projection video, which helps to reduce or eliminate the hole in the second three-dimensional dynamic world model, thereby helping to reduce the hole in the target three-dimensional dynamic world model after fusion, and further helping to improve the integrity of the target three-dimensional dynamic world model and the picture effect of the target three-dimensional dynamic world model.
[0025] In another possible implementation, the method further includes: obtaining a plurality of candidate videos according to a plurality of supplementary camera poses and a plurality of supplementary pictures; determining a target candidate video from the plurality of candidate videos, the hole in the target candidate video being greater than that in at least part of the candidate videos other than the target candidate video; providing a fifth page, the fifth page being used to render a picture of a target view angle obtained by the initial three-dimensional dynamic world model, the picture of the target view angle being a picture of the target candidate video; adjusting the view angle of the picture obtained by rendering the initial three-dimensional dynamic world model, and obtaining at least one picture of a supplementary view angle; and rendering the at least one picture of the supplementary view angle to obtain at least one supplementary video.
[0026] In the implementation, the target candidate video is obtained, and the view angle of the picture obtained by rendering the initial three-dimensional dynamic world model is adjusted from the starting point of the picture of the target view angle corresponding to the target candidate video, thereby obtaining at least one supplementary candidate video. This helps to obtain a video of a view angle with a hole from the initial three-dimensional dynamic world model, and further helps to optimize the hole in the initial three-dimensional dynamic world model.
[0027] In another possible implementation manner, the method further includes: constructing a distribution field of a space body of the initial three-dimensional dynamic world model, the space body including a plurality of three-dimensional points of the initial three-dimensional dynamic world model, and the distribution field being used to indicate a distribution of the plurality of three-dimensional points; determining a density of the three-dimensional points in a plurality of directions centered on a position of the camera pose of the initial video in the distribution field; obtaining the target camera pose according to a target direction and the camera pose of the initial video, the density of the three-dimensional points in the target direction being less than a density of the three-dimensional points in at least part of the directions other than the target direction in the plurality of directions; inputting the target camera pose and an initial picture rendered by the target camera pose into the neural network model to obtain at least one supplementary camera pose output by the neural network model and at least one supplementary picture rendered by the at least one supplementary camera pose; and generating at least one supplementary video according to the at least one supplementary camera pose and the at least one supplementary picture.
[0028] In this implementation manner, the target camera pose is obtained through the initial three-dimensional dynamic world model, and the supplementary video is obtained based on the target camera pose. Since the target camera pose is a camera pose in a direction in which a hole in the initial three-dimensional dynamic world model is relatively large, when other camera poses are predicted based on the target camera pose through the neural network model, the camera pose in the direction in which the hole in the initial three-dimensional dynamic world model is relatively large is predicted, so that the hole in the initial three-dimensional dynamic world model is repaired, and the presentation effect of the initial three-dimensional dynamic world model is improved.
[0029] In another possible implementation manner, the method further includes: providing a sixth page, the sixth page being used to display a picture rendered by the target camera pose of the initial three-dimensional dynamic world model; adjusting a viewing angle of the picture rendered by the initial three-dimensional dynamic world model, and obtaining a picture of at least one supplementary viewing angle; and rendering the picture of the at least one supplementary viewing angle to obtain at least one supplementary video.
[0030] In this implementation manner, the target candidate video is obtained, and the viewing angle of the picture rendered by the initial three-dimensional dynamic world model is adjusted based on a starting point of the picture corresponding to the target viewing angle of the target candidate video, so that at least one supplementary candidate video is obtained. In this way, the video of the viewing angle with the hole in the initial three-dimensional dynamic world model is obtained, and the hole in the initial three-dimensional dynamic world model is optimized.
[0031] In another possible implementation manner, the initial video is generated according to the prompt information, and the initial three-dimensional dynamic world model is generated according to the initial video, including: inputting the prompt information into a video generation model to obtain the initial video output by the video generation model; and inputting the initial video into a dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model.
[0032] In another possible implementation, the initial video is input into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model, including: performing target processing on the initial video to obtain auxiliary information of the initial video, the target processing including at least one of video depth estimation or dynamic information estimation; and inputting the initial video and the auxiliary information of the initial video into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model.
[0033] In this implementation, the initial video is processed to obtain auxiliary information, and then the initial three-dimensional dynamic world model is generated based on the auxiliary information and the initial video, which helps to improve the comprehensiveness of information used to generate the initial three-dimensional dynamic world model, and thus helps to improve the picture quality and picture accuracy of the generated initial three-dimensional dynamic world model.
[0034] In another possible implementation, the initial video is input into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model, including: performing processing on a hole of the initial video to obtain a target initial video; and inputting the target initial video into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model.
[0035] In this implementation, the hole of the initial video is processed to obtain a target initial video, which helps to reduce or eliminate the hole on the target initial video, and then the initial three-dimensional dynamic world model is generated based on the target initial video, which helps to reduce or eliminate the hole on the initial three-dimensional dynamic world model, and thus helps to improve the integrity of the initial three-dimensional dynamic world model, and further helps to improve the picture effect of the initial three-dimensional dynamic world model.
[0036] In another possible implementation, the initial video is input into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model, including: performing processing on a hole of the initial video to obtain a target initial video; performing target processing on the target initial video to obtain auxiliary information of the target initial video; and inputting the target initial video and the auxiliary information of the target initial video into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model.
[0037] In another possible implementation, at least one supplementary three-dimensional dynamic world model is generated based on at least one supplementary video, including: performing target processing on a first supplementary video in the at least one supplementary video to obtain auxiliary information of the first supplementary video, the target processing including at least one of video depth estimation or dynamic information estimation; and inputting the first supplementary video and the auxiliary information of the first supplementary video into the dynamic generation model to obtain a first supplementary three-dimensional dynamic world model in the at least one supplementary three-dimensional dynamic world model output by the dynamic generation model.
[0038] In the implementation, the target processing is performed on the supplementary video to obtain the auxiliary information, and the three-dimensional dynamic world model is generated by using the auxiliary information and the supplementary video, so that the comprehensiveness of the information used to generate the supplementary three-dimensional dynamic world model is improved, the picture quality and picture accuracy of the generated supplementary rendered three-dimensional dynamic world model are improved, and the picture quality and picture accuracy of the target three-dimensional dynamic world model after fusion are improved.
[0039] In another possible implementation, the generating the at least one supplementary three-dimensional dynamic world model according to the at least one supplementary video includes: performing target processing on a hole in a first supplementary video in the at least one supplementary video to obtain a target first supplementary video; performing target processing on the target first supplementary video to obtain auxiliary information of the target first supplementary video; and inputting the target first supplementary video and the auxiliary information of the target first supplementary video into a dynamic generation model to obtain the first supplementary three-dimensional dynamic world model output by the dynamic generation model.
[0040] In a second aspect, a three-dimensional model generation apparatus is provided, which includes functional units for performing any of the methods provided in the first aspect, and actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software. For example, the three-dimensional model generation apparatus can include an information module, a generation module, an optimization module, and a fusion module. The information module is configured to obtain prompt information input by a user, where the prompt information is used to indicate at least one element and an action performed by a target element in the at least one element; the generation module is configured to generate an initial video according to the prompt information, and generate an initial three-dimensional dynamic world model according to the initial video, where the initial three-dimensional dynamic world model includes a three-dimensional model of the at least one element and a three-dimensional model of the action performed by the target element; the optimization module is configured to obtain at least one supplementary video according to target initial information, where the target initial information includes at least one of the initial three-dimensional dynamic world model or the initial video; the generation module is further configured to generate at least one supplementary three-dimensional dynamic world model according to the at least one supplementary video; and the fusion module is configured to fuse the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model to obtain a target three-dimensional dynamic world model.
[0041] In a third aspect, a processor is provided, which can be used to execute any of the methods provided in the first aspect.
[0042] In a fourth aspect, a chip is provided, which includes a processor and a power supply circuit. The power supply circuit can be used to supply power to the chip, and the processor can be used to execute any of the methods provided in the first aspect.
[0043] In a fifth aspect, a computing device is provided, comprising: a processor, a memory, and a computer program / instruction stored on the memory; the processor executes the computer program / instruction to cause the computing device to perform any of the methods provided in the first aspect.
[0044] In a sixth aspect, a computing device cluster is provided, comprising: at least one computing device, each of the at least one computing device comprising a processor, a memory, and a computer program / instruction stored on the memory; the processor of each computing device executes the computer program / instruction to cause the computing device cluster to implement any of the methods provided in the first aspect.
[0045] In a seventh aspect, a computer program product is provided, the computer program product comprising a computer program / instruction, the computer program / instruction being executed by a computing device to implement any of the methods provided in the first aspect.
[0046] In an eighth aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program / instruction, the computer program / instruction being executed by a computing device to implement any of the methods provided in the first aspect.
[0047] The technical effects brought by any of the implementation manners of the second aspect to the eighth aspect can be referred to the technical effects brought by the different implementation manners of the first aspect, which will not be described herein. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 One of the schematic diagrams of a system architecture provided by the present application;
[0049] Figure 2 Another of the schematic diagrams of a system architecture provided by the present application;
[0050] Figure 3 Another of the schematic diagrams of a system architecture provided by the present application;
[0051] Figure 4 A flowchart of a three-dimensional model generation method provided by the present application;
[0052] Figure 5 One of the schematic diagrams of a three-dimensional model generation provided by the present application;
[0053] Figure 6 Another of the schematic diagrams of a three-dimensional model generation provided by the present application;
[0054] Figure 7 Another of the schematic diagrams of a three-dimensional model generation provided by the present application;
[0055] Figure 8A schematic diagram of a distributed field provided for the present application;
[0056] Figure 9 A schematic diagram of a three-dimensional model generation device provided for the present application;
[0057] Figure 10 A schematic diagram of a computing device provided for the present application;
[0058] Figure 11 A schematic diagram of a computing device cluster provided for the present application;
[0059] Figure 12 A schematic diagram of a connection of a computing device cluster provided for the present application. DETAILED DESCRIPTION
[0060] The technical solutions provided by the present application will be described in detail below with reference to the accompanying drawings.
[0061] Currently, when making a three-dimensional dynamic scene, not only does it need to be manually designed, but also a professional needs to model it through modeling software. Among them, scene design includes scenery design and environment design, and a rigorous scene design usually includes a plan view, a structural exploded view, a color atmosphere view, etc. so as to deconstruct the scene from various perspectives such as plan, internal structure, color, etc. and design them respectively, and then combine them in the subsequent process. Therefore, the designer of the three-dimensional dynamic scene needs to be proficient in various technologies, and the learning cost is very high, which also leads to a very high labor cost of the three-dimensional dynamic scene.
[0062] For example, making a three-dimensional dynamic scene usually includes: (a) designing a scene in a professional software, for example, designing multiple aspects such as plan, structure, color, etc.; (b) modeling elements in the scene through modeling software; (c) mapping and binding the modeling; (d) for elements that need dynamic effects, using various special effect modules to calculate; (e) processing and rendering the light; (f) generating a rendering or effect video through the software, and modifying steps (a) to (e) according to the effect presented by the rendering or effect video until the effect presented by the rendering or effect video meets the requirements.
[0063] In general, the entire production process requires complex design and editing operations to define various structures and local details. Moreover, in order to ensure that the dynamic effects are both beautiful and realistic, not only does it need to carefully perform each of the above steps, but the entire production process also needs to be constantly modified.
[0064] Since each scene in a three-dimensional dynamic scene needs to be built from scratch, and there are many elements in the scene, the entire production process is time-consuming and labor-intensive, and the production cycle is very long. In addition, during the production process, it is necessary to accurately simulate physical effects such as object collision, fluid motion, cloth deformation, etc. through simulation of a physical engine algorithm, so as to present better dynamic effects. However, the simulation of the physical engine algorithm not only has a high learning threshold and great implementation difficulty, but also has a very large amount of calculation, resulting in very high calculation cost. Dynamic effect rendering requires a device with high computing power to implement, and the production threshold is high. Due to the limitation of the computing power of the device and software, it is impossible to implant dynamic effects on all elements in the scene, and only some elements can be selected for dynamic processing. In addition, the professional software and professional hardware required in the entire production process need to be learned by professional personnel before they can be operated. However, the learning process is complex, and it is difficult to complete the design work without years of accumulation, especially the design of dynamic effects, which needs to follow the physical engine of the real world, otherwise the realism of the finished product will be greatly discounted. The professional software, professional personnel, etc. have a high cost, so the cost of the entire design process is relatively high. In addition, since the entire production process needs to be constantly iterated and modified, and each time an effect picture or an effect video is generated, a large amount of time is consumed, resulting in a particularly high time cost.
[0065] In related technologies, the objects in the real world are also scanned and reconstructed to assist modeling. However, when scanning dynamic objects (such as vehicles, animals, etc.), not only is multi-frame synchronization required, making the scanning difficult, but motion blur or data loss is also likely to occur. When scanning static objects (such as rocks, buildings, etc.), not only does it need to be split, but it also results in a lack of internal structure data of the static objects, and the dynamic effects of the static elements cannot be scanned. In addition, scanning and reconstruction also require professional equipment and additional personnel to handle, for example, scanning and reconstructing a city scene requires a drone / laser radar device to implement multi-angle collection, which not only has a large amount of data and requires a lot of time to process, but also has a very high cost.
[0066] As can be seen, the entire production process has very low generation efficiency, very high generation cost, and requires a large amount of time cost, labor cost, and software and hardware cost.
[0067] Therefore, the three-dimensional model generation method provided in the application can be applied to the fields of games, film and television design, automatic driving, embodied intelligence, etc. For example, in the fields of games and film and television design, the three-dimensional model generation method can be used to design a three-dimensional dynamic scene, etc. For another example, in the fields of automatic driving and embodied intelligence, the three-dimensional model generation method can be used to generate a dynamic scene in simulation data, etc.
[0068] For example, the three-dimensional model generation method provided in the application can be applied to the fields of games, film and television design, automatic driving, embodied intelligence, etc. For example, in the fields of games and film and television design, the three-dimensional model generation method can be used to design a three-dimensional dynamic scene, etc. For another example, in the fields of automatic driving and embodied intelligence, the three-dimensional model generation method can be used to generate a dynamic scene in simulation data, etc.
[0069] It should be noted that the application of the three-dimensional model generation method provided in the application is not limited, and the above is only an exemplary introduction.
[0070] It should be noted that the various implementation manners of the three-dimensional model generation method provided in the application will be described in detail in the Figure 4 The implementation manners of the three-dimensional model generation method provided in the application will be described in detail in the
[0071] Next, the system architecture related to the technical solutions provided in the application will be further introduced in combination with the accompanying drawings.
[0072] The application provides a device applied to the three-dimensional model generation method, which can be referred to as a three-dimensional model generation device.
[0073] In one example, as Figure 1As shown, the three-dimensional model generation apparatus can be deployed on a terminal device of a user, and the user can implement the three-dimensional model generation method provided by the present application through the three-dimensional model generation apparatus on the terminal device of the user.
[0074] Optionally, the terminal device can be a device with a display screen, which can be used to display a picture obtained by rendering the three-dimensional dynamic world model, for example, a picture obtained by rendering the initial three-dimensional dynamic world model, a picture obtained by rendering the target three-dimensional dynamic world model, and the like. Exemplarily, the terminal device can be a tablet computer, a handheld computer, a personal computer (PC), a personal digital assistant (PDA), an ultra-mobile personal computer (UMPC), a notebook computer, a netbook, a desktop computer, or an all-in-one computer, and the like.
[0075] It should be noted that the present application does not limit the device form of the terminal device, and the above is only an exemplary introduction.
[0076] Exemplarily, the three-dimensional dynamic world model can be a four-dimensional (4D) world model.
[0077] It should be noted that the present application does not limit the type of the three-dimensional dynamic world model, and the above is only an exemplary introduction.
[0078] In another example, as shown, Figure 2 The three-dimensional model generation apparatus includes a server and a client, the client can be installed on the terminal device of the user, and the server can be deployed on a computing device of a cloud service system. The terminal device and the cloud service system communicate through a network. For example, it can be directly installed on the computing device, or it can also be installed on a virtual machine running on the computing device. Exemplarily, the cloud service system includes a cloud management platform and an infrastructure managed by the cloud management platform, and the infrastructure includes the computing device.
[0079] Method 1, the cloud management platform is used to provide an access interface (such as an interface or an API), a tenant can operate the client on the terminal device to remotely access the cloud management platform to register a cloud account and a password, and log in to the cloud management platform. After the cloud management platform authenticates the cloud account and the password successfully, the tenant can further pay for a virtual machine of a specific specification (processor, memory, disk) on the cloud management platform, and purchase the virtual machine, which is deployed on the computing device of the infrastructure. After the payment and purchase are successful, the cloud management platform provides a remote login account and password of the purchased virtual machine, and the client can remotely log in to the virtual machine, and install and run an application in the virtual machine. For example, the server of the three-dimensional model generation apparatus can be installed.
[0080] Exemplarily, a client is installed on a terminal device of a user. The cloud management platform can receive a generation request from the client, and send prompt information carried by the generation request to a server of the three-dimensional model generation apparatus. The server can generate an initial three-dimensional dynamic world model according to the prompt information carried by the generation request. The cloud management platform can also generate a supplementary three-dimensional dynamic world model, and fuse the initial three-dimensional dynamic world model and the supplementary three-dimensional dynamic world model to obtain a target three-dimensional dynamic world model. Then, the cloud management platform can return the target three-dimensional dynamic world model to the client, so as to display a picture obtained by rendering the target three-dimensional dynamic world model to the user through the client.
[0081] It should be noted that the present application does not limit the process of implementing the three-dimensional model generation method provided by the present application by the client and the server, and the above is only an exemplary introduction.
[0082] Alternatively, the computing device can be an ultra-mobile personal computer (UMPC), a notebook computer, a netbook, a desktop computer, an all-in-one computer, etc. Or, the computing device can be a network device, etc.
[0083] The network device can include a server, a bare-metal server, etc. The server can be a physical server, or two or more physical servers sharing different responsibilities and cooperating with each other to implement the functions of the server. Exemplarily, the server can be a blade server, a high-density server, a rack server, or a tower server, etc.
[0084] It should be noted that the present application does not limit the device form of the computing device, and the above is only an exemplary description.
[0085] Option 2, the cloud service system can provide a three-dimensional dynamic scene generation service for tenants. Exemplarily, a server of the three-dimensional model generation apparatus can be deployed on a computing device in the infrastructure, or a server of the three-dimensional model generation apparatus can be deployed in a virtual machine on the computing device. The tenant can purchase a three-dimensional dynamic scene generation service through the cloud management platform, and install a client of the three-dimensional model generation apparatus on his own terminal device. Then, the cloud service system can provide a three-dimensional dynamic scene generation service for the tenant through the server and the client of the three-dimensional model generation apparatus.
[0086] It should be noted that other related descriptions of embodiment 2 can refer to the descriptions of embodiment 1 above, which will not be repeated here.
[0087] Exemplarily, as shown in Figure 3 The three-dimensional model generation apparatus can include an information module, a generation module, an optimization module, a fusion module, and a display module.
[0088] The information module can be configured to receive prompt information input by a user, various instructions (e.g., a first instruction, a second instruction, etc.), and the like, and can process the prompt information to convert the prompt information into input information that can be recognized by the video generation model.
[0089] The generation module can be configured to generate an initial video, a three-dimensional dynamic world model, and the like. For example, the three-dimensional dynamic world model can be an initial three-dimensional dynamic world model, a supplemented three-dimensional dynamic world model, and the like. For example, the generation module includes a video generation model (also referred to as a video generation algorithm), which is configured to generate a corresponding video according to input information, such as a video of a scene described by the input information. For example, the generation module further includes a dynamic generation model (also referred to as a 3D dynamic generation model), which is configured to generate a three-dimensional dynamic world model according to a video, and the like. For example, the dynamic generation model can generate an initial three-dimensional dynamic world model according to an initial video output by the video generation model.
[0090] For example, the video generation model (also referred to as a video large model) can be a DIT (diffusion transformer) model, and the like.
[0091] For example, the video generation model can be obtained by iteratively training a base model using training data. For example, the training data is configured to describe elements and actions in a three-dimensional dynamic scene to be generated. After inputting the training data 1 into the base model, a video 1 output by the base model is obtained. Then, the base model is iteratively optimized according to the video 1 and a loss function, and a video generation model is obtained.
[0092] It should be noted that the type of the video generation model is not limited in the present application, and the above is only an example.
[0093] For example, the dynamic generation model can be a 4D Gaussian model, and the like. In the case where the generation module generates a three-dimensional dynamic world model by using the 4D Gaussian model, the generated three-dimensional dynamic world model can be a Gaussian world model.
[0094] It should be noted that the type of the dynamic generation model is not limited in the present application, and the above is only an example.
[0095] For example, the dynamic generation model can be obtained by iteratively training a base model using training data. For example, the training data is a video. After inputting the training data 1 into the base model, a three-dimensional dynamic world model 1 output by the base model is obtained. Then, the base model is iteratively optimized according to the three-dimensional dynamic world model 1 and a loss function, and a video generation model is obtained.
[0096] The generation module can further include a video completion unit, which can be configured to complete the holes of the video.
[0097] The holes of the video: due to the different depth of field and occlusion relationship of the objects in the pictures of the video, the occluded objects or backgrounds in the current viewpoint can be exposed in the virtual viewpoint, resulting in lack of effective information, and these areas lacking effective information are referred to as holes.
[0098] The optimization module can be configured to obtain the supplementary video to optimize the initial three-dimensional dynamic world model, etc. For example, the optimization module can optimize the initial three-dimensional dynamic world model by completing the holes, supplementing the camera poses, etc. For example, the generation module can further generate a supplementary three-dimensional dynamic world model according to the supplementary video, so as to optimize the initial three-dimensional dynamic world model through the supplementary three-dimensional dynamic world model.
[0099] The fusion module is configured to fuse the supplementary three-dimensional dynamic world model into the initial three-dimensional dynamic world model. For example, the fusion module includes a fusion model, which can be configured to fuse the initial three-dimensional dynamic world model and the supplementary three-dimensional dynamic world model.
[0100] For example, the fusion model can be a 4D Gaussian model, etc.
[0101] It should be noted that the type of the fusion model is not limited in the present application, and the above is only an example.
[0102] The display module can be configured to display the three-dimensional dynamic scene obtained by rendering the three-dimensional dynamic world model, etc. For example, the display module displays the three-dimensional dynamic scene obtained by rendering the three-dimensional dynamic world model, so that the user can perform 3D roaming.
[0103] For example, in combination with the system architecture shown in Figure 2 The client of the three-dimensional model generation device can include the display module and the acquisition unit in the information module, and the acquisition unit in the information module is configured to acquire the prompt information input by the user, various instructions, etc. The server of the three-dimensional model generation device includes the conversion unit of the information module, the generation module, the optimization module, and the fusion module, etc. The conversion unit of the information module is configured to convert the prompt information into input information that can be recognized by the video module, etc.
[0104] It should be noted that the names of the various modules of the three-dimensional model generation device are not limited in the present application, and the above is only an example. For example, the generation module can be referred to as a Gaussian reconstruction module, and the display module can also be referred to as a world roaming module. The information module can also be referred to as an input processing module. The optimization module can also be referred to as a trajectory planning module / user interaction module, etc.
[0105] It should be noted that the above division of the three-dimensional model generation apparatus is only exemplary, for example, the three-dimensional model generation apparatus can also be divided into more modules or fewer modules, or the three-dimensional model generation apparatus can also be as a whole, to realize the three-dimensional model generation method provided by the present application.
[0106] It should be noted that Figures 1 to 3 The system architecture shown does not constitute a limitation on the system architecture for executing the three-dimensional model generation method provided by the present application.
[0107] It should be noted that the system architecture and application scenarios described in the present application are for more clearly illustrating the technical solutions of the present application, and do not constitute a limitation on the technical solutions provided by the present application. Those skilled in the art can know that with the evolution of system architecture and the appearance of new application scenarios, the technical solutions provided by the present application are also applicable to similar technical problems.
[0108] For ease of understanding, the three-dimensional model generation method provided by the present application is introduced below in combination with the above system architecture and the accompanying drawings.
[0109] Figure 4 A flowchart of a three-dimensional model generation method provided by the present application is provided. Exemplarily, the three-dimensional model generation method can include the following steps 401-405. In the present application, "step" can be abbreviated as "S", and the following will not be described in detail.
[0110] Exemplarily, in the three-dimensional model generation method provided by the present application, Figure 1 The system architecture shown is taken as an example for exemplary introduction of the present application. Among them, Figure 2 The system architecture shown implements the process of the three-dimensional model generation method provided by the present application, which can be referred to in the following embodiments, and the following will not be described in detail.
[0111] S401: Obtain the prompt information input by the user. The prompt information is used to indicate at least one element and an action performed by a target element in the at least one element.
[0112] Exemplarily, the prompt information is used to indicate a three-dimensional dynamic scene to be generated, and the at least one element is an element in the three-dimensional dynamic scene to be generated. The action performed by the target element is a dynamic effect of the three-dimensional dynamic scene to be generated.
[0113] Exemplarily, as Figure 5As shown in (a) in FIG. 4, the terminal device of the user displays a target interface of the three-dimensional model generation apparatus, and the target interface is used for inputting prompt information. The user can input the prompt information on the target interface, and the prompt information is used for describing the three-dimensional dynamic scene to be generated, that is, used for describing the three-dimensional dynamic scene that the user wants to generate. For example, the prompt information can be used to indicate at least one element in the three-dimensional dynamic scene to be generated and a dynamic effect of the three-dimensional dynamic scene.
[0114] For example, the prompt information can be "a man sits on a sofa in a living room, watches a sports event played on a projection screen, and the living room has an audio-visual device, a humanoid intelligent agent and a pet penguin". At least one element can include a man, a living room, a sofa, a projection screen, an audio-visual device, a humanoid intelligent agent and a pet penguin. The action performed by the target element can include the projection screen playing the sports event, and the dynamic effect includes the projection screen playing the sports event.
[0115] For example, the form of the prompt information can include at least one of a text form, a text form, an audio form, an image form and a video form.
[0116] It should be noted that the form of the prompt information is not limited in the present application, and the above is only an example. Hereinafter, the prompt information is taken as a text form as an example to introduce the present application.
[0117] In this example, by setting multiple forms of prompt information, a three-dimensional dynamic world model can be generated by triggering multiple modal information, so as to realize the dynamic effect specified by the user, which helps to improve the user experience.
[0118] Optionally, the prompt information can include text information (i.e. information in the form of text) and pictures (i.e. information in the form of images), and the text information is used to indicate at least one element and an action performed by a target element in the at least one element. For example, the at least one element can be an element in the picture included in the prompt information.
[0119] In this embodiment, by setting that the prompt information includes text information and pictures, it helps to improve the comprehensiveness of the content provided by the prompt information, thereby helping to improve the accuracy and reliability of generating the initial three-dimensional dynamic world model according to the prompt information.
[0120] S402: generating an initial video according to the prompt information, and generating an initial three-dimensional dynamic world model according to the initial video.
[0121] The initial three-dimensional dynamic world model includes a three-dimensional model of at least one element and a three-dimensional model of a target element performing the action.
[0122] The initial three-dimensional dynamic world model may include, for example, a first 3D model of at least one element and a second 3D model of a dynamic effect. When rendering the initial three-dimensional dynamic world model, the at least one element may be obtained by rendering the first 3D model, and the target element performing the action may be obtained by rendering the second 3D model, so as to obtain the at least one element in the three-dimensional dynamic scene and the dynamic effect of the three-dimensional dynamic scene. That is, the three-dimensional dynamic scene obtained by rendering the target three-dimensional dynamic world model includes the three-dimensional dynamic scene to be generated indicated by the prompt information, the at least one element in the three-dimensional dynamic scene to be generated, and the dynamic effect of the three-dimensional dynamic scene to be generated, and the like. In other words, the initial three-dimensional dynamic world model may be used to render the at least one element in the three-dimensional dynamic scene and the dynamic effect of the three-dimensional dynamic scene.
[0123] For example, rendering the initial three-dimensional dynamic world model may obtain a plurality of pictures, and the plurality of pictures may constitute the three-dimensional dynamic scene indicated by the prompt information.
[0124] It should be noted that the related descriptions of the supplementary three-dimensional dynamic world model and the target three-dimensional dynamic world model may refer to the initial three-dimensional dynamic world model, and will not be described in detail hereinafter.
[0125] For example, the initial video includes at least one element indicated by the prompt information, a dynamic effect, and the like.
[0126] For example, as shown in Figure 6 After the three-dimensional model generation apparatus obtains the prompt information input by the user, the three-dimensional model generation apparatus automatically generates an initial video in response to the user input prompt information. Then, the three-dimensional model generation apparatus generates an initial three-dimensional dynamic world model according to the initial video.
[0127] Hereinafter, the various implementation manners of S402 will be exemplarily introduced through mode 1 and mode 2.
[0128] In mode 1, after the three-dimensional model generation apparatus obtains the initial video, the three-dimensional model generation apparatus automatically generates an initial three-dimensional dynamic world model according to the initial video. Hereinafter, mode 1 will be exemplarily introduced through S1-S2.
[0129] S1: input the prompt information into the video generation model of the three-dimensional model generation apparatus to obtain the initial video output by the video generation model.
[0130] For example, the generation module of the three-dimensional model generation apparatus inputs the prompt information into the video generation model, and the video generation model generates a video according to the prompt information and outputs the initial video.
[0131] It should be noted that the application does not limit the manner of generating the initial video, and the above is only an exemplary introduction.
[0132] It should be noted that the number of initial videos generated according to the prompt information is not limited in the present application. For example, one initial video can be generated according to the prompt information, or a plurality of initial videos can be generated according to the prompt information.
[0133] S2: input the initial video into the dynamic generation model to obtain an initial three-dimensional dynamic world model output by the dynamic generation model.
[0134] For example, after the generation module of the three-dimensional model generation apparatus obtains the initial video, the initial video is input into the dynamic generation model, the dynamic generation model generates a three-dimensional dynamic world model according to the initial video, and outputs the initial three-dimensional dynamic world model.
[0135] In this way, after the three-dimensional model generation apparatus generates the initial video, the next step is automatically started directly, that is, the initial three-dimensional dynamic world model is generated, thereby helping to improve the efficiency of generating the initial three-dimensional dynamic world model.
[0136] In the following, the plurality of implementation manners of S2 are exemplarily introduced through mode 1A and mode 1B.
[0137] In mode 1A, the three-dimensional model generation apparatus can process the holes of the initial video, for example, fill the holes of the initial video, and then generate the initial three-dimensional dynamic world model. In the following, mode 1A is exemplarily introduced through S2a and S2b.
[0138] S2a: process the holes of the initial video to obtain a target initial video.
[0139] In this way, the holes of the target initial video are smaller than the holes of the initial video.
[0140] For example, after the three-dimensional model generation apparatus obtains the initial video, it can be determined whether the initial video has holes. In the case that the initial video has holes, the three-dimensional model generation apparatus can process the holes of the video (for example, fill the holes), thereby obtaining the target initial video.
[0141] S2b: input the target initial video into the dynamic generation model to obtain an initial three-dimensional dynamic world model output by the dynamic generation model.
[0142] For example, after the three-dimensional model generation apparatus obtains the target initial video, the target initial video is input into the dynamic generation model, the dynamic generation model generates a three-dimensional dynamic world model according to the target initial video, and outputs the generated initial three-dimensional dynamic world model.
[0143] In this way, by processing the holes of the initial video to obtain the target initial video, it is helpful to reduce or eliminate the holes on the target initial video, so as to generate the initial three-dimensional dynamic world model according to the target initial video, and it is helpful to reduce or eliminate the holes in the initial three-dimensional dynamic world model, thereby helping to improve the integrity of the rendered initial three-dimensional dynamic world model, and further helping to improve the picture effect of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model.
[0144] It should be noted that the holes in the three-dimensional dynamic world model in the present application can be considered as holes in the picture obtained by rendering the three-dimensional dynamic world model, and the following will not be described in detail.
[0145] In mode 1B, the three-dimensional model generation device can obtain auxiliary information of the initial video, and generate the initial three-dimensional dynamic world model in combination with the auxiliary information. Hereinafter, mode 1B will be exemplarily introduced through S2c and S2d.
[0146] S2c: performing target processing on the initial video to obtain auxiliary information of the initial video.
[0147] Exemplarily, after the three-dimensional model generation device obtains the initial video, the target processing can be performed on the initial video. For example, the target processing includes at least one of video depth estimation or dynamic information estimation, thereby obtaining the auxiliary information of the initial video.
[0148] In one example, the target processing includes video depth estimation. As shown in Figure 6 , the three-dimensional model generation device performs video depth estimation on the initial video to obtain first auxiliary information of the initial video.
[0149] Exemplarily, the first auxiliary information can include a depth map of the initial video. The depth map is used to indicate the depth value of each pixel point in the video. The depth value of each pixel point is used to represent the distance from each pixel point in the scene to the camera. The depth information of the three-dimensional model can be constructed through the depth map.
[0150] It should be noted that the manner of video depth estimation of the video in the present application is not limited.
[0151] In another example, the target processing includes dynamic information estimation. As shown in Figure 6 , the three-dimensional model generation device performs dynamic information estimation on the initial video to obtain second auxiliary information of the initial video.
[0152] Exemplarily, the second auxiliary information can include frame rate, resolution, code rate and motion information of the initial video. The motion information includes the trajectory, speed and direction of the movement of the elements in the initial video.
[0153] It should be noted that the application does not limit the manner of dynamic information estimation on the video.
[0154] In yet another example, the target processing includes video depth estimation and dynamic information estimation.
[0155] For example, the three-dimensional model generation apparatus performs video depth estimation on the initial video first, and then performs dynamic information estimation on the initial video. Alternatively, the three-dimensional model generation apparatus can perform dynamic information estimation on the initial video first, and then perform video depth estimation on the initial example.
[0156] It should be noted that the related description of this example can refer to the above two examples, which will not be repeated here.
[0157] S2d: input the initial video and the auxiliary information of the initial video into the dynamic generation model to obtain an initial three-dimensional dynamic world model.
[0158] For example, the three-dimensional model generation apparatus inputs the initial video and the auxiliary information (e.g., the first auxiliary information and / or the second auxiliary information, etc.) of the initial video into the dynamic generation model. The dynamic generation model generates a three-dimensional dynamic world model according to the initial video and the auxiliary information of the initial video, and outputs the generated initial three-dimensional dynamic world model.
[0159] In this implementation, the target processing is performed on the initial video to obtain auxiliary information, so that the three-dimensional dynamic world model is generated by the auxiliary information and the initial video. In this way, it is helpful to improve the comprehensiveness of the information used to generate the initial three-dimensional dynamic world model, thereby helping to improve the picture quality and picture accuracy of the three-dimensional dynamic scene obtained by rendering the generated initial three-dimensional dynamic world model.
[0160] It should be noted that the scheme of S2a-S2b can be used in combination with the scheme of S2c-S2d, or can be used alone, and the application does not limit this.
[0161] It should be noted that the application does not limit the manner of generating the initial three-dimensional dynamic world model, and the above is only an exemplary introduction. For example, in addition to the above manner 1A and manner 1B, the initial video can also be inputted into the dynamic generation model to generate the initial three-dimensional dynamic world model.
[0162] In mode 2, the three-dimensional model generation apparatus generates the initial video and presents the initial video to the user to facilitate the user to view the effect of the initial video, etc. Hereinafter, mode 2 is exemplarily introduced by S3-S5.
[0163] S3: input the prompt information into the video generation model of the three-dimensional model generation apparatus to obtain an initial video output by the video generation model.
[0164] It should be noted that the related description of S3 can refer to the description of S1 above, which will not be repeated here.
[0165] S4: provide a third page, wherein the third page is used to display the initial video.
[0166] For example, as shown in (a) of FIG. 6, after the three-dimensional model generation apparatus obtains the initial video, the third page for displaying the initial video can be provided, so that the user can view the dynamic effect presented by the initial video. For example, after the video generation model of the three-dimensional model generation apparatus outputs the initial video, the initial video can be sent to the display module to be displayed by the display module, so as to present the initial video to the user. Figure 7
[0167] S5: in response to a second instruction input by the user, input the initial video into the dynamic generation model to obtain an initial three-dimensional dynamic world model output by the dynamic generation model. The second instruction is used to indicate confirmation of the initial video or to indicate any one of the generation of the three-dimensional model.
[0168] For example, after the three-dimensional model generation apparatus displays the initial video through the third page, the user can determine whether the initial video achieves the effect he needs according to the presentation effect of the initial video, for example, whether the elements in the initial video are correct, whether the dynamic effect is correct, etc. In the case where the user determines that the dynamic effect he needs is achieved, the user can input the second instruction, wherein the second instruction can be used to indicate that the user confirms that the content of the initial video is correct, or can be used to indicate that the generation of the three-dimensional dynamic world model is continued. After the three-dimensional model generation apparatus receives the second instruction, in response to the second instruction, the initial three-dimensional dynamic world model is generated according to the initial video.
[0169] For example, as shown in (b) of FIG. 6, the interface of the initial video displayed by the three-dimensional model generation apparatus includes a "confirm" control, and the user can input the second instruction to the three-dimensional model generation apparatus by clicking the "confirm" control. Figure 7
[0170] It should be noted that the manner in which the user inputs the second instruction is not limited by the present application, and the above is only an exemplary introduction.
[0171] It should be noted that other related descriptions of S5 can refer to the description of S2 above, which will not be repeated here.
[0172] In this way, after the three-dimensional model generation apparatus generates the initial video, the initial video is presented to the user, so that the user can view the dynamic effect of the initial video presentation, thereby helping to ensure that the initial three-dimensional dynamic world model is regenerated in the case that the dynamic effect of the initial video presentation meets the user's demand, and further helping to ensure that the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model meets the user's demand, so as to help improve the user's use experience.
[0173] S403: Obtain at least one supplementary video according to the target initial information, wherein the target initial information includes at least one of the initial three-dimensional dynamic world model or the initial video.
[0174] For example, after the three-dimensional model generation apparatus obtains the initial three-dimensional dynamic world model, at least one supplementary video can be obtained according to the target initial information, and the at least one supplementary video is fused into the initial three-dimensional dynamic world model, so as to optimize the initial three-dimensional dynamic world model through the at least one supplementary video.
[0175] In the following, the various implementation manners of S403 are exemplarily introduced through mode 3 and mode 4.
[0176] Mode 3, after the three-dimensional model generation apparatus obtains the initial three-dimensional dynamic world model, at least one supplementary video can be directly obtained according to the target initial information.
[0177] In this way, by setting the three-dimensional model generation apparatus to obtain the initial three-dimensional dynamic world model, and then directly obtaining at least one supplementary video according to the target initial information, the efficiency of obtaining the supplementary video is improved, thereby helping to improve the efficiency of obtaining the supplementary three-dimensional dynamic world model.
[0178] Mode 4, after the three-dimensional model generation apparatus obtains the initial three-dimensional dynamic world model, the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model can be presented to the user, so that the user can view the effect of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model.
[0179] In the following, mode 4 is exemplarily introduced through S6-S7.
[0180] S6: Provide a first page, and the first page is used to display a picture rendered by the initial three-dimensional dynamic world model.
[0181] For example, after the three-dimensional model generation apparatus obtains the initial three-dimensional dynamic world model, the three-dimensional model generation apparatus can display a picture of a three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, so that the user can check whether the three-dimensional dynamic scene presented by the initial three-dimensional dynamic world model is correct. For example, after the dynamic model generation module of the three-dimensional model generation apparatus outputs the initial three-dimensional dynamic world model, the three-dimensional model generation apparatus can send the initial three-dimensional dynamic world model to the display module, so that the display module displays a three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, thereby presenting the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model to the user.
[0182] S7: In response to a first instruction input by the user, at least one supplementary video is obtained according to the target initial information. The first instruction is used to confirm the picture obtained by rendering the initial three-dimensional dynamic world model or to indicate any one of the supplementary three-dimensional dynamic scene.
[0183] For example, after the three-dimensional model generation apparatus displays the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, the user can determine whether the content in the initial three-dimensional dynamic scene is correct according to the presentation effect of the three-dimensional dynamic scene, for example, whether at least one element in the three-dimensional dynamic scene, the dynamic effect of the three-dimensional dynamic scene, etc. is correct. In the case of correct determination, the user can input a second instruction to the three-dimensional model generation apparatus. After the three-dimensional model generation apparatus receives the second instruction, at least one supplementary video is obtained according to the target initial information in response to the second instruction.
[0184] It should be noted that other related descriptions of S6-S7 can refer to the descriptions of SS4-S5 above, which will not be described here.
[0185] In this way, after the three-dimensional model generation apparatus generates the initial three-dimensional dynamic world model, the three-dimensional model generation apparatus presents the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model to the user, so that the user can check the presentation effect of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, thereby helping to ensure that the initial three-dimensional dynamic world model is generated and optimized again only in the case that the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model is correct, thereby helping to ensure the necessity of the optimization process and the accuracy of the optimization result.
[0186] For example, as shown in Figure 6 After the three-dimensional model generation apparatus obtains the initial three-dimensional dynamic world model, the three-dimensional model generation apparatus can obtain the supplementary video through trajectory planning / user interaction, etc. Hereinafter, the process of "obtaining at least one supplementary video" is exemplarily introduced through mode 5 and mode 6.
[0187] In the manner 5, the three-dimensional model generation apparatus can acquire at least one supplementary video according to the pictures of the respective perspectives of the initial three-dimensional dynamic world model. In the following, the manner 5 is exemplarily introduced through S8-S9.
[0188] S8: providing a fourth page, the fourth page being configured to display the picture of at least one perspective rendered by the initial three-dimensional dynamic world model.
[0189] Exemplarily, after the three-dimensional model generation apparatus obtains the initial three-dimensional dynamic world model, the three-dimensional model generation apparatus displays the picture rendered by the initial three-dimensional dynamic world model through the fourth page. Then, the three-dimensional model generation apparatus adjusts the perspective of the picture rendered by the initial three-dimensional dynamic world model, thereby displaying the picture of at least one perspective rendered by the initial three-dimensional dynamic world model.
[0190] It should be noted that the manner of adjusting the perspective of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model is not limited in the present application. For example, the user can manually adjust the perspective of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model, or the three-dimensional model generation apparatus can automatically adjust the perspective of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model.
[0191] S9: acquiring the picture of at least one perspective, and rendering the picture of at least one perspective to obtain at least one candidate video, wherein the at least one candidate video is the at least one supplementary video.
[0192] Exemplarily, when the picture of at least one perspective is displayed on the fourth page, the three-dimensional model generation apparatus acquires the picture of at least one perspective. Then, the three-dimensional model generation apparatus can render the picture of at least one perspective to obtain at least one candidate video, for example, the three-dimensional model generation apparatus can render the picture of at least one perspective through the optimization module. For example, one perspective picture can be rendered into one candidate video. The three-dimensional model generation apparatus can take the at least one candidate video as the at least one supplementary video.
[0193] In this manner, by adjusting the perspective of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model, at least one supplementary video is acquired, which not only helps to improve the convenience of acquiring the supplementary video, but also helps to take the picture of a perspective with a large empty hole in the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model as a supplementary video, thereby realizing the completion of the empty hole of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model, and further realizing the optimization of the three-dimensional dynamic scene rendered by the final target three-dimensional dynamic world model.
[0194] In manner 5, after the three-dimensional model generation apparatus obtains the plurality of candidate videos, the three-dimensional model generation apparatus can further obtain a supplementary video based on the perspective corresponding to the target candidate video in the plurality of candidate videos. The following is exemplarily introduced through S10 to S13.
[0195] S10: From the plurality of candidate videos, determine a target candidate video, and the target candidate video has a larger hole than at least part of the candidate videos other than the target candidate video in the plurality of candidate videos.
[0196] Exemplarily, after the three-dimensional model generation apparatus obtains the plurality of candidate videos, the three-dimensional model generation apparatus determines the size of the hole of each candidate video in the plurality of candidate videos. For example, the size of the hole can be represented by the number of pixel points of the hole. The three-dimensional model generation apparatus sorts the holes of the plurality of candidate videos in descending order to obtain a sorting result.
[0197] In one example, the three-dimensional model generation apparatus can take the candidate video ranked first as the target candidate video. In this way, based on the perspective corresponding to the target candidate video, when the supplementary candidate video is obtained, it is helpful to improve the probability of obtaining a video with a hole in the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, and further to optimize the hole in the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model.
[0198] In another example, the plurality of candidate videos is K candidate videos, and the three-dimensional model generation apparatus can take the candidate video ranked Nth as the target candidate video, where N is less than K, and N is a sorting threshold.
[0199] It should be noted that the size of the sorting threshold is not limited in the present application, and the sorting threshold can be dynamically set in the actual scene.
[0200] It should be noted that at least part of the present application can be all, for example, the at least part of the candidate videos can include K-1 candidate videos in the K candidate videos other than the target candidate video. Alternatively, at least part of the present application can be part of all. For example, the at least part of the candidate videos can include Q candidate videos in the K candidate videos other than the target candidate video, and Q is less than K-1.
[0201] S11: Provide a fifth page, the fifth page is used to render the picture of the target perspective obtained by the initial three-dimensional dynamic world model, and the picture of the target perspective is the picture of the target candidate video.
[0202] Exemplarily, after obtaining the target candidate video, the three-dimensional model generation apparatus determines a picture of a target view angle of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, and obtains the target candidate video. Based on this, the three-dimensional model generation apparatus displays the picture of the target view angle of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model through the fifth page.
[0203] S12: adjusting a view angle of the picture obtained by rendering the initial three-dimensional dynamic world model, and obtaining a picture of at least one supplementary view angle.
[0204] Exemplarily, in the process of providing the fifth page, the three-dimensional model generation apparatus takes the picture of the target view angle as a starting point, adjusts the view angle of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, obtains the picture after adjusting the view angle, and takes the picture as the picture of at least one supplementary view angle.
[0205] S13: rendering the picture of at least one supplementary view angle to obtain at least one supplementary video.
[0206] Exemplarily, the three-dimensional model generation apparatus renders the picture of at least one supplementary view angle through the optimization module to obtain at least one supplementary candidate video, and takes the at least one supplementary candidate video as the supplementary video.
[0207] It should be noted that other related descriptions of S11 can refer to the description of S8 above, which will not be described here.
[0208] In an embodiment, by obtaining the target candidate video, and taking the picture of the target view angle corresponding to the target candidate video as a starting point, adjusting the view angle of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, at least one supplementary candidate video is obtained. In this way, it is not only helpful to obtain a video with a hollow view angle from the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, thereby helping to optimize the hollow of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, but also helpful to improve the richness and comprehensiveness of the supplementary video.
[0209] In mode 6, the three-dimensional model generation apparatus can obtain at least one supplementary video according to the camera pose of the initial video. Hereinafter, S14 and S15 are exemplarily introduced.
[0210] S14: inputting a target camera pose and an initial picture rendered by the target camera pose into a neural network model to obtain at least one supplementary camera pose output by the neural network model and at least one supplementary picture rendered by the at least one supplementary camera pose.
[0211] The target camera pose is the camera pose of the initial video.
[0212] Exemplarily, after obtaining the initial three-dimensional dynamic world model, the three-dimensional model generation apparatus can render a target camera pose (e.g., a camera pose of an initial video) to obtain an initial picture. Then, the three-dimensional model generation apparatus inputs the target camera pose and the initial picture rendered by the target camera pose into a neural network model, and the neural network model predicts a next camera pose according to the target camera pose and the initial picture, and outputs at least one supplementary camera pose and at least one supplementary picture rendered by the at least one supplementary camera pose. For example, one supplementary camera pose can render one supplementary picture.
[0213] Exemplarily, the neural network model can be a pre-trained trajectory prediction model, which can be used to predict a next camera pose according to a camera pose and a picture rendered by the camera pose. For example, the next camera pose (e.g., the at least one supplementary camera pose) output by the trajectory prediction model can be an optimal camera pose relative to other camera poses.
[0214] It should be noted that the next camera pose refers to a camera pose that can appear after the target camera pose, and the camera pose that can appear after the target camera pose can include various cases, so the next camera pose can include at least one camera pose, and each camera pose in the at least one camera pose can be a next camera pose of the target camera pose.
[0215] It should be noted that the type of the trajectory prediction model is not limited in the present application. For example, it can be a sequential network, a graph neural network, a generative model, etc.
[0216] S15: generating at least one supplementary video according to the at least one supplementary camera pose and the at least one supplementary picture.
[0217] Exemplarily, after obtaining the at least one supplementary camera pose and the at least one supplementary picture, the three-dimensional model generation apparatus can input the at least one supplementary camera pose and the at least one supplementary picture into a video generation model, and the video generation model generates a video according to the supplementary camera pose and the supplementary picture, and outputs the generated at least one supplementary video. For example, the video generation model can generate one supplementary video according to one supplementary camera pose and one supplementary picture rendered by the one supplementary camera pose.
[0218] In the implementation, based on the camera poses of the initial video, other camera poses are predicted through the neural network model, and the supplementary video is generated through the predicted camera poses. In this way, not only the diversity of the manner of obtaining the supplementary video is improved, but also the camera poses that are missing in the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model are obtained, thereby the richness and comprehensiveness of the camera poses of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model are improved.
[0219] In addition, predicting other camera poses based on the camera poses of the initial video also helps to ensure that the elements on the supplementary video corresponding to the predicted camera poses are the same as those on the initial video, thereby the accuracy of the elements on the three-dimensional dynamic scene of the supplementary rendered three-dimensional dynamic world model corresponding to the supplementary video is ensured, and further the accuracy of the elements of the fused target three-dimensional dynamic world model is ensured, avoiding missing elements in the initial video or adding elements that are not in the initial video.
[0220] In mode 6, the target camera pose can also be a camera pose determined according to the initial three-dimensional dynamic world model. Hereinafter, the process of determining the target camera pose according to the initial three-dimensional dynamic world model is exemplarily introduced through S16 to S18.
[0221] S16: Construct a distribution field of a space body of the initial three-dimensional dynamic world model, the space body including a plurality of three-dimensional points of the initial three-dimensional dynamic world model, and the distribution field being used to indicate the distribution of the plurality of three-dimensional points.
[0222] Exemplarily, after obtaining the initial three-dimensional dynamic world model, the three-dimensional model generation device can determine a space body of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model, the space body including a plurality of 3D points of the initial three-dimensional dynamic world model, for example, including a plurality of 3D points of the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic scene model. Then, as shown in Figure 8 The three-dimensional model generation device constructs a distribution field of the space body, which can be used to indicate the distribution of the plurality of 3D points, for example, can be used to indicate the distribution rule of the plurality of 3D points.
[0223] It should be noted that the present application does not limit the representation method of the space body. For example, the space body can be represented by coordinates in a three-dimensional space. It should be noted that the present application does not limit the method of constructing the distribution field. For example, the coordinates of each 3D point of the space body can be placed in the coordinate system of the three-dimensional space, thereby obtaining the distribution field of the space body.
[0224] S17: Determine the density of the three-dimensional points in a plurality of directions centered on the position of the camera pose of the initial video in the distribution field.
[0225] Exemplarily, as shown in Figure 8As shown, after determining the distribution field of the spatial body, the three-dimensional model generation apparatus determines the position (e.g., target position) of the camera pose of the initial video in the distribution field. Then, the three-dimensional model generation apparatus determines the density of 3D points in multiple directions centered on the target position. For example, the density 1 of 3D points in the X-axis direction is determined, the density 2 of 3D points in the Y-axis direction is determined, the density 3 of 3D points in the Z-axis direction is determined, and so on.
[0226] The density of 3D points in each of the multiple directions can be used to indicate the size of the hole in each direction. The smaller the density of 3D points, the larger the hole. Conversely, the larger the density of 3D points, the smaller the hole.
[0227] For example, the three-dimensional model generation apparatus can sort the densities of 3D points in multiple directions in ascending order to obtain a sorting result. The density of 3D points in the target direction is the density ranked first, for example, the density 1 of 3D points in the X-axis direction is the smallest in the multiple directions, that is, the density 1 of 3D points in the X-axis direction is ranked first, and the target direction can be the X-axis direction.
[0228] For example, the target direction can be the direction with the second smallest density. For example, the density 2 of 3D points in the Y-axis direction is the second smallest in the multiple directions, that is, the density 1 of 3D points in the Y-axis direction is ranked second, and the target direction can be the Y-axis direction.
[0229] It should be noted that the specific value of the sorting threshold is not limited in the present application, and the sorting threshold can be dynamically set according to the actual scene.
[0230] For example, the target direction can be the direction with the second smallest density. For example, the density 2 of 3D points in the Y-axis direction is the second smallest in the multiple directions, that is, the density 1 of 3D points in the Y-axis direction is ranked second, and the target direction can be the Y-axis direction.
[0231] It should be noted that the specific value of the sorting threshold is not limited in the present application, and the sorting threshold can be dynamically set according to the actual scene.
[0232] S18: obtaining a target camera pose according to the target direction and the camera pose of the initial video, the density of 3D points in the target direction being smaller than the density of 3D points in at least part of the multiple directions other than the target direction.
[0233] For example, the three-dimensional model generation apparatus determines the camera pose of the target direction according to the camera pose of the initial video and the relationship between the target direction and the target position, thereby obtaining the target camera pose.
[0234] In this implementation, the target camera pose is obtained by the initial three-dimensional dynamic world model, which helps to improve the diversity of the target camera pose. In addition, since the target camera pose is a camera pose in a direction in which the hole in the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model is relatively large, when predicting other camera poses based on the target camera pose by using the neural network model, the camera pose in which the hole in the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model is relatively large can be predicted, thereby helping to repair the hole in the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model, and further helping to improve the presentation effect of the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model.
[0235] Optionally, the target camera pose corresponds to a target view angle. The target view angle corresponding to the target camera pose refers to a view angle corresponding to the target camera pose in the three-dimensional dynamic scene obtained by rendering the initial three-dimensional dynamic world model. For example, the camera pose of the first view angle obtained by rendering the initial three-dimensional dynamic world model is the target camera pose, and based on this, the first view angle can be used as the target view angle, that is, the picture of the first view angle can be used as the picture of the target view angle.
[0236] For example, the three-dimensional model generation apparatus provides a sixth page, and the sixth page is used to display a picture obtained by rendering the target camera pose of the initial three-dimensional dynamic world model. Then, the three-dimensional model generation apparatus adjusts the view angle of the picture obtained by rendering the initial three-dimensional dynamic world model, and obtains a picture of at least one supplementary view angle. Then, the three-dimensional model generation apparatus renders the picture of the at least one supplementary view angle, and obtains at least one supplementary video.
[0237] It should be noted that other related descriptions of this scheme can be referred to the description of S11 above, and will not be repeated here.
[0238] In this embodiment, by using the view angle corresponding to the target camera pose as the target view angle, the diversity of the selection manner of the target view angle is improved.
[0239] S404: generating at least one supplementary three-dimensional dynamic world model according to the at least one supplementary video.
[0240] It should be noted that the related description of the supplementary three-dimensional dynamic world model can be referred to the description of the initial three-dimensional dynamic world model above, and will not be repeated here.
[0241] For example, the at least one supplementary three-dimensional dynamic world model can be used to expand the initial three-dimensional dynamic world model.
[0242] Exemplarily, the three-dimensional model generation apparatus can generate a supplementary three-dimensional dynamic world model according to one supplementary video. Alternatively, the three-dimensional model generation apparatus can also generate a supplementary three-dimensional dynamic world model according to multiple supplementary videos.
[0243] It should be noted that the number of supplementary videos used to generate a supplementary three-dimensional dynamic world model is not limited in the present application, and the above is only an exemplary introduction. Hereinafter, the generation of a supplementary three-dimensional dynamic world model from one supplementary video is taken as an example to exemplarily introduce the present application.
[0244] Optionally, the hole of the first picture in the three-dimensional dynamic scene rendered by the supplementary three-dimensional dynamic world model is smaller than the hole of the first picture in the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model. In this way, the hole of the first picture in the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model can be optimized by the supplementary three-dimensional dynamic world model.
[0245] Optionally, the supplementary three-dimensional dynamic world model includes a picture of a first camera pose that is missing in the three-dimensional dynamic scene rendered by the initial three-dimensional dynamic world model. In this way, it is helpful to supplement the picture of the first camera pose for the initial three-dimensional dynamic world model by the supplementary three-dimensional dynamic world model.
[0246] Hereinafter, the multiple implementation manners of S404 are exemplarily introduced through the first implementation manner and the second implementation manner.
[0247] Exemplarily, the at least one supplementary video includes a first supplementary video, and the first supplementary video can be any one of the at least one supplementary video. The at least one supplementary three-dimensional dynamic world model includes a first supplementary three-dimensional dynamic world model, and the first supplementary three-dimensional dynamic world model is a three-dimensional dynamic world model generated based on the first supplementary video.
[0248] In the first implementation manner, as shown in Figure 6 After the three-dimensional model generation apparatus obtains the supplementary video, the hole of the supplementary video is video-completed. Hereinafter, the first implementation manner is exemplarily introduced through S404a-S404b.
[0249] S404a: processing the hole of the first supplementary video to obtain a target first supplementary video.
[0250] S404b: inputting the target first supplementary video into the dynamic generation model to obtain a first supplementary three-dimensional dynamic world model output by the dynamic generation model.
[0251] It should be noted that the related description of S404a-S404b can refer to the description of S2a-S2b above, which will not be repeated here.
[0252] In the implementation, by processing the holes of the supplementary video, the holes of the supplementary video are eliminated or reduced, the holes of the three-dimensional dynamic scene of the supplementary rendered three-dimensional dynamic world model are eliminated or reduced, and the holes of the three-dimensional dynamic scene of the target three-dimensional dynamic world model obtained by fusing are eliminated or reduced.
[0253] In the second implementation, as shown in Figure 6 After the three-dimensional model generation apparatus obtains the supplementary video, the supplementary video is subjected to video depth estimation and the like. The second implementation is exemplarily introduced through S404c-S404d below.
[0254] S404c: target processing is performed on the first supplementary video to obtain auxiliary information of the first supplementary video, and the target processing includes at least one of video depth estimation or dynamic information estimation.
[0255] S404d: the first supplementary video and the auxiliary information of the first supplementary video are input into a dynamic generation model to obtain a first supplementary three-dimensional dynamic world model output by the dynamic generation model.
[0256] In the implementation, the target processing is performed on the supplementary video to obtain the auxiliary information, and the three-dimensional dynamic world model is generated by using the auxiliary information and the supplementary video, so that the comprehensiveness of the information used to generate the supplementary three-dimensional dynamic world model is improved, the picture quality and the picture accuracy of the three-dimensional dynamic scene of the supplementary rendered three-dimensional dynamic world model are improved, and the picture quality and the picture accuracy of the three-dimensional dynamic scene of the target three-dimensional dynamic world model obtained by fusing are improved.
[0257] It should be noted that the related descriptions of S404c-S404d can refer to the descriptions of S2c-S2d, which will not be described here.
[0258] It should be noted that the first implementation and the second implementation can be combined or used alone, and the present application does not limit this.
[0259] S405: fusing at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model to obtain a target three-dimensional dynamic world model.
[0260] Exemplarily, after the three-dimensional model generation apparatus obtains at least one three-dimensional dynamic world model, the at least one three-dimensional dynamic world model and the initial three-dimensional dynamic world model are fused to optimize the initial three-dimensional dynamic world model and obtain a target three-dimensional dynamic world model.
[0261] Exemplarily, the rendering target three-dimensional dynamic world model can obtain a plurality of pictures, which can constitute the three-dimensional dynamic scene indicated by the prompt information. Wherein, the plurality of pictures obtained by the rendering target three-dimensional dynamic world model can have a smaller hole than the plurality of pictures obtained by the rendering initial three-dimensional dynamic world model.
[0262] Exemplarily, the initial three-dimensional dynamic world model can be used to render a smaller range (i.e. a local area) of the three-dimensional dynamic scene indicated by the prompt information. The initial three-dimensional dynamic world model can be gradually expanded by at least one supplementary three-dimensional dynamic world model, so as to optimize the initial three-dimensional dynamic world model into a complete target three-dimensional dynamic world model.
[0263] In one example, the dynamic scene generation system can first obtain a first supplementary three-dimensional dynamic world model, and fuse the first supplementary three-dimensional dynamic world model with the initial three-dimensional dynamic world model to obtain an intermediate three-dimensional dynamic world model 1. Then, the dynamic scene generation system iteratively performs the foregoing steps until the obtained intermediate three-dimensional dynamic world model k meets the user's demand, and takes the intermediate three-dimensional dynamic world model k as the target three-dimensional dynamic world model.
[0264] In another example, the dynamic scene generation system can directly obtain at least one supplementary three-dimensional dynamic world model, and fuse the at least one supplementary three-dimensional dynamic world model with the initial three-dimensional dynamic world model to obtain the target three-dimensional dynamic world model.
[0265] In the following, the various implementation manners of S405 are exemplarily introduced by way A and way B.
[0266] In way A, the three-dimensional model generation device can take the initial fusion result as the target three-dimensional dynamic world model. In the following, way A is exemplarily introduced by S405a.
[0267] In way A, S405 can include the following S405a.
[0268] S405a: Place the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model in the world coordinate system to obtain a first three-dimensional dynamic world model. Wherein, the first three-dimensional dynamic world model is the target three-dimensional dynamic world model.
[0269] In one example, the fusion module of the three-dimensional model generation apparatus may, for example, place the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model in a world coordinate system, fuse the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model through the world coordinate system, and obtain a first three-dimensional dynamic world model. The three-dimensional model generation apparatus may take the first three-dimensional dynamic world model as the target three-dimensional dynamic world model. In this example, the three-dimensional model generation apparatus directly fuses the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model, which helps to improve the diversity of the fusion mode.
[0270] In another example, the three-dimensional model generation apparatus inputs the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model into a fusion model, the fusion model places the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model in a world coordinate system, and obtains a first three-dimensional dynamic world model. In this example, the fusion model fuses the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model, which helps to improve the fusion efficiency and simplify the fusion process.
[0271] In this implementation, by taking the first three-dimensional dynamic world model in the world coordinate system as the target three-dimensional dynamic world model, the efficiency of obtaining the target three-dimensional dynamic world model can be improved.
[0272] In mode B, after obtaining the first three-dimensional dynamic world model, the three-dimensional model generation apparatus may optimize the first three-dimensional dynamic world model and take the optimized first three-dimensional dynamic world model as the target three-dimensional dynamic world model.
[0273] In the following, mode B is exemplarily introduced through S405b to S405d.
[0274] S405b: Place the target three-dimensional dynamic world model and the three-dimensional dynamic world model in a world coordinate system to obtain a first three-dimensional dynamic world model.
[0275] It should be noted that the related description of S405b can refer to the description of S405a above, which will not be repeated here.
[0276] S405c: Project the first three-dimensional dynamic world model to obtain at least one projection video of the first three-dimensional dynamic world model.
[0277] In one example, after obtaining the first three-dimensional dynamic world model, the three-dimensional model generation apparatus projects the first three-dimensional dynamic world model to obtain a plurality of initial projection videos. Then, according to the camera pose of the target video and the camera pose of the plurality of initial projection videos, the three-dimensional model generation apparatus obtains the camera pose of at least one projection video from the plurality of initial projection videos. The camera pose of the at least one projection video is the same as the camera pose of the target video. The target video can include one or more videos of the initial video and the supplementary video. For example, the at least one projection video includes projection video 1 and projection video 2, and the target video includes the initial video and supplementary video 1, wherein the camera pose of the projection video 1 is the same as the camera pose of the initial video, and the camera pose of the projection video 2 is the same as the camera pose of the supplementary video 1.
[0278] In another example, after obtaining the first three-dimensional dynamic world model, the three-dimensional model generation apparatus can project the first three-dimensional dynamic world model according to the camera pose of the target video to obtain at least one projection video.
[0279] S405d: According to the difference between the at least one projection video and the at least one target video, the first three-dimensional dynamic world model is optimized to obtain a target three-dimensional dynamic world model.
[0280] For example, the three-dimensional model generation apparatus determines the difference between two videos with the same camera pose, and optimizes the first three-dimensional dynamic world model according to the difference to obtain a target three-dimensional dynamic world model. For example, the two videos with the same camera pose can include a projection video and a target video with the same camera pose. For example, the two videos with the same camera pose include projection video 1 and initial video, and the two videos with the same camera pose can also include projection video 2 and supplementary video 1.
[0281] For example, the hole of the projection video 2 is larger than that of the supplementary video 1, and the three-dimensional model generation apparatus inputs the supplementary video 1 into the dynamic generation model to obtain a dynamic generation model output second three-dimensional dynamic world model 1. The three-dimensional model generation apparatus fuses the second three-dimensional dynamic world model 1 and the first three-dimensional dynamic world model in the world coordinate system to obtain a fused three-dimensional dynamic world model 1, which can be used as a target three-dimensional dynamic world model. In this way, it is helpful to reduce or eliminate the hole in the target three-dimensional dynamic world model.
[0282] Optionally, the three-dimensional model generation method can further include S405e-405g. Through S405e-405g, the three-dimensional model generation apparatus can optimize the hole of the three-dimensional dynamic scene of the first rendered three-dimensional dynamic world model.
[0283] Exemplarily, the at least one projection video includes a first projection video, the at least one target video includes a first target video, the camera pose of the first projection video is same as the camera pose of the first target video, and there is a difference between the first projection video and the first target video. For example, the difference can be a difference in size of the hole, a difference in number of elements, etc.
[0284] Hereinafter, S405e-S405g are exemplarily introduced through the first projection video.
[0285] S405e: processing the hole of the first projection video to obtain a target first projection video.
[0286] S405f: inputting the target first projection video into the dynamic generation model to obtain a third three-dimensional dynamic world model output by the dynamic generation model.
[0287] It should be noted that the related description of S405e-S405f can refer to the description of S2a-S2b above, and will not be repeated here.
[0288] S405g: fusing the third three-dimensional dynamic world model and the first three-dimensional dynamic world model to obtain a target three-dimensional dynamic world model.
[0289] Exemplarily, the three-dimensional model generation apparatus can place the third three-dimensional dynamic world model and the first three-dimensional dynamic world model in a world coordinate system, thereby fusing the third three-dimensional dynamic world model and the first three-dimensional dynamic world model to obtain the target three-dimensional dynamic world model.
[0290] In this implementation manner, the hole of the first projection video is processed to obtain a target first projection video, and the third three-dimensional dynamic world model generated through the target first projection video is used to repair the first three-dimensional dynamic world model to obtain a target three-dimensional dynamic world model, thereby helping to reduce the hole in the target three-dimensional dynamic world model.
[0291] Optionally, the three-dimensional model generation method can further include S405h-S405i. Through S405h-S405i, the three-dimensional model generation apparatus can improve the comprehensiveness of the information for generating the supplementary three-dimensional dynamic world model.
[0292] S405h: target processing is performed on the first projection video to obtain auxiliary information of the first projection video, and the target processing includes at least one of video depth estimation or dynamic information estimation.
[0293] S405i: inputting the first projection video and the auxiliary information of the first projection video into the dynamic generation model to obtain a third three-dimensional dynamic world model output by the dynamic generation model.
[0294] It should be noted that the related descriptions of S405h-S405i can refer to the descriptions of S2c-S2d, which will not be repeated here.
[0295] It should be noted that the scheme of S405h-S405i and the scheme of S405e-S405f can be used in combination, or can also be used alone, and the present application does not limit this.
[0296] Optionally, in the case that the hole of the first picture of the three-dimensional dynamic scene rendered by supplementally rendering the three-dimensional dynamic world model is smaller than the hole of the first picture of the three-dimensional dynamic scene rendered by rendering the initial three-dimensional dynamic world model, the hole of the first picture of the three-dimensional dynamic scene rendered by rendering the target three-dimensional dynamic world model is smaller than the hole of the first picture of the three-dimensional dynamic scene rendered by rendering the initial three-dimensional dynamic world model.
[0297] In this implementation, by setting the hole of the first picture of the three-dimensional dynamic scene rendered by supplementally rendering the three-dimensional dynamic world model to be smaller than the hole of the first picture of the three-dimensional dynamic scene rendered by rendering the initial three-dimensional dynamic world model, the hole of the first picture of the three-dimensional dynamic scene rendered by rendering the target three-dimensional dynamic world model can be made to be smaller than the hole of the first picture of the three-dimensional dynamic scene rendered by rendering the initial three-dimensional dynamic world model, which helps to reduce the hole on the three-dimensional dynamic scene of the target three-dimensional dynamic world model.
[0298] Optionally, in the case that the three-dimensional dynamic scene rendered by supplementally rendering the three-dimensional dynamic world model includes the first camera pose that is missing in the three-dimensional dynamic scene rendered by rendering the initial three-dimensional dynamic world model, the three-dimensional dynamic scene rendered by rendering the target three-dimensional dynamic world model includes the picture of the first camera pose.
[0299] In this implementation, by setting the three-dimensional dynamic scene rendered by supplementally rendering the three-dimensional dynamic world model to include the first camera pose that is missing in the three-dimensional dynamic scene rendered by rendering the initial three-dimensional dynamic world model, the three-dimensional dynamic scene rendered by rendering the target three-dimensional dynamic world model can be made to include the picture of the first camera pose, which helps to improve the richness and comprehensiveness of the camera poses of the three-dimensional dynamic scene rendered by rendering the target three-dimensional dynamic world model, thereby helping to improve the visual effect of the three-dimensional dynamic scene rendered by rendering the target three-dimensional dynamic world model.
[0300] In this way, by fusing the supplementary three-dimensional dynamic world model and the initial three-dimensional exchange scene, a target three-dimensional dynamic world model is obtained. This helps to reduce the voids on the three-dimensional dynamic scene rendered from the target three-dimensional dynamic world model, ensure the accuracy of the elements on the three-dimensional dynamic scene rendered from the target three-dimensional dynamic world model, improve the content integrity and camera pose diversity of the three-dimensional dynamic scene rendered from the target three-dimensional dynamic world model, thereby helping to improve the display effect of the three-dimensional dynamic scene rendered from the target three-dimensional dynamic world model, and further helping to improve the user's viewing experience.
[0301] It should be noted that the present application does not limit the method for obtaining the target three-dimensional dynamic world model, and the above is only an exemplary introduction.
[0302] Optionally, the three-dimensional model generation method may further include: providing a second page for displaying the three-dimensional dynamic scene rendered from the target three-dimensional dynamic world model. The three-dimensional dynamic scene rendered from the target three-dimensional dynamic world model includes the to-be-generated three-dimensional dynamic scene indicated by the prompt information, at least one element in the to-be-generated three-dimensional dynamic scene, and the dynamic effect of the to-be-generated three-dimensional dynamic scene.
[0303] Exemplarily, as shown in (b) of Figure 5 After the three-dimensional model generation device obtains the target three-dimensional dynamic world model, it can display the image rendered from the target three-dimensional dynamic world model through the second page, so as to present the three-dimensional dynamic scene rendered from the target three-dimensional dynamic world model to the user.
[0304] In this embodiment, the three-dimensional model generation device displays the three-dimensional dynamic scene rendered from the target three-dimensional dynamic world model through the second page, and presents the final three-dimensional dynamic scene to the user, so as to facilitate viewing the effect of the three-dimensional dynamic scene.
[0305] Exemplarily, the three-dimensional model generation device can iteratively optimize the initial three-dimensional dynamic world model through the above solution until the three-dimensional dynamic scene of the optimized rendered three-dimensional dynamic world model meets the user's requirements. The three-dimensional model generation device can determine the three-dimensional dynamic world model that meets the user's requirements as the target three-dimensional dynamic world model.
[0306] For example, the three-dimensional model generation device obtains at least one supplementary video a according to the initial video and the initial three-dimensional dynamic world model, and generates at least one supplementary three-dimensional dynamic world model a1 through the at least one supplementary video a, and optimizes the initial three-dimensional dynamic world model to obtain an optimized supplementary three-dimensional dynamic world model a2. Then, the three-dimensional model generation device can obtain at least one supplementary video b according to the at least one supplementary video a and the supplementary three-dimensional dynamic world model a2, and generate at least one supplementary three-dimensional dynamic world model b1 through the at least one supplementary video b, and optimize the supplementary three-dimensional dynamic world model a2 to obtain an optimized supplementary three-dimensional dynamic world model b2. The three-dimensional model generation device repeats the foregoing steps until the optimized three-dimensional dynamic world model meets the user's requirements.
[0307] In the above embodiment, the initial video is generated based on the prompt information input by the user, and the initial three-dimensional dynamic world model is generated according to the initial video, so that the complete three-dimensional dynamic scene is rendered through the three-dimensional dynamic world model, and the user is provided with the scheme of quickly designing the three-dimensional dynamic scene, the scheme of generating the effect video, and the like in the three-dimensional dynamic scene production process. In addition, the user can obtain the supplementary video through the automatic intelligent trajectory planning or manual roaming of the three-dimensional model generation device, generate the supplementary three-dimensional dynamic scene, and iteratively optimize the initial three-dimensional dynamic scene through the supplementary three-dimensional dynamic scene, so as to further improve the initial three-dimensional dynamic world model, for example, to complete the holes of the initial three-dimensional dynamic world model, and to realize the iterative three-dimensional dynamic scene based on the video generation. In this way, the user's scheme of quickly modifying the three-dimensional dynamic scene in the three-dimensional dynamic scene production process is improved, so that the design can be quickly modified to obtain intuitive actual effects, and the design threshold can be greatly reduced, the design process can be accelerated, the design period can be reduced, and the design cost can be saved. In addition, since the three-dimensional dynamic world model capable of rendering the three-dimensional dynamic scene is automatically generated based on the video, even a user without professional design capability can also design the three-dimensional dynamic scene.
[0308] In addition, since the process of generating the three-dimensional dynamic scene is simple, dynamic effects can be generated for all elements in the three-dimensional scene, rather than being limited to a small number of objects and characters, thereby improving the richness of the three-dimensional dynamic scene. Moreover, the generated dynamic effects are diverse and not limited to basic physical models. In addition, since the initial three-dimensional dynamic world model is optimized through the supplementary three-dimensional dynamic world model, the optimized target three-dimensional dynamic world model can be arbitrarily expanded, and operations such as adjusting the perspective and moving do not affect the integrity of the three-dimensional dynamic scene.
[0309] The above describes the solutions provided by the embodiments of the present application from the method perspective. To implement the above functions, the three-dimensional model generation apparatus includes hardware structures and / or software modules corresponding to the functions. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0310] The embodiments of the present application can divide the three-dimensional model generation apparatus into functional modules according to the above method. For example, the three-dimensional model generation apparatus can include functional modules corresponding to each function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware or software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical functional division. Actual implementation can have another division manner.
[0311] For example, Figure 9 A possible schematic diagram of the three-dimensional model generation apparatus (i.e., the three-dimensional model generation apparatus 900) involved in the above embodiments is shown. The actions performed by the three-dimensional model generation apparatus 900 can be performed by a computing device or by a corresponding software implemented by a computing device. For example, the three-dimensional model generation apparatus can include an information module 901, a generation module 902, an optimization module 903, and a fusion module 904. The information module 901 is configured to obtain prompt information input by a user, wherein the prompt information is used to indicate at least one element and an action performed by a target element in the at least one element. For example, as shown in S401. Figure 4 The generation module 902 is configured to generate an initial video according to the prompt information, and generate an initial three-dimensional dynamic world model according to the initial video, wherein the initial three-dimensional dynamic world model includes a three-dimensional model of the at least one element and a three-dimensional model of the action performed by the target element. For example, as shown in S402. Figure 4 The optimization module 903 is configured to obtain at least one supplementary video according to target initial information, wherein the target initial information includes at least one of the initial three-dimensional dynamic world model or the initial video. For example, as shown in S403. Figure 4 The generation module 902 is further configured to generate at least one supplementary three-dimensional dynamic world model according to the at least one supplementary video. For example, as shown in S404. Figure 4S404. The fusion module 904 is configured to fuse the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model to obtain a target three-dimensional dynamic world model. For example, as shown in S404. Figure 4 S405.
[0312] Optionally, the three-dimensional model generation apparatus further includes a display module 905. The display module 905 is configured to provide a first page, where the first page is configured to display a picture rendered from the initial three-dimensional dynamic world model. The optimization module 903 is specifically configured to: in response to a first instruction input by a user, obtain at least one supplementary video according to target initial information, where the first instruction is used to confirm the initial three-dimensional dynamic world model or is used to indicate any of the initial three-dimensional dynamic world model.
[0313] Optionally, the display module 905 is further configured to: provide a second page, where the second page is configured to display a picture rendered from the target three-dimensional dynamic world model.
[0314] Optionally, the generation module 902 is specifically configured to: input the prompt information into a video generation model to obtain an initial video output by the video generation model. The display module 905 is further configured to: provide a third page, where the third page is configured to display the initial video. The generation module 902 is specifically configured to: in response to a second instruction input by a user, input the initial video into a dynamic generation model to obtain an initial three-dimensional dynamic world model output by the dynamic generation model, where the second instruction is used to indicate confirmation of the initial video or is used to indicate any of generation of the three-dimensional dynamic world model.
[0315] Optionally, the optimization module 903 is specifically configured to: provide a fourth page, where the fourth page is configured to display a picture of at least one view angle rendered from the initial three-dimensional dynamic world model. The optimization module 903 is specifically configured to: obtain the picture of the at least one view angle and render the picture of the at least one view angle to obtain the at least one supplementary video.
[0316] Optionally, the optimization module 903 is specifically configured to: input a camera pose of the initial video and an initial picture rendered by the camera pose of the initial video into a neural network model to obtain at least one supplementary camera pose output by the neural network model and at least one supplementary picture rendered by the at least one supplementary camera pose. The optimization module 903 is specifically configured to: generate at least one supplementary video according to the at least one supplementary camera pose and the at least one supplementary picture.
[0317] Optionally, the generation module 902 is specifically configured to: process a hole in a first supplementary video of the at least one supplementary video to obtain a target first supplementary video. The generation module 902 is specifically configured to: input the target first supplementary video into a dynamic generation model to obtain a first supplementary three-dimensional dynamic world model of the at least one supplementary three-dimensional dynamic world model output by the dynamic generation model.
[0318] Optionally, the fusion module 904 is specifically configured to: input the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model into a fusion model to obtain the target three-dimensional dynamic world model.
[0319] Optionally, the fusion module 904 is specifically configured to: place the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model in a world coordinate system to obtain a first three-dimensional dynamic world model, and the first three-dimensional dynamic world model is the target three-dimensional dynamic world model.
[0320] Optionally, the fusion module 904 is specifically configured to: place the target three-dimensional dynamic world model and the three-dimensional dynamic world model in a world coordinate system to obtain a first three-dimensional dynamic world model; project the first three-dimensional dynamic world model to obtain at least one projection video of the first three-dimensional dynamic world model; optimize the first three-dimensional dynamic world model according to a difference between the at least one projection video and at least one target video to obtain the target three-dimensional dynamic world model; wherein a camera pose of the at least one projection video is the same as a camera pose of the at least one target video, and the at least one target video includes at least one of the at least one supplementary video and the initial video.
[0321] Optionally, the fusion module 904 is specifically configured to: process a hole in a first projection video of the at least one projection video to obtain a target first projection video; input the target first projection video into a dynamic generation model to obtain at least one second three-dimensional dynamic world model output by the dynamic generation model; and fuse the second three-dimensional dynamic world model and the first three-dimensional dynamic world model to obtain the target three-dimensional dynamic world model.
[0322] Optionally, the optimization module 903 is specifically configured to: obtain a plurality of candidate videos according to a plurality of supplementary camera poses and a plurality of supplementary pictures; determine a target candidate video from the plurality of candidate videos, and a hole of the target candidate video is larger than at least part of candidate videos other than the target candidate video in the plurality of candidate videos; provide a fifth page, the fifth page being used to render a picture of a target view angle obtained by the initial three-dimensional dynamic world model, and the picture of the target view angle being a picture of the target candidate video; adjust a view angle of the picture obtained by rendering the initial three-dimensional dynamic world model, and obtain a picture of at least one supplementary view angle; and render the picture of the at least one supplementary view angle to obtain at least one supplementary video.
[0323] Optionally, the optimization module 903 is specifically configured to: construct a distribution field of space bodies of the initial three-dimensional dynamic world model, the space bodies comprising a plurality of three-dimensional points of the initial three-dimensional dynamic world model, the distribution field being configured to indicate a distribution of the plurality of three-dimensional points; determine densities of three-dimensional points in a plurality of directions centered on a position of the camera pose of the initial video in the distribution field; obtain the target camera pose according to a target direction and the camera pose of the initial video, the density of the three-dimensional points in the target direction being less than densities of the three-dimensional points in at least part of the directions other than the target direction; input the target camera pose and an initial picture rendered by the target camera pose into the neural network model to obtain at least one supplementary camera pose and at least one supplementary picture rendered by the at least one supplementary camera pose output by the neural network model; and generate at least one supplementary video according to the at least one supplementary camera pose and the at least one supplementary picture.
[0324] Optionally, the optimization module 903 is specifically configured to: provide a sixth page, the sixth page being configured to display a picture rendered by the target camera pose of the initial three-dimensional dynamic world model; adjust a viewing angle of the picture rendered by the initial three-dimensional dynamic world model and obtain a picture of at least one supplementary viewing angle; and render the picture of the at least one supplementary viewing angle to obtain at least one supplementary video.
[0325] Optionally, the generation module 902 is specifically configured to: input the prompt information into the video generation model to obtain an initial video output by the video generation model; and input the initial video into the dynamic generation model to obtain an initial three-dimensional dynamic world model output by the dynamic generation model.
[0326] Optionally, the generation module 902 is specifically configured to: perform target processing on the initial video to obtain auxiliary information of the initial video, the target processing comprising at least one of video depth estimation or dynamic information estimation; and input the initial video and the auxiliary information of the initial video into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model.
[0327] Optionally, the generation module 902 is specifically configured to: perform processing on a hole of the initial video to obtain a target initial video; and input the target initial video into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model.
[0328] Optionally, the generation module 902 is specifically configured to: perform processing on a hole of the initial video to obtain a target initial video; perform target processing on the target initial video to obtain auxiliary information of the target initial video; and input the target initial video and the auxiliary information of the target initial video into the dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model.
[0329] Optionally, the generating module 902 is specifically configured to: performing target processing on a first supplementary video in the at least one supplementary video to obtain auxiliary information of the first supplementary video, the target processing including at least one of video depth estimation or dynamic information estimation; and inputting the first supplementary video and the auxiliary information of the first supplementary video into the dynamic generation model to obtain a first supplementary three-dimensional dynamic world model in the at least one supplementary three-dimensional dynamic world model output by the dynamic generation model.
[0330] Optionally, the generating module 902 is specifically configured to: processing a hole of a first supplementary video in the at least one supplementary video to obtain a target first supplementary video; performing target processing on the target first supplementary video to obtain auxiliary information of the target first supplementary video; and inputting the target first supplementary video and the auxiliary information of the target first supplementary video into the dynamic generation model to obtain the first supplementary three-dimensional dynamic world model output by the dynamic generation model.
[0331] For specific description of the above optional manners, refer to the foregoing method embodiments, which will not be described here again. In addition, the foregoing explanations and beneficial effect descriptions of any one of the three-dimensional model generation apparatuses 900 provided above can refer to the corresponding method embodiments described above, which will not be described here again.
[0332] In the present application, the information module 901, the generating module 902, the optimizing module 903, the fusing module 904, and the display module 905 can be implemented by software or by hardware. For example, the implementation of the generating module 902 is described as follows. Similarly, the implementation of the information module 901, the optimizing module 903, the fusing module 904, and the display module 905 can refer to the implementation of the generating module 902.
[0333] As an example of a software functional unit, the generating module 902 can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the generating module 902 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region) or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple data centers with similar geographical locations. Generally, one region can include multiple AZs.
[0334] Similarly, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, usually one VPC is set in one region, and communication between two VPCs in the same region and between VPCs in different regions needs to set a communication gateway in each VPC to realize the interconnection between VPCs through the communication gateway.
[0335] As an example of a hardware functional unit, the generation module 902 can include at least one computing device, such as a server, etc. Alternatively, the generation module 902 can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. Among them, the above-mentioned PLD can be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0336] The multiple computing devices included in the generation module 902 can be distributed in the same region or in different regions. The multiple computing devices included in the generation module 902 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the generation module 902 can be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs, etc.
[0337] It should be noted that in other embodiments, the information module 901 can be configured to perform any of the steps of the three-dimensional model generation method, the generation module 902 can be configured to perform any of the steps of the three-dimensional model generation method, the optimization module 903 can be configured to perform any of the steps of the three-dimensional model generation method, the fusion module 904 can be configured to perform any of the steps of the three-dimensional model generation method, and the display module 905 can be configured to perform any of the steps of the three-dimensional model generation method. The steps responsible for the information module 901, the generation module 902, the optimization module 903, the fusion module 904, and the display module 905 can be specified as needed, and the entire function of the three-dimensional model generation device can be achieved by the information module 901, the generation module 902, the optimization module 903, the fusion module 904, and the display module 905 respectively implementing different steps of the three-dimensional model generation method.
[0338] The present application also provides a computing device 1000. As shown in Figure 10 The computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate through the bus 1002. The computing device 1000 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1000.
[0339] The bus 1002 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 In the present application, only one line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus. The bus 1002 can include a path for transmitting information between various components (e.g., the memory 1006, the processor 1004, the communication interface 1008) of the computing device 1000.
[0340] The processor 1004 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0341] The memory 1006 can include volatile memory (volatile memory), such as random access memory (RAM). The processor 1004 can also include non-volatile memory (non-volatile memory), such as read-only memory (ROM), flash memory, a mechanical hard disk drive (HDD) or a solid state drive (SSD).
[0342] The executable program code stored in the memory 1006 is executed by the processor 1004 to realize the functions of the aforementioned information module 901, generation module 902, optimization module 903, fusion module 904, display module 905, respectively, so as to realize the three-dimensional model generation method. That is, the memory 1006 has instructions for executing the three-dimensional model generation method.
[0343] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card, a transceiver, to realize the communication between the computing device 1000 and other devices or communication networks.
[0344] For example, the above-mentioned computing device 1000 can be Figure 1 The terminal device shown, Figure 2 The computing device and the terminal device shown.
[0345] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a desktop computer, a notebook computer, or a terminal device such as a smart phone.
[0346] As Figure 11 shown, the computing device cluster 1100 includes at least one computing device 1000. The memory 1006 in one or more computing devices 1000 in the computing device cluster 1100 can store the same instructions for executing the three-dimensional model generation method.
[0347] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster can also store partial instructions for executing the three-dimensional model generation method, respectively. In other words, the combination of one or more computing devices 1000 can collectively execute the instructions for executing the three-dimensional model generation method.
[0348] It should be noted that the memories 1006 in different computing devices 1000 in the computing device cluster can store different instructions for respectively performing part of the functions of the cloud scheduling platform. That is, the memories 1006 in different computing devices 1000 store instructions that can implement the functions of one or more of the information module 901, the generation module 902, the optimization module 903, the fusion module 904, and the display module 905.
[0349] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc. Figure 12 A possible implementation manner is shown. As shown in Figure 12 Two computing devices 1000A and 1000B are connected through a network. Specifically, the computing devices are connected to the network through the communication interfaces in the computing devices.
[0350] In this type of possible implementation manner, the memory 1006 in the computing device 1000A stores instructions for performing the functions of the information module 901, the generation module 902, and the optimization module 903. Meanwhile, the memory 1006 in the computing device 1000B stores instructions for performing the functions of the fusion module 904 and the display module 905.
[0351] Figure 12 The connection manner between the computing device cluster shown in
[0352] It should be understood that Figure 12 The functions of the computing device 1000A shown in
[0353] The functions of the computing device 1000B can also be completed by multiple computing devices 1000. Figure 11 Figure 12 The present application also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection manner of the computing device cluster shown in
[0354] In some possible implementation manners, the memory 1006 of one or more computing devices 1000 in the computing device cluster can also respectively store partial instructions for performing the three-dimensional model generation method. In other words, the combination of one or more computing devices 1000 can collectively execute the instructions for performing the three-dimensional model generation method.
[0355] The present application also provides a processor, which can be used to execute the above method.
[0356] The present application also provides a chip, which includes a processor and a power supply circuit; the power supply circuit can be used to supply power for the processor; and the processor can be used to execute the above method.
[0357] The present application also provides a computing device, which can include a processor, a memory, and a computer program / instruction stored on the memory; the processor executes the computer program / instruction to enable the computing device to implement the above method.
[0358] The present application also provides a computing device cluster, which includes at least one computing device; each of the at least one computing device includes a processor, a memory, and a computer program / instruction stored on the memory; the processor of each computing device executes the computer program / instruction stored in the memory of each computing device to enable each computing device to implement the above method.
[0359] The present application also provides a computer program product. The computer program product includes computer program / instruction, which can be run on a computing device or stored in any available medium or software or program product. When the computer program / instruction is executed on at least one computing device, the at least one computing device can execute the above method.
[0360] The present application also provides a computer readable storage medium, which can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The computer readable storage medium stores computer program / instruction, which, when executed on at least one computing device, enables the at least one computing device to execute the above method.
[0361] For example, the available medium can be a magnetic medium (for example, a floppy disk, a magnetic disk, a magnetic tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid state drive (SSD)), etc.
[0362] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the same; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can still be modified, or some technical features therein can be replaced by equivalents; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A three-dimensional model generation method characterized by comprising: The method comprises: obtaining prompt information input by a user, wherein the prompt information is used to indicate at least one element and an action performed by a target element in the at least one element; generating an initial video according to the prompt information, and generating an initial three-dimensional dynamic world model according to the initial video, wherein the initial three-dimensional dynamic world model comprises a three-dimensional model of the at least one element and a three-dimensional model of the target element performing the action; obtaining at least one supplementary video according to target initial information, wherein the target initial information comprises at least one of the initial three-dimensional dynamic world model or the initial video; generating at least one supplementary three-dimensional dynamic world model according to the at least one supplementary video; fusing the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model to obtain a target three-dimensional dynamic world model.
2. The method of claim 1, wherein, The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises:
3. The method according to claim 1 or 2, characterized in that, providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises:
4. The method according to any one of claims 1 to 3, characterized in that, providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises:
5. The method according to any one of claims 1-4, characterized in that, providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model.
6. The method according to any one of claims 1-4, characterized in that, The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises:
7. The method according to any one of claims 1 to 6, characterized in that, providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: providing a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model. The method further comprises: provid process a hole in a first supplementary video of the at least one supplementary video to obtain a target first supplementary video; input the target first supplementary video into a dynamic generation model to obtain a first supplementary three-dimensional dynamic world model of the at least one supplementary three-dimensional dynamic world model output by the dynamic generation model.
8. A three-dimensional model generation apparatus characterized by comprising: The three-dimensional model generation apparatus comprises: an information module configured to acquire prompt information input by a user, wherein the prompt information is used to indicate at least one element and an action performed by a target element in the at least one element; a generation module configured to generate an initial video according to the prompt information, and generate an initial three-dimensional dynamic world model according to the initial video, wherein the initial three-dimensional dynamic world model comprises a three-dimensional model of the at least one element and a three-dimensional model of the action performed by the target element; an optimization module configured to acquire at least one supplementary video according to target initial information, wherein the target initial information comprises at least one of the initial three-dimensional dynamic world model or the initial video; the generation module is further configured to generate at least one supplementary three-dimensional dynamic world model according to the at least one supplementary video; a fusion module configured to fuse the at least one supplementary three-dimensional dynamic world model and the initial three-dimensional dynamic world model to obtain a target three-dimensional dynamic world model.
9. The apparatus of claim 8, wherein, The three-dimensional model generation apparatus further comprises a display module; the display module is further configured to provide a first page, wherein the first page is used to display a picture rendered by the initial three-dimensional dynamic world model; the optimization module is specifically configured to acquire the at least one supplementary video according to the target initial information in response to a first instruction input by the user, wherein the first instruction is used to confirm the initial three-dimensional dynamic world model or to indicate supplementing any one of the initial three-dimensional dynamic world model.
10. The apparatus of claim 8 or 9, wherein, The three-dimensional model generation apparatus further comprises a display module; the display module is further configured to provide a second page, wherein the second page is used to display a picture rendered by the target three-dimensional dynamic world model.
11. The apparatus of any one of claims 8-10, wherein, The three-dimensional model generation apparatus further comprises a display module; the generation module is specifically configured to input the prompt information into a video generation model to obtain the initial video output by the video generation model; the display module is further configured to provide a third page, wherein the third page is used to display the initial video; the generation module is specifically configured to input the initial video into a dynamic generation model to obtain the initial three-dimensional dynamic world model output by the dynamic generation model in response to a second instruction input by the user, wherein the second instruction is used to indicate confirming the initial video or to indicate generating a three-dimensional dynamic world model.
12. The apparatus of any one of claims 8-11, wherein, The three-dimensional model generation apparatus further comprises a display module; the display module is further configured to provide a fourth page, wherein the fourth page is used to display a picture of at least one view angle rendered by the initial three-dimensional dynamic world model; the optimization module is specifically configured to acquire the picture of the at least one view angle, and render the picture of the at least one view angle to obtain the at least one supplementary video.
13. The apparatus of any one of claims 8-11, wherein, The optimization module is specifically used for: inputting the camera pose of the initial video and the initial picture rendered by the camera pose of the initial video into a neural network model, to obtain at least one supplementary camera pose output by the neural network model and at least one supplementary picture rendered by the at least one supplementary camera pose; generating the at least one supplementary video according to the at least one supplementary camera pose and the at least one supplementary picture.
14. The apparatus of any one of claims 8-13, wherein, The generation module is specifically used for: processing a hole in a first supplementary video in the at least one supplementary video to obtain a target first supplementary video; inputting the target first supplementary video into a dynamic generation model to obtain a first supplementary three-dimensional dynamic world model in the at least one supplementary three-dimensional dynamic world model output by the dynamic generation model.
15. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device; each computing device in the at least one computing device includes a processor, a memory, and a computer program / instruction stored on the memory; the processor of each computing device executes the computer program / instruction stored in the memory of each computing device, so that each computing device implements the method in any one of claims 1-7.
16. A computer program product, characterized in that, the computer program product includes a computer program / instruction, and when the computer program / instruction is executed by a computing device, the computing device implements the method in any one of claims 1-7.
17. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program / instruction, and when the computer program / instruction is executed by a computing device, the computing device implements the method in any one of claims 1-7.