Method and device for updating three-dimensional image

By acquiring images of two-dimensional images and three-dimensional models, and using a video generation model to process and generate a surround video with an updated three-dimensional model, the problem of low efficiency in three-dimensional image updates in existing technologies is solved, and efficient three-dimensional image updates are achieved.

WO2025260712A1PCT designated stage Publication Date: 2025-12-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/070689
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-27
Filing Date
2025-01-06
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

In existing technologies, updating 3D images requires professionals to use specialized equipment to collect data on changed areas, resulting in low update efficiency.

Method used

By acquiring images of two-dimensional and three-dimensional models, and processing them using a video generation model, an updated surround video of the three-dimensional model is generated, thus updating the three-dimensional model and avoiding the need to rebuild a new three-dimensional model.

Benefits of technology

It reduces the workload and complexity of updating 3D images, improves update efficiency, and enhances the reliability and realism of updates through auxiliary information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070689_26122025_PF_FP_ABST
    Figure CN2025070689_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a method and device for updating a three-dimensional image, for use in improving the efficiency of updating three-dimensional images. The method comprises: acquiring a first two-dimensional image and a second image of a three-dimensional model, the first image comprising a region to be updated in the three-dimensional model; acquiring a surround video of the three-dimensional model on the basis of the first image and the second image; inputting the first image and the surround video of the three-dimensional model into a video generation model to obtain an updated surround video of the three-dimensional model; and updating the second image on the basis of the updated surround video of the three-dimensional model.
Need to check novelty before this filing date? Find Prior Art

Description

A method and apparatus for updating three-dimensional images

[0001] This application claims priority to Chinese Patent Application No. 202410788946.2, filed on June 18, 2024, entitled "A Method and Apparatus for Updating a 3D Urban Base Plate", and Chinese Patent Application No. 202410788946.4, filed on August 27, 2024, entitled "A Method and Apparatus for Updating a Three-Dimensional Image", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence (AI), and more particularly to a method and apparatus for updating three-dimensional images. Background Technology

[0003] Real-world 3D refers to the virtual 3D digital representation of real-world scenes using technologies such as remote sensing mapping, big data, cloud computing, and intelligent sensing, reflecting the real world through three-dimensional images. Based on real-world 3D, the previously abstract two-dimensional image representation is transformed into a three-dimensional and intuitive one, realizing digital twin reality, with broad application scenarios.

[0004] As the real world changes, real-world 3D images also need to be modified accordingly. Related technical solutions require professionals to use specialized equipment (such as drones) to collect data on the changed areas, and then reconstruct updated 3D images of the scene based on the new data.

[0005] This technical solution requires professionals to use specialized equipment to collect data on the changing areas. The data collection requirements are high and the difficulty is great, which reduces the efficiency of updating the 3D image. Summary of the Invention

[0006] This application provides a method and apparatus for updating a three-dimensional image. In the method for updating a three-dimensional image, the data that needs to be re-acquired when updating the three-dimensional image is the two-dimensional image of the area to be updated. By processing the two-dimensional image and the image of the original three-dimensional model, the image of the three-dimensional model can be updated without having to start from scratch and reconstruct the image of the new three-dimensional model. This reduces the workload and complexity of updating three-dimensional images and improves the efficiency of updating three-dimensional images.

[0007] In a first aspect, this application provides a method for updating a three-dimensional image, comprising:

[0008] A 2D first image and a 3D model second image are acquired. The first image includes the region to be updated in the 3D model. In other words, the second image is a 3D image showing the 3D model, and it contains the region to be updated. The first and second images are processed to obtain a surround video of the 3D model. The first image and the surround video of the 3D model are input into a video generation model to obtain an updated surround video of the 3D model. In the updated surround video, the region to be updated has been updated; this is referred to as the updated region. The updated surround video can be understood as a video captured around the updated region as the center, taking 360° shots. Based on the updated surround video of the 3D model, the second image is updated. In the updated second image, the region to be updated has been updated.

[0009] In this application, when updating a 3D image, the data that needs to be re-acquired is the 2D image of the area to be updated. By processing the 2D image and the original 3D model image, the 3D model image can be updated without having to start from scratch and reconstruct a new 3D model image. This reduces the workload and complexity of 3D image updates and improves the efficiency of 3D image updates.

[0010] In some optional implementations of the first aspect, the video generation model includes multiple neural networks. These neural networks process the first image and the surround video of the 3D model input to the video generation model to obtain the surround video updated with the 3D model. Specifically, the first image and the surround video of the 3D model are input into the first neural network of the video generation model to generate image features corresponding to the first image and video features corresponding to the surround video of the 3D model. The image features and video features are then input into the attention network of the video generation model to obtain fused features. The fused features are then input into the second neural network of the video generation model to obtain the surround video updated with the 3D model.

[0011] In some optional implementations of the first aspect, textual information can also be obtained. This textual information describes the attribute information of at least one of the first image, the second image, or the 3D model. The attribute information includes the style of the image, the style of the 3D model, and update information of the 3D model, serving as auxiliary information to make the updated surround video of the 3D model more realistic. In the scheme of obtaining textual information, the input to the video generation model includes textual information in addition to the first image, the surround video of the 3D model, and the textual information itself.

[0012] In this application, the input to the video generation model also includes text information, which serves as auxiliary information to describe the attribute information of at least one of the first image, the second image, or the three-dimensional model, making the surround video of the updated three-dimensional model closer to the real situation and improving the reliability of the three-dimensional image update.

[0013] In some optional implementations of the first aspect, in the scheme for obtaining text information, the first image, the surround video of the 3D model, and the text information are input into a video generation model to obtain a surround video updated with the 3D model. Specifically, the first image, the surround video of the 3D model, and the text information are input into the first neural network of the video generation model to generate image features corresponding to the first image, video features corresponding to the surround video of the 3D model, and text features corresponding to the text information. Then, the image features, video features, and text features are input into the attention network of the video generation model to obtain fused features. Finally, the fused features are input into the second neural network of the video generation model to obtain a surround video updated with the 3D model.

[0014] In some optional implementations of the first aspect, obtaining a surround video of the 3D model based on the first image and the second image includes: determining the first pose of the first image relative to the 3D model based on the first image and the second image; and then generating a surround video of the 3D model based on the first pose and the first image. Specifically, the first pose is used as the pose of a certain frame, and the surround video of the 3D model is constructed with the region to be updated as the surround center.

[0015] In this application, a two-dimensional first image and a three-dimensional second image are processed to obtain the first pose of the first image relative to the three-dimensional model, thus realizing pose transformation. Therefore, the surround video of the three-dimensional model generated based on the first pose and the first image more closely resembles the actual situation of a real three-dimensional model, further improving the reliability of the technical solution.

[0016] In several possible implementations of the first aspect, determining the first pose of the first image relative to the 3D model based on the first image and the second image can be done in several ways. Optionally, multiple rendered images corresponding to multiple viewpoints can be obtained for the second image. Then, based on the similarity between the first image and the multiple rendered images, the target rendered image with the highest similarity to the first image is determined from the multiple rendered images. The second pose of the first image relative to the target rendered image is then determined. Since the pose of the target rendered image in the 3D model is fixed, the first pose of the first image relative to the 3D model can be determined based on the second pose.

[0017] In this application, the first pose of the first image relative to the target rendered image is determined based on the second pose of the first image relative to the target rendered image. Since the first image is the image with the highest similarity to the second image among multiple rendered images from multiple viewpoints, the second pose is close to the first pose, making the calculation of the first pose simpler.

[0018] In some alternative implementations of the first aspect, the first pose of the first image relative to the 3D model can also be determined in other ways. Optionally, the first image and the second image can be input into a third neural network to obtain the first pose of the first image relative to the 3D model.

[0019] In this application, there are multiple possibilities for determining the first pose of the first image relative to the 3D model, which enriches the implementation methods and application scenarios of the technical solution of this application and improves the flexibility of the solution.

[0020] Among some alternative implementations of the first aspect, the three-dimensional model includes: a city three-dimensional model, a rural three-dimensional model, a scenic area three-dimensional model, or a natural landscape three-dimensional model.

[0021] In this application, the three-dimensional model has multiple possibilities, which enriches the implementation methods and application scenarios of the technical solution and enhances the practical value of the solution.

[0022] Secondly, embodiments of this application provide a method for training a video generation model, including:

[0023] The first image and the surround video of the 3D model are input into the first neural network of the video generation model to generate image features corresponding to the first image and video features corresponding to the surround video of the 3D model. The first image is a 2D image that includes the region of the 3D model to be updated. The surround video of the 3D model is the surround video before the 3D model is updated. The image features and video features are input into the attention network of the video generation model to obtain fused features. The fused features are input into the second neural network of the video generation model to obtain the surround video after the 3D model is updated. The video generation model is iteratively trained, updating at least one of the parameters of the first neural network, the attention network, or the second neural network, until the video generation model converges or reaches a preset number of iterations. In the surround video after the 3D model is updated, the region of the 3D model to be updated has been updated. The surround video after the 3D model is updated is used to update the second image of the 3D model.

[0024] In this application, the output of the video generation model is used to update 3D images. The input of the video generation model is a 2D image of the region to be updated and a surround video of the original 3D model. Based on this, an updated surround video of the 3D model can be obtained without having to reconstruct a new 3D model image from scratch, thus reducing the workload and complexity of 3D image updates and improving the efficiency of 3D image updates.

[0025] In some optional implementations of the second aspect, textual information can also be obtained, which is used to describe the attribute information of at least one of the first image, the second image, or the 3D model. The input to the first neural network also includes textual information, and textual features are obtained after passing through the first neural network. The input to the attention network also includes textual features, that is, the attention network processes textual features, image features, and video features to obtain fused features.

[0026] In this application, the input to the video generation model also includes text information, which serves as auxiliary information to describe the attribute information of at least one of the first image, the second image, or the three-dimensional model, making the surround video of the updated three-dimensional model closer to the real situation and improving the reliability of the three-dimensional image update.

[0027] Thirdly, this application provides a three-dimensional image updating device, comprising:

[0028] The acquisition unit is used to acquire a two-dimensional first image and a three-dimensional model second image, the first image including the region to be updated in the three-dimensional model. Based on the first and second images, a surround video of the three-dimensional model is acquired.

[0029] The processing unit is used to input the first image and the surround video of the 3D model into the video generation model to obtain the surround video updated with the 3D model. Based on the surround video updated with the 3D model, the second image is updated.

[0030] The three-dimensional image updating device is used to implement the method shown in the first aspect or any possible implementation of the first aspect, as detailed above, and will not be repeated here.

[0031] Fourthly, this application provides a training device, comprising:

[0032] The processing unit is configured to input a 2D first image and a surround video of a 3D model into a first neural network of a video generation model, generating image features corresponding to the first image and video features corresponding to the surround video of the 3D model. The image features and video features are then input into an attention network of the video generation model to obtain fused features. The fused features are then input into a second neural network of the video generation model to obtain the surround video updated with the 3D model. The video generation model is iteratively trained, updating at least one of the parameters of the first neural network, the attention network, or the second neural network, until the video generation model converges or reaches a preset number of iterations. The first image includes the region of the 3D model to be updated.

[0033] The training device is used to implement the method shown in the second aspect above, or any possible implementation of the second aspect, as detailed above, and will not be repeated here.

[0034] Fifthly, this application provides a computer device including a processor and a memory, wherein the processor stores instructions that, when executed on the processor, implement the methods shown in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.

[0035] In a sixth aspect, this application provides a computer program product containing instructions that, when executed on a processor, implement the methods shown in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect; or, when the instructions are run by a cluster of computer devices, cause the cluster of computer devices to implement the methods shown in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.

[0036] In a seventh aspect, this application provides a computer-readable storage medium storing computer program instructions that, when executed on a processor, implement the methods shown in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect; or, when executed by a cluster of computer devices, cause the cluster of computer devices to implement the methods shown in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.

[0037] Eighthly, this application provides a chip including at least one processor and a communication interface, the communication interface and the at least one processor being interconnected via a circuit, the at least one processor being used to execute computer programs or instructions to perform the methods shown in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.

[0038] In some possible implementations of the eighth aspect, the chip described above in this application further includes at least one memory storing instructions. This memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).

[0039] The beneficial effects shown in any of the fifth to eighth aspects are similar to those of the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect, and will not be repeated here. Attached Figure Description

[0040] Figure 1 is a schematic diagram of the artificial intelligence architecture provided in an embodiment of this application;

[0041] Figure 2 is a schematic diagram of a system architecture provided in an embodiment of this application;

[0042] Figure 3 is a schematic diagram of another system architecture provided in an embodiment of this application;

[0043] Figure 4 is a flowchart illustrating a training method for a video generation model provided in an embodiment of this application;

[0044] Figure 5 is a schematic diagram of a video generation model provided in an embodiment of this application;

[0045] Figure 6 is a flowchart illustrating a method for updating a three-dimensional image provided in an embodiment of this application;

[0046] Figure 7 is another flowchart illustrating the three-dimensional image updating method provided in an embodiment of this application;

[0047] Figure 8 is another flowchart illustrating the three-dimensional image updating method provided in an embodiment of this application;

[0048] Figure 9 is a structural schematic diagram of a training device provided in an embodiment of this application;

[0049] Figure 10 is a schematic diagram of a three-dimensional image updating device provided in an embodiment of this application;

[0050] Figure 11 is a schematic diagram of a computing device provided in an embodiment of this application;

[0051] Figure 12 is a schematic diagram of a computing device cluster provided in an embodiment of this application;

[0052] Figure 13 is another schematic diagram of the computing device cluster provided in an embodiment of this application;

[0053] Figure 14 is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation

[0054] This application provides a method and apparatus for updating a three-dimensional image. In the method for updating a three-dimensional image, the data that needs to be re-acquired when updating the three-dimensional image is the two-dimensional image of the area to be updated. By processing the two-dimensional image and the image of the original three-dimensional model, the image of the three-dimensional model can be updated without having to start from scratch and rebuild the image of the new three-dimensional model. This reduces the workload and complexity of updating the three-dimensional image and improves the efficiency of updating the three-dimensional image.

[0055] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0056] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses. Additionally, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can be expressed as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0057] First, please refer to Figure 1, which is a schematic diagram of the artificial intelligence architecture provided in the embodiment of this application. The schematic diagram describes the overall workflow of the artificial intelligence system and is applicable to general artificial intelligence field needs.

[0058] The above-mentioned artificial intelligence framework will be elaborated from two dimensions: "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis).

[0059] The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it could be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom."

[0060] The "IT value chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence, information (provided and processed by technology) to the industrial ecosystem of systems.

[0061] (1) Infrastructure.

[0062] The infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, including central processing units (CPUs), neural network processing units (NPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs) and other hardware acceleration chips. The basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.

[0063] (2) Data.

[0064] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0065] (3) Data processing.

[0066] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.

[0067] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.

[0068] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.

[0069] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.

[0070] (4) General ability.

[0071] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0072] (5) Smart products and industry applications.

[0073] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They are the encapsulation of overall artificial intelligence solutions, productizing intelligent information decision-making and realizing practical applications. Their application areas mainly include: intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, smart cities, and intelligent terminals.

[0074] Please refer to Figure 2 below, which is a system architecture diagram provided in an embodiment of this application. In the embodiment shown in Figure 2, the image processing system includes an execution device 210, a training device 220, a database 230, a terminal device 240, a data storage system 250, and a data acquisition device 260. The execution device 210 includes a computing module 211 and an input / output (I / O) interface 212.

[0075] Data acquisition device 260 is used to acquire a large-scale training dataset and store the training dataset in database 230. Training device 220 trains the model 201 constructed in this application based on the training dataset maintained in database 230. The trained model 201 runs on execution device 210. Execution device 210 can access data, code, etc. in data storage system 250, and can also store data, code, etc. in data storage system 250.

[0076] Regarding the three-dimensional image update method in this embodiment, the data acquired by the data acquisition device 260 can be a two-dimensional first image and a surround video of the three-dimensional model. The execution device 210 can acquire the first image and the surround video of the three-dimensional model stored in the database 230 or the data storage system 250, or input by the terminal device 240 through the I / O interface 212, and process them to obtain the surround video of the three-dimensional model updated. Then, the execution device 210 can send the surround video of the three-dimensional model updated to the database 230 through the I / O interface 212, or it can send the surround video of the three-dimensional model updated to the terminal device 240 through the I / O interface 212, and the terminal device 240 can send the feature map to the database 230.

[0077] Training device 220 can acquire the updated surround video of the 3D model from database 230 as training data for training model 201 until model 201 converges or reaches a preset number of training iterations, thus completing training. Execution device 210 can also update the second image based on the updated surround video of the 3D model. The updated second image can be output to terminal device 240 via I / O interface 212, or directly displayed on the display interface of terminal device 240.

[0078] It should be noted that Figure 2 is merely a schematic diagram of a system architecture provided in an embodiment of this application, and does not constitute any limitation on the positional relationship between various devices, components, modules, etc. For example, in the embodiment shown in Figure 2, the data storage system 250 is an external memory relative to the execution device 210. In some optional embodiments, the data storage system 250 may also be integrated into the execution device 210. In the embodiment shown in Figure 2, the terminal device 240 is an external device relative to the execution device 210. In some optional embodiments, the terminal device 240 may also be integrated into the execution device 210; the specifics are not limited here.

[0079] In some optional implementations, the three-dimensional image updating method provided in this application can also be applied to the field of cloud computing. Referring to Figure 3, the system architecture of cloud computing is illustrated. Figure 3 is a schematic diagram of the system architecture provided in an embodiment of this application.

[0080] As shown in Figure 3, the tenant logs into the cloud service platform 303 via client 301 through the Internet 302 using the account and password registered on the cloud service platform 303. The cloud service platform 303 manages the infrastructure, which includes multiple data centers located in different regions. For example, region 1 in Figure 1 includes cloud data center 1 and cloud data center 2, and region 2 includes cloud data center 3 and cloud data center 4. Each cloud data center has multiple servers running business instances (including at least one of virtual machines, containers, and dedicated servers).

[0081] In this embodiment, a 3D image update service is deployed in the business instance. Tenants can purchase or use this cloud service through a client on the cloud service platform 303. Specifically, the tenant sends a call request to the cloud service platform 303, which requests the cloud service, i.e., requests an update to the 3D image. The specific implementation of this service will be described later.

[0082] Please refer to Figure 4, which is a flowchart illustrating the training method of the video generation model provided in this embodiment. In this embodiment, the video generation model is trained by a training device, including:

[0083] 401. Input the first image and the surround video of the 3D model into the first neural network of the video generation model to generate image features corresponding to the first image and video features corresponding to the surround video of the 3D model.

[0084] In this embodiment, the video generation model includes a first neural network, a second neural network, and an attention model. The video generation model is used to generate a surround video with an updated 3D model. This updated surround video is used to update a second image, which is an image of the unupdated 3D model.

[0085] The input to the video generation model, which is also the input to the first neural network, includes a first image and a surround video of the 3D model. The first image is a two-dimensional image and includes the region of the 3D model that needs to be updated. In other words, the first image indicates the region in the 3D model that needs updating. The surround video of the 3D model refers to the surround video of the 3D model before it has been updated.

[0086] The first neural network includes a neural network image encoder and a neural network video encoder. The neural network image encoder encodes the first image into image features and encodes the surround video of the 3D model into video features. The image features and video features are the outputs of the first neural network.

[0087] 402. Input the image features and video features into the attention network of the video generation model to obtain fused features.

[0088] The output of the first neural network serves as the input to the attention network. The attention network processes image and video features using an attention mechanism to obtain fused features. In other words, the output of the attention network is the fused feature.

[0089] The attention network can be of various types, including cross attention networks, self attention networks, and other types of attention mechanisms. No specific restrictions are made here.

[0090] 403. Input the fused features into the second neural network of the video generation model to obtain the surround video after the 3D model is updated.

[0091] The output of the attention network serves as the input to the second neural network. The second neural network processes the fused features to obtain the surround video with the updated 3D model. In the surround video with the updated 3D model, the 3D model is the updated version.

[0092] Optionally, the fused features can be input into the noise of the second neural network, and the image features or video features can be input into the control condition network of the second neural network.

[0093] Optionally, image features or video features can be input into the noise of the second neural network, and the fused features can be input into the control condition network of the second neural network.

[0094] Optionally, the second neural network can be a diffusion transformer (DiT) network or other neural networks capable of generating surround video; no specific limitation is made here.

[0095] 404. Iteratively train the video generation model, updating the parameters of the first neural network, the attention network, and the second neural network until the video generation model converges or reaches the preset number of iterations.

[0096] The training device iteratively trains the video generation model so that the surround video updated by the 3D model output by the video generation model has the same shooting pose as the surround video of the input 3D model, and the surround video output by the video generation model is identical to or approximately similar to the real surround video to a threshold. This objective can be achieved by updating the parameters of the first neural network, the attention network, and the second neural network, i.e., achieving convergence of the video generation model or reaching a preset number of iterations.

[0097] Optionally, during training, it's also necessary to ensure that the surround video of the updated 3D model output by the video generation model has the same or similar visual style as the surround video of the input 3D model. Visual style can be evaluated from multiple perspectives, such as color, lighting, and shooting season. In simpler terms, the output video should have a consistent style with the input video, and the output video should not appear jarring compared to the input video.

[0098] In the embodiment shown in Figure 4, the input to the first neural network includes a first image and a surround video of the 3D model. In practical applications, the input to the first neural network may also include text information. Please refer to Figure 5, which is a schematic diagram of the video generation model provided in an embodiment of this application.

[0099] As shown in Figure 5, the first image, the surround video of the 3D model, and text information are input into the first neural network of the video generation model. The text information is used to describe the attribute information of at least one of the first image, the 3D model, or the second image, such as the style of the image or the 3D model, the update information of the 3D model, etc. The text information serves as auxiliary information, making the surround video updated by the 3D model closer to reality. The surround video of the first image and the 3D model is similar to step 401 above, and will not be repeated here.

[0100] In addition to the neural network image encoder and neural network video encoder mentioned earlier, the first neural network can also include a neural network text encoder. The neural network text encoder encodes text information into text features. Therefore, the output of the first neural network includes not only image features and video features, but also text features.

[0101] An attention network processes image features, video features, and text features using an attention mechanism to obtain fused features. These fused features are then input into a second neural network to obtain a surround video with an updated 3D model.

[0102] Similarly, the training device iteratively trains the video generation model shown in Figure 5, updating at least one of the parameters of the first neural network, the attention network, or the second neural network, until the video generation model converges or reaches the preset number of iterations.

[0103] The video generation model described in the embodiments shown in Figure 4 or Figure 5 above is applied in the three-dimensional image updating method provided in the embodiments of this application. Please refer to Figure 6 below, which is a flowchart illustrating the three-dimensional image updating method provided in the embodiments of this application.

[0104] 601. Obtain a first two-dimensional image and a second three-dimensional image of the model, wherein the first image includes the region to be updated in the three-dimensional model.

[0105] The real-world scene is constructed in 3D to obtain a 3D model and a second image. In other words, the second image is 3D, providing a three-dimensional and intuitive representation of the real-world scene. The first image, on the other hand, is 2D, displaying the area of ​​the 3D model to be updated, or, in other words, an image obtained after capturing the updated real-world scene. The first image can be captured using ordinary image acquisition devices, such as cameras, terminals equipped with camera lenses, or road traffic cameras, or it can be captured using devices with image acquisition capabilities, such as drones or satellites; the specific method is not limited here.

[0106] For example, suppose the second image shows a 3D model of a real city scene. When the second image was constructed, there was no school in the eastern part of the city. Later, with urban planning, a school was built in the eastern part of the city. Therefore, the first image is an image taken of that school. The first image could be taken from a bird's-eye view or other angles; no specific limitation is made here. Thus, the first image indicates that the area to be updated in the 3D model of the second image is the eastern part of the city.

[0107] In other words, the first image shows the updated cityscape, while the second image shows the 3D model of the cityscape before the update. The purpose of the 3D image updating method provided in this application is to update the 3D model image based on the updated cityscape image.

[0108] The 3D image updating device can acquire the first image in response to user operations, such as user-uploaded first image, or it can acquire the first image through external devices, such as hard drives or USB flash drives. The second image of the 3D model can be understood as the initial image or a historical image of the 3D model. Correspondingly, after the 3D model is constructed or updated, it is stored internally or externally within the 3D image updating device. The 3D image updating device can obtain the storage location of the second image, thereby acquiring the second image.

[0109] It should be noted that, in the embodiments of this application, the 3D model can be of various types, including: a city 3D model, a rural 3D model, a scenic area 3D model, or a natural landscape 3D model. Accordingly, the 3D image updating method provided in the embodiments of this application can be applied to various scenarios, such as digital twin cities, cloud-based scenic area tours, and digital rural construction, etc., and is not specifically limited here.

[0110] In this application, the three-dimensional model has multiple possibilities, which enriches the implementation methods and application scenarios of the technical solution and enhances the practical value of the solution.

[0111] 602. Based on the first image and the second image, obtain a surround video of the 3D model.

[0112] After acquiring the first and second images, the 3D image updating device processes them to obtain a surround video of the 3D model. The surround video of the 3D model refers to the surround video before the 3D model is updated; that is, in this surround video, the area of ​​the 3D model to be updated has not yet been updated.

[0113] In summary, the 3D image updating device determines the first pose of the first image relative to the 3D model based on the first image and the second image. The first pose, that is, the shooting pose of the first image obtained by capturing the area to be updated, using the pose of the 3D model as the coordinate system, can be the pose of a real image acquisition device or a virtual image acquisition device. Based on the first pose and the first image, a surround video of the 3D model is generated.

[0114] There are several possible ways for the 3D image updating device to determine the first pose, which will be explained below:

[0115] In some optional implementations, the 3D image updating device can determine the first pose based on the rendered image of the second image. Specifically, the 3D image updating device renders the second image from multiple viewpoints to obtain multiple rendered videos. The similarity between the first image and each of the multiple rendered images is compared, and a target rendered image with the highest similarity to the first image is determined from the multiple rendered images. The highest similarity between the first image and the target rendered image means that the viewpoint of the first image is closest to that of the target rendered image, and therefore the pose of the target rendered image is the most relevant. The 3D image updating device determines the second pose of the first image relative to the target rendered image. Since the target rendered image is obtained by rendering the second image, which is either the initial image or a historical image of the 3D model, the poses of the second image and the target rendered image are known, and this pose is a pose with the 3D model as the coordinate system. Therefore, based on the second pose of the first image relative to the target rendered image, the 3D image updating device can calculate the first pose of the first image relative to the 3D model.

[0116] For example, suppose the 3D image updating device renders 10 rendered images from 10 different viewpoints for the second image. The scale-invariant feature transform (SIFT) features of the first image and each rendered image are extracted. The SIFT features of the first image are compared with the SIFT features of the 10 rendered images, and the rendered image with the shortest feature distance is identified as the target rendered image.

[0117] Optionally, the above example uses SIFT features. In practical applications, other features can also be used to determine the target rendering image, such as super point features, or other features that can represent the correspondence between pixels. No specific limitations are made here.

[0118] In this application, the first pose of the first image relative to the target rendered image is determined based on the second pose of the first image relative to the target rendered image. Since the first image is the image with the highest similarity to the second image among multiple rendered images from multiple viewpoints, the second pose is close to the first pose, making the calculation of the first pose simpler.

[0119] In some alternative implementations, the three-dimensional image updating device may also input the first image and the second image into a third neural network to obtain the first pose of the first image relative to the three-dimensional model.

[0120] In this embodiment, a two-dimensional first image and a three-dimensional second image are processed to obtain the first pose of the first image relative to the three-dimensional model, thus realizing pose transformation. Therefore, the surround video of the three-dimensional model generated based on the first pose and the first image more closely resembles the actual situation of a real three-dimensional model, further improving the reliability of the technical solution.

[0121] After determining the first pose, the first pose of the first image is used as the pose of a certain frame, and the center pose of the region to be updated is used as the surround center to construct a surround video of the 3D model.

[0122] Optionally, the first pose can be used as the pose for the first frame. In this method, the construction of the orbital trajectory is simpler.

[0123] Optionally, the center pose of the area to be updated can be manually marked, or it can be the default setting. Alternatively, the user can be prompted whether to modify it after setting it. No specific restrictions are imposed here.

[0124] 603. Input the first image and the surround video of the 3D model into the video generation model to obtain the surround video after the 3D model is updated.

[0125] After acquiring the first image and the surround video of the 3D model, the 3D image updating device inputs it into the video generation model to obtain the surround video updated with the 3D model.

[0126] In this embodiment, the video generation model includes a first neural network, an attention network, and a second neural network. The input to the video generation model is also the input to the first neural network. That is, the input to the first neural network is a first image and a surround video of the 3D model.

[0127] The first neural network processes the first image and the surround video of the 3D model to obtain image features corresponding to the first image and video features corresponding to the surround video of the 3D model. Specifically, the first neural network includes a neural network image encoder and a neural network video encoder. The neural network image encoder encodes the first image to obtain image features. The neural network video encoder encodes the surround video of the 3D model to obtain video features.

[0128] The input to the attention network is the output of the first neural network, that is, the image features and video features are input into the attention network of the video generation model to obtain fused features. The attention network can be a self-attention network, a cross-attention network, etc., and is not limited here.

[0129] The input to the second neural network is the output of the attention network, and the output of the second neural network is the output of the video generation model. That is, the fused features are input into the second neural network to obtain the surround video with the updated 3D model. In the surround video with the updated 3D model, the regions of the 3D model that were to be updated have already been updated.

[0130] For example, the second neural network can be a DiT network. Then, inputting the fused features into the second neural network can be done by inputting the fused features into the noise of the DiT network, and inputting the image features and / or video features into the control condition network of the DiT network. Alternatively, it can be done by inputting the image features and / or video features into the noise of the DiT network, and inputting the fused features into the control condition network of the DiT network.

[0131] In some optional implementations, the 3D image updating device can also acquire text information, which describes attribute information of at least one of the first image, the second image, or the 3D model. For example, image style, 3D model update information, etc., are used to assist in generating a surround video with an updated 3D model.

[0132] In the scheme for acquiring text information, the input to the video generation model can include text information in addition to the first image and the surround video of the 3D model. Specifically, the 3D image updating device inputs the first image, the surround video of the 3D model, and the text information into a first neural network to obtain image features, video features, and text features. The image features, video features, and text features are then input into an attention network to obtain fused features. The fused features are then input into a second neural network to obtain the surround video updated with the 3D model.

[0133] In this scheme, the first neural network further includes a neural network text encoder for encoding text information into text features. For example, the second neural network can be a DiT network. Then, inputting the fused features into the second neural network can be done by adding the fused features to the noise of the DiT network, or by inputting at least one of image features, video features, or text features into the control condition network of the DiT network. Alternatively, it can be done by inputting at least one of image features, video features, or text features into the noise of the DiT network, and then inputting the fused features into the control condition network of the DiT network.

[0134] In this embodiment of the application, the input to the video generation model also includes text information. The text information serves as auxiliary information, describing the attribute information of at least one of the first image, the second image, or the three-dimensional model, so that the surround video of the updated three-dimensional model is closer to the real situation and the reliability of the three-dimensional image update is improved.

[0135] 604. Update the second image based on the surround video updated from the 3D model.

[0136] The 3D image updating device acquires the surround video with the updated 3D model and then updates the second image. Specifically, this includes reconstructing the 3D model of the region to be updated based on the updated surround video, thus obtaining a locally reconstructed 3D model. Then, the 3D model of the region to be updated in the second image is replaced with the locally reconstructed 3D model.

[0137] In this application, when updating a 3D image, the data that needs to be re-acquired is the 2D image of the area to be updated. By processing the 2D image and the original 3D model image, the 3D model image can be updated without having to start from scratch and reconstruct a new 3D model image. This reduces the workload and complexity of 3D image updates and improves the efficiency of 3D image updates.

[0138] The following description uses an example where the 3D model is a 3D city base and the second image is the initial image of the 3D city base to further illustrate the 3D image updating method provided in this application. Please refer to Figures 7 and 8, which are schematic flowcharts of the 3D image updating method provided in this application.

[0139] As shown in Figure 7, the 3D image updating device includes: an image acquisition module, an image registration module, a 3D model video rendering module, a new scene video generation module, and a 3D model updating module.

[0140] The image acquisition module acquires new scene images (i.e., the first image) using image acquisition devices such as satellites and digital cameras. The image registration module registers the acquired image to the original 3D city model, solving for the image's pose (the first pose). The 3D model video rendering module, based on the original 3D city model, uses the first image as the pose of one frame in the rendering trajectory to render a surround video of the original scene (i.e., a surround video of the 3D model). The new scene video generation module uses the captured first image and the surround video rendered from the original scene to generate a surround video of the changed scene, i.e., a new scene surround video. The 3D model update module uses the surround video of the changed scene to perform 3D reconstruction and update the city's 3D model.

[0141] As shown in Figure 8, the new scene is based on the initial scene with the addition of buildings. The 3D image updating device acquires a 2D image of the new scene, i.e., the first image. In the embodiment shown in Figure 8, the first image is obtained by a camera taking a top-down view of the new scene.

[0142] Based on a visual positioning algorithm, the first image is registered into the 3D model to obtain the first pose of the first image relative to the city's 3D model. This first pose is a six-degree-of-freedom (6DoF) pose. Specifically, this can be achieved by using positioning information (such as GPS prior information) to segment the region related to the area captured by the first image from the initial complete city 3D model, thus obtaining the second image. Then, based on the first and second images, the first pose of the first image relative to the city 3D model is determined.

[0143] For example, multiple images are rendered from multiple perspectives for the second image. Image features are extracted from the rendered images, and the image features of the first image are compared with the image features of the multiple rendered images to determine the image pair with the highest similarity. The relative pose between the first image and the most similar rendered image is calculated. Since the pose of the rendered image in the 3D base plate coordinate system is already known, the first pose of the first image relative to the 3D base plate can also be determined.

[0144] The first pose is taken as the pose of the first frame in the trajectory rendering process. Using the center pose of the manually annotated scene update region as the orbital center, trajectory rendering is performed to obtain the video trajectory. The video trajectory can be understood as an orbital trajectory, which is a sequence of camera poses. Then, video rendering is performed on the video trajectory to obtain the orbital video of the original scene. The region corresponding to the original scene here is the same as or similar to the region corresponding to the new scene.

[0145] The first image and the surround video of the original scene are then processed by a video generation model to generate a surround video of the new scene. The specific process is detailed above and will not be repeated here. In this embodiment, during training, the video generated by the video generation model maintains the same shooting pose trajectory as the input video and a similar visual style to the original 3D city model. Therefore, in the 3D image update method, the surround video of the new scene output by the video generation model is consistent with or similar in style to the original 3D city model.

[0146] After obtaining the surround video of the new scene, a local 3D model is reconstructed from it, and the content of the 3D model in the original scene within the corresponding area of ​​the new scene is replaced, that is, the second image is updated. The updated second image displays the 3D model of the new scene.

[0147] The relevant equipment involved in the embodiments of this application will be described below.

[0148] Please refer to Figure 9, which is a schematic diagram of the structure of the training device provided in an embodiment of this application. As shown in Figure 9, the training device 900 includes an acquisition unit 901 and a processing unit 902.

[0149] In some optional embodiments, the processing unit 902 is specifically configured to: input a two-dimensional first image and a surround video of a three-dimensional model into a first neural network of a video generation model to generate image features corresponding to the first image and video features corresponding to the surround video of the three-dimensional model; input the image features and video features into an attention network of the video generation model to obtain fused features; input the fused features into a second neural network of the video generation model to obtain a surround video updated with the three-dimensional model; and iteratively train the video generation model, updating at least one of the parameters of the first neural network, the parameters of the attention network, or the parameters of the second neural network until the video generation model converges or reaches a preset number of iterations. The first image includes the region of the three-dimensional model to be updated.

[0150] In some optional implementations, the acquisition unit 901 is used to acquire text information, which is used to describe the attribute information of at least one of the first image, the second image, or the three-dimensional model.

[0151] Processing unit 902 is specifically used to input the two-dimensional first image, the surround video of the three-dimensional model, and text information into the first neural network of the video generation model to generate image features corresponding to the first image, video features corresponding to the surround video of the three-dimensional model, and text features corresponding to the text information. The image features, video features, and text features are then input into the attention network of the video generation model to obtain fused features.

[0152] It should be noted that the acquisition unit 901 and the processing unit 902 respectively implement different steps in the training method of the video generation model, thereby realizing all the functions of the training device and the cloud service platform. The training device 900 is used to execute the operations performed by the execution device in the embodiment shown in Figure 2, the cloud service platform in the embodiment shown in Figure 3, and the training device in the embodiments shown in Figures 4 and 5, so as to realize the training method of the video generation model provided in this application embodiment, as detailed above, and will not be repeated here.

[0153] Please refer to Figure 10, which is a schematic diagram of the structure of a three-dimensional image updating device provided in an embodiment of this application. As shown in Figure 10, the three-dimensional image updating device 1000 includes an acquisition unit 1001 and a processing unit 1002.

[0154] In some optional implementations, the acquisition unit 1001 is used to acquire a first image and a second image of the 3D model, the first image including the region to be updated in the 3D model. Based on the first and second images, a surround video of the 3D model is acquired.

[0155] Processing unit 1002 is used to input the first image and the surround video of the 3D model into the video generation model to obtain the surround video updated with the 3D model. Based on the surround video updated with the 3D model, the second image is updated.

[0156] In some optional embodiments, the processing unit 1002 is specifically configured to: input the first image and the surround video of the 3D model into the first neural network of the video generation model to generate image features corresponding to the first image and video features corresponding to the surround video of the 3D model; input the image features and video features into the attention network of the video generation model to obtain fused features; and input the fused features into the second neural network of the video generation model to obtain the surround video updated with the 3D model.

[0157] In some optional implementations, the acquisition unit 1001 is also used to acquire text information, which is used to describe the attribute information of at least one of the first image, the second image, or the three-dimensional model.

[0158] The processing unit 1002 is specifically used to input the first image, the surround video of the 3D model, and text information into the video generation model to obtain the surround video updated by the 3D model.

[0159] In some optional implementations, the acquisition unit 1001 is specifically used to: determine a first pose of the first image relative to the 3D model based on the first image and the second image; and generate a surround video of the 3D model based on the first pose and the first image.

[0160] In some optional implementations, the acquisition unit 1001 is specifically used to: acquire multiple rendered images corresponding to multiple viewpoints of the second image.

[0161] Based on a first image and multiple rendered images, a second pose of the first image relative to a target rendered image is determined. The target rendered image is the image among the multiple rendered images that has the highest similarity to the first image. Based on the second pose, a first pose of the first image relative to the 3D model is determined.

[0162] In some optional implementations, the acquisition unit 1001 is specifically used to input the first image and the second image into the third neural network to obtain the first pose of the first image relative to the three-dimensional model.

[0163] In some alternative implementations, the 3D model includes: a city 3D model, a rural 3D model, a scenic area 3D model, or a natural landscape 3D model.

[0164] It should be noted that by implementing different steps in the three-dimensional image update method through the acquisition unit 1001 and the processing unit 1002, all the functions of the aforementioned execution device, cloud service platform, or three-dimensional image update apparatus are realized. The three-dimensional image update apparatus 1000 is used to execute the operations performed by the execution device in the embodiment shown in FIG2, the cloud service platform in the embodiment shown in FIG3, and the three-dimensional image update apparatus in the embodiments shown in FIG6 to 8, so as to realize the three-dimensional image update method or video generation model training method provided in the embodiments of this application, as detailed above, and will not be repeated here.

[0165] The acquisition unit 901, processing unit 902, acquisition unit 1001, and processing unit 1002 can all be implemented in software or in hardware. For example, the implementation of processing unit 1002 will be described below. Similarly, the implementation of acquisition unit 1001 can refer to the implementation of processing unit 1002.

[0166] As an example of a software functional unit, processing unit 1002 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, processing unit 1002 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0167] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0168] As an example of a hardware functional unit, the processing unit 1002 may include at least one computing device, such as a server. Alternatively, the processing unit 1002 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0169] The processing unit 1002 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the processing unit 1002 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the processing unit 1002 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0170] Please refer to Figure 11, which is a schematic diagram of a computing device provided in an embodiment of this application. The computing device 1100 includes a processor 1101, a communication interface 1102, a bus 1103, and a memory 1104. The processor 1101, the communication interface 1102, and the memory 1104 communicate with each other via the bus 1103. In practical applications, communication can also be achieved through other means such as wireless transmission; the specifics are not limited here.

[0171] The computing device 1100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memory in the computing device 1100.

[0172] Processor 1101 may include any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).

[0173] The communication interface 1102 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1100 and other devices or communication networks.

[0174] Bus 1103 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 11, but this does not imply that there is only one bus or one type of bus. Bus 1103 can include pathways for transmitting information between various components of computing device 1100 (e.g., memory 1104, processor 1101, communication interface 1102).

[0175] Memory 1104 may include volatile memory, such as random access memory (RAM). Memory 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0176] Optionally, the memory 1104 stores executable program code, and the processor 1101 executes the executable program code to implement the functions of the aforementioned acquisition unit 901 and processing unit 902, thereby realizing the training method of the video generation model. That is, the memory 1104 stores instructions for executing the training method of the video generation model.

[0177] Optionally, the memory 1104 stores executable program code, and the processor 1101 executes the executable program code to implement the functions of the aforementioned acquisition unit 1001 and processing unit 1002, thereby realizing the three-dimensional image update method. That is, the memory 1104 stores instructions for executing the three-dimensional image update method.

[0178] This application also provides a computing device cluster, which includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some optional embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0179] Please refer to Figures 12 and 13, which are schematic diagrams of the structure of the computing device cluster provided in the embodiments of this application.

[0180] As shown in Figure 12, the computing device cluster includes at least one computing device 1100. The memory 1104 of one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the training method of the video generation model or the update method of the three-dimensional image provided in the embodiments of this application.

[0181] In some possible implementations, the memory 1104 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the training method of the video generation model. In other words, a combination of one or more computing devices 1104 can jointly execute the instructions for executing the training method of the video generation model.

[0182] It should be noted that the memory 1104 in different computing devices 1100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the training device. That is, the instructions stored in the memory 1104 of different computing devices 1100 can implement the functions of one or more units among the acquisition unit 901 and the processing unit 902.

[0183] In some possible implementations, the memory 1104 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the method of updating the three-dimensional image. In other words, a combination of one or more computing devices 1104 can jointly execute the instructions for executing the method of updating the three-dimensional image.

[0184] It should be noted that the memory 1104 in different computing devices 1100 within the computing device cluster can store different instructions, which are used to execute parts of the functions of the 3D image updating device. That is, the instructions stored in the memory 1104 of different computing devices 1100 can implement the functions of one or more units among the acquisition unit 1001 and the processing unit 1002.

[0185] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 13 illustrates one possible implementation. As shown in Figure 13, two computing devices 1100A and 1100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1104 in computing device 1100A stores instructions for executing the functions of the acquisition unit 1001. Simultaneously, the memory 1104 in computing device 1100B stores instructions for executing the functions of the processing unit 1002.

[0186] The connection method between the computing device clusters shown in Figure 13 can be considered in the three-dimensional image update method provided in this application, in which the processing operation and the operation other than the processing operation are executed separately. That is, it is considered that the function of the acquisition unit 1001 is executed by the computing device 1100A, and the function of the processing unit 1002 is executed by the computing device 1100B.

[0187] In some optional embodiments, the memory 1104 in computing device 1100A stores instructions for performing the functions of acquisition unit 901. Simultaneously, the memory 1104 in computing device 1100B stores instructions for performing the functions of processing unit 902. This connection method can be considered in the training method of the video generation model provided in this application, where processing operations and non-processing operations are executed separately; that is, the functions of acquisition unit 901 are considered to be performed by computing device 1100A, and the functions of processing unit 902 are considered to be performed by computing device 1100B.

[0188] It should be understood that the functions of computing device 1100A shown in Figure 13 can also be performed by multiple computing devices 1100. Similarly, the functions of computing device 1100B can also be performed by multiple computing devices 1100.

[0189] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device cluster described in Figures 12 and 13, and will not be repeated here.

[0190] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computer device, it causes the at least one computer device to execute the above-described video generation model training method or the 3D image updating method.

[0191] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-described video generation model training method or the 3D image updating method.

[0192] Please refer to Figure 14, which illustrates an embodiment of the chip in this application. This chip can be represented as a neural network processor 1400. The neural network processor 1400 is mounted as a coprocessor on the main CPU, and tasks are assigned by the host CPU. The core of the neural network processor 1400 is the arithmetic circuit 1403. The controller 1404 can control the arithmetic circuit 1403 to extract data from the weight memory 1402 or the input memory 1401 and perform calculations.

[0193] In some implementations, the arithmetic circuit 1403 internally includes multiple process engines (PEs). In some implementations, the arithmetic circuit 1403 can be a two-dimensional pulsating array. The arithmetic circuit 1403 can also be a one-dimensional pulsating array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1403 is a general-purpose matrix processor.

[0194] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 1402 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 1401 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is ​​stored in the accumulator 1408.

[0195] Unified memory 1406 is used to store input and output data. Weight data is directly transferred to weight memory 1402 via direct memory access controller (DMAC) 1405. Input data is also transferred to unified memory 1406 via DMAC.

[0196] The bus interface unit 1410 (BIU) is used to enable interaction between the main CPU, DMAC, and instruction fetch buffer 1409 (IFB) via a bus. The instruction fetch buffer 1409 is used to store instructions used by the controller 1404.

[0197] The vector computation unit 1407 includes multiple arithmetic processing units that further process the output of the computation circuit as needed, such as vector multiplication, vector addition, exponential operations, logarithmic operations, size comparisons, etc. It is mainly used for computations in non-convolutional / fully connected layers of neural networks, such as pixel-level summation and upsampling of feature planes.

[0198] In the aforementioned embodiments, the operations of each layer in the neural network can be performed by the operation circuit 1403 or the vector calculation unit 1407.

[0199] The processor mentioned above can be a central processing unit, a microprocessor, or one or more integrated circuits used to control the execution of a program in the first aspect of the method.

[0200] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0201] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.

Claims

1. A method for updating a three-dimensional image, characterized in that, include: Acquire a first two-dimensional image and a second three-dimensional model, wherein the first image includes the region to be updated in the three-dimensional model; Based on the first image and the second image, obtain a surround video of the 3D model; The first image and the surround video of the 3D model are input into the video generation model to obtain the surround video updated by the 3D model; The second image is updated based on the surround video updated from the 3D model.

2. The method according to claim 1, characterized in that, The step of inputting the first image and the surround video of the 3D model into the video generation model to obtain the surround video updated by the 3D model includes: The first image and the surround video of the 3D model are input into the first neural network of the video generation model to generate image features corresponding to the first image and video features corresponding to the surround video of the 3D model. The image features and the video features are input into the attention network of the video generation model to obtain fused features; The fused features are input into the second neural network of the video generation model to obtain the surround video updated by the 3D model.

3. The method according to claim 1, characterized in that, The method further includes: Obtain text information, wherein the text information is used to describe the attribute information of at least one of the first image, the second image, or the three-dimensional model; The step of inputting the first image and the surround video of the 3D model into the video generation model to obtain the surround video updated by the 3D model includes: The first image, the surround video of the 3D model, and the text information are input into the video generation model to obtain the surround video updated by the 3D model.

4. The method according to any one of claims 1 to 3, characterized in that, The step of obtaining a surround video of the 3D model based on the first image and the second image includes: Based on the first image and the second image, determine the first pose of the first image relative to the 3D model; Based on the first pose and the first image, a surround video of the 3D model is generated.

5. The method according to claim 4, characterized in that, Determining the first pose of the first image relative to the 3D model based on the first image and the second image includes: Obtain multiple rendered images corresponding to multiple viewpoints for the second image; Based on the first image and the plurality of rendered images, a second pose of the first image relative to a target rendered image is determined, wherein the target rendered image is the image among the plurality of rendered images that has the highest similarity to the first image; Based on the second pose, the first pose of the first image relative to the three-dimensional model is determined.

6. The method according to claim 4, characterized in that, Determining the first pose of the first image relative to the 3D model based on the first image and the second image includes: The first image and the second image are input into a third neural network to obtain the first pose of the first image relative to the three-dimensional model.

7. The method according to any one of claims 1 to 6, characterized in that, The three-dimensional models include: urban three-dimensional models, rural three-dimensional models, scenic area three-dimensional models, or natural landscape three-dimensional models.

8. A three-dimensional image updating device, characterized in that, include: An acquisition unit is used to acquire a first image and a second image of a three-dimensional model, wherein the first image includes the region to be updated in the three-dimensional model; The acquisition unit is further configured to acquire a surround video of the three-dimensional model based on the first image and the second image; The processing unit is configured to input the first image and the surround video of the three-dimensional model into the video generation model to obtain the surround video updated by the three-dimensional model; The processing unit is also used to update the second image based on the surround video updated by the three-dimensional model.

9. The apparatus according to claim 8, characterized in that, The processing unit is specifically used for: The first image and the surround video of the 3D model are input into the first neural network of the video generation model to generate image features corresponding to the first image and video features corresponding to the surround video of the 3D model. The image features and the video features are input into the attention network of the video generation model to obtain fused features; The fused features are input into the second neural network of the video generation model to obtain the surround video updated by the 3D model.

10. The apparatus according to claim 8, characterized in that, The acquisition unit is further configured to acquire text information, which is used to describe the attribute information of at least one of the first image, the second image, or the three-dimensional model; The processing unit is specifically used to input the first image, the surround video of the three-dimensional model, and the text information into the video generation model to obtain the surround video updated by the three-dimensional model.

11. The apparatus according to any one of claims 8 to 10, characterized in that, The acquisition unit is specifically used for: Based on the first image and the second image, determine the first pose of the first image relative to the 3D model; Based on the first pose and the first image, a surround video of the 3D model is generated.

12. The apparatus according to claim 11, characterized in that, The acquisition unit is specifically used for: Obtain multiple rendered images corresponding to multiple viewpoints for the second image; Based on the first image and the plurality of rendered images, a second pose of the first image relative to a target rendered image is determined, wherein the target rendered image is the image among the plurality of rendered images that has the highest similarity to the first image; Based on the second pose, the first pose of the first image relative to the three-dimensional model is determined.

13. The apparatus according to claim 12, characterized in that, The acquisition unit is specifically used to input the first image and the second image into a third neural network to obtain the first pose of the first image relative to the three-dimensional model.

14. A computing device, characterized in that, Includes a processor, which is coupled to a memory; The memory stores instructions that, when executed on the processor, cause the computer device to perform the method of any one of claims 1 to 7.

15. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 7.

16. A computer program product containing instructions, characterized in that, When the instructions are executed by the computing device, the computing device performs the method as described in any one of claims 1 to 7; Alternatively, when the instruction is run by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 7.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer program instructions that, when executed by a computing device, cause the computing device to perform the method as described in any one of claims 1 to 7; Alternatively, when the computer program instructions are executed by the computing device cluster, the computing device cluster is caused to perform the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for making static photo into three-dimensional effect video

    CN111612878A

  • Three-dimensional model updating method and device and storage medium

    CN115457202A

  • Three-dimensional reconstruction model generation method, three-dimensional reconstruction method, three-dimensional reconstruction device and related equipment

    CN117974904A

  • Pixel-region-to-be-changed extraction device, image processing system, pixel-region-to-be-changed extraction method, image processing method, and program

    JP2020057038A