Object generation method and related device

By using an interactive editing method of reference view diagrams, the problems of low efficiency and resource waste in existing object generation algorithms are solved, enabling the efficient generation of digital objects that meet user expectations.

WO2026158292A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2026-01-20
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing object generation algorithms based on large models need to repeat the entire process when the generated objects do not meet user expectations, resulting in low efficiency and wasted resources.

Method used

Through interactive editing capabilities, computing devices can acquire and edit multiple reference viewpoints of digital objects to generate 3D object resources that meet user expectations, reducing repetitive steps and improving generation efficiency and resource utilization.

Benefits of technology

It enables the efficient generation of controllable digital objects, reduces algorithm runtime, saves computing resources, and improves generation efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026073659_30072026_PF_FP_ABST
    Figure CN2026073659_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of AI, and discloses an object generation method and a related device. The method comprises: acquiring N reference view images of a digital object respectively corresponding to different viewpoints, N being an integer greater than 1; on the basis of an editing operation of a user for the N reference view images, obtaining M reference view images of the digital object, M being an integer greater than or equal to N; and on the basis of the M reference view images of the digital object, generating a three-dimensional object resource of the digital object. The present application is applied to efficiently generate three-dimensional object resources of controllable digital objects.
Need to check novelty before this filing date? Find Prior Art

Description

An object generation method and related equipment

[0001] This application claims priority to Chinese Patent Application No. 202510098847.6, filed with the State Intellectual Property Office of China on January 21, 2025, entitled “An Object Generation Method and Related Equipment”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence (AI), and more particularly to an object generation method and related equipment. Background Technology

[0003] With the rapid development of the digital media and entertainment industry, the production of digital content has become a growing market demand. In this field, the generation of three-dimensional (3D) objects, characters, and motion is a crucial step in constructing high-quality digital content. To enhance the flexibility and innovation of digital content creation, using artificial intelligence to generate digital objects has become a hot research topic in the field of digital content production.

[0004] However, current object generation algorithms based on large models can generate digital objects through text or multimodal conditions. Moreover, these object generation algorithms are mostly end-to-end outputs. When the final generated result does not meet the user's expectations, it often takes a lot of time to repeatedly run the entire generation process. This not only reduces the efficiency of object generation but also causes a significant waste of resources.

[0005] Therefore, how to efficiently generate controllable 3D object resources is a technical problem that urgently needs to be solved by those in this field. Summary of the Invention

[0006] This application provides an object generation method and related equipment, which can reduce the waste of time and computing power in the digital object generation process, thereby efficiently generating controllable three-dimensional object resources.

[0007] To achieve the above objectives, this application adopts the following technical solution.

[0008] In a first aspect, embodiments of this application provide an object generation method, the method comprising: a computing device acquiring N reference viewpoint images corresponding to a digital object from different perspectives, where N is an integer greater than 1; the computing device obtaining M reference viewpoint images of the digital object based on user editing operations on the N reference viewpoint images, where M is an integer greater than or equal to N; and the computing device generating a three-dimensional object resource of the digital object based on the M reference viewpoint images of the digital object.

[0009] In this context, a reference viewpoint corresponds to a viewpoint parameter, which (e.g., camera parameter) can be described using spherical coordinate angles. Specifically, it can include polar angle, pitch angle, and the distance from the viewpoint (e.g., camera center point) to the center point of the digital object.

[0010] For example, the digital object here can be any of the following: a digital human figure (e.g., a game character or an anime character), a digital object (e.g., a bridge, an airplane or a vehicle), a digital animal (e.g., a cat, a fish or a bird) or a digital plant (e.g., a flower or grass).

[0011] In the above-described solution, the object generation algorithm provided in this application embodiment has interactive editing capabilities. That is, by responding to the user's editing operations on N reference viewpoints, the computing device can regenerate more accurate M reference viewpoints. In other words, if the three-dimensional object resources of the final generated digital object are not accurate enough, i.e. do not meet the user's expectations, it is not necessary to spend a lot of time repeating the entire process like traditional object generation algorithms. Instead, it can directly repeat one step, i.e., edit some or all of the reference viewpoints in multiple reference viewpoints, so as to facilitate the subsequent generation of controllable three-dimensional object resources for digital objects. This not only reduces the algorithm's running time and improves the object generation efficiency, but also effectively saves computing resources.

[0012] In one possible implementation, the editing operation includes one or more of the following: redraw operation, deletion operation, addition operation, and regeneration operation.

[0013] In the above implementation, the editing operation can include multiple methods, so that when users participate in editing N reference viewpoint images, they can flexibly choose one or more methods. For example, if the digital object is a digital human image, the editing operation here may include redrawing the color of a certain part (e.g., hair color), adding accessories (e.g., glasses), etc., so that the final M reference viewpoint images are more accurate, so as to facilitate the subsequent generation of controllable 3D object resources of digital objects.

[0014] In one possible implementation, the N reference viewpoint images include a first reference image; the computing device obtains M reference viewpoint images of the digital object based on the user's editing operation on the N reference viewpoint images, including: the computing device responding to the user's editing operation on the first reference image to obtain an edited first reference image; the computing device calling a multi-view image generation model, and obtaining the M reference viewpoint images of the digital object based on the edited first reference image and the multi-view image generation model.

[0015] In the above implementation, the computing device can edit some of the reference viewpoints in the N reference viewpoints through interaction with the user, and then, based on the edited reference viewpoints, can more quickly regenerate M reference viewpoints that meet the user's expectations. This avoids the waste of time and computing power caused by the uncontrollable generation algorithm. In other words, the efficiency and accuracy of the 3D object resources of digital objects generated by this implementation are higher.

[0016] In one possible implementation, the computing device obtains M reference viewpoints of a digital object based on editing operations on N reference viewpoints, including: the computing device acquiring M viewpoint parameters, where M is an integer greater than N, based on the editing operations on the N reference viewpoints; and the computing device obtaining M reference viewpoints of the digital object in response to a generation operation on the M viewpoint parameters.

[0017] In the above implementation, the computing device can increase the number of viewpoint parameters through user interaction. For example, the number of viewpoint parameters can be increased from the original 4 to 24, resulting in a denser reference viewpoint map. Compared to 3D object resources generated using a sparse reference viewpoint map, 3D object resources generated using a denser reference viewpoint map can improve the quality of object generation.

[0018] In one possible implementation, the M viewpoint parameters include K viewpoint parameters, where K is an integer less than or equal to M, and the K viewpoint parameters are collected within the viewpoint range corresponding to the key parts of the digital object.

[0019] For example, if the digital object is a digital human figure, and the digital object has accessories (e.g., a necklace or scarf) around its neck and a sword at its waist, then the M viewpoint parameters here may include at least one viewpoint parameter collected within the viewpoint range corresponding to the neck and at least one viewpoint parameter collected within the viewpoint range corresponding to the waist. This method of collecting viewpoint parameters can obtain more reference viewpoint images for the key parts that the user is concerned about, thereby making the three-dimensional object resources of the subsequently generated digital object more controllable.

[0020] In one possible implementation, the M reference viewpoints are determined based on M viewpoint parameters, N reference images, and initial information used to describe the digital object.

[0021] The above implementation method can be applied to reconstruction scenarios. For example, the computing device can generate images by calling a multi-view image generation model. Since M view parameters are used to control the reference view images of digital objects from different viewpoints, the dependence on camera calibration can be reduced, which not only improves the efficiency of object reconstruction but also effectively saves computing resources.

[0022] In one possible implementation, the M reference viewpoint maps are determined based on M viewpoint parameters, N reference maps, first information describing the digital object, and M depth maps; the M depth maps are obtained by rendering the geometry of the digital object based on the M viewpoint parameters.

[0023] In this system, each depth map corresponds to a viewpoint parameter. A depth map is a grayscale image whose pixels record the distance from the viewpoint to the encoded occluded object. Depth maps are a commonly used image representation method in computer vision. The above implementation can be applied to reprojection scenarios, effectively controlling the geometric consistency between the M reference viewpoint maps and the digital object, resulting in more accurate 3D object resources for the subsequently generated digital object.

[0024] In one possible implementation, the computing device acquires N reference view images corresponding to the digital object from different perspectives, including: the computing device acquiring a second reference image based on user input; and the computing device acquiring N reference view images corresponding to the digital object from different perspectives based on a multi-view generation operation of the second reference image.

[0025] In the above implementation, the 3D object resource of the digital object is not directly generated based on the first information used to describe the digital object. Instead, it generates reference viewpoint images corresponding to the digital object from multiple viewpoints based on a reference viewpoint image (i.e., a second reference image) of the digital object from a certain perspective. Then, the 3D object resource of the digital object is generated based on these multiple reference viewpoint images. That is, the object generation method of this application embodiment is a two-stage generation route. The generation process of the second reference image requires interaction between the computing device and user a. The participation of user a can generate not only real-world objects but also virtual objects (i.e., objects without actual reference objects), which can effectively improve the flexibility and innovation of digital content creation.

[0026] In one possible implementation, the second reference diagram is obtained by transforming the first information used to describe the digital object based on the object style of the digital object. The first information is determined in response to user input.

[0027] For example, the first information may be one or more of text, voice, images, or videos.

[0028] In the above implementation, the object style of the digital object can include one of the styles such as realistic, comic, or ink painting. This means that the embodiments of this application can provide the ability to generate digital objects of different styles. By using the stylized image generation method, it is possible to generate three-dimensional object resources of stylized digital objects, thereby effectively improving the flexibility and innovation of digital content creation.

[0029] In one possible implementation, the computing device acquires a second reference image based on user input, including: the computing device acquiring second information corresponding to a digital object in response to the user input; the second information including first information describing the digital object, the geometry of the digital object, and N viewpoint parameters; the computing device acquiring N depth maps corresponding to the geometry in response to user rendering operations on the second information, each depth map corresponding to one of the N viewpoint parameters; and the computing device acquiring the second reference image in response to generation operations on the N depth maps and the first information.

[0030] The above implementation method can be applied to reprojection scenarios. The depth map can effectively control the geometric consistency between the second reference map and the digital object, making the 3D object resources of the subsequently generated digital object more accurate.

[0031] Secondly, embodiments of this application provide an object generation apparatus, comprising: an acquisition module, configured to acquire N reference viewpoint images corresponding to a digital object from different viewpoints, where N is an integer greater than 1; a determination module, configured to obtain M reference viewpoint images of the digital object based on user editing operations on the N reference viewpoint images, where M is an integer greater than or equal to N; and a generation module, configured to generate a three-dimensional object resource of the digital object based on the M reference viewpoint images of the digital object.

[0032] In one possible implementation, the editing operation includes one or more of the following: redraw operation, deletion operation, addition operation, and regeneration operation.

[0033] In one possible implementation, the N reference viewpoint images include a first reference image; a determining module is used to obtain the edited first reference image in response to the user's editing operation on the first reference image; the determining module is also used to call a multi-view image generation model to obtain M reference viewpoint images of the digital object based on the edited first reference image and the multi-view image generation model.

[0034] In one possible implementation, the determining module is used to obtain M viewpoint parameters based on the editing operation for N reference viewpoint images, where M is an integer greater than N; the determining module is also used to obtain M reference viewpoint images of the digital object in response to the generation operation for the M viewpoint parameters.

[0035] In one possible implementation, the M viewpoint parameters include K viewpoint parameters, where K is an integer less than or equal to M, and the K viewpoint parameters are collected within the viewpoint range corresponding to the key parts of the digital object.

[0036] In one possible implementation, the M reference viewpoints are determined based on M viewpoint parameters, N reference images, and initial information used to describe the digital object.

[0037] In one possible implementation, the M reference viewpoint maps are determined based on M viewpoint parameters, N reference maps, first information describing the digital object, and M depth maps; the M depth maps are obtained by rendering the geometry of the digital object based on the M viewpoint parameters.

[0038] In one possible implementation, the acquisition module is used to acquire a second reference image based on user input; the acquisition module is also used to acquire N reference view images corresponding to the digital object from different viewpoints based on the multi-viewpoint generation operation for the second reference image.

[0039] In one possible implementation, the second reference diagram is obtained by transforming the first information used to describe the digital object based on the object style of the digital object, the first information being determined in response to user input.

[0040] In one possible implementation, the acquisition module is used to acquire second information corresponding to the digital object in response to user input; the second information includes first information describing the digital object, the geometry of the digital object, and N view parameters; the acquisition module is also used to acquire N depth maps corresponding to the geometry in response to user rendering operation on the second information, with each depth map corresponding to one of the N view parameters; the acquisition module is also used to acquire a second reference map in response to the generation operation on the N depth maps and the first information.

[0041] Thirdly, embodiments of this application provide a computing device, which includes a chip system, a processor, and a power supply circuit. The power supply circuit supplies power to the processor, and the processor performs the method described in the first aspect or any possible implementation thereof.

[0042] The processor can be implemented through a GPU, or through computing devices such as a DPU, NPU, XPU, SoC, offload card, or accelerator card.

[0043] Fourthly, embodiments of this application provide a computing device cluster, including at least one computing device. Each computing device includes a chip system, the chip system including a processor and a power supply circuit. The power supply circuit is used to supply power to the processor, and the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.

[0044] The processor can be implemented through a GPU, or through computing devices or AI chips such as a DPU, NPU, XPU, SoC, offload card, or accelerator card.

[0045] Fifthly, embodiments of this application provide a computer-readable storage medium including computer program instructions, which, when executed by a cluster of computing devices, enable the cluster of computing devices to perform a method as described in the first aspect or any implementation thereof.

[0046] In a sixth aspect, embodiments of this application provide a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform a method as described in the first aspect or any implementation thereof.

[0047] The technical effects produced by any of the above-mentioned second to sixth aspects and any of the above-mentioned implementation methods can be referred to the above-mentioned first aspect and the corresponding implementation methods in the first aspect. The repeated parts will not be repeated here. Attached Figure Description

[0048] Figure 1 is a schematic diagram of a network architecture provided in an embodiment of this application;

[0049] Figure 2 is a schematic diagram of an interactive object generation interface provided in an embodiment of this application;

[0050] Figure 3 is a schematic diagram of a process for generating a three-dimensional object resource for digital objects according to an embodiment of this application;

[0051] Figure 4 is a schematic diagram of an interface for displaying N reference viewpoint images provided in an embodiment of this application;

[0052] Figure 5 is a schematic diagram of an interface for editing a first reference drawing according to an embodiment of this application;

[0053] Figure 6 is an interactive schematic diagram of a three-dimensional object resource for reconstructing digital objects provided in an embodiment of this application;

[0054] Figure 7 is a schematic diagram of an interface for obtaining a second reference image provided in an embodiment of this application;

[0055] Figure 8 is an interactive schematic diagram of a three-dimensional object resource for reprojecting digital objects provided in an embodiment of this application;

[0056] Figure 9 is a schematic diagram of the structure of an object generation device provided in an embodiment of this application;

[0057] Figure 10 is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0058] Figure 11 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;

[0059] Figure 12 is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation

[0060] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. To facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" are not necessarily different. Furthermore, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being better or more advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.

[0061] Unless otherwise stated, " / " means "or". For example, A / B can mean A or B. In this document, "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "more than one" means two or more. "One or more of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent: a, b, c; a and b; a and c; b and c; or a and b and c. Here, a, b, and c can be single or multiple.

[0062] It is understood that in this application, "when," "if," and "if" all refer to the device performing a corresponding action under certain objective circumstances, and are not time-limited, nor do they require the device to perform a judgment action when it is implemented, nor do they imply any other limitations. The device performing a corresponding action under certain objective circumstances includes: satisfying the objective circumstances, i.e., being able to perform the corresponding action; or satisfying both the objective circumstances and other circumstances, in order to perform the corresponding action.

[0063] In this application, "simultaneous" can be understood as "parallel", or at the same point in time, or within a period of time, or within the same cycle. The specific meaning can be understood in conjunction with the context.

[0064] In this application, the use of singular designations for elements is intended to represent "one or more" rather than "one and only one," unless otherwise specified.

[0065] It should be understood that the object generation method provided in this application can be applied to the field of artificial intelligence. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0066] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision (CV), speech processing, natural language processing (NLP), machine learning (ML) / deep learning, autonomous driving, and intelligent transportation.

[0067] Computer vision is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or 3D data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0068] Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0069] To facilitate understanding of the technical solutions provided in the embodiments of this application, the AI ​​models involved in the embodiments of this application will first be introduced:

[0070] 1. Image generation model

[0071] Image generation models can be obtained using a large model fine-tuning technique (e.g., Low-Rank Adaptation, LoRA), hence the name LoRA model. LoRA fine-tunes a large model by freezing its pre-trained weights and training only the biases, enabling it to generate images of specific styles (e.g., realistic, comic, ink painting) or specific objects, such as specific figures or painting styles. Furthermore, due to the generalization ability of large models, LoRA can be transferred to different large models to achieve similar results.

[0072] For example, the image generation model in this application embodiment may include two types, specifically including a first type of image generation model and a second type of image generation model. The first type of image generation model may include multiple branches, one branch being used to convert digital objects of a certain object style. The second type of image generation model is an image generation model corresponding to the object style, that is, it is used to convert digital objects of a specific object style.

[0073] 2. Multi-view image generation model

[0074] A large-scale modeling technique can generate images of the same object from different viewpoints, guided by text or multimodal methods, where the generated viewpoints are controllable. To represent points on a 2D image (an image from a specific viewpoint) within a 3D scene, a camera coordinate system is introduced. For example, a 3D direct coordinate system is established with the camera's optical center as the origin, where the x and y axes are parallel to the x' and y' axes of the imaging plane, and the z-axis is the camera's optical axis, perpendicular to the imaging plane. Multi-view image generation models can typically be combined with camera parameters to generate images from corresponding viewpoints. These camera parameters are primarily used to estimate the size of digital objects; common camera parameters include translation, rotation, and focal length.

[0075] The network architecture of the object generation method provided in the embodiments of this application is described below:

[0076] Please refer to Figure 1, which is a schematic diagram of a network architecture provided in an embodiment of this application. As shown in Figure 1, the network architecture may include a server 100 and a cluster of terminal devices. The cluster of terminal devices may include one or more terminal devices, and the number of terminal devices is not limited here. As shown in Figure 1, the cluster of terminal devices may specifically include terminal devices 111, 112, 113, ..., 114. As shown in Figure 1, terminal devices 111, 112, 113, ..., 114 can respectively connect to the server 100 via the network, so that each terminal device can interact with the server 100 through the network connection. The network connection method is not limited here; it can be directly or indirectly connected via wired communication, directly or indirectly connected via wireless communication, or other methods, which are not limited here.

[0077] Each terminal device in this terminal device cluster can include: smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, in-vehicle terminals, smart TVs, and other smart terminals with data processing capabilities. It should be understood that each terminal device in the terminal device cluster shown in Figure 1 can have an application client installed for generating digital objects. When this application client runs on each terminal device, it can interact with the server 100 shown in Figure 1. This application client can include social clients, multimedia clients (e.g., video clients), entertainment clients (e.g., game clients), information stream clients, educational clients, live streaming clients, etc. This application client can be a standalone client or an embedded sub-client integrated into a client (e.g., social clients, educational clients, and multimedia clients, etc.), and is not limited here.

[0078] As shown in Figure 1, the server 100 in this embodiment can be the server corresponding to the application client. The server 100 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. This embodiment does not limit the number of servers.

[0079] In this embodiment of the application, the computing device with object generation function can be a server or any terminal device in the terminal device cluster shown in Figure 1. The specific form of the computing device will not be limited here.

[0080] It is understood that the object generation algorithm provided in this application embodiment has interactive editing capabilities, meaning that users can participate in the object generation process, enabling computing devices to efficiently generate controllable 3D object resources for digital objects. Specifically, the computing device can obtain N reference viewpoint images corresponding to the digital object from different perspectives, where N is an integer greater than 1. If one or more of these N reference viewpoint images do not meet the user's expectations, the user can interact with the computing device to edit them. At this time, the computing device can obtain M reference viewpoint images of the digital object based on the user's editing operations on the N reference viewpoint images, where M is an integer greater than or equal to N. The editing operations here can include one or more of the following: redrawing, deleting, adding, and regenerating. In this way, the 3D object resources of the digital object subsequently generated by the computing device based on these M reference viewpoint images will be more accurate and better meet the user's expectations. Therefore, this application embodiment can adjust some reference viewpoint images during the object generation algorithm's execution without repeatedly running the entire object generation process, greatly reducing the algorithm's running time, improving object generation efficiency, and saving computing resources.

[0081] For ease of understanding, the computing device in this application embodiment can be a terminal device (e.g., terminal device 111 shown in FIG1) to solve the problem of uncontrollable object generation results caused by end-to-end algorithms in existing object generation algorithms. For example, please refer to FIG2, which is a schematic diagram of an interactive object generation interface provided in an embodiment of this application. Interfaces 20J1, 20J2, and 20J3 can all be display interfaces provided by the computing device corresponding to user a.

[0082] In one possible implementation, the computing device can acquire N reference viewpoint images of a digital object (e.g., object 2D) at different viewpoints, as shown in interface 20J1 of Figure 2. Here, N can be 4, specifically including reference viewpoint image P1, reference viewpoint image P2, reference viewpoint image P3, and reference viewpoint image P4. Reference viewpoint image P1 is the reference viewpoint image of the digital object at viewpoint 1 (i.e., the viewpoint corresponding to viewpoint parameter 1); reference viewpoint image P2 is the reference viewpoint image of the digital object at viewpoint 2 (i.e., the viewpoint corresponding to viewpoint parameter 2); reference viewpoint image P3 is the reference viewpoint image of the digital object at viewpoint 3 (i.e., the viewpoint corresponding to viewpoint parameter 3); and reference viewpoint image P4 is the reference viewpoint image of the digital object at viewpoint 4 (i.e., the viewpoint corresponding to viewpoint parameter 4).

[0083] As shown in Figure 2, the interface 20J1 may also include a control K1 (e.g., a "re-edit" control). To generate a 3D object resource that matches the user's expectations for the digital object, user a can perform a trigger operation on control K1 to edit one or more of the four reference viewpoints. This trigger operation can include contact operations such as clicking and long-pressing, as well as non-contact operations such as voice and gestures; these are not limited here. After editing, the computing device can obtain M reference viewpoints of the object in 2D, as shown in interface 20J2 of Figure 2. Here, M can be 6, specifically including reference viewpoints P1, P2, P3, P4, P5, and P6.

[0084] It is understandable that, in interface 20J2, the reference view P1 displayed in area Q1 can be understood as a new reference view obtained by user a after modifying the original reference view P1, under the same view; the reference view P5 and reference view P6 displayed in area Q2 can be understood as two views added by user a; while reference view P2, reference view P3 and reference view P4 are the reference view images that user a chose to keep, that is, these three reference view images were not edited.

[0085] Furthermore, user a can perform a trigger operation on control K2 (e.g., the "object generation" control) in interface 20J2, so that the computing device can more accurately generate the three-dimensional object resource of object 2D based on the six reference viewpoints of object 2D. For example, the computing device can display the three-dimensional object resource of object 2D in interface 20J3.

[0086] Therefore, the object generation algorithm provided in this application embodiment has interactive editing capabilities. That is, the computing device can intervene in the object generation process by responding to the relevant triggering operation of user a. Based on the editing operation of user a on part of the reference view map, a more accurate reference view map can be obtained, so that the subsequent 2D object three-dimensional object resources can meet the user's needs. This not only reduces the algorithm running time and improves the object generation efficiency, but also saves computing resources.

[0087] It should be noted that the interfaces and controls shown in the embodiments of this application are merely some forms of representation for reference. In actual business scenarios, developers can make relevant designs according to product requirements. The embodiments of this application do not limit the specific forms of the interfaces and controls involved.

[0088] The specific implementation method of the computing device generating controllable digital object 3D object resources more efficiently through interaction with the user can be found in the embodiment methods corresponding to Figures 3-8 below.

[0089] Further, please refer to Figure 3, which is a flowchart illustrating a method for generating a three-dimensional object resource for digital objects according to an embodiment of this application. As shown in Figure 3, this method can be executed by a computing device, which can be the terminal device shown in Figure 1 above. The method may include at least steps S301-S303:

[0090] Step S301: Obtain N reference view images corresponding to the digital object from different viewpoints, where N is an integer greater than 1.

[0091] Each viewpoint corresponds to a viewpoint parameter, which can be predefined or configured by the user corresponding to the computing device; no limitation will be imposed here. For example, this viewpoint parameter (e.g., camera parameter) can be described using spherical coordinates, specifically including polar angle, pitch angle, and the distance from the viewpoint (e.g., camera center point) to the center point of the digital object. The digital object can be any of the following: a digital human figure (e.g., a game character or anime character), a digital object (e.g., a bridge, airplane, or vehicle), a digital animal (e.g., a cat, fish, or bird), or a digital plant (e.g., a flower or grass).

[0092] The N reference viewpoint images can be determined by the computing device in response to the user's image upload operation (i.e., the trigger operation for uploading reference viewpoint images). For example, these N reference viewpoint images can be drawn by the user using an image editing tool (e.g., a painting tool or an AI tool), selected by the user from the computing device's storage space (e.g., a photo album), downloaded by the user using the computing device's search function from a website, or taken by the user using the computing device's camera function for digital objects. There are no limitations on these types of images.

[0093] Optionally, these N reference viewpoint images can also be generated by the computing device based on a reference viewpoint image (also known as a second reference image) from a certain perspective. For example, the computing device can obtain the second reference image based on user input, and then, based on a multi-viewpoint generation operation on the second reference image, obtain N reference viewpoint images corresponding to the digital object from different perspectives. Here, the multi-viewpoint generation operation refers to the triggering operation used to generate multiple reference viewpoint images corresponding to different perspectives. For example, when the digital object is a digital human figure, the second reference image can be a reference viewpoint image (also known as a frontal reference image) corresponding to the digital object from a frontal perspective.

[0094] For ease of understanding, please further refer to Figure 4, which is a schematic diagram of an interface for displaying N reference viewpoint images provided in an embodiment of this application. Interfaces 40J1, 40J2, and 40J3 can all be display interfaces provided by the computing device corresponding to user a.

[0095] It should be understood that the computing device can acquire a second reference image (e.g., reference view image P1 in interface 40J1) based on the input operation of user a. Interface 40J1 may further include control K3 (e.g., an "edit" control) and control K4 (e.g., a "generate reference view image" control). Control K3 can be used to edit the reference view image P1, such as redrawing a portion of the reference view image P1 (e.g., adding a new pattern), or transforming the reference view image P1 based on the object style of the digital object.

[0096] In one possible implementation, the computing device can respond to a user's trigger operation on control K4 and display the interface 40J2 shown in Figure 4. This interface 40J2 may include control K5 (a configuration control, such as a "camera parameter configuration" control) and control K6 (such as a "generate" control). Control K5 can be used to configure the viewpoint parameters of a digital object. For example, control K5 can be used to configure the number N of viewpoint parameters and the specific value corresponding to each viewpoint parameter. Control K6 can be used to generate multiple reference viewpoint images of the digital object at different viewpoints. It is understood that the aforementioned multi-viewpoint generation operation can refer to the trigger operation performed by user A on control K6.

[0097] In other words, if the number of viewpoint parameters N for the digital object is 4, then after user a performs a trigger operation on control K6, the computing device can display 4 reference viewpoint images in interface 40J3, specifically including reference viewpoint image P1, reference viewpoint image P2, reference viewpoint image P3, and reference viewpoint image P4. These N reference viewpoint images may include a second reference image.

[0098] In another possible implementation, the viewpoint parameters of the digital object may have been pre-configured. Therefore, in response to user a's trigger operation on control K4, the computing device can directly display the interface 40J3 shown in Figure 4. In this case, the aforementioned multi-viewpoint generation operation can refer to the trigger operation performed by user a on control K4.

[0099] Therefore, the object generation method provided in this application embodiment is a two-stage generation approach. Instead of directly generating a 3D object resource based on the first information describing the digital object, it first generates a reference viewpoint image (i.e., a second reference image) of the digital object from a certain perspective. Then, based on the second reference image, it generates reference viewpoint images corresponding to the digital object from multiple perspectives. Finally, it generates the 3D object resource of the digital object based on these multiple reference viewpoint images. Since the generation process of the second reference image requires interaction between the computing device and user a, user a's participation not only enables the generation of digital objects that do not exist in reality but also provides the ability to generate digital objects in different styles. This effectively enhances the flexibility and innovation of digital content creation.

[0100] In one alternative approach, the aforementioned second reference image may be determined by a computing device based on first information used to describe a digital object. This first information may be determined by the computing device in response to a user's input operation, and may be one or more of text, voice, images, or videos. If the first information is an image, then embodiments of this application may determine the first information as the second reference image. For example, if the first information is one or more of text, voice, images, or videos, the computing device may further perform conversion processing on the first information to obtain the second reference image.

[0101] The conversion process described here can be used to convert first information from other modalities (e.g., text or voice) into an image modality, or it can be used to convert the first information into an image (e.g., a comic) that matches the object style of the digital object; this is not limited here. In other words, the second reference image can be obtained by converting the first information based on the object style of the digital object. For example, the computing device can first obtain the object style of the digital object, and then use the image generation model described above and the object style of the digital object to convert the first information to obtain at least one initial reference image. When the computing device responds to a user's selection operation for at least one initial reference image, it can determine the initial reference image corresponding to the selection operation as the second reference image.

[0102] If the image generation model invoked by the computing device is the first type of image generation model described above, then the input of the first type of image generation model may include the first information and the object style of the digital object. If the image generation model invoked by the computing device is the second type of image generation model described above (i.e., the image generation model corresponding to the object style of the digital object), then the input of the second type of image generation model may include the first information.

[0103] In an alternative approach, the second reference image can also be drawn by the user using drawing tools or AI tools. For example, it can be drawn locally using AI algorithms based on the user-selected area and line drawing reference.

[0104] In another alternative approach, the second reference image may be selected by the user from a plurality of images corresponding to a certain perspective. For example, the plurality of images may include nine frontal reference images centered on the digital object, and the second reference image may be any one of these nine frontal reference images.

[0105] As can be seen from the above implementation method, the digital object generated by the second reference image can be a real object or a virtual object (i.e., without a real reference object), which will effectively expand the flexibility and innovation of digital content creation.

[0106] Step S302: Based on the user's editing operations on N reference viewpoints, M reference viewpoints of the digital object are obtained, where M is an integer greater than or equal to N.

[0107] The editing operations here may include one or more of the following: redrawing, deleting, adding, and regenerating.

[0108] In one possible implementation, the computing device can maintain the original number of viewpoint parameters and regenerate N reference viewpoint images. Specifically, the computing device can determine the reference viewpoint image to be edited (i.e., the first reference image) from the N reference viewpoint images, and obtain the edited first reference image in response to an editing operation on the first reference image. After editing, the computing device can obtain M reference viewpoint images of the digital object, where M equals N. These M reference viewpoint images may include the edited first reference image, as well as the other reference viewpoint images from the N reference viewpoint images besides the first reference image.

[0109] For example, the computing device can directly call the multi-view image generation model to regenerate N reference view images. As shown in Figure 2, when user a is dissatisfied with the four reference view images in interface 20J1, they can perform a trigger operation on control K1. At this time, the computing device can determine all four reference view images as the first reference images and regenerate the four reference view images through the multi-view image generation model. The input of the multi-view image generation model can include at least one reference view image corresponding to a viewpoint (e.g., the four reference view images in interface 20J1), first information describing the digital object, and four viewpoint parameters.

[0110] Optionally, the computing device can also directly edit unsatisfactory reference view images through user interaction. This editing can be either manual or intelligent, selected through user interaction; neither is limited here.

[0111] For ease of understanding, please refer to Figure 5, which is a schematic diagram of an interface for editing a first reference image according to an embodiment of this application. As shown in Figure 5, interface 50J1 can be an interface displayed by a computing device in response to a trigger operation on control K1 in Figure 2. This interface 50J1 may include four selectable reference view images. It is understood that these four reference view images are the same as the reference view images in interface 20J1 shown in Figure 2.

[0112] As shown in Figure 5, the interface 50J1 may include an editing control, which may include at least one of the following controls: Control K 71 (For example, the "Edit 1" control,) or control K 72 (For example, the "Edit 2" control), where control K 71 Control K is used to represent manual editing, such as manual editing using drawing tools or AI tools. 72 This is used to indicate intelligent editing, such as automatic editing by calling a multi-view image generation model.

[0113] After user a selects a reference view (e.g., reference view P1) in interface 50J1 as the reference view to be edited, they can target control K in interface 50J1. 71 Upon execution of the trigger operation, the computing device can determine the reference viewpoint P1 as the first reference image and display it in interface 50J2. User a, corresponding to the computing device, can directly edit the first reference image using drawing tools or AI tools. After editing, a trigger operation is executed on control K8 (e.g., the "Complete" control) in interface 50J2 to obtain the edited first reference image (i.e., the new reference viewpoint P1). Then, the computing device can display interface 50J3, which shows: the new reference viewpoint P1, the original reference viewpoint P2, the original reference viewpoint P3, and the original reference viewpoint P4.

[0114] Of course, to reduce the number of interface switching times, user a can also select H reference viewpoints in interface 50J1 as the reference viewpoints to be edited (i.e., the first reference viewpoint), where H is less than or equal to N. Based on this, interface 50J2 can include H editable areas, with one editable area used to edit one first reference viewpoint. For example, if user a selects reference viewpoints P1 and P2, and applies control K... 71 When the trigger operation is executed, the interface 50J2 can include two editable areas: one editable area for editing the reference view P1 and the other editable area for editing the reference view P2.

[0115] Optionally, the computing device can also replace the second reference image through interaction with the user. For example, the four reference view images displayed by the computing device in the interface 50J1 are generated based on the original second reference image (e.g., reference view image P1). Now the computing device can determine a new second reference image (i.e., a reference view image from another perspective) and then regenerate N reference view images based on the new second reference image.

[0116] As shown in Figure 5, user a can reselect the second reference image (e.g., reference view image P2) in interface 50J1, and then user a can target control K in interface 50J1. 72 Upon triggering the operation, the computing device can identify the three reference viewpoint images (excluding reference viewpoint image P2) in interface 50J1 as the first reference images. Further, the computing device can invoke a multi-view image generation model and, based on the model and reference viewpoint image P2 in interface 50J1, regenerate four reference viewpoint images. Then, the computing device can display interface 50J3, which shows the original reference viewpoint image P2 and the edited first reference images (including new reference viewpoint images P1, P3, and P4). Compared to the manual editing method described above, this eliminates the need for time-consuming waiting for the user to edit the three reference viewpoint images sequentially; instead, the multi-view image generation model can generate the images more efficiently, thus improving editing efficiency.

[0117] In another possible implementation, the computing device can increase the number of viewpoint parameters through user interaction to obtain a denser reference viewpoint map. For example, the computing device can obtain M viewpoint parameters, where M is an integer greater than N, based on an editing operation on N reference viewpoint maps. Further, the computing device can, in response to a user's generation operation on the M viewpoint parameters, obtain M reference viewpoint maps of the digital object.

[0118] For example, if user a in Figure 5 is not satisfied with the 3D object resources of the digital objects generated based on the N reference viewpoint images displayed in step S301, user a can choose to generate denser reference viewpoint images. That is, user a can first increase the number of viewpoint parameters from N to M based on the configuration controls displayed on interface 50J1 (e.g., the "Camera Parameter Configuration" control), and then adjust the control K. 72 The triggering operation (i.e., the generation operation for M viewpoint parameters) is executed. Then, the computing device can respond to this generation operation by invoking a multi-view image generation model to regenerate M reference viewpoint images. The input to the multi-view image generation model can include the M viewpoint parameters, first information describing the digital object, and a reference viewpoint image corresponding to at least one viewpoint parameter.

[0119] For ease of explanation, in this embodiment, the center position of the digital object can be defined as the zero point of the coordinate system, and a sphere can be formed with the zero point as the center and the distance between the camera and the zero point as the radius. The aforementioned M viewpoint parameters can be viewpoint parameters uniformly collected on the sphere, or viewpoint parameters corresponding to a specified viewpoint on the sphere; no limitation will be imposed here.

[0120] For example, the M viewpoint parameters here may include K viewpoint parameters, which are collected within the viewpoint range corresponding to key parts of the digital object, where K is an integer less than or equal to M. For instance, if the digital object is a digital human figure, and the digital object has accessories (e.g., a necklace or scarf) around its neck and carries a sword at its waist, in order to more accurately generate the 3D object resource of the digital object, these M viewpoint parameters may include at least one viewpoint parameter collected within the viewpoint range corresponding to the neck, and at least one viewpoint parameter collected within the viewpoint range corresponding to the waist.

[0121] In another possible implementation, the computing device can also edit a portion of the N reference viewpoint images through user interaction, and then regenerate M reference viewpoint images based on the edited reference viewpoint images. For example, the computing device can respond to a user's editing operation on a first reference image (i.e., the reference viewpoint image to be edited), obtain an edited first reference image, and then invoke a multi-view image generation model to obtain M reference viewpoint images of the digital object based on the edited first reference image and the multi-view image generation model.

[0122] For example, if the first reference image determined by the computing device is the reference view image P1 in the interface 50J1 shown in Figure 5 above, then user a can first target control K. 71A trigger operation is executed, and the reference viewpoint image P1 is edited in interface 50J2 to obtain the edited reference viewpoint image P1. After editing, user a can continue to execute trigger operations on control K8. When the computing device responds to the trigger operation, it can call the multi-view image generation model to regenerate N reference viewpoint images. The input of the multi-view image generation model can include N viewpoint parameters, first information describing the digital object, and at least one reference viewpoint image (e.g., the edited reference viewpoint image P1, the original reference viewpoint image P2, the original reference viewpoint image P3, and the original reference viewpoint image P4).

[0123] For example, if user a, as shown in Figure 5, is dissatisfied with the 3D object resources of the digital object generated based on the N reference viewpoint images displayed in step S301, then user a not only needs to edit the first reference image, but also needs to select to generate a denser set of reference viewpoint images. That is, user a can increase the number of viewpoint parameters from N to M based on the configuration controls displayed on interface 50J1 (e.g., the "Camera Parameter Configuration" control). Then, after the computing device determines the edited first reference image (e.g., reference viewpoint image P1 in interface 50J2), it can call the multi-view image generation model to regenerate M reference viewpoint images. The input of the multi-view image generation model here can include M viewpoint parameters, first information describing the digital object, and reference viewpoint images (e.g., the edited reference viewpoint image P1, the original reference viewpoint image P2, the original reference viewpoint image P3, and the original reference viewpoint image P4).

[0124] Step S303: Generate a 3D object resource of the digital object based on M reference viewpoint images of the digital object.

[0125] Here, the 3D object resource of the digital object can be reconstructed based on M reference viewpoints of the digital object, or it can be reprojected based on M reference viewpoints of the digital object; there is no limitation on this.

[0126] Therefore, the object generation algorithm provided in this application embodiment has interactive editing capabilities. That is, by responding to the user's editing operation on N reference viewpoints, the computing device can regenerate more accurate M reference viewpoints. In other words, if the three-dimensional object resources of the final generated digital object are not accurate enough, i.e. do not meet the user's expectations, it is not necessary to spend a lot of time repeating the entire process like traditional object generation algorithms. Instead, it can directly repeat one step, i.e., edit some or all of the reference viewpoints in multiple reference viewpoints. This not only reduces the algorithm's running time and improves the object generation efficiency, but also effectively saves computing resources.

[0127] It is understood that the object generation method provided in this application embodiment can be applied not only to reconstruction scenarios but also to reprojection scenarios, and may also be applied to other object generation scenarios, which will not be limited here. The following will describe the different application scenarios:

[0128] To facilitate understanding of the object generation algorithm in the reconstruction scenario, please refer to Figure 6, which is an interactive schematic diagram of a 3D object resource for reconstructing digital objects provided in an embodiment of this application. As shown in Figure 6, the method can be executed by a computing device, which can be the terminal device shown in Figure 1 above. The method can include at least steps S601-S606:

[0129] Step S601: The computing device acquires first information for describing the digital object.

[0130] For example, the first information here may be obtained by the computing device in response to the user's input operation (the triggering operation for inputting information), and the first information may be one or more of text, voice, images or videos.

[0131] In step S602, the computing device obtains the object style of the digital object, and based on the object style, performs conversion processing on the first information to obtain the third reference image.

[0132] For example, the computing device can first obtain the object style of a digital object, and then, using the aforementioned image generation model and the object style of the digital object, perform transformation processing on the first information to obtain at least one initial reference image. When the computing device responds to a user's selection operation on at least one initial reference image, it can determine the initial reference image corresponding to the selection operation as the third reference image.

[0133] Step S603: The computing device acquires the second reference image.

[0134] The second reference image here can be the third reference image mentioned above. Optionally, the second reference image can also be obtained by the computing device in response to the user's editing operation on the third reference image (i.e., the first editing operation).

[0135] To facilitate understanding of the specific implementation of steps S601-S603 above, please further refer to Figure 7, which is a schematic diagram of an interface for obtaining a second reference diagram provided by an embodiment of this application. Interfaces 70J1 and 70J2 can both be display interfaces provided by the computing device corresponding to user a.

[0136] As shown in Figure 7, the interface 70J1 may include an information input area Q3 (i.e., the first area) and an object style configuration area Q4 (i.e., the second area). The information input area Q3 may include areas for inputting information in at least one modality; specifically, it may include an area for inputting text and an area for inputting images. The object style configuration area Q4 may include at least one candidate object style, specifically object style 1 (e.g., comic) and object style 2 (e.g., realistic).

[0137] It should be understood that after user a performs an input operation on the information input area Q3 in interface 70J1, the computing device can display the first information describing the digital object in the information input area Q3. After user a performs a selection operation on a certain object style (e.g., object style 2) in the object style configuration area Q4, the computing device can determine object style 2 as the object style of the digital object.

[0138] At this time, user a can perform a trigger operation on control K9 (e.g., the "conversion" control) in interface 70J1 to cause the computing device to convert the first information based on object style 2 and display reference view diagram 71P1.

[0139] For example, the computing device can invoke the image generation model corresponding to object style 2, input the first information into the image generation model, and have the image generation model transform and process the first information to obtain at least one initial reference image. For instance, the at least one initial reference image may include nine frontal reference images centered on the digital object. Then, the user can perform a selection operation on one of these nine frontal reference images (e.g., reference view image 71P1). In response to the selection operation, the computing device can display reference view image 71P1 in interface 70J2.

[0140] If user a is satisfied with reference viewpoint 71P1, then reference viewpoint 71P1 in interface 70J2 can be understood as the second reference viewpoint.

[0141] If user a is not satisfied with reference viewpoint 71P1, then reference viewpoint 71P1 in interface 70J2 can be understood as the third reference view. Then user a needs to perform the first editing operation on reference viewpoint 71P1. After the editing is completed, the computing device can respond to the first editing operation and obtain the edited reference viewpoint 71P1 (i.e., the second reference view).

[0142] Step S604: The computing device acquires N reference view images corresponding to the digital object from different viewpoints.

[0143] These N reference viewpoint images can be determined by the computing device based on the multi-view generation operation of the second reference image. For example, in response to the multi-view generation operation of the second reference image, the computing device can call a first multi-view image generation model (i.e., a multi-view image generation model applied in the reconstructed scene), input N viewpoint parameters, first information, and the second reference image into the first multi-view image generation model, and have the first multi-view image generation model perform generation processing to obtain N reference viewpoint images.

[0144] Understandably, when the user is satisfied with these N reference viewpoints, the computing device can directly generate a 3D object resource of the digital object based on these N reference viewpoints; when the user is not satisfied with these N reference viewpoints, the computing device can continue to execute the following step S605, that is, it can respond to the user's editing operation (i.e., the second editing operation) on these N reference viewpoints. The second editing operation can specifically include one or more of the following: redrawing operation, deletion operation, addition operation, or regeneration operation.

[0145] Step S605: The computing device acquires M reference view images of the digital object.

[0146] For example, the M reference viewpoint images can be determined based on M viewpoint parameters, N reference images, and the first information.

[0147] For example, these M reference viewpoint images can be regenerated by the computing device calling the first multi-view image generation model; they can also be obtained by directly editing unsatisfactory reference viewpoint images through user interaction; they can also be regenerated based on the replaced second reference image through user interaction; they can also be generated by increasing the number of viewpoint parameters through user interaction to create a denser reference viewpoint image; or they can be regenerated based on the edited reference viewpoint images after editing some of the N reference viewpoint images through user interaction. For details, please refer to the description of step S302 in the embodiment corresponding to Figure 3 above, which will not be repeated here.

[0148] Step S606: The computing device reconstructs the M reference viewpoint images to obtain the three-dimensional object resources of the digital object.

[0149] For example, the computing device can use M reference viewpoint images and the viewpoint parameters corresponding to the M reference viewpoint images as input to the AI ​​reconstruction algorithm, and finally obtain the three-dimensional object resources (including geometry and material) of the digital object. The specific implementation of the AI ​​reconstruction algorithm can be found in existing reconstruction technologies, and will not be described in detail in this application embodiment.

[0150] To facilitate understanding of the object generation algorithm in a reprojection scenario, please refer to Figure 8, which is an interactive schematic diagram of a 3D object resource for reprojecting digital objects provided in an embodiment of this application. As shown in Figure 8, the method can be executed by a computing device, which can be the terminal device shown in Figure 1 above. The method can include at least steps S801-S807:

[0151] Step S801: The computing device acquires the second information of the digital object.

[0152] The second information may be acquired by the computing device in response to the user's input operation, and may specifically include first information describing the digital object, the geometry (white model) of the digital object, and N viewpoint parameters (N is an integer greater than 1). The first information may be one or more of text, voice, images, or videos, and the N viewpoint parameters may be predefined or configured by the user; there are no restrictions on them here.

[0153] In step S802, the computing device displays N depth maps corresponding to the geometry.

[0154] Here, the N depth maps are obtained by the computing device in response to the user's rendering operation on the second information. In other words, the computing device can render depth maps of geometry from N viewpoints using N viewpoint parameters. Rendering refers to the process by which the computing device converts the geometry of a digital object into its corresponding actual drawing on the screen. The most commonly used technique in rendering is rasterization, which converts data into visible pixels.

[0155] Understandably, a depth map corresponds to a viewpoint parameter. A depth map is a grayscale image whose pixels record the distance from the viewpoint to the encoded occlusion. Depth maps are a commonly used image representation in computer vision, primarily used to control the consistency between the subsequently generated second reference map and the geometry.

[0156] In step S803, the computing device generates at least one initial reference map based on N depth maps.

[0157] Wherein, at least one initial reference map is an initial reference map of a digital object from the same viewpoint, and the at least one initial reference map is generated by the computing device in response to a generation operation for N depth maps.

[0158] For example, a computing device can generate at least one initial reference image using an image generation model, combined with N depth maps as controls. The input to this image generation model can include first information and the N depth maps. Alternatively, if the computing device obtains the object style of a user-selected digital object, the input to the image generation model can also include the first information, the object style of the digital object, and the N depth maps.

[0159] Step S804: The computing device acquires the second reference image.

[0160] It is understood that the second reference image here can be a third reference image determined by the user from at least one of the aforementioned initial reference images. Optionally, the second reference image can also be a reference image obtained by editing the third reference image (also known as an edited third reference image). The edited third reference image is obtained by the computing device in response to an editing operation (i.e., a first editing operation) on the second reference image. For example, the edited second reference image is drawn by the user using an image editing tool (e.g., a drawing tool or an AI tool), for example, by using an AI algorithm to perform local drawing based on a user-selected area and a line drawing reference.

[0161] Step S805: The computing device acquires N reference view images corresponding to the digital object from different viewpoints.

[0162] These N reference viewpoint images can be obtained by the computing device based on the multi-view generation operation of the second reference image. For example, in response to the multi-view generation operation of the second reference image, the computing device can call the second multi-view image generation model (i.e., the multi-view image generation model applied in the reprojection scene), input N viewpoint parameters, the first information, N depth maps, and the second reference image into the second multi-view image generation model, and the second multi-view image generation model can perform generation processing to obtain N reference viewpoint images.

[0163] Understandably, when the user is satisfied with these N reference viewpoints, the computing device can directly generate a 3D object resource of the digital object based on these N reference viewpoints; when the user is not satisfied with these N reference viewpoints, the computing device can continue to support the following step S806, that is, it can respond to the editing operation (i.e., the second editing operation) performed by the user on these N reference viewpoints. The second editing operation can specifically include one or more of the following: redrawing operation, deletion operation, addition operation, or regeneration operation.

[0164] Step S806: The computing device acquires M reference view images of the digital object.

[0165] The M reference viewpoint images are determined based on M viewpoint parameters (i.e., the viewpoint parameters corresponding to the M reference viewpoint images respectively), reference viewpoint images (at least one reference viewpoint image corresponding to a viewpoint parameter, for example, the aforementioned N reference viewpoint images), the first information, and M depth images. The M depth images are obtained by rendering the geometry of the digital object based on the M viewpoint parameters. It is understood that the generation method of the M reference viewpoint images can be specifically referred to in the description of step S302 in the embodiment corresponding to Figure 3 above, or the description of step S605 in the embodiment corresponding to Figure 6, and will not be repeated here.

[0166] In step S807, the computing device reprojects the M reference viewpoint images to obtain the three-dimensional object resources of the digital object.

[0167] For example, a computing device can map the M reference viewpoints into a texture space based on the viewpoint parameters corresponding to the M reference viewpoints respectively, obtain the material corresponding to the digital object, and then perform rendering processing on the material and the geometry of the digital object to generate a three-dimensional object resource of the digital object.

[0168] Here, the texture space can be UV space, where UV is short for material map coordinates. Each vertex in the geometry (e.g., mesh representation) corresponds to a two-dimensional UV coordinate, linking it to information such as color on the material map. Because of this correspondence, a perfect circle in UV space may not necessarily be a perfect circle when mapped (rendered) geometrically. For example, this computing device can use a rasterization reprojection algorithm to map M reference viewpoints into UV space to obtain the corresponding material.

[0169] The foregoing details the method provided in this application. To facilitate the implementation of the above-described solutions in the embodiments of this application, corresponding apparatus or devices are also provided in the embodiments of this application.

[0170] This application divides a computing device into functional modules according to the above-described method embodiments. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. The computing device of this application embodiment will be described in detail below with reference to Figures 9 to 12.

[0171] Further, please refer to Figure 9, which is a schematic diagram of an object generation apparatus provided in an embodiment of this application. As shown in Figure 9, the object generation apparatus 1 may include at least one of an acquisition module 91, a determination module 92, and a generation module 93. These modules can perform the corresponding functions of the devices in the above method embodiments.

[0172] In one possible implementation, the object generation device 1 can be used to implement the functions of the computing device described above. The computing device can be the server shown in FIG1, or any terminal device in the terminal device cluster shown in FIG1. ​​The specific form of the computing device will not be limited here.

[0173] Specifically, the acquisition module 91 is used to acquire N reference viewpoint images corresponding to the digital object from different perspectives, where N is an integer greater than 1; the determination module 92 is used to obtain M reference viewpoint images of the digital object based on the user's editing operations on the N reference viewpoint images, where M is an integer greater than or equal to N; and the generation module 93 is used to generate a 3D object resource of the digital object based on the M reference viewpoint images of the digital object.

[0174] In one possible implementation, the editing operation includes one or more of the following: redraw operation, deletion operation, addition operation, and regeneration operation.

[0175] In one possible implementation, the N reference viewpoint images include a first reference image; the determining module 92 is used to obtain the edited first reference image in response to the user's editing operation on the first reference image; the determining module 92 is also used to call a multi-view image generation model to obtain M reference viewpoint images of the digital object based on the edited first reference image and the multi-view image generation model.

[0176] In one possible implementation, the determining module 92 is used to obtain M viewpoint parameters based on the editing operation for N reference viewpoint images, where M is an integer greater than N; the determining module 92 is also used to obtain M reference viewpoint images of the digital object in response to the generation operation for the M viewpoint parameters.

[0177] In one possible implementation, the M viewpoint parameters include K viewpoint parameters, where K is an integer less than or equal to M, and the K viewpoint parameters are collected within the viewpoint range corresponding to the key parts of the digital object.

[0178] In one possible implementation, the M reference viewpoints are determined based on M viewpoint parameters, N reference images, and initial information used to describe the digital object.

[0179] In one possible implementation, the M reference viewpoint maps are determined based on M viewpoint parameters, N reference maps, first information describing the digital object, and M depth maps; the M depth maps are obtained by rendering the geometry of the digital object based on the M viewpoint parameters.

[0180] In one possible implementation, the acquisition module 91 is used to acquire a second reference image based on user input; the acquisition module 91 is also used to acquire N reference view images corresponding to the digital object from different viewpoints based on the multi-viewpoint generation operation for the second reference image.

[0181] In one possible implementation, the second reference diagram is obtained by transforming the first information used to describe the digital object based on the object style of the digital object, the first information being determined in response to user input.

[0182] In one possible implementation, the acquisition module 91 is used to acquire second information corresponding to the digital object in response to user input; the second information includes first information describing the digital object, the geometry of the digital object, and N view parameters; the acquisition module 91 is also used to acquire N depth maps corresponding to the geometry in response to user rendering operation on the second information, with each depth map corresponding to one of the N view parameters; the acquisition module 91 is also used to acquire a second reference map in response to generation operation on the N depth maps and the first information.

[0183] The specific implementation methods of the acquisition module 91, the determination module 92, and the generation module 93 can be found in the description of steps S301-S303 in the embodiment corresponding to Figure 3 above, and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated here.

[0184] The acquisition module 91, determination module 92, and generation module 93 can all be implemented in software or hardware. For example, the implementation of the acquisition module 91 will be described below. Similarly, the implementation of the determination module 92 and generation module 93 can refer to the implementation of the acquisition module 91.

[0185] As an example of a software functional unit, module 91 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, module 91 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0186] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0187] As an example of a hardware functional unit, the acquisition module 91 may include at least one computing device, such as a server. Alternatively, the acquisition module 91 may be implemented using a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system-on-chip (SoC), an offload card, an accelerator card, or any combination thereof.

[0188] The multiple computing devices included in the acquisition module 91 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 91 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 91 can be distributed in the same VPC or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offloading cards, and accelerator cards.

[0189] It should be noted that, in other embodiments, the acquisition module 91 can be used to execute any step in the object generation method, the determination module 92 can be used to execute any step in the object generation method, and the generation module 93 can be used to execute any step in the object generation method. The steps implemented by the acquisition module 91, the determination module 92, and the generation module 93 can be specified as needed. By implementing different steps in the object generation method through the acquisition module 91, the determination module 92, and the generation module 93, all functions of the object generation device 1 can be realized.

[0190] This application also provides a chip system including a processor and a power supply circuit. The power supply circuit supplies power to the processor, which executes the operation steps corresponding to the object generation method. For simplicity, further details are omitted here. The processor can be implemented using a GPU, or it can be implemented using computing devices such as a DPU, NPU, XPU, SoC, offloading card, or accelerator card.

[0191] Further, please refer to Figure 10, which is a schematic diagram of the structure of a computing device provided in an embodiment of this application. As shown in Figure 10, the computing device 2 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other via the bus 1002. The computing device 2 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 2.

[0192] Bus 1002 can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The Unified Bus is also known as the Lingqu Bus. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one line is used in Figure 5, but this does not imply that there is only one bus or one type of bus. Bus 1002 can include pathways for transmitting information between various components of computing device 2 (e.g., memory 1006, processor 1004, communication interface 1008). The Unified Bus can also be referred to as the Lingqu Bus.

[0193] The processor 1004 may include any one or more of the following computing devices: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP) or digital signal processor (DSP), ASIC, FPGA, CPLD, NPU, SoC, offload card, accelerator card, etc.

[0194] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 1004 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD). Furthermore, the memory 1006 may also be implemented using storage class memory (SCM), phase change memory (PCM), or other types of storage media.

[0195] It is worth noting that the same type of storage medium can be configured in the same computing device to realize the function of memory 1006, or two or more types of storage media can be configured to realize the function of memory 1006. This application does not limit this.

[0196] The memory 1006 stores executable program code, which the processor 1004 executes to implement the functions of the aforementioned acquisition module 91, determination module 92, and generation module 93, thereby realizing the object generation method. That is, the memory 1006 stores instructions for executing the object generation method.

[0197] The communication interface 1008 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 2 and other devices or communication networks.

[0198] As one possible implementation, computing device 2 may also include a chip system, which includes a processor and a power supply circuit. The power supply circuit supplies power to the processor, and the processor executes the operation steps corresponding to the object generation method. For simplicity, further details are omitted here. The processor can be implemented using a GPU, or it can be implemented using computing devices or AI chips such as a DPU, NPU, XPU, SoC, offloading card, or accelerator card.

[0199] As one possible implementation, the computing device 2 may include multiple types of processors 1004, i.e., the computing device 2 is a heterogeneous device. For example, the computing device 2 may include a CPU and a GPU, and the operation steps corresponding to the object generation method can be executed by at least one of the processors 1004. For the sake of brevity, further details will not be provided here.

[0200] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0201] Furthermore, please refer to Figure 11, which is a schematic diagram of a computing device cluster provided in an embodiment of this application. The computing device cluster includes at least one computing device 2. The memory 1006 of one or more computing devices 2 in the computing device cluster may store the same instructions for executing object generation methods.

[0202] In some possible implementations, the memory 1006 of one or more computing devices 2 in the computing device cluster may also store partial instructions for executing the object generation method. In other words, a combination of one or more computing devices 2 can jointly execute the instructions for executing the object generation method.

[0203] It should be noted that the memory 1006 in different computing devices 2 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the object generation device. That is, the instructions stored in the memory 1006 of different computing devices 2 can implement the functions of one or more modules among the acquisition module 91, the determination module 92, and the generation module 93.

[0204] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 12 illustrates one possible implementation. As shown in Figure 12, two computing devices 2A and 2B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1006 in computing device 2A stores instructions for executing the functions of the acquisition module 91 and the determination module 92. Simultaneously, the memory 1006 in computing device 2B stores instructions for executing the function of the generation module 93. Specifically, the acquisition module 91 in computing device 2A acquires N reference viewpoint images corresponding to the digital object from different perspectives; the determination module 92 in computing device 2A generates M reference viewpoint images of the digital object based on user editing operations on the N reference viewpoint images; and the generation module 93 in computing device 2B generates a 3D object resource of the digital object based on the M reference viewpoint images.

[0205] It should be understood that the functions of computing device 2A shown in Figure 12 can also be performed by multiple computing devices 2. Similarly, the functions of computing device 2B can also be performed by multiple computing devices 2.

[0206] In some possible implementations, the memory 1006 of one or more computing devices 2 in the computing device cluster may also store partial instructions for executing the object generation method. In other words, a combination of one or more computing devices 2 can jointly execute the instructions for executing the object generation method.

[0207] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform an object generation method.

[0208] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform an object generation method.

[0209] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0210] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0211] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical business division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0212] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0213] Those skilled in the art will recognize that, in one or more of the examples above, the services described in this application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these services can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of computer programs from one place to another. Storage media can be any available medium accessible to general-purpose or special-purpose computers.

[0214] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application.

[0215] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

A method of object generation, characterized in that The method includes: Obtain N reference view images corresponding to the digital object from different perspectives, where N is an integer greater than 1; Based on the user's editing operations on the N reference viewpoint images, M reference viewpoint images of the digital object are obtained, where M is an integer greater than or equal to N; Based on M reference viewpoint images of the digital object, a three-dimensional object resource of the digital object is generated. The method of claim 1, wherein The editing operations include one or more of the following: redrawing, deleting, adding, and regenerating. The method according to any one of claims 1 or 2, characterized in that, The N reference view images include the first reference image; The process of obtaining M reference viewpoints for the digital object based on user editing operations on the N reference viewpoints includes: In response to the user's editing operation on the first reference image, the edited first reference image is obtained; The multi-view image generation model is invoked, and based on the edited first reference image and the multi-view image generation model, M reference view images of the digital object are obtained. The method according to any one of claims 1 or 2, characterized in that, The process of obtaining M reference viewpoints for the digital object based on the editing operations on the N reference viewpoints includes: Based on the editing operations performed on the N reference viewpoint images, M viewpoint parameters are obtained, where M is an integer greater than N; In response to the generation operation for the M viewpoint parameters, M reference viewpoint images of the digital object are obtained. The method according to claim 4, characterized in that The M viewpoint parameters include K viewpoint parameters, where K is an integer less than or equal to M. The K viewpoint parameters are collected within the viewpoint range corresponding to the key parts of the digital object. The method according to claim 4 or 5, characterized in that The M reference viewpoint images are determined based on the M viewpoint parameters, the N reference images, and the first information used to describe the digital object. The method according to claim 4 or 5, characterized in that The M reference viewpoint images are determined based on the M viewpoint parameters, the N reference images, the first information used to describe the digital object, and the M depth images; the M depth images are obtained by rendering the geometry of the digital object based on the M viewpoint parameters. The method according to any one of claims 1 to 6, characterized in that The acquisition of N reference viewpoint images corresponding to the digital object from different viewpoints includes: Based on the user's input, a second reference image is obtained; Based on the multi-view generation operation of the second reference image, N reference view images corresponding to the digital object under different viewpoints are obtained. The method of claim 8, wherein The second reference diagram is obtained by transforming the first information used to describe the digital object based on the object style of the digital object. The first information is determined in response to the user's input operation. The method of claim 8, wherein The step of obtaining the second reference image based on the user's input operation includes: In response to the user's input operation, second information corresponding to the digital object is obtained; the second information includes first information describing the digital object, the geometry of the digital object, and N viewpoint parameters. In response to the user's rendering operation on the second information, N depth maps corresponding to the geometry are obtained, and each depth map corresponds to one of the N view parameters. In response to the generation operation of the N depth maps and the first information, a second reference map is obtained. An object generation device characterized by comprising: include: The acquisition module is used to acquire N reference view images corresponding to digital objects from different perspectives, where N is an integer greater than 1; The determining module is used to obtain M reference view images of the digital object based on the user's editing operations on the N reference view images, where M is an integer greater than or equal to N; The generation module is used to generate a three-dimensional object resource of the digital object based on M reference viewpoint images of the digital object. The apparatus of claim 11, wherein The editing operations include one or more of the following: redrawing, deleting, adding, and regenerating. The apparatus of any one of claims 11 or 12, wherein The N reference view images include the first reference image; The determining module is used to obtain the edited first reference image in response to the user's editing operation on the first reference image; The determining module is further configured to call a multi-view image generation model to obtain M reference view images of the digital object based on the edited first reference image and the multi-view image generation model. The apparatus according to any one of claims 11 or 12 is characterized in that, The determining module is used to obtain M view parameters based on the editing operations performed on the N reference view images, where M is an integer greater than N; The determining module is further configured to obtain M reference view images of the digital object in response to the generation operation for the M view parameters. The apparatus of claim 14, wherein The M viewpoint parameters include K viewpoint parameters, where K is an integer less than or equal to M. The K viewpoint parameters are collected within the viewpoint range corresponding to the key parts of the digital object. The apparatus according to claim 14 or 15, characterized in that The M reference viewpoint images are determined based on the M viewpoint parameters, the N reference images, and the first information used to describe the digital object. The apparatus according to claim 14 or 15, characterized in that The M reference viewpoint images are determined based on the M viewpoint parameters, the N reference images, the first information used to describe the digital object, and the M depth images; the M depth images are obtained by rendering the geometry of the digital object based on the M viewpoint parameters. The apparatus according to any one of claims 11-16 is characterized in that, The acquisition module is used to acquire a second reference image based on the user's input operation; The acquisition module is further configured to acquire N reference view images corresponding to the digital object from different viewpoints based on the multi-view generation operation for the second reference image. The apparatus of claim 18, wherein The second reference diagram is obtained by transforming the first information used to describe the digital object based on the object style of the digital object, the first information being determined in response to user input. The apparatus according to claim 18 is characterized in that, The acquisition module is used to acquire second information corresponding to the digital object in response to the user's input operation; the second information includes first information describing the digital object, the geometry of the digital object, and N viewpoint parameters. The acquisition module is further configured to, in response to the user's rendering operation on the second information, acquire N depth maps corresponding to the geometry, wherein one depth map corresponds to one of the N view parameters; The acquisition module is further configured to acquire a second reference image in response to the generation operation of the N depth maps and the first information. A cluster of computing devices, characterized in that, The method includes at least one computing device, one of which includes a chip system, the chip system including a processor and a power supply circuit, the power supply circuit being used to supply power to the processor, the processor being used to perform the operation steps of the method as described in any one of claims 1 to 10. A computer-readable storage medium, characterized by It includes computer program instructions, which, when executed by a cluster of computing devices, perform the operational steps of the method as described in any one of claims 1 to 10. A computer program product comprising instructions, characterized in that When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the operation steps of the method as described in any one of claims 1 to 10. A chip system, characterized by The chip system includes a processor and a power supply circuit, the power supply circuit being used to supply power to the processor, the processor being used to perform the operation steps of the method as described in any one of claims 1 to 10.