Electronic device, method, and non-transitory computer-readable storage medium for generating three-dimensional mesh model

WO2026205777A1PCT designated stage Publication Date: 2026-10-01NCSOFT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/002815
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-21
Filing Date
2026-02-13
Publication Date
2026-10-01

Smart Images

  • Figure KR2026002815_01102026_PF_FP_ABST
    Figure KR2026002815_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device comprises: a memory storing instructions and including one or more storage media; and at least one processor including a processing circuit, wherein the instructions, when executed individually or collectively by the at least one processor, may instruct the electronic device to: provide information about an object to a trained model; obtain a generated image representing the object viewed from a plurality of directions, generated by the trained model using the information; and generate a three-dimensional (3D) mesh model of the object using the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transient computer-readable storage medium for generating a 3D mesh model

[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for generating a 3D mesh model.

[0002] Artificial intelligence is a technology for simulating human (or biological) neural activities such as perception and / or inference, and can be implemented as hardware, software, or a combination thereof designed to perform calculations for simulating neural activities.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0004] An electronic device is described. The electronic device may include at least one processor comprising a memory that stores instructions and includes one or more storage media, and a processing circuit. The instructions may cause the electronic device to provide information about an object to a trained model when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain a generated image created by the trained model using the information, which represents the object viewed from multiple directions when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to create a three-dimensional mesh model of the object using the generated image when executed individually or collectively by the at least one processor.

[0005] A method is described. The method may include an operation of providing information about an object to a trained model. The method may include an operation of acquiring a generated image, which is generated by the trained model using the information, representing the object as viewed from multiple directions. The method may include an operation of generating a three-dimensional mesh model of the object using the generated image.

[0006] A non-transient computer-readable storage medium is described. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to provide information about an object to a trained model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain a generated image, which is generated by the trained model using the information, representing the object as viewed from multiple directions when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to generate a three-dimensional mesh model of the object using the generated image when executed by the electronic device.

[0007] Figure 1 illustrates an example of photogrammetry.

[0008] Figure 2 is a simplified block diagram of an exemplary electronic device.

[0009] FIG. 3 is a flowchart illustrating exemplary operations of an electronic device for generating a 3D mesh model of an object.

[0010] Figure 4 illustrates an example of a trained model that generates a generated image.

[0011] FIG. 5 illustrates an example of a video representing an object seen in multiple directions as the object is continuously rotated.

[0012] Figure 6 illustrates an example of a user interface (UI) that asks whether to generate a 3D mesh model using a generated image.

[0013] Figure 7 illustrates an example of generating a 3D mesh model of an object based on the point cloud of the object.

[0014] Figure 8 illustrates an example of performing post-processing on a 3D mesh model of an object.

[0015] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0016] Figure 1 illustrates an example of photogrammetry.

[0017] Referring to FIG. 1, an electronic device (e.g., the electronic device (200) of FIG. 2) may be described as a device available for generating a three-dimensional mesh model. For example, the electronic device may be one of various types of mobile devices, such as at least one server device including circuits (or circuitry) for generating a three-dimensional mesh model, a personal computer (PC) (e.g., a laptop and / or desktop), smartphones having various form factors (e.g., a bar-type smartphone, a foldable-type smartphone, or a rollable-type smartphone), a tablet, a wearable device, a cellular phone, and / or other similar computing devices.

[0018] The electronic device can generate a 3D mesh model of the object (105) based on photogrammetry of the object (105). For example, photogrammetry can be described as generating a 3D mesh model by obtaining a point cloud using multiple images representing the object (105) taken from multiple angles through a camera, and converting the point cloud into a 3D mesh and texture. To generate a 3D mesh model of the object (105) based on photogrammetry, multiple images representing the object (105) taken from multiple angles may be required.

[0019] State (100) can be described as a state for performing 3D scanning for photogrammetry. For example, a camera (110) (e.g., an image sensor) may be used to perform 3D scanning. In state (100), the camera (110) may acquire multiple images (115) each containing the object (105) as seen from multiple angles by photographing the object (105) from multiple angles. For example, the camera (110) may be composed of multiple cameras surrounding the object (105) or may photograph the object (105) while moving around the object (105).

[0020] The electronic device can acquire a plurality of images (115) captured through a camera (110). For example, the electronic device can acquire a point cloud using the plurality of images (115). For example, the point cloud may include points corresponding to the outer surface of an object (105) in 3D space. The electronic device can generate a 3D mesh model of the object (105) by connecting the points in the point cloud. The 3D mesh model of the object (105) can be generated based on the lines and surfaces to which the points in the point cloud are connected.

[0021] For example, in order to create a 3D mesh model of an object (105) using multiple images (115), it may be required that the object (105) viewed from multiple directions be sufficiently represented by the multiple images (115). For example, relatively large costs and time may be required to obtain multiple images (115) that sufficiently represent the object (105) viewed from multiple directions.

[0022] For example, since multiple images (115) are acquired through the camera (110), the multiple images (115) may be affected by the shooting environment of the camera (110) (e.g., weather or lighting). For example, if the size of the object (105) is relatively large or if there is an error in approaching the object (105), the object (105) viewed from multiple directions may not be sufficiently represented by the multiple images (115). For example, if the object (105) is a virtual object (or an unreal object), there may be an error in photographing the object (105).

[0023] For example, if an object (105) viewed from multiple directions is not sufficiently represented by multiple images (115), the 3D mesh model of the object (105) may be partially missing or have noise. For example, generating a 3D mesh model of the object (105) based on 3D scanning for photogrammetry may cause such inconvenience to the user. A solution may be required to resolve the inconvenience to the user caused by generating a 3D mesh model of the object (105) based on 3D scanning for photogrammetry.

[0024] To resolve this inconvenience, the electronic device may use a trained model. The electronic device may use the trained model to obtain a generated image representing an object (105) viewed from multiple directions. The electronic device may use the generated image to generate a 3D mesh model of the object (105). For example, photogrammetry may be performed to generate a 3D mesh model of the object (105) using the generated image. For example, the term “generated image” may be used interchangeably with the term “video.” Here, the generated image may refer to a video and / or a single frame image included in the video. To generate a 3D mesh model of the object (105), the electronic device may perform operations exemplified within the description of FIGS. 3 through 8. The electronic device may include components for performing said operations. said components may be exemplified within the description of FIG. 2.

[0025] Figure 2 is a simplified block diagram of an exemplary electronic device.

[0026] Referring to FIG. 2, the electronic device (200) may be one of various types of mobile devices, such as at least one server device, a personal computer (PC) (e.g., a laptop and / or a desktop), smartphones having various form factors (e.g., a bar-type smartphone, a foldable-type smartphone, or a rollable-type smartphone), a tablet, a wearable device, a cellular phone, and / or other similar computing devices. For example, the electronic device (200) may include or correspond to the electronic device of FIG. 1. For example, the electronic device (200) may include at least one processor (210) and memory (220). For example, the electronic device (200) may further include a display (230).

[0027] At least one processor (210) may include a processing circuit. For example, at least one processor (210) may include a CPU (central processing unit) (e.g., including a processing circuit). For example, at least one processor (210) may include a GPU (graphic processing unit) (e.g., including a processing circuit) and / or an NPU (neural processing unit) (e.g., including a processing circuit). For example, at least one processor (210) may be described as an application processor. For example, at least one processor (210) may be configured to control a memory (220) and / or a display (230). At least one processor (210) may be configured to execute instructions stored in the memory (220) individually or collectively to cause the electronic device (200) to perform at least some of the operations illustrated in the description of FIG. 1. At least one processor (210) may be configured to execute instructions stored in memory (220) individually or collectively to cause the electronic device (200) to perform at least some of the operations to be illustrated in the description of FIGS. 3 through 8.

[0028] For example, the term “processor” as used herein, including in the claims, may include various processing circuits comprising at least one processor, and one or more of said at least one processor may be configured to perform the various functions described below in a distributed manner, individually and / or collectively. As used below, where “processor,” “at least one processor,” and “one or more processors” are described as being configured to perform various functions, these terms encompass, for example, but not limited to, situations where one processor performs some of the cited functions and another processor(s) perform other parts of the cited functions, and also situations where one processor can perform all of the cited functions. Additionally, said at least one processor may include a combination of processors that perform the enumerated / disclosed various functions, for example, in a distributed manner. At least one processor may execute program instructions to achieve or perform the various functions.

[0029] The memory (220) may include one or more storage media. For example, the memory (220) may store various data used by at least one component of the electronic device (200) (e.g., at least one processor (210) and / or display (230)). For example, the data may include input data or output data for software and related commands. The memory (220) may include volatile memory or non-volatile memory.

[0030] A display (230) can output visualized information under the control of at least one processor (210). For example, the display (230) may include a flat panel display (FPD) and / or electronic paper. The FPD may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). For example, the LEDs may include organic LEDs (OLEDs). The display (230) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by the touch. For example, the display (230) may be configured to display a UI. For example, the display (230) may be configured to receive user input. For example, a display (230) that supports touch functions may be referred to as a touchscreen. The display (230) may further include a structure capable of detecting input using a stylus pen, such as EMR (electro-magnetic resonance) or AES (active electrostatic solution).

[0031] The electronic device (200) illustrated in the description of FIG. 2 can perform at least some of the operations illustrated in the descriptions of FIG. 3 through 8. For example, the operations illustrated in the descriptions of FIG. 3 through 8 can be caused by (or in) the electronic device (200) under the control of at least one processor (210).

[0032] FIG. 3 is a flowchart illustrating exemplary operations of an electronic device for generating a 3D mesh model of an object.

[0033] Referring to Fig. 3,

[0034] In operation 300, at least one processor (210) may provide (or input) information about an object (e.g., information (405) of FIG. 4) to a trained model (e.g., trained model (400) of FIG. 4). For example, the trained model may include a machine learning model, a deep learning model, and an artificial intelligence (AI) model. For example, the trained model may include a generative artificial intelligence model. For example, the trained model may include a diffusion model.

[0035] For example, the trained model may be located within an electronic device (200) or within an external electronic device (e.g., a server). For example, at least one processor (210) may provide (or input) a prompt to the trained model located within the electronic device (200) or provide (or input) a prompt to the trained model located within the external electronic device (e.g., a server) by transmitting a prompt to the external electronic device. However, it is not limited thereto.

[0036] For example, information about an object can be described as a prompt. For example, the prompt can be described as natural language text to be provided (or input) to a trained model. For example, the prompt can be generated based on user input (e.g., text input) or by a prompt generator. For example, the prompt generator can be composed of a trained model.

[0037] For example, information about an object may be described as information about an object to be created as a 3D mesh model. For example, information about an object may include the object's appearance, the object's shape, the object's pose, the object's state, the object's color, the object's properties, the object's texture, and / or the object's size. However, it is not limited thereto. For example, information about an object may include text describing the object, text indicating the object, and / or text representing the object.

[0038] For example, the prompt may further include information regarding the rotation of an object to generate a generated image representing an object viewed from multiple directions. For example, the information regarding the rotation of an object may include text requesting the rotation of the object. For example, the information regarding the rotation of an object may include the direction of rotation, the speed of rotation, the continuity of rotation, and / or the angle of rotation, but is not limited thereto.

[0039] In operation 310, at least one processor (210) can acquire a generated image generated by a trained model using information representing an object viewed from multiple directions. Generating a generated image using a trained model is exemplified in the description of FIG. 4.

[0040] Figure 4 illustrates an example of a trained model that generates a generated image.

[0041] Referring to FIG. 4, the trained model (400) may include a generative artificial intelligence model for generating a generated image. For example, the trained model (400) may be configured to generate a generated image using text, images, and / or videos. For example, the trained model (400) may include an artificial neural network model comprising a plurality of layers and / or operations (or computations). As an example without limitation, the trained model (400) may include one of a feedforward neural network (FNN), a deep neural network (DNN), a convolutional neural network (CNN), a region with convolutional neural network (R-CNN), a region proposal network (RPN), a recurrent neural network (RNN), a stacking-based deep neural network (S-DNN), a state-space dynamic neural network (S-SDNN), a deconvolution network, a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, a fully convolutional network, a long short-term memory (LSTM) network, a classification network, or a combination of two or more of these, but is not limited thereto.

[0042] At least one processor (210) may provide information (405) about an object to a trained model (400). For example, information about the object may include the appearance of the object (e.g., ancient stone sculpture bust), the state of the object (e.g., floating in mid-air), the shape of the object (e.g., consistent shape), the color of the object (e.g., full color or no black-and-white), and / or the texture of the object (e.g., ancient relic texture).

[0043] For example, the information (405) may include a prompt. For example, the prompt may include information about the rotation of the object. For example, the information about the rotation of the object may include seamless rotation, the angle of rotation (e.g., rotation of 360° or more), the direction of rotation (e.g., rotation in one direction), and / or the speed of rotation (e.g., slow rotation).

[0044] For example, the prompt may include additional information regarding contrast. For example, the information regarding contrast may include low contrast, no harsh contrast, and / or low contrast brightness. For example, the prompt may include additional information regarding lighting (e.g., soft diffused lighting or frontal lighting). For example, the prompt may include additional information regarding shadows (e.g., minimal shadow or subtle shadow only). For example, the prompt may include additional information regarding the object's gloss (e.g., non-gloss). For example, the prompt may include additional information regarding the object's background (e.g., transparent background or plain background). For example, the prompt may include additional information regarding resolution (e.g., high-resolution), but is not limited thereto.

[0045] For example, at least one processor (210) can acquire a generated image (415) generated by using information (405) by a trained model (400). For example, the generated image (415) can represent an object viewed from multiple directions.

[0046] As another example, at least one processor (210) may further provide the trained model (400) with a reference image (410) containing an object and / or a reference video containing an object. For example, the reference image (410) containing an object and / or a reference video containing an object may be referred to as reference data. For example, the trained model (400) may generate a generated image (415) containing an object in the reference image (410) and / or reference video using the reference image (410) containing an object and / or a reference video containing an object. For example, an object in the reference image (410) and / or reference video may correspond to an object in the generated image (415) generated by the trained model (400). For example, at least one processor (210) can obtain a generated image (415) containing an object intended by the user by further providing the trained model (400) with a reference image (410) containing an object and / or a reference video containing an object.

[0047] For example, the term “generated image” may be used interchangeably with the term “video.” Here, the generated image may refer to a video and / or a single frame image contained in said video. For example, the generated image (415) may be a video representing an object based on at least one of information about smooth rotation and / or information about the direction of rotation. A video generated by the trained model (400) is exemplified in the description of FIG. 5.

[0048] FIG. 5 illustrates an example of a video representing an object seen in multiple directions as the object is continuously rotated.

[0049] Referring to FIG. 5, the states (500), (510), and (515) can be described as states representing objects (505) that are viewed from different directions as the video (502) is played. For example, an object (505) in the video (502) can be described as an object indicated by information (e.g., information (405) in FIG. 4). For example, the video (502) may include an object (505) having an appearance, shape, pose, state, color, characteristics, texture, and / or size indicated by a prompt, but is not limited thereto. For example, an object (505) in the video (502) can be determined based on a prompt (or prompting).

[0050] For example, as the video (502) is played, an object (505) within the video (502) may be rotated. For example, the rotation of the object (505) within the video (502) may be indicated by the prompt. For example, the object (505) within the video (502) may be rotated in a direction, speed, continuity, and / or angle indicated by the prompt. For example, the object (505) within the video (502) may be rotated smoothly (or continuously). For example, the object (505) within the video (502) may be rotated in one direction. For example, the object (505) within the video (502) may be rotated by a reference angle (e.g., 360°). For example, the rotation of the object (505) within the video (502) may be determined based on the prompt (or prompting).

[0051] For example, while the object (505) in the video (502) is rotated, the object (505) in the video (502) may have a consistent appearance (or shape). For example, within the frames of the video (502), the object (505) may not be physically deformed or distorted. For example, the object (505) in the video (502) may have a contrast (e.g., de-lighting) indicated by a prompt. For example, depending on the contrast indicated by a prompt, the object (505) in the video (502) may be displayed in detail. For example, the video (502) may have a resolution (e.g., 1080p (progressive)) indicated by a prompt.

[0052] As the video (502) is played, the object (505) can be switched from state (500) to state (510). As the video (502) is played, the object (505) can be switched from state (510) to state (515). For example, in state (500), the object (505) can be viewed from the front direction. For example, in state (510), the object (505) can be viewed from the side direction. For example, in state (515), the object (505) can be viewed from the rear direction. For example, the video (502) may include a first frame representing the object (505) viewed from the front direction, a second frame representing the object (505) viewed from the side direction, and a third frame representing the object (505) viewed from the rear direction.

[0053] For example, the video (502) may include consecutive frames between the first frame and the second frame. For example, because the object (505) in the video (502) rotates smoothly (or continuously), the consecutive frames between the first frame and the second frame may represent the object (505) as seen in consecutive directions between the frontal direction and the lateral direction.

[0054] For example, the video (502) may include consecutive frames between the second frame and the third frame. For example, because the object (505) in the video (502) rotates smoothly (or continuously), the consecutive frames between the second frame and the third frame may represent the object (505) being viewed in consecutive directions between the side direction and the rear direction.

[0055] For example, since the video (502) represents an object (505) that is seen in multiple directions (e.g., consecutive directions between the front direction and the side direction and consecutive directions between the side direction and the rear direction), the 3D mesh model of the object (505) created using the video (502) can represent the object (505) that is seen in multiple directions.

[0056] Referring again to FIG. 3, in operation 330, at least one processor (210) can generate a 3D mesh model of an object using video. For example, the 3D mesh model of the object can be generated based on photogrammetry. At least one processor (210) can generate a point cloud of the object by matching representations of the object within consecutive frames of the video. At least one processor (210) can generate a 3D mesh model of the object based on the point cloud of the object.

[0057] As another example, at least one processor (210) may inquire whether to generate a 3D mesh model of an object (505) using the generated image (or video (502)) based on acquiring the generated image (or video (502)). A UI for inquiring whether to generate a 3D mesh model of an object (505) using the generated image (or video (502)) is exemplified in the description of FIG. 6.

[0058] Figure 6 illustrates an example of a user interface (UI) that asks whether to generate a 3D mesh model using a generated image.

[0059] Referring to FIG. 6, the state (600) can be described as a state in which a generated image (415) is displayed. In the state (600), at least one processor (210) may display a UI (or window) (605) through a display (230) based on acquiring the generated image (415) and inquiring whether to generate a 3D mesh model using the generated image (415). For example, the generated image (415) and the UI (605) may be displayed simultaneously. For example, at least one processor (210) may provide an object that is viewed from multiple directions by displaying the generated image (415) through the display (230).

[0060] For example, the UI (605) may include text asking whether to create a 3D mesh model using an object in the generated image (415) (e.g., “Would you like to create a 3D model of the above object?”). For example, the UI (605) may include a UI object (610) indicating not to create a 3D mesh model using the generated image (415) and a UI object (620) indicating to create a 3D mesh model using the generated image (415).

[0061] At least one processor (210) can identify user input (615) for a UI object (610). For example, user input (615) may include touch input having a contact point on the UI object (610). For example, user input (615) may be received via a display (230) (e.g., a touchscreen). For another example, user input (615) may include voice input (or verbal input) received via a microphone (not shown). For another example, user input (615) may include input received via an external electronic device (e.g., a mouse). However, it is not limited thereto.

[0062] At least one processor (210) may refrain from (or stop, or skip, or bypass, or not generate) a 3D mesh model using a generated image (415) based on user input (615). For example, at least one processor (210) may obtain a different generated image by providing a prompt (e.g., information (405) in FIG. 4) to a trained model (e.g., trained model (400) in FIG. 4) based on user input (615). As another example, at least one processor (210) may display a field for prompt input based on user input (615). At least one processor (210) may generate a different prompt based on the input identified through said field. At least one processor (210) may obtain a different generated image by providing a different prompt to the trained model. At least one processor (210) may generate a 3D mesh model of an object using the different generated image.

[0063] At least one processor (210) can identify user input (625) for a UI object (620). For example, user input (625) may include touch input having a contact point on the UI object (620). For example, user input (625) may be received via a display (230) (e.g., a touchscreen). For another example, user input (625) may include voice input (or verbal input) received via a microphone (not shown). For another example, user input (625) may include input received via an external electronic device (e.g., a mouse). However, it is not limited thereto.

[0064] At least one processor (210) can generate a 3D mesh model of an object within a generated image (415) using a generated image (415) based on user input (625).

[0065] As another example, at least one processor (210) can acquire an intermediate image by providing information (e.g., information (405) of FIG. 4) to a trained model. For example, the intermediate image may be described as a reference image containing an object. At least one processor (210) may display, through a display (230), an intermediate image and a UI object inquiring whether to generate a generated image using the intermediate image. For example, the UI object inquiring whether to generate a generated image using the intermediate image may correspond to a UI object (605). At least one processor (210) can acquire (or generate) a generated image (415) using the intermediate image based on user input to the UI object inquiring whether to generate a generated image using the intermediate image.

[0066] Creating a 3D mesh model of an object using a generated image (415) is exemplified in the description of FIG. 7.

[0067] Figure 7 illustrates an example of generating a 3D mesh model of an object based on the point cloud of the object.

[0068] Referring to FIG. 7, the state (700) can be described as a state for generating a point cloud (705) of an object. For example, the point cloud (705) may include points corresponding to the outer surface of an object in 3D space. In the state (700), at least one processor (210) may generate the point cloud (705) of an object using a generated image (or consecutive frames of a video). For example, at least one processor (210) may identify the direction in which the object in the generated image (or consecutive frames of a video) is viewed (e.g., shooting position). For example, the direction may be stored in association with the generated image (or frames of a video). For example, the direction may be stored as metadata in the electronic device (200).

[0069] For example, at least one processor (210) can set coordinate values ​​of an object within a generated image (or frames of a video) based on the direction. The coordinate values ​​of the object can be set as x, y, and z values ​​in 3D space. For example, at least one processor (210) can generate a point cloud (705) of an object by connecting representations of the object included in the generated image (or consecutive frames of a video) based on the coordinate values ​​of the object. For example, representations of the object can be connected by matching feature points of the object.

[0070] As another example, at least one processor (210) can determine one or more frames containing representations of an object having a consistent shape among the frames by comparing consecutive frames of video. For example, one or more frames can be described as images refined on a frame-by-frame basis. At least one processor (210) can generate a point cloud (705) using one or more frames. For example, a point cloud (705) generated using one or more frames (e.g., including representations of an object having a consistent shape) can be distributed over an object having a consistent appearance (or shape).

[0071] The electronic device can transition from state (700) to state (710) by generating a 3D mesh model (715) of an object based on a point cloud (705). For example, the point cloud (705) may be distributed over the object (or the outer surface of the object). In state (710), at least one processor (210) may obtain lines connecting the points by connecting the points within the point cloud (705). At least one processor (210) may use the lines to generate surfaces enclosed by the lines. For example, the surfaces generated by connecting the points within the point cloud (705) may be composed of polygons (e.g., triangles). For example, the surfaces may be defined as polygons. For example, the surfaces may constitute the appearance (or shape) of the object. At least one processor (210) can obtain (or generate) a 3D mesh model (715) of an object based on the faces.

[0072] For example, since the 3D mesh model (715) of an object is generated using a generated image (or video) created by a trained model, it may not be necessary to photograph the object (105). For example, at least one processor (210) can generate the 3D mesh model (715) of an object without photographing the object through a camera. For example, at least one processor (210) can reduce the cost and time required to acquire multiple images representing the object as viewed from multiple directions. For example, at least one processor (210) can generate the 3D mesh model (715) of a virtual object (or unreal object, or object that is difficult to photograph) by generating the 3D mesh model (715) of an object using a generated image (or video) created by a trained model.

[0073] For example, at least one processor (210) may perform post-processing on the 3D mesh model (715) of an object to enhance the detail (or quality) of the 3D mesh model (715) of the object. Performing post-processing on the 3D mesh model (715) of an object is exemplified in the description of FIG. 8.

[0074] Figure 8 illustrates an example of performing post-processing on a 3D mesh model of an object.

[0075] Referring to FIG. 8, the state (800) can be described as a state for performing post-processing on the 3D mesh model (805) of an object. In the state (800), at least one processor (210) can perform retopology on the 3D mesh model (805) of an object. For example, retopology can be described as optimizing the topology of the 3D mesh model (805). For example, the 3D mesh model (805) of an object may be composed of polygons. For example, the topology of the 3D mesh model (805) of an object may be described as the connection relationships of the polygons of the 3D mesh model (805) of an object. For example, at least one processor (210) can provide natural motion of the 3D mesh model (805) of an object by optimizing the topology of the 3D mesh model (805) of an object. For example, retopology can be performed by a software application (e.g., Zbrush).

[0076] At least one processor (210) can perform physically based rendering (PBR) on the 3D mesh model (805) of an object. For example, by performing physically based rendering, at least one processor (210) can represent light reflected by the 3D mesh model (805) of an object (e.g., diffuse reflection or specular reflection). For example, the light reflected by the 3D mesh model (805) of an object may be determined by the roughness (or glossiness) of the surface of the 3D mesh model (805) of the object. For example, texture processing of the 3D mesh model (805) of an object may be performed based on physically based rendering. For example, physically based rendering may be performed by a software application (e.g., substance painter).

[0077] At least one processor (210) can perform UV mapping (or UV unwrap) on the 3D mesh model (805) of the object. For example, UV mapping can be described as mapping the 3D mesh model (805) of the object onto a plane. For example, at least one processor (210) can obtain 2D (two-dimensional) images (810) of the object by performing UV mapping on the 3D mesh model (805) of the object. For example, the 2D images (810) of the object may be referred to as islands. For example, at least one processor (210) can perform texture processing (e.g., delighting) on ​​the 2D images (810) of the object.

[0078] The electronic device (200) can transition from state (800) to state (815) by performing post-processing on the 3D mesh model (805) of an object. In state (815), at least one processor (210) can obtain the 3D mesh model (820) of the object that has undergone post-processing. For example, the 3D mesh model (820) of the object that has undergone post-processing can be optimized for use in a game engine. For example, at least one processor (210) can provide natural motion of the object through the 3D mesh model (820) of the object that has undergone post-processing. For example, at least one processor (210) can provide enhanced user experiences (UX) through the 3D mesh model (820) of the object that has undergone post-processing.

[0079] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.

[0080] The electronic device described above (e.g., the electronic device (200) of FIG. 2) may include a memory (e.g., the memory (220) of FIG. 2) that stores instructions and includes one or more storage media, and at least one processor (e.g., at least one processor (210) of FIG. 2) that includes a processing circuit. The instructions may cause the electronic device to provide information about an object to a trained model when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain a generated image created by the trained model using the information, which represents the object viewed from multiple directions when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to create a three-dimensional mesh model of the object using the generated image when executed individually or collectively by the at least one processor.

[0081] For example, the information may include at least one of information about the consistent shape of the object, information about low contrast, information about seamless rotation, information about the rotation of the object, and / or information about the direction of the rotation of the object.

[0082] For example, the generated image may be a video representing the object rotating based on at least one of the information regarding the smooth rotation and / or the information regarding the direction of the rotation.

[0083] For example, the instructions may cause the electronic device to generate a point cloud of the object by matching representations of the object within consecutive frames of the video when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to generate a 3D mesh model of the object based on the point cloud of the object when executed individually or collectively by the at least one processor.

[0084] For example, the above instructions may cause the electronic device to further provide the trained model with a reference image containing the object when executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to obtain a video representing the object of the reference image using the reference image and the information when executed individually or collectively by the at least one processor.

[0085] For example, the instructions may cause the electronic device to identify consecutive frames of the video when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to determine one or more frames containing a representation of the object having a consistent shape among the frames by comparing the frames when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to generate the 3D mesh model of the object using one or more frames among the frames when executed individually or collectively by the at least one processor.

[0086] For example, the above instructions may cause the electronic device to perform post-processing on the 3D mesh model of the object when executed individually or collectively by the at least one processor. The post-processing may include retopology and / or physically based rendering (PBR).

[0087] For example, the electronic device may further include a display. The instructions may cause the electronic device to obtain an intermediate image generated by the trained model using the information when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to display, through the display, the intermediate image and a UI object inquiring whether to generate a generated image using the intermediate image when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to generate the generated image of the object using the intermediate image based on user input instructing to generate a generated image using the intermediate image identified through the UI object when executed individually or collectively by the at least one processor.

[0088] For example, the electronic device may further include a display. The instructions may cause the electronic device to obtain the generated image created by the trained model using the information when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to display, through the display, the generated image and a UI object inquiring whether to create a 3D mesh model using the generated image when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to create the 3D mesh model of the object using the generated image based on user input instructing to create a 3D mesh model using the generated image identified through the UI object when executed individually or collectively by the at least one processor.

[0089] The above-described method may be performed within an electronic device. The method may include an operation of providing information about an object to a trained model. The method may include an operation of obtaining a generated image created by the trained model using the information, which represents the object viewed from multiple directions. The method may include an operation of creating a three-dimensional mesh model of the object using the generated image.

[0090] For example, the information may include at least one of information about the consistent shape of the object, information about low contrast, information about seamless rotation, information about the rotation of the object, and / or information about the direction of the rotation of the object.

[0091] For example, the generated image may be a video representing the object rotating based on at least one of the information regarding the smooth rotation and / or the information regarding the direction of the rotation.

[0092] For example, the above method may include the operation of generating a point cloud of the object by matching representations of the object within consecutive frames of the video. The above method may include the operation of generating a 3D mesh model of the object based on the point cloud of the object.

[0093] For example, the above method may include the operation of further providing a reference image containing the object to the trained model. The above method may include the operation of obtaining a video representing the object of the reference image using the reference image and the information.

[0094] For example, the method may include an operation of identifying consecutive frames of the video. The method may include an operation of determining one or more frames among the frames that include a representation of the object having a consistent shape by comparing the frames. The method may include an operation of generating the 3D mesh model of the object using the one or more frames among the frames.

[0095] For example, the above method may include an operation to perform post-processing on the 3D mesh model of the object. The post-processing may include retopology and / or physically based rendering (PBR).

[0096] For example, the electronic device may further include a display. The method may include an operation of obtaining an intermediate image generated by the trained model using the information. The method may include an operation of displaying, through the display, the intermediate image and a UI object inquiring whether to generate a generated image using the intermediate image. The method may include an operation of generating the generated image of the object using the intermediate image based on user input instructing to generate a generated image using the intermediate image identified through the UI object.

[0097] For example, the electronic device may further include a display. The method may include an operation of obtaining the generated image created by the trained model using the information. The method may include an operation of displaying the generated image and a UI object inquiring whether to create a 3D mesh model using the generated image through the display. The method may include an operation of creating the 3D mesh model of the object using the generated image based on user input instructing to create a 3D mesh model using the generated image identified through the UI object.

[0098] The above-described non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to provide information about an object to a trained model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain a generated image created by the trained model using the information, which represents the object viewed from multiple directions when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to generate a three-dimensional mesh model of the object using the generated image when executed by the electronic device.

[0099] For example, the information may include at least one of information about the consistent shape of the object, information about low contrast, information about seamless rotation, information about the rotation of the object, and / or information about the direction of the rotation of the object.

[0100] For example, the generated image may be a video representing the object rotating based on at least one of the information regarding the smooth rotation and / or the information regarding the direction of the rotation.

[0101] For example, the one or more programs may include instructions that cause the electronic device to generate a point cloud of the object by matching representations of the object within consecutive frames of the video when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to generate a 3D mesh model of the object based on the point cloud of the object when executed by the electronic device.

[0102] For example, the one or more programs may include instructions that cause the electronic device to further provide the trained model with a reference image containing the object when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain a video representing the object of the reference image using the reference image and the information when executed by the electronic device.

[0103] For example, the one or more programs may include instructions that cause the electronic device to identify consecutive frames of the video when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to determine, when executed by the electronic device, one or more frames containing a representation of the object having a consistent shape among the frames by comparing the frames. The one or more programs may include instructions that cause the electronic device to generate the 3D mesh model of the object using the one or more frames among the frames when executed by the electronic device.

[0104] For example, the one or more programs may include instructions that cause the electronic device to perform post-processing on the 3D mesh model of the object when executed by the electronic device. The post-processing may include retopology and / or physically based rendering (PBR).

[0105] For example, the electronic device may further include a display. The one or more programs may include instructions that cause the electronic device to obtain an intermediate image generated by the trained model using the information when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display, through the display, the intermediate image and a UI object inquiring whether to generate a generated image using the intermediate image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to generate the generated image of the object using the intermediate image based on user input instructing to generate a generated image using the intermediate image identified through the UI object when executed by the electronic device.

[0106] For example, the electronic device may further include a display. The one or more programs may include instructions that cause the electronic device to obtain the generated image created by the trained model using the information when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display, through the display, the generated image and a UI object inquiring whether to create a 3D mesh model using the generated image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to create the 3D mesh model of the object using the generated image based on user input instructing to create a 3D mesh model using the generated image identified through the UI object when executed by the electronic device.

[0107] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.

Claims

1. In an electronic device, Memory that stores instructions and includes one or more storage media; and It includes at least one processor comprising a processing circuit, and When the above instructions are executed individually or collectively by the at least one processor: Provide information about the object to the trained model; Acquiring a generated image, which is generated by the trained model using the information representing the object viewed from multiple directions; and To generate a 3D (three-dimensional) mesh model of the object using the generated image above, causing the above electronic device, Electronic device.

2. In Claim 1, The above information is, at least one of information regarding the consistent shape of the object, information regarding low contrast, information regarding seamless rotation, information regarding the rotation of the object, and / or information regarding the direction of the rotation of the object. Electronic device.

3. In Claim 2, The generated image above is, A video representing the object rotating based on at least one of the information regarding the smooth rotation and / or the information regarding the direction of the rotation, Electronic device.

4. In Claim 3, When the above instructions are executed individually or collectively by the at least one processor: By matching representations of the object within consecutive frames of the video, a point cloud of the object is generated; and To generate the 3D mesh model of the object based on the point cloud of the object, causing the above electronic device, Electronic device.

5. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor: Further providing a reference image including the above object to the trained model; and To obtain a video representing the object of the reference image using the reference image and the information above, causing the above electronic device, Electronic device.

6. In Claim 3, When the above instructions are executed individually or collectively by the at least one processor: Identify consecutive frames of the above video; By comparing the above frames, one or more frames containing a representation of the object having a consistent shape among the above frames are determined; and To generate the 3D mesh model of the object using one or more of the frames among the above frames, causing the above electronic device, Electronic device.

7. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor: To perform post-processing on the 3D mesh model of the above object, Causing the above electronic device, The above post-processing is, including retopology and / or physically based rendering (PBR), Electronic device.

8. In Claim 1, Includes additional displays; When the above instructions are executed individually or collectively by the at least one processor: Using the above information, obtain an intermediate image generated by the above-mentioned trained model; Displaying a UI object through the above display that inquires whether to generate a generated image using the intermediate image and the intermediate image; and Based on user input instructing to generate a generated image using the intermediate image identified through the UI object, to generate the generated image of the object using the intermediate image, causing the above electronic device, Electronic device.

9. In Claim 1, Includes additional displays; When the above instructions are executed individually or collectively by the at least one processor: Using the above information, obtain the generated image created by the above-mentioned trained model; Displaying, through the above display, a UI object that inquires whether to generate a 3D mesh model using the generated image and the generated image; and Based on user input instructing to generate a 3D mesh model using the generated image identified through the UI object, to generate the 3D mesh model of the object using the generated image, causing the above electronic device, Electronic device.

10. A method executed within an electronic device, wherein the method comprises: The action of providing information about an object to a trained model; An operation to acquire a generated image, generated by the trained model using the information representing the object viewed from multiple directions; and A method comprising the operation of generating a 3D (three-dimensional) mesh model of the object using the generated image, method.

11. In Claim 10, The above information is, at least one of information regarding the consistent shape of the object, information regarding low contrast, information regarding seamless rotation, information regarding the rotation of the object, and / or information regarding the direction of the rotation of the object. method.

12. In Claim 11, The generated image above is, A video representing the object rotating based on at least one of the information regarding the smooth rotation and / or the information regarding the direction of the rotation, method.

13. In claim 12, the method comprises: The operation of generating a point cloud of the object by matching representations of the object within consecutive frames of the video; and The operation of generating the 3D mesh model of the object based on the point cloud of the object, method.

14. In claim 10, the method comprises: The operation of further providing a reference image including the above object to the trained model; and The operation of obtaining a video representing the object of the reference image using the reference image and the information, method.

15. In claim 12, the method comprises: An operation to identify consecutive frames of the above video; The operation of determining one or more frames containing a representation of the object having a consistent shape among the frames by comparing the frames; and The operation of generating the 3D mesh model of the object using one or more of the frames among the above frames, method.

16. In claim 10, the method comprises: It includes an operation to perform post-processing on the 3D mesh model of the above object, and The above post-processing is, including retopology and / or physically based rendering (PBR), method.

17. In Claim 10, The above electronic device is, Includes more displays, The above method is: An operation to acquire an intermediate image generated by the trained model using the above information; An operation to display a UI object through the above display that inquires whether to generate a generated image using the intermediate image and the intermediate image; and Based on user input instructing to generate a generated image using the intermediate image identified through the UI object, the operation of generating the generated image of the object using the intermediate image, method.

18. In Claim 10, The above electronic device is, Includes more displays, The above method is: An operation to acquire the generated image created by the trained model using the above information; An operation to display, through the above display, a UI object that inquires whether to generate a 3D mesh model using the generated image and the generated image; and A method comprising an operation to generate the 3D mesh model of the object using the generated image based on user input instructing to generate the 3D mesh model using the generated image identified through the UI object. method.

19. In a non-transient computer-readable storage medium storing one or more programs, When one or more of the above programs are executed by an electronic device: Provide information about the object to the trained model; Acquiring a generated image, which is generated by the trained model using the information representing the object viewed from multiple directions; and To generate a 3D (three-dimensional) mesh model of the object using the generated image above, Including instructions that cause the above electronic device, Non-transient computer-readable storage media.

20. In Claim 19, The above information is, at least one of information regarding the consistent shape of the object, information regarding low contrast, information regarding seamless rotation, information regarding the rotation of the object, and / or information regarding the direction of the rotation of the object. Non-transient computer-readable storage media.