Modeling method, related electronic equipment and storage medium
By using a regular RGB camera to capture images on a terminal device and generating 3D models using image correlation, the problem of high hardware requirements for existing 3D modeling is solved, achieving an efficient and low-cost 3D modeling process and improving the user experience.
Patent Information
- Application Number
- CN202510890245.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-26
- Publication Date
- 2025-11-21
AI Technical Summary
Existing 3D modeling methods involve complex data acquisition processes and require high-end hardware, such as LIDAR sensors or RGB-D cameras. Furthermore, the modeling process is cumbersome and results in a poor user experience.
By using a regular RGB camera to capture multiple frames of images on a terminal device, and utilizing the correlation between the images to generate a 3D model, the data can be processed on the terminal device or in the cloud, reducing hardware requirements and simplifying the operation process.
It enables 3D modeling with low hardware requirements, improves modeling efficiency and user experience, and is suitable for terminal devices with weak computing power.
Smart Images

Figure CN120997269A_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 202110715044.2 and the original application date is June 26, 2021. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This application relates to the field of three-dimensional reconstruction, and more particularly to a modeling method and related electronic devices and storage media. Background Technology
[0003] 3D reconstruction applications / software can be used to create 3D models of objects. Currently, when creating 3D models, users need to first use mobile tools (such as mobile phones, cameras, etc.) to collect the data required for 3D modeling (such as images, depth information, etc.). Then, the 3D reconstruction application can use the collected data to reconstruct the object in 3D and obtain the corresponding 3D model.
[0004] However, current 3D modeling methods involve complex processes, including collecting the data required for 3D modeling and reconstructing the object using that data. These processes also place high demands on the hardware. For example, collecting the data requires the acquisition device (such as the aforementioned mobile tool) to be equipped with specialized hardware like a LiDAR (light detection and ranging) sensor or an RGB depth (RGB-D) camera. Furthermore, reconstructing the object using the acquired data requires the processing device running the 3D reconstruction application to be equipped with a high-performance dedicated graphics card. Summary of the Invention
[0005] This application provides a modeling method, related electronic devices, and storage media, which simplifies the process of collecting data required for 3D modeling and the process of 3D reconstruction of objects based on the collected data, and has low requirements for device hardware.
[0006] In a first aspect, embodiments of this application provide a modeling method, which is applied to a terminal device, and the method includes:
[0007] The terminal device displays a first interface, which includes the image captured by the terminal device. Responding to the acquisition operation, the terminal device acquires multiple frames of images corresponding to the target object to be modeled, and obtains the correlation relationships between the multiple frames. Based on the multiple frames and the correlation relationships between them, the terminal device obtains the 3D model corresponding to the target object. The terminal device then displays the 3D model corresponding to the target object.
[0008] During the acquisition of multiple frames of images corresponding to the target object, the terminal device displays a first virtual bounding volume; the first virtual bounding volume includes multiple facets. In response to the acquisition operation, the terminal device acquires multiple frames of images corresponding to the target object to be modeled, and obtains the correlation relationships between the multiple frames of images, including:
[0009] When the terminal device is in the first pose, the terminal device acquires the first image and changes the display effect of the corresponding patch of the first image; when the terminal device is in the second pose, the terminal device acquires the second image and changes the display effect of the corresponding patch of the second image; after changing the display effect of the multiple patches of the first virtual bounding body, the terminal device obtains the correlation relationship between the multiple frames of images based on the multiple patches.
[0010] For example, for each keyframe, the terminal device can determine the matching information of the keyframe based on the association between the patch corresponding to the keyframe and other patches. The association between the patch corresponding to the keyframe and other patches can include: in the patch model, which patches correspond to the four cardinal directions (up, down, left, right) of the patch corresponding to the keyframe.
[0011] For example, taking a patch model with two layers, each containing 20 patches, and assuming that the patch corresponding to a certain keyframe is patch 1 in the first layer, the association between patch 1 and other patches can include: the patch below patch 1 is patch 21, the patch to the left of patch 1 is patch 20, and the patch to the right of patch 1 is patch 2. Based on the aforementioned association between patch 1 and other patches, the mobile phone can determine that other keyframes associated with this keyframe include the keyframe corresponding to patch 21, the keyframe corresponding to patch 20, and the keyframe corresponding to patch 2. Therefore, the mobile phone can obtain the matching information of this keyframe, including the identification information of the keyframe corresponding to patch 21, the identification information of the keyframe corresponding to patch 20, and the identification information of the keyframe corresponding to patch 2.
[0012] Understandably, for mobile phones, the relationships between different patches in the patch model are known quantities.
[0013] In this modeling method (or 3D modeling method), the terminal device only needs a regular RGB camera to collect the data required for 3D modeling. The data acquisition process does not require specialized hardware such as a LiDAR sensor or RGB-D camera on the terminal device. Based on multiple frames and the relationships between them, the terminal device obtains the 3D model of the target object, effectively reducing the computational load and improving the efficiency of the 3D modeling process.
[0014] Furthermore, in this method, users only need to perform the data acquisition operations required for 3D modeling on the terminal device, and then view or preview the final 3D model on the terminal device. For users, all operations are completed on the terminal device, making the operation simpler and the user experience better.
[0015] In one possible design, the terminal device includes a first application, and before the terminal device displays a first interface, the method further includes: the terminal device displaying a second interface in response to an operation of opening the first application.
[0016] The terminal device displays a first interface, including: the terminal device displays the first interface in response to an operation that launches the 3D modeling function of the first application on the second interface.
[0017] For example, the second interface may include a function control for activating the 3D modeling function. The user can click or touch this function control on the second interface, and the phone can respond to this action by activating the 3D modeling function of the first application. In other words, the user's action of clicking or touching this function control on the second interface is equivalent to activating the 3D modeling function of the first application on the second interface.
[0018] In some embodiments, the first virtual bounding body includes one or more layers, with the plurality of patches distributed in the one or more layers.
[0019] For example, in one implementation, the structure of the patch model may include two layers, top and bottom, with each layer including multiple patches. In another implementation, the structure of the patch model may include three layers, top, middle, and bottom, with each layer including multiple patches. In yet another implementation, the structure of the patch model may be a single layer composed of multiple patches. No restrictions are imposed here.
[0020] Optionally, the method further includes: the terminal device displaying a first prompt message, the first prompt message being used to remind the user to place the target object in the center of the captured image.
[0021] For example, the first prompt could be "Please place the target object in the center of the screen".
[0022] Optionally, the method further includes: the terminal device displaying a second prompt message; the second prompt message is used to remind the user to adjust one or more of the following: the shooting environment of the target object, the shooting method of the target object, and the screen ratio of the target object.
[0023] For example, the second prompt could be "Place the object on a solid-color plane with soft lighting, take a picture around the object, and make sure the object has a large and complete screen ratio."
[0024] In this embodiment, after the user adjusts the shooting environment and screen ratio of the target object according to the prompts in the second prompt information, the subsequent data collection process can be faster and the quality of the collected data can be better.
[0025] Optionally, before the terminal device obtains the 3D model corresponding to the target object based on the multi-frame images and the correlation between the multi-frame images, the method further includes: the terminal device detecting the operation of generating the 3D model; and the terminal device displaying a third prompt message in response to the operation of generating the 3D model, the third prompt message being used to prompt the user that the target object is being modeled.
[0026] For example, the third prompt message could be "Modeling in progress".
[0027] Optionally, after the terminal device obtains the three-dimensional model corresponding to the target object based on the multi-frame images and the correlation between the multi-frame images, the method further includes: the terminal device displays a fourth prompt message, which is used to prompt the user that the modeling of the target object has been completed.
[0028] For example, the fourth prompt could be "Modeling completed".
[0029] Optionally, the terminal device displaying the three-dimensional model corresponding to the target object further includes: the terminal device responding to an operation of changing the display angle of the three-dimensional model corresponding to the target object, changing the display angle of the three-dimensional model corresponding to the target object; the operation of changing the display angle of the three-dimensional model corresponding to the target object includes dragging the three-dimensional model corresponding to the target object to rotate clockwise or counterclockwise along a first direction.
[0030] The first direction can be any direction, such as horizontal or vertical. The terminal device responds to operations that change the display angle of the 3D model corresponding to the target object. Changing the display angle of the 3D model corresponding to the target object allows for the presentation of the 3D model to the user from different angles.
[0031] Optionally, the terminal device displaying the three-dimensional model corresponding to the target object further includes: the terminal device responding to an operation of changing the display size of the three-dimensional model corresponding to the target object, changing the display size of the three-dimensional model corresponding to the target object; the operation of changing the display size of the three-dimensional model corresponding to the target object includes an operation of enlarging or shrinking the three-dimensional model corresponding to the target object.
[0032] For example, zooming out can be done by a user sliding two fingers inwards (in opposite directions) on the 3D model preview interface, while zooming in can be done by a user sliding two fingers outwards (in opposite directions) on the 3D model preview interface. The 3D model preview interface is the interface on the terminal device that displays the 3D model corresponding to the target object.
[0033] In other implementations, the zoom-in or zoom-out operation performed by the user on the 3D model corresponding to the target object can be a double-click operation, a long-press operation, or the 3D model preview interface can also include a function control that can perform zoom-in or zoom-out operations, etc., without any restrictions.
[0034] In some embodiments, the association between the multiple frames of images includes matching information for each frame of the multiple frames of images; the matching information for each frame of the image includes identification information of other images associated with the image in the multiple frames of images; the matching information for each frame of the image is obtained based on the association between each frame of the image and the corresponding patch, as well as the association between the multiple patches.
[0035] For example, for a keyframe with image number 18, the identification information of other keyframes associated with keyframe with image number 18 is the image number of the other keyframes associated with keyframe with image number 18, such as 26, 45, 59, 78, 89, 100, 449, etc.
[0036] Optionally, the terminal device responds to the acquisition operation by acquiring multiple frames of images corresponding to the target object to be modeled, and obtains the correlation between the multiple frames of images, which further includes: the terminal device determining the target object based on the captured image; when the terminal device acquires multiple frames of images, the position of the target object in the captured image is the center position of the captured image.
[0037] Optionally, the terminal device acquires multiple frames of images corresponding to the target object to be modeled, including: during the process of shooting the target object, the terminal device performs blur detection on each frame of the captured image and acquires the image with a sharpness greater than a first threshold as the image corresponding to the target object.
[0038] For each shooting position (one shooting position can correspond to one patch), the terminal device can obtain some high-quality keyframe images by performing blur detection on the image captured at that shooting position. The keyframe images are the images corresponding to the target object, and the number of keyframe images corresponding to each patch can be one or more.
[0039] Optionally, the terminal device displays a three-dimensional model corresponding to the target object, including: the terminal device displays the three-dimensional model corresponding to the target object in response to an operation of previewing the three-dimensional model corresponding to the target object.
[0040] For example, a terminal device can display a "view" button. When a user clicks this button, the terminal device can respond by displaying a 3D model of the target object. Clicking the "view" button allows the user to preview the 3D model of the target object.
[0041] Optionally, the 3D model corresponding to the target object includes the basic 3D model of the target object and the texture of the target object's surface.
[0042] The texture of the target object's surface can be a texture map of the target object's surface. A 3D model of the target object can be generated based on its basic 3D model and its surface texture. Mapping the target object's surface texture onto the surface of its basic 3D model in a specific way can more realistically reproduce the target object's surface, making the target object appear more lifelike.
[0043] In one possible design, the terminal device is connected to a server; the terminal device obtains a 3D model corresponding to the target object based on the multi-frame images and the correlation between the multi-frame images, including: the terminal device sending the multi-frame images and the correlation between the multi-frame images to the server; and the terminal device receiving the 3D model corresponding to the target object sent from the server.
[0044] In this design, the process of generating a 3D model of the target object based on the multiple frames of images and the relationships between them can be completed on the server side. That is, this design can utilize the server's computing resources to achieve 3D modeling. This design is applicable to scenarios where some terminal devices have limited computing power, improving the versatility of this 3D modeling method.
[0045] Optionally, the method further includes: the terminal device sending to the server the camera intrinsic parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to the multiple frames of images.
[0046] Optionally, the method further includes: the terminal device receiving an indication message from the server, the indication message being used to indicate to the terminal device that the server has completed modeling the target object.
[0047] For example, after receiving the instruction message, the terminal device can display the aforementioned fourth prompt message.
[0048] Optionally, before the terminal device receives the 3D model corresponding to the target object sent by the server, the method further includes: the terminal device sending a download request message to the server, the download request message being used to request the server to download the 3D model corresponding to the target object.
[0049] After receiving the download request message, the server can send the 3D model of the target object to the terminal device.
[0050] Optionally, in some embodiments, when the terminal device collects the data required for 3D modeling of the target object, it can also display the scanning progress on the first interface. For example, the first interface may include a scan button, and the terminal device can display the scanning progress through the circular black fill effect of the scan button on the first interface.
[0051] It is understandable that the UI presentation of the scan button may differ, and the way the phone displays the scan progress on the first screen may also differ; therefore, no restrictions are imposed here.
[0052] In other embodiments, the terminal device may not need to display the scanning progress; the user can understand the scanning progress by observing the lit-up areas in the first virtual bounding volume.
[0053] Secondly, embodiments of this application provide a modeling apparatus that can be applied to a terminal device to implement the modeling method described in the first aspect. The function of this apparatus can be implemented in hardware or by executing corresponding software. The hardware or software includes one or more modules or units corresponding to the aforementioned functions. For example, the apparatus may include a display unit and a processing unit. The display unit and processing unit can be used in conjunction to implement the modeling method described in the first aspect.
[0054] For example, a display unit is used to display a first interface, which includes the image captured by the terminal device.
[0055] The processing unit is configured to, in response to an acquisition operation, acquire multiple frames of images corresponding to the target object to be modeled, and obtain the correlation relationships between the multiple frames of images. Based on the multiple frames of images and the correlation relationships between them, a 3D model corresponding to the target object is obtained.
[0056] The display unit is also used to display the three-dimensional model corresponding to the target object.
[0057] During the acquisition of multiple frames of images corresponding to the target object, the display unit is also used to display a first virtual bounding volume; the first virtual bounding volume includes multiple facets. Specifically, the processing unit is used to acquire a first image and change the display effect of the facets corresponding to the first image when the terminal device is in a first pose; acquire a second image and change the display effect of the facets corresponding to the second image when the terminal device is in a second pose; and after changing the display effect of the multiple facets of the first virtual bounding volume, obtain the correlation relationship between the multiple frames of images based on the multiple facets.
[0058] Optionally, the display unit and the processing unit are also used to implement other display functions and processing functions described in the first aspect above, which will not be elaborated here.
[0059] Optionally, for the implementation method described in the first aspect above, whereby the terminal device sends the multi-frame images and the association between the multi-frame images to the server, and the server generates a three-dimensional model corresponding to the target object based on the multi-frame images and the association between the multi-frame images, the modeling device may further include a sending unit and a receiving unit. The sending unit is used to send the multi-frame images and the association between the multi-frame images to the server, and the receiving unit is used to receive the three-dimensional model corresponding to the target object sent by the server.
[0060] Thirdly, embodiments of this application provide an electronic device, including: a processor; a memory; and a computer program; wherein the computer program is stored in the memory, and when the computer program is executed by the processor, it causes the electronic device to implement the method as described in the first aspect and any possible implementation thereof.
[0061] Fourthly, embodiments of this application provide a computer-readable storage medium comprising a computer program that, when executed on an electronic device, causes the electronic device to implement the method described in the first aspect and any possible implementation thereof.
[0062] Fifthly, embodiments of this application also provide a computer program product, including computer-readable code, which, when executed in an electronic device, causes the electronic device to implement the method described in the first aspect and any possible implementation thereof.
[0063] The beneficial effects of the second to fifth aspects mentioned above can be referred to in the first aspect, and will not be repeated here.
[0064] Sixthly, embodiments of this application also provide a modeling method, the method being applied to a server, the server being connected to a terminal device; the method comprising: the server receiving multiple frames of images corresponding to a target object and the association relationships between the multiple frames of images sent from the terminal device; the server generating a three-dimensional model corresponding to the target object based on the multiple frames of images and the association relationships between the multiple frames of images; and the server sending the three-dimensional model corresponding to the target object to the terminal device.
[0065] This method can utilize server computing resources to achieve 3D modeling. The process of generating a 3D model of the target object based on multiple frames of images and the relationships between them can be completed on the server side. By combining the relationships between the multiple frames of images for 3D modeling, the server's computing load can be effectively reduced, and modeling efficiency can be improved.
[0066] For example, when performing 3D modeling, the server can combine the relationships between the multiple frames of images and perform feature detection and matching on each frame and other images associated with it, without needing to perform feature detection and matching on the image with all other images. This method allows for rapid comparison of two adjacent frames, effectively reducing the server's computational load and improving the efficiency of 3D modeling.
[0067] For example, after determining the mapping relationship between the texture of the first frame image and the surface of the basic 3D model of the target object, the server can combine the matching information of the first frame image to quickly and accurately determine the mapping relationship between the texture of other images associated with the first frame image and the surface of the basic 3D model of the target object. Similarly, for each subsequent frame image, the server can combine the matching information of that image to quickly and accurately determine the mapping relationship between the texture of other images associated with that image and the surface of the basic 3D model of the target object.
[0068] This method can be applied to scenarios where terminal devices have limited computing power, thus improving the universality of this 3D modeling method.
[0069] In addition, this method also possesses the other beneficial effects described in the first aspect above, such as: the process of acquiring the data required for 3D modeling does not rely on special hardware such as LiDAR sensors or RGB-D cameras on the terminal device. The server obtains the 3D model corresponding to the target object based on multiple frames of images and the relationships between them, which can effectively reduce the computational load of the 3D modeling process and improve the efficiency of 3D modeling, etc., which will not be elaborated further here.
[0070] Optionally, the method further includes: the server receiving camera intrinsic parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to the multiple frames of images sent from the terminal device.
[0071] The server generates a 3D model of the target object based on the multiple frames of images and the relationships between them. This includes: the server generating a 3D model of the target object based on the multiple frames of images and the relationships between them, the camera intrinsic parameters, gravity direction information, image name, image number, camera pose information, and timestamps corresponding to the multiple frames of images.
[0072] Optionally, the association between the multiple frames of images includes matching information for each frame of the multiple frames of images; the matching information for each frame of the image includes identification information of other images associated with the image in the multiple frames of images; the matching information for each frame of the image is obtained based on the association between each frame of the image and the corresponding patch, as well as the association between the multiple patches.
[0073] In a seventh aspect, embodiments of this application provide a modeling apparatus that can be applied to a server to implement the modeling method described in the sixth aspect above. The functionality of this apparatus can be implemented in hardware or by executing corresponding software. The hardware or software includes one or more modules or units corresponding to the aforementioned functions. For example, the apparatus may include a receiving unit, a processing unit, and a transmitting unit. The receiving unit, processing unit, and transmitting unit can be used in conjunction to implement the modeling method described in the sixth aspect above.
[0074] For example, the receiving unit can be used to receive multiple frames of images corresponding to the target object and the correlation relationships between the multiple frames of images sent from the terminal device. The processing unit can be used to generate a three-dimensional model corresponding to the target object based on the multiple frames of images and the correlation relationships between the multiple frames of images. The sending unit can be used to send the three-dimensional model corresponding to the target object to the terminal device.
[0075] Optionally, the receiving unit, processing unit, and sending unit can be used to implement all the functions that the server in the method described in the sixth aspect above can perform, and will not be described in detail here.
[0076] Eighthly, embodiments of this application provide an electronic device, including: a processor; a memory; and a computer program; wherein the computer program is stored in the memory, and when the computer program is executed by the processor, causes the electronic device to implement the method as described in the sixth aspect and any possible implementation thereof.
[0077] Ninthly, embodiments of this application provide a computer-readable storage medium comprising a computer program that, when executed on an electronic device, causes the electronic device to implement the method described in the sixth aspect and any possible implementation thereof.
[0078] In a tenth aspect, embodiments of this application also provide a computer program product, including computer-readable code, which, when executed in an electronic device, causes the electronic device to implement the method described in the sixth aspect and any possible implementation thereof.
[0079] The beneficial effects described in aspects seven through ten above can be found in aspect six, and will not be repeated here.
[0080] Eleventhly, embodiments of this application also provide an edge-cloud collaborative system, including: a terminal device and a server, the terminal device being connected to the server; the terminal device displays a first interface, the first interface including the image captured by the terminal device; the terminal device, in response to a acquisition operation, acquires multiple frames of images corresponding to a target object to be modeled, and obtains the correlation between the multiple frames of images; wherein, during the acquisition of the multiple frames of images corresponding to the target object, the terminal device displays a first virtual bounding volume; the first virtual bounding volume includes multiple facets; the terminal device, in response to a acquisition operation, acquires multiple frames of images corresponding to the target object to be modeled, and obtains the correlation between the multiple frames of images, including: when the terminal device is in the first pose, the terminal device... The system acquires a first image and modifies the display effect of the corresponding facets of the first image; when the terminal device is in a second pose, the terminal device acquires a second image and modifies the display effect of the corresponding facets of the second image; after modifying the display effect of the multiple facets of the first virtual bounding body, the terminal device obtains the correlation relationship between the multiple frames of images based on the multiple facets; the terminal device sends the multiple frames of images and the correlation relationship between the multiple frames of images to the server; the server obtains the 3D model corresponding to the target object based on the multiple frames of images and the correlation relationship between the multiple frames of images; the server sends the 3D model corresponding to the target object to the terminal device; and the terminal device displays the 3D model corresponding to the target object.
[0081] The beneficial effects of the eleventh aspect mentioned above can be referred to in the first and sixth aspects, and will not be repeated here.
[0082] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0083] Figure 1 This is a schematic diagram of the composition of the end-to-cloud collaborative system provided in the embodiments of this application;
[0084] Figure 2 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application;
[0085] Figure 3 A schematic diagram of the main interface of a mobile phone provided in an embodiment of this application;
[0086] Figure 4 A schematic diagram of the main interface of the first application provided in this application embodiment;
[0087] Figure 5 A schematic diagram of a 3D modeling data acquisition interface provided in an embodiment of this application;
[0088] Figure 6 Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0089] Figure 7A Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0090] Figure 7B Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0091] Figure 7C Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0092] Figure 7D Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0093] Figure 7EAnother schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0094] Figure 7F Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0095] Figure 8 This is a schematic diagram of the structure of the patch model provided in the embodiments of this application;
[0096] Figure 9 Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0097] Figure 10 Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0098] Figure 11 Another schematic diagram of the 3D modeling data acquisition interface provided in the embodiments of this application;
[0099] Figure 12 A schematic diagram of a 3D model preview interface provided in an embodiment of this application;
[0100] Figure 13 Another schematic diagram of the 3D model preview interface provided in the embodiments of this application;
[0101] Figure 14 Another schematic diagram of the 3D model preview interface provided in the embodiments of this application;
[0102] Figure 15 This is a schematic diagram illustrating a user performing a counter-clockwise rotation operation along the horizontal direction on a 3D model of a toy car, as provided in an embodiment of this application.
[0103] Figure 16 Another schematic diagram of the 3D model preview interface provided in the embodiments of this application;
[0104] Figure 17 This is an illustration of a user performing a scaling-down operation on a 3D model of a toy car, as provided in an embodiment of this application.
[0105] Figure 18 Another schematic diagram of the 3D model preview interface provided in the embodiments of this application;
[0106] Figure 19 A flowchart illustrating the 3D modeling method provided in the embodiments of this application;
[0107] Figure 20 A logical schematic diagram illustrating the 3D modeling method implemented in the end-to-cloud collaborative system provided in this application embodiment;
[0108] Figure 21This is a schematic diagram of the modeling apparatus provided in the embodiments of this application;
[0109] Figure 22 Another schematic diagram of the modeling apparatus provided in the embodiments of this application;
[0110] Figure 23 This is another schematic diagram of the modeling apparatus provided in the embodiments of this application. Detailed Implementation
[0111] The terminology used in the following embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to also include expressions such as “one or more,” unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, “at least one” and “one or more” refer to one or more (including two). The character “ / ” generally indicates that the preceding and following objects are in an “or” relationship.
[0112] References to "one embodiment" or "some embodiments" as used in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes both direct and indirect connections, unless otherwise stated.
[0113] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0114] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0115] 3D reconstruction technology is widely used in virtual reality, augmented reality, extended reality (XR), mixed reality (MR), gaming, film and television, education, and healthcare. For example, 3D reconstruction technology can be used to model characters, props, and vegetation in games, or to model human figures in films and television shows. It can also be used in education for modeling chemical structures and in healthcare for modeling human anatomy.
[0116] Currently, most 3D reconstruction applications / software that can be used for 3D modeling require a personal computer (PC) to implement, while a small number can perform 3D modeling on mobile devices (such as smartphones). When using PC-based 3D reconstruction applications, users first need to use mobile tools (such as smartphones, cameras, etc.) to collect the necessary data (such as images, depth information, etc.) for 3D modeling, and then upload the collected data to the PC. The PC-based 3D reconstruction application can then perform 3D modeling processing based on the uploaded data. When using mobile 3D reconstruction applications, users can directly collect the necessary data using their mobile devices, and the mobile 3D reconstruction application can directly perform 3D modeling processing based on the data collected on the mobile device.
[0117] However, both of the above-mentioned 3D modeling methods rely on specialized hardware such as LiDAR (light detection and ranging) sensors or RGB depth (RGB-D) cameras on mobile devices when users collect the data required for 3D modeling. The data acquisition process for 3D modeling requires high hardware specifications. PC / mobile 3D reconstruction applications also have high hardware requirements for the PC / mobile devices used for 3D modeling; for example, they may require a high-performance dedicated graphics card.
[0118] In addition, the above-mentioned method of implementing 3D modeling on the PC is still quite cumbersome. For example, after the user performs relevant data collection operations on the mobile device, the user not only needs to copy the collected data or transmit it to the PC via the network, but also needs to perform relevant modeling operations on the 3D reconstruction application on the PC.
[0119] Against this background, embodiments of this application provide a 3D modeling method applicable to an end-to-cloud collaborative system comprised of a terminal device and a cloud platform. Here, "end" refers to the terminal device, and "cloud" refers to the cloud platform, which can also be called a cloud server or cloud platform. In this method, the terminal device can collect data required for 3D modeling, preprocess the data, and then upload the preprocessed data to the cloud. The cloud platform can then perform 3D modeling based on the received preprocessed data. The terminal device can download the 3D model obtained from the cloud modeling and provide a preview function for the 3D model.
[0120] In this 3D modeling method, the terminal device only needs a regular RGB camera to collect the data required for 3D modeling. The process of collecting the data for 3D modeling does not require the terminal device to have special hardware such as a LiDAR sensor or an RGB-D camera. The 3D modeling process is completed in the cloud, and it does not require the terminal device to be equipped with a high-performance dedicated graphics card. In other words, this 3D modeling method has low hardware requirements for the terminal device.
[0121] Furthermore, compared to the PC-based 3D modeling method described above, this method only requires users to perform data acquisition operations on the terminal device, and then view or preview the final 3D model on the terminal device. For users, all operations are completed on the terminal device, making the process simpler and providing a better user experience.
[0122] For example, Figure 1 This is a schematic diagram illustrating the composition of the end-to-cloud collaborative system provided in an embodiment of this application. Figure 1 As shown, the end-to-cloud collaborative system provided in this application embodiment may include: cloud 100 and terminal device 200, and the terminal device 200 can connect to cloud 100 through a wireless network.
[0123] In this context, cloud 100 refers to the server. For example, in some embodiments, cloud 100 may be a single server or a server cluster consisting of multiple servers. This application does not limit the implementation architecture of cloud 100.
[0124] Optionally, in this embodiment, the terminal device 200 may be an interactive electronic whiteboard with shooting function, a mobile phone, a wearable device (such as a smartwatch, smart bracelet, etc.), a tablet computer, a laptop computer, a desktop computer, a portable electronic device (such as a laptop computer), an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart TV (such as a smart screen), an in-vehicle computer, a smart speaker, an augmented reality (AR) device, a virtual reality (VR) device, and other smart devices with a display screen. Alternatively, it may be a digital camera, a DSLR camera / mirrorless camera, an action camera, a gimbal camera, a drone, or other professional shooting equipment. This embodiment does not limit the specific type of terminal device.
[0125] It should be understood that when the terminal device is a gimbal camera, drone, or other shooting device, it will also include a display device that provides a shooting interface, used to display the data acquisition interface required for 3D modeling, the 3D model preview interface, etc. For example, the display device for a gimbal camera can be a mobile phone, and the display device for an aerial drone can be a remote control device, etc.
[0126] It should be noted that, Figure 1 An example of a terminal device 200 is provided. However, it should be understood that the terminal device 200 in this end-to-cloud collaborative system may include one or more devices, and the multiple terminal devices 200 may be the same, different, or partially the same, without limitation. The 3D modeling method provided in this application embodiment is a process of achieving 3D modeling through interaction between each terminal device 200 and the cloud 100.
[0127] For example, taking a mobile phone as an example, terminal device 200, Figure 2 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 2As shown, a mobile phone may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.
[0128] Processor 210 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0129] The controller can serve as the nerve center and command center of the mobile phone. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0130] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0131] In some embodiments, the processor 210 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM interface, and / or a USB interface, etc.
[0132] The external storage interface 220 can be used to connect an external storage card, such as a Micro SD card, to expand the phone's storage capacity. The external storage card communicates with the processor 210 through the external storage interface 220 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0133] Internal memory 221 can be used to store computer executable program code, which includes instructions. Processor 210 executes various functions and data processing of the mobile phone by running the instructions stored in internal memory 221.
[0134] The internal memory 221 may further include a program storage area and a data storage area. The program storage area may store the operating system, at least one application required for a function (such as the first application described in this embodiment), etc. The data storage area may store data created during the use of the mobile phone (such as image data, phonebook data, etc.). Furthermore, the internal memory 221 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0135] The charging management module 240 receives charging input from the charger. While charging the battery 242, the charging management module 240 can also power the mobile phone via the power management module 241. The power management module 241 connects to the battery 242, the charging management module 240, and the processor 210. The power management module 241 can also receive input from the battery 242 to power the mobile phone.
[0136] The wireless communication function of a mobile phone can be implemented through antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the mobile phone can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0137] Mobile phones can perform audio functions, such as music playback and recording, through an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, and an application processor.
[0138] The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an accelerometer sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.
[0139] Camera 293 can include various types. For example, camera 293 can include telephoto cameras, wide-angle cameras, or ultra-wide-angle cameras with different focal lengths. Telephoto cameras have a small field of view, suitable for capturing distant objects within a small area; wide-angle cameras have a large field of view; and ultra-wide-angle cameras have a larger field of view than wide-angle cameras, and can be used to capture panoramic or other large-scale scenes. In some embodiments, the telephoto camera with a small field of view can be rotated, thereby allowing it to capture objects within different ranges.
[0140] The mobile phone can capture raw images (also known as RAW images or digital negatives) through camera 293. For example, camera 293 includes at least a lens and a sensor. When taking a photo or recording video, the shutter is opened, and light is transmitted through the lens of camera 293 to the sensor. The sensor converts the light signal passing through the lens into an electrical signal, and then performs analog-to-digital (A / D) conversion on the electrical signal to output the corresponding digital signal. This digital signal is the RAW image. Afterwards, the mobile phone can use a processor (such as an ISP or DSP) to perform subsequent ISP processing and YUV domain processing on the RAW image, converting it into an image that can be displayed, such as a JPEG image or a high-efficiency image file format (HEIF) image. The JPEG image or HEIF image can be transmitted to the mobile phone's display screen for display, and / or transmitted to the mobile phone's memory for storage. Thus, the mobile phone can perform the shooting function.
[0141] In one possible design, the sensor's photosensitive element can be a charge-coupled device (CCD), and the sensor also includes an A / D converter. In another possible design, the sensor's photosensitive element can be a complementary metal-oxide-semiconductor (CMOS).
[0142] For example, ISP processing may include: bad pixel correction (DPC), RAW domain noise reduction, black level correction (BLC), lens shading correction (LSC), auto white balance (AWB), demosica color interpolation, color correction matrix (CCM), dynamic range compression (DRC), gamma, 3D lookup table (LUT), YUV domain noise reduction, sharpening, and detail enhancement. YUV domain processing may include: multi-frame registration, fusion, and noise reduction of high-dynamic range (HDR) images, as well as super-resolution (SR) algorithms, skin smoothing algorithms, distortion correction algorithms, and bokeh algorithms to improve sharpness.
[0143] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone may include one or N displays 294, where N is a positive integer greater than 1. For example, display screen 294 can be used to display an application interface.
[0144] The mobile phone implements its display function through a GPU, a display screen 294, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0145] Understandable Figure 2The structure shown does not constitute a specific limitation on the mobile phone. In some embodiments, the mobile phone may also include... Figure 2 This could mean having more or fewer components, combining some components, separating some components, or having different component arrangements. Or, Figure 2 Some of the components shown can be implemented in hardware, software, or a combination of both.
[0146] Additionally, when the terminal device 200 is an interactive whiteboard, wearable device, tablet computer, laptop computer, desktop computer, portable electronic device, UMPC, netbook, PDA, smart TV, in-vehicle computer, smart speaker, AR device, VR device, or other smart devices with a display screen, or other forms of terminal devices such as digital camera, SLR camera / mirrorless camera, action camera, gimbal camera, or drone, the specific structure of these other forms of terminal devices can also refer to... Figure 2 As shown. Exemplarily, other forms of terminal devices may be... Figure 2 The components were added or removed based on the given structure, which will not be elaborated here.
[0147] It should also be understood that the terminal device 200 (such as a mobile phone) may run one or more applications capable of acquiring data required for 3D modeling, preprocessing the data, and previewing the 3D model. These applications may be called 3D modeling applications or 3D reconstruction applications. When the terminal device 200 runs the aforementioned 3D modeling or 3D reconstruction application, the application can, according to user input, call the terminal device 200's camera to capture data required for 3D modeling and preprocess that data. Additionally, the application can display a preview interface of the 3D model on the terminal device 200's screen, allowing the user to view and preview the 3D model.
[0148] The following is based on the above. Figure 1 Taking a mobile phone as an example of the terminal device 200 in the cloud-edge collaborative system shown, and combining the scenario of a user using a mobile phone to perform 3D modeling, the 3D modeling method provided in this application embodiment will be illustrated by way of example.
[0149] It should be noted that although this application embodiment uses a mobile phone as an example of terminal device 200 for illustration, it should be understood that the 3D modeling method provided in this application embodiment is also applicable to other terminal devices with shooting functions, and this application does not limit the specific type of terminal device.
[0150] Taking the 3D modeling of a target object as an example, the 3D modeling method provided in this application embodiment may include the following three parts:
[0151] The first part involves users using their mobile phones to collect the data needed to create a 3D model of the target object.
[0152] In the second part, the mobile phone preprocesses the data collected for 3D modeling of the target object and uploads the preprocessed data to the cloud.
[0153] The third part involves cloud-based 3D modeling based on data uploaded from the mobile phone, resulting in a 3D model of the target object.
[0154] By following the steps one through three described above, a 3D model of the target object can be created. After obtaining the 3D model of the target object, the mobile phone can download the 3D model from the cloud for the user to preview.
[0155] The following sections will provide specific explanations of Parts 1 through 3.
[0156] For Part One:
[0157] In this embodiment, a first application may be installed on the mobile phone, which is the 3D modeling application or 3D reconstruction application described in the foregoing embodiments. For example, the name of the first application may be "3D Magic Cube," and there is no limitation on the name of the first application. When a user wants to perform 3D modeling of a target object, they can launch and run the first application on the mobile phone. After the mobile phone launches and runs the first application, the main interface of the first application may include a function control for launching the 3D modeling function. The user can click or touch the function control on the main interface of the first application. The mobile phone can launch the 3D modeling function of the first application in response to the user's click or touch operation of the function control. After the 3D modeling function of the first application is launched, the mobile phone can switch the display interface from the main interface of the first application to the 3D modeling data acquisition interface, and launch the camera's shooting function to display the image captured by the camera on the 3D modeling data acquisition interface. During the process of the mobile phone displaying the 3D modeling data acquisition interface, the user can hold the mobile phone to collect the data required for 3D modeling of the target object.
[0158] In some embodiments, the mobile phone can display functional controls for launching the first application on the main interface (or desktop), such as the application icon (or button) of the first application. When a user wants to use the first application to perform 3D modeling of a target object, they can click or touch the application icon of the first application. After receiving the user's click or touch operation on the application icon, the mobile phone can respond to the user's click or touch operation by launching and running the first application and displaying the main interface of the first application.
[0159] For example, Figure 3 This is a schematic diagram of the main interface of a mobile phone provided in an embodiment of this application. Figure 3As shown, the main interface 301 of the mobile phone may include an application icon 302 for the first application. The main interface 301 of the mobile phone may also include application icons for other applications such as application A, application B, and application C. The user can click or touch the application icon 302 on the main interface 301 of the mobile phone to trigger the mobile phone to launch and run the first application and display the main interface of the first application.
[0160] In other embodiments, the mobile phone may also display the function controls for launching the first application on other display interfaces such as the pull-down interface or the negative one screen. In the pull-down interface or the negative one screen, the function controls for the first application may be presented in the form of an application icon or other function buttons, without limitation.
[0161] The pull-down screen is the interface that appears after swiping down from the top of the phone's main screen. This screen displays frequently used function buttons, such as WLAN and Bluetooth, allowing users to quickly access these functions. For example, if the phone is currently displaying the home screen, a user can swipe down from the top of the screen to switch the display to the pull-down screen (or have the pull-down screen overlaid on the home screen). The negative one screen is the interface that appears after swiping right from the phone's main screen (or home screen). This screen displays frequently used applications, functions, subscribed services, and information for quick browsing and use. For example, if the phone is currently displaying the home screen, a user can swipe right from the screen to switch the display to the negative one screen.
[0162] It is understood that "negative one screen" is just a term used in the embodiments of this application, and its meaning has been recorded in the embodiments of this application, but its name does not constitute any limitation on the embodiments of this application; in addition, in some other embodiments, "negative one screen" may also be called other names such as "desktop assistant", "shortcut menu", "widget collection interface" etc., which are not limited here.
[0163] In some embodiments, when a user wants to use the first application to perform 3D modeling of a target object, they can also control the phone to launch and run the first application through a voice assistant. This application does not limit the method of launching the first application.
[0164] For example, Figure 4 This is a schematic diagram of the main interface of the first application provided in an embodiment of this application. Figure 4As shown, the main interface 401 of the first application may include a function control: "Start Modeling" 402. "Start Modeling" 402 is the function control used to initiate the 3D modeling function. Users can click or touch "Start Modeling" 402 on the main interface 401 of the first application. In response to the user's click or touch of "Start Modeling" 402, the mobile phone will initiate the 3D modeling function of the first application, switch the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface, and activate the camera's shooting function to display the image captured by the camera on the 3D modeling data acquisition interface.
[0165] For example, Figure 5 This is a schematic diagram of a 3D modeling data acquisition interface provided in an embodiment of this application. Figure 5 As shown, the 3D modeling data acquisition interface 501 displayed on the mobile phone may include functional controls: a scan button 502, and the image captured by the mobile phone's camera. For example, please refer to... Figure 5 As shown, assuming a user wants to create a 3D model of a toy car placed on a table, the user can point their phone's camera at the toy car. The image captured by the camera on the 3D modeling data acquisition interface 501 will include the toy car 503 and the table 504. The user can move the phone's camera angle to center the toy car 503 on the phone screen (i.e., the 3D modeling data acquisition interface 501) and tap or click the scan button 502 on the 3D modeling data acquisition interface 501. The phone will respond to the user's tap or click of the scan button 502 and begin collecting the data needed to create a 3D model of the target object (i.e., the toy car 503) located in the center of the phone screen.
[0166] That is, in this embodiment of the application, when the mobile phone responds to the user's click or touch scan button 502 and starts to collect the data required for 3D modeling of the target object, the object located in the center of the mobile phone screen can be used as the target object.
[0167] In this embodiment of the application, the 3D modeling data acquisition interface can be referred to as the first interface, and the main interface of the first application can be referred to as the second interface.
[0168] Optionally, Figure 6 Another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 6 As shown, after the mobile phone receives the user's click or touch scan button 502, it can also display a prompt message on the modeling data acquisition interface 501: "Please place the target object in the center of the screen" 505. Here, the target object is the target object, and this prompt message can be used to remind the user to place the target object in the center of the mobile phone screen.
[0169] For example, the display position of "Please place the target object in the center of the screen" 505 in the modeling data acquisition interface 501 can be above the scan button 502, or in a position slightly below the center of the screen, etc. There is no limitation on the display position of "Please place the target object in the center of the screen" 505 in the modeling data acquisition interface 501.
[0170] Additionally, it should be noted that the prompt message "Please place the target object in the center of the screen" 505 is merely an illustrative example. In some other embodiments, other textual identifiers may be used to remind the user to place the target object in the center of the phone screen. This application does not limit the content of this prompt message. In this embodiment, the prompt message used to remind the user to place the target object in the center of the phone screen can be referred to as the first prompt message.
[0171] In this embodiment, when the mobile phone responds to the user's click or touch of the scan button 502 and begins collecting data required for 3D modeling of the target object, the user can hold the phone and circle around the target object to take pictures. During this circling process, the phone can collect 360-degree panoramic data of the target object. The data collected by the phone for 3D modeling of the target object may include: images of the target object captured by the phone during the circling process. These images can be in JPG / JPEG format. For example, the phone can capture a RAW image of the target object through its camera, and then the phone's processor can perform ISP processing and JPEG encoding on the RAW image to obtain a corresponding JPG / JPEG image of the target object.
[0172] Figure 7A This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 7A As shown, in one possible design, when the mobile phone responds to the user's click or touch of the scan button 502 and begins collecting data required for 3D modeling of a target object (taking a toy car as an example), a surface model 701 (or bounding volume or virtual bounding volume) can be displayed around the target object on the 3D modeling data acquisition interface 501. The surface model 701 can be positioned around the center of the target object. The surface model 701 can include two layers, each containing multiple surfaces. The upper layer can be called the first layer, and the lower layer can be called the second layer. Each surface in each layer can correspond to an angle range within a 360-degree radius around the target object. For example, assuming there are 20 surfaces in each layer, each surface in each layer corresponds to an angle range of 18 degrees.
[0173] When a user holds their phone to take a picture of a target object, they need to circle the object twice. The first circle involves shooting around the object with the phone's camera looking down (e.g., at a 30-degree angle, no restriction). The second circle involves shooting around the object with the phone's camera facing directly at it. When the user shoots around the object with the camera looking down during the first circle, the phone will illuminate the first layer of panels sequentially as it moves around the object. When the user shoots around the object with the camera facing directly at it during the second circle, the phone will illuminate the second layer of panels sequentially as it moves around the object.
[0174] For example, suppose a user looks down at a target object with their phone's camera and takes a first round of shots around the object in a clockwise or counter-clockwise direction. Assuming the initial shooting angle is 0 degrees and each layer has 20 facets, when the phone captures an image of the target object within an angle range of 0 to 18 degrees, the phone can light up the first facet in the first layer. When the phone captures an image of the target object within an angle range of 18 to 36 degrees, the phone can light up the second facet. Similarly, when the phone captures an image of the target object within an angle range of 342 to 360 degrees, the phone can light up the 20th facet. In other words, after the user takes a first round of shots around the target object looking down at it, all 20 facets in the first layer will be lit. Likewise, after the user takes a second round of shots around the target object looking directly at it, all 20 facets in the second layer will be lit.
[0175] In this embodiment, the image corresponding to the first patch can be referred to as the first image, and the image corresponding to the second patch can be referred to as the second image. The posture of the mobile phone when capturing the first image can be referred to as the first pose, and the posture of the mobile phone when capturing the second image can be referred to as the second pose.
[0176] For example, please continue to refer to Figure 7A As shown, when a user points their phone camera downwards at the target object and takes the first shot around the object counter-clockwise, the phone can illuminate the first panel in the first layer, as shown in the image. Figure 7A As shown in 702, the illuminated surface can display a different pattern or color (i.e., change the display effect) compared to other unilluminated surfaces.
[0177] Figure 7B This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 7B As shown above, in the above Figure 7ABased on the example shown, when the user looks down at the target object with their phone camera and takes the first round of shots around the target object in a counter-clockwise direction, the phone can continue to light up the second and third panels in the first layer as the user holds the phone and rotates it counter-clockwise around the target object by a certain angle (moves a certain distance).
[0178] Figure 7C This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 7C As shown above, in the above Figure 7B Based on the above, when the user looks down at the target object with the phone camera and takes the first round of shots around the target object in a counterclockwise direction, when the user moves the phone counterclockwise around the target object to the other side of the target object, the number of the facets in the first layer that the phone can light up can reach half or more.
[0179] Figure 7D This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 7D As shown above, in the above Figure 7C Based on the example shown, the user positions their phone camera downwards at the target object and takes the first round of shots around the object counter-clockwise. As the user moves the phone counter-clockwise around the target object... Figure 7A When the phone is in the initial shooting position described in the description (or when the user holds the phone and rotates it counterclockwise around the target object), the phone can light up all the faces in the first layer.
[0180] After the phone illuminates all the faces in the first layer, the user can adjust the phone's shooting position relative to the target object, lower the phone to a certain distance so that the phone camera is facing the target object, and take a second round of shots around the target object in a counter-clockwise direction.
[0181] For example, Figure 7E This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 7E As shown, when the user looks directly at the target object with the phone's camera and takes a second shot around the target object in a counter-clockwise direction, the phone can light up the first panel in the second layer.
[0182] Similarly, as the user holds the phone's camera directly in front of the target object and takes a second round of shots around the object counter-clockwise, the phone gradually illuminates the second layer of pixels as the user moves the phone counter-clockwise around the object. For example, Figure 7F This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 7FAs shown, when the user points their phone camera directly at the target object and takes a second shot around the object counter-clockwise, the phone moves counter-clockwise around the target object until... Figure 7E At the initial shooting position described above (or when the user holds the phone and rotates it counterclockwise around the target object), the phone can illuminate all the facets in the second layer. The process from illuminating the first facet of the second layer to illuminating all the facets is similar to the illumination process of the first layer, as described above. Figures 7A to 7D The process shown is not described in detail here.
[0183] Optionally, in this embodiment of the application, the rule for the mobile phone to light up each surface can be as follows:
[0184] 1) When a mobile phone takes a picture of a target object at a certain location, it performs blur detection and keyframe selection on each frame of the captured image (which can be called the input frame) to obtain an image with the required clarity and image features.
[0185] For example, in this embodiment, when the mobile phone takes a picture of a target object, it can acquire a preview stream corresponding to the target object, which includes multiple frames of images. For instance, the mobile phone can take pictures of the target object at frame rates such as 24 frames per second or 30 frames per second, without limitation. The mobile phone can perform blur detection on each captured frame to obtain images with a sharpness greater than a first threshold. The first threshold can be determined according to requirements and the blur detection algorithm, and its size is not limited. If the sharpness of the current frame does not meet the requirements (e.g., less than or equal to the first threshold), the next frame is acquired. Then, the mobile phone can perform keyframe selection (or keyframe filtering) on the images with a sharpness greater than the first threshold to obtain images whose image features meet the requirements. The requirements for image features may include: the image contains relatively clear and rich features; the features contained in the image are easy to extract; and the image contains less redundant information. The algorithm and specific requirements for keyframe selection are not limited here.
[0186] For each shooting location (one shooting location can correspond to one patch), the mobile phone can perform blur detection and keyframe selection on the image captured at that shooting location to obtain some high-quality keyframe images. The number of keyframe images corresponding to each patch can be one or more.
[0187] 2) The camera pose information (i.e., the pose information of the mobile phone camera) corresponding to the image obtained in 1) is calculated by the mobile phone.
[0188] For example, when a mobile phone supports AR engine capabilities, AR core capabilities, or AR KIT capabilities, the mobile phone can call the aforementioned capabilities to directly obtain the camera pose information corresponding to the image.
[0189] For example, camera pose information may include qw, qx, qy, qz, tx, ty, and tz. Here, qw, qx, qy, and qz represent rotation matrices composed of unit quaternions, and tx, ty, and tz can form translation matrices. These rotation and translation matrices represent the relative position and angle between the camera (mobile phone camera) and the target object. The mobile phone can use the aforementioned rotation and translation matrices to transform the coordinates of the target object from the world coordinate system to the camera coordinate system, obtaining the target object's coordinates in the camera coordinate system. The world coordinate system can refer to a coordinate system with the center of the target object as its origin, and the camera coordinate system can refer to a coordinate system with the center of the camera as its origin.
[0190] 3) Based on the camera pose information corresponding to the image obtained in 2), the mobile phone determines the relationship between the image obtained in 1) and each facet in the facet model, and obtains the facet corresponding to the image.
[0191] For example, as shown in 2) above, the mobile phone can transform the coordinates of the target object from the world coordinate system to the camera coordinate system based on the camera pose information (rotation matrix and translation matrix) corresponding to the image, thus obtaining the target object's coordinates in the camera coordinate system. Then, the mobile phone can determine the line connecting the camera coordinates and the target object's coordinates based on the target object's coordinates and the camera coordinates. The intersection of this line with the surface model is the surface corresponding to that frame of the image. Here, the camera coordinates are known parameters for the mobile phone.
[0192] 4) Store the image in a frame sequence file and illuminate the corresponding facets. The frame sequence file includes images corresponding to each illuminated facet, which can be used as data for 3D modeling of the target object.
[0193] The images included in the frame sequence file can be in JPG format. For example, the images saved in the frame sequence file can be numbered sequentially as 001.jpg, 002.jpg, 003.jpg, and so on.
[0194] Understandably, once a certain area is lit up, the user can move the phone to take a picture from the next angle, and the phone can continue to light up the next area according to the above rules.
[0195] Each frame in the aforementioned frame sequence file can be called a keyframe. These keyframes serve as the data collected by the mobile phone in the first part for 3D modeling of the target object. The frame sequence file can also be called a keyframe sequence file.
[0196] Optionally, in this embodiment of the application, the user can hold the mobile phone and take pictures of the target object within a range of 1.5 meters. When the shooting distance is too close (such as when the 3D modeling data acquisition interface cannot display the full view of the target object), the mobile phone can turn on the wide-angle camera to take pictures.
[0197] In the first part above, as the user holds the mobile phone and circles around the target object to take pictures, the mobile phone lights up the facets in the facet model in sequence as it moves around the target object. This can guide the user to collect the data required for 3D modeling of the target object. The dynamic UI guidance through the 3D guidance interface (i.e. the 3D modeling data collection interface that displays the facet model above) enhances the user interactivity and allows the user to intuitively perceive the data collection process.
[0198] It should be noted that the description of the patch model (or bounding body) in the first part above is merely illustrative. In other embodiments, the number of layers in the patch model may include more or fewer layers, and the number of patches in each layer may be greater than 20 or less than 20. This application does not limit the number of layers in the patch model or the number of patches in each layer.
[0199] For example, Figure 8 This is a schematic diagram of the structure of the patch model provided in an embodiment of this application. Please refer to... Figure 8 As shown, in one implementation, the structure of the patch model can be as follows: Figure 8 As shown in (a) above, it includes two layers, each of which may include multiple facets (i.e., the structure described in the preceding embodiments). In another implementation, the structure of the facet model can be as follows: Figure 8 As shown in (b) above, it includes three layers: top, middle, and bottom, and each layer can include multiple facets. In another implementation, the structure of the facet model can be as follows: Figure 8 As shown in (c), this is a single-layer structure composed of multiple facets. In another implementation, the facet model structure can also be as follows: Figure 8 As shown in (d), it includes two layers, and each layer can include multiple facets.
[0200] Figure 8 The structures of the patch models shown are illustrative. This application does not limit the structure of the patch models, nor the tilt angle (tilt angle relative to the central axis) of each layer in the patch models.
[0201] It should be understood that the patch model described in the embodiments of this application is a virtual model. The patch model can be pre-installed in the mobile phone, such as by configuring it in the file directory of the first application in the form of a configuration file.
[0202] In some embodiments, the mobile phone may have multiple pre-installed patch models. When the mobile phone collects data required for 3D modeling of a target object, it can recommend a target patch model that matches the shape of the target object, or select a target patch model from multiple patch models based on the user's selection operation, and use the target patch model to implement the guidance function described in the foregoing embodiments. The target patch model can also be referred to as the first virtual bounding volume.
[0203] Optionally, Figure 9 This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 9 As shown, after the phone switches the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface 501, when a target object is detected in the image, a prompt message can be displayed on the modeling data acquisition interface 501: "Target object detected, click the button to start scanning" 506. The button is also the scan button 502, and this prompt message reminds the user to click the scan button 502 to start the phone collecting the data required for 3D modeling of the target object.
[0204] Optionally, in this embodiment of the application, the mobile phone can respond to the user's operation of clicking or touching "Start Modeling" 402, start the 3D modeling function of the first application, and switch the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface. It can also first display relevant prompt information on the 3D modeling data acquisition interface to remind the user to adjust the shooting environment of the target object, the shooting method of the target object, and the screen ratio of the object.
[0205] For example, Figure 10 This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 10 As shown, the mobile phone can respond to the user's click or touch of "Start Modeling" 402, launching the 3D modeling function of the first application. After switching the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface, it can also display a prompt message 1001 on the 3D modeling data acquisition interface. The content of prompt message 1001 can be "Place the object statically on a solid-color plane, with soft lighting, and take photos around the object, ensuring the object has a large and complete screen ratio." This serves to remind the user to adjust the shooting environment of the target object to be statically placed on a solid-color plane with soft lighting; to take photos of the target object by circling the object; and to ensure the object has a large and complete screen ratio. (Continue to refer to...) Figure 10As shown, the mobile phone can also first display function control 1002 on the 3D modeling data acquisition interface, such as the function space for "OK" or "Confirm". After the user clicks on function control 1002, the mobile phone no longer displays the prompt message 1001 and function control 1002, and presents the interface as described above. Figure 5 The 3D modeling data acquisition interface shown.
[0206] In this embodiment, after the user adjusts the shooting environment and screen ratio of the target object according to the prompts in prompt 1001, the subsequent data acquisition process can be faster and the quality of the acquired data can be better. Prompt 1001 can also be referred to as the second prompt.
[0207] In some embodiments, after the mobile phone switches the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface, it can also display the prompt message 1001 on the 3D modeling data acquisition interface for a preset duration. When the preset duration is reached, the mobile phone can automatically stop displaying the prompt message 1001 and present the above-described... Figure 5 The 3D modeling data acquisition interface is shown. The preset duration can be 20 seconds, 30 seconds, etc., and is not limited here.
[0208] In some embodiments, after the mobile phone has collected the data required for 3D modeling of the target object (each frame image included in the frame sequence file) in the manner described in the first part above, the second part can be executed automatically.
[0209] In other embodiments, after the mobile phone collects the data required for 3D modeling of the target object (each frame image included in the frame sequence file) in the manner described in Part 1 above, it can display a function control for uploading to the cloud for 3D modeling on the 3D modeling data acquisition interface. The user can click on this function control for uploading to the cloud for 3D modeling. The mobile phone can respond to the user's click on the function control for uploading to the cloud for 3D modeling and execute Part 2.
[0210] For example, Figure 11 This is another schematic diagram of the 3D modeling data acquisition interface provided in an embodiment of this application. For example... Figure 11As shown, after the mobile phone collects the data required for 3D modeling of the target object (each frame image included in the frame sequence file) in the manner described in Part 1 above, it can display the function control "Upload to Cloud Modeling" 1101 on the 3D modeling data acquisition interface. "Upload to Cloud Modeling" 1101 is the function control used to upload the data to the cloud for 3D modeling. The user can click "Upload to Cloud Modeling" 1101, and the mobile phone can respond to the user's click on "Upload to Cloud Modeling" 1101 and execute Part 2. In this embodiment of the application, the user's click on "Upload to Cloud Modeling" 1101 can be referred to as the operation of generating a 3D model.
[0211] Alternatively, please continue to refer to Figure 11 As shown, after the mobile phone collects the data required for 3D modeling of the target object (each frame image included in the frame sequence file) in the manner described in Part 1 above, it can also display a prompt message on the 3D modeling data acquisition interface to inform the user that the mobile phone has collected the data required for 3D modeling of the target object, such as "Scan completed" 1102.
[0212] Alternatively, please refer to the foregoing. Figures 5 to 7E and 9 to Figure 11 As shown, the mobile phone can also display an exit button 1103 in the 3D modeling data acquisition interface (only when...). Figure 11 (As indicated by the winning bid), during the execution of the first part of the process, the user can click the exit button 1103 at any time. The phone can respond to the user's click of the exit button 1103 and exit the execution process of the first part. After exiting the execution process of the first part, the phone can switch the display interface from the 3D modeling data acquisition interface to... Figure 4 The main interface of the first application is shown.
[0213] For Part Two:
[0214] As described in Part One above, the data collected by the mobile phone in Part One for 3D modeling the target object consists of each frame (i.e., keyframe) in the frame sequence file mentioned in Part One. Part Two, the preprocessing of the collected data for 3D modeling the target object by the mobile phone, refers to the preprocessing of the keyframes included in the frame sequence file collected in Part One, as follows:
[0215] 1) The phone calculates the matching information of each keyframe and saves the matching information of each keyframe in a first file. For example, the first file can be a JavaScript object notation (JSON) file.
[0216] For each keyframe, the matching information can include the identification information of other keyframes associated with it. For example, the matching information of a keyframe can include the identification information of the keyframes corresponding to its four cardinal directions (up, down, left, and right) (such as the nearest keyframe). These keyframes are the other keyframes associated with the keyframe. The keyframe identification information can be the image number of the keyframe. The matching information of this keyframe can be used to indicate which images in the frame sequence file are associated with it.
[0217] For example, the matching information for each keyframe is obtained based on the association between each keyframe and the corresponding patches, as well as the association between the multiple patches. That is, for each keyframe, the mobile phone can determine the matching information based on the association between the patch corresponding to that keyframe and other patches. The association between the patch corresponding to that keyframe and other patches can include: in the patch model, which patches correspond to the four cardinal directions (up, down, left, right) of the patch corresponding to that keyframe.
[0218] For example, with the above Figure 7A Taking the patch model shown as an example, assuming that the patch corresponding to a certain keyframe is patch 1 in the first layer, the association between patch 1 and other patches can include: the patch below patch 1 is patch 21, the patch to the left of patch 1 is patch 20, and the patch to the right of patch 1 is patch 2. Based on the aforementioned association between patch 1 and other patches, the mobile phone can determine that other keyframes associated with this keyframe include the keyframe corresponding to patch 21, the keyframe corresponding to patch 20, and the keyframe corresponding to patch 2. Therefore, the mobile phone can obtain the matching information of this keyframe, including the identification information of the keyframe corresponding to patch 21, the identification information of the keyframe corresponding to patch 20, and the identification information of the keyframe corresponding to patch 2.
[0219] Understandably, for mobile phones, the relationships between different patches in the patch model are known quantities.
[0220] Optionally, the first file may also include: camera intrinsics, gravity information, image name, image index, camera pose information, timestamp, etc., corresponding to each keyframe.
[0221] For example, the first file includes three parts: "intrinsics", "keyframes", and "matching_list". The "intrinsics" part contains camera intrinsic parameters; the "keyframes" part contains information such as gravity direction, image name, image number, camera pose information, and timestamp for each keyframe; and the "matching_list" part contains matching information for each keyframe.
[0222] For example, the contents of the first file can be as follows:
[0223]
[0224]
[0225] In the first file provided in the above example, cx, cy, fx, fy, height, k1, k2, k3, p1, p2, and width are all camera intrinsic parameters. Specifically, cx and cy represent the offset of the optical axis from the center of the projection plane coordinate system; fx and fy represent the focal lengths in the x and y directions, respectively, when the camera is shooting; k1, k2, and k3 represent radial distortion coefficients; p1 and p2 represent tangential distortion coefficients; and height and width represent the camera resolution when shooting.
[0226] x, y, and z represent the direction of gravity. This direction of gravity information can be obtained by the phone using its built-in gyroscope and can represent the offset angle when the phone takes a picture.
[0227] 18.jpg represents the image name, and 18 is the image index (18.jpg is used as an example here). That is, the above example contains the camera intrinsics, gravity information, image name, image index, camera pose information, timestamp, and matching information corresponding to 18.jpg.
[0228] qw, qx, qy, qz, tx, ty, and tz represent camera pose information. qw, qx, qy, and qz represent rotation matrices composed of unit quaternions, while tx, ty, and tz can form translation matrices. These rotation and translation matrices represent the relative position and angle between the camera (mobile phone camera) and the target object. The mobile phone can use these rotation and translation matrices to transform the target object's coordinates from the world coordinate system to the camera coordinate system, obtaining the target object's coordinates in the camera coordinate system. The world coordinate system can refer to a coordinate system with the center of the target object as its origin, and the camera coordinate system can refer to a coordinate system with the center of the camera as its origin.
[0229] The timestamp represents the time when the camera captured the keyframe.
[0230] `src_id` represents the image number of each keyframe. For example, in the first file content given in the example above, image number 18, and the "matching_list" part contains the matching information for keyframe with image number 18. `tgt_id` represents the image numbers of other keyframes associated with keyframe with image number 18 (i.e., the identification information of other keyframes associated with keyframe with image number 18). For example, in the first file content given in the example above, the image numbers of other keyframes associated with keyframe with image number 18 include: 26, 45, 59, 78, 89, 100, 449, etc. That is, other keyframes associated with keyframe with image number 18 include: keyframe with image number 26, keyframe with image number 45, keyframe with image number 59, keyframe with image number 78, keyframe with image number 89, keyframe with image number 100, keyframe with image number 449, etc.
[0231] It should be noted that the above example only shows a portion of the content of the first file using the keyframe with image number 18 as an example, and is not intended to limit the content of the first file.
[0232] 2) The mobile phone packages the first file with all the keyframes in the frame sequence file (i.e., all the frame images included in the frame sequence file).
[0233] The result obtained by packaging all keyframes in the first file and the frame sequence file (such as a packaged file or data packet) is the preprocessed data obtained in the second part by preprocessing the data collected by the mobile phone in the first part for 3D modeling of the target object.
[0234] That is, in this embodiment of the application, the data after the mobile phone preprocesses the data required for 3D modeling of the target object collected in the first part may include: each key frame image saved to the frame sequence file during the process of the mobile phone taking pictures around the target object, and a first file including the matching information of each key frame.
[0235] After receiving the preprocessed data, the mobile phone can send (i.e. upload) the preprocessed data to the cloud. The cloud can then execute the third part, performing 3D modeling based on the data uploaded by the mobile phone to obtain a 3D model of the target object.
[0236] For Part Three:
[0237] The process of cloud-based 3D modeling based on data uploaded from a mobile phone can be as follows:
[0238] 1) The cloud decompresses the data packets (including frame sequence files and the first file) received from the mobile phone and extracts the frame sequence files and the aforementioned first file.
[0239] 2) The cloud performs 3D modeling processing based on the keyframe images included in the frame sequence file and the first file mentioned above to obtain the 3D model of the target object.
[0240] For example, the steps of performing 3D modeling processing in the cloud based on the keyframe images included in the frame sequence file and the aforementioned first file may include at least: key target extraction, feature detection and matching, global optimization and fusion, sparse point cloud computing, dense point cloud computing, surface reconstruction, and texture generation.
[0241] Key target extraction refers to the process of separating the target object of interest from the background in a keyframe image, identifying and interpreting meaningful object entities from the image, and extracting different image features.
[0242] Feature detection and matching refers to: detecting unique pixels in keyframe images as feature points of the keyframe images; describing the salient feature points in different keyframe images; and comparing the similarity between the two descriptions to determine whether the feature points in different keyframe images are the same feature.
[0243] In this embodiment of the application, when the cloud performs feature detection and matching, for each key frame, the cloud can determine other key frames associated with the key frame based on the matching information of the key frame included in the first file (i.e., the identification information of other key frames associated with the key frame), and perform feature detection and matching on the key frame and other key frames associated with the key frame.
[0244] For example, taking the content of the first file exemplified by the keyframe with image number 18 as an example, the cloud can determine other keyframes associated with keyframe 18 based on the first file, including: keyframes with image number 26, 45, 59, 78, 89, 100, and 449. Then, the cloud can perform feature detection and matching between keyframe 18 and keyframes with image numbers 26, 45, 59, 78, 89, 100, and 449, without needing to perform feature detection and matching between keyframe 18 and all other keyframes in the frame sequence file.
[0245] As can be seen in this embodiment, for each keyframe, the cloud can combine the matching information of that keyframe included in the first file and perform feature detection and matching on that keyframe and other keyframes associated with it. It is not necessary to perform feature detection and matching on that keyframe and all other keyframes in the frame sequence file. This approach effectively reduces the computational load on the cloud and improves the efficiency of 3D modeling.
[0246] Global optimization and fusion refers to using global optimization and fusion algorithms to optimize and fuse the matching results of feature detection and matching. The results of global optimization and fusion can be used to generate basic 3D models.
[0247] Sparse point cloud computing and dense point cloud computing refer to generating 3D point cloud data corresponding to a target object based on the results of global optimization and fusion. Compared to images, point clouds have an irreplaceable advantage—depth. 3D point cloud data directly provides data in 3D space, while images require perspective geometry to infer 3D data.
[0248] Surface reconstruction refers to the process of accurately restoring the three-dimensional surface shape of an object using three-dimensional point cloud data, thereby obtaining a basic 3D model of the target object.
[0249] Texture generation refers to generating a texture (also called texture mapping) on the surface of a target object based on keyframe images or their features. After obtaining the texture of the target object's surface, the texture is mapped onto the surface of the target object's basic 3D model in a specific way, which can more realistically reproduce the target object's surface and make the target object look more lifelike.
[0250] In this embodiment, the cloud can also quickly and accurately determine the mapping relationship between the texture and the surface of the basic 3D model of the target object based on the matching information of each key frame included in the first file, which can further improve modeling efficiency and effect.
[0251] For example, after determining the mapping relationship between the texture of the first keyframe and the surface of the basic 3D model of the target object, the cloud can combine the matching information of the first keyframe to quickly and accurately determine the mapping relationship between the texture of other keyframes associated with the first keyframe and the surface of the basic 3D model of the target object. Similarly, for each subsequent keyframe, the cloud can combine the matching information of that keyframe to quickly and accurately determine the mapping relationship between the texture of other keyframes associated with that keyframe and the surface of the basic 3D model of the target object.
[0252] After obtaining the basic 3D model of the target object and its surface texture, the cloud can generate a 3D model of the target object based on these elements. The cloud can then save the basic 3D model and surface texture of the target object for download to a mobile phone.
[0253] As can be seen, in the process of cloud-based 3D modeling based on data uploaded from mobile phones, the matching information of each key frame included in the first file can effectively improve the processing speed of 3D modeling, reduce the computational load on the cloud, and improve the efficiency of cloud-based 3D modeling.
[0254] For example, the basic 3D model of the target object can be stored in OBJ format, and the texture of the target object's surface can be stored in JPG format (such as a texture map). For instance, the basic 3D model of the target object can be an OBJ file, and the texture of the target object's surface can be a JPG file.
[0255] Optionally, the cloud can save the basic 3D model of the target object and the texture of its surface for a certain period of time (e.g., 7 days). After this period, the cloud can automatically delete the basic 3D model of the target object and the texture of its surface. Alternatively, the cloud can permanently retain the basic 3D model of the target object and the texture of its surface; there is no limitation on this.
[0256] Through the first to third parts described above, 3D modeling of the target object can be achieved, resulting in a 3D model of the target object. As can be seen from the first to third parts, in the 3D modeling method provided in this application embodiment, the mobile phone only needs a regular RGB camera (camera) to collect the data required for 3D modeling to achieve 3D modeling. The process of collecting the data required for 3D modeling does not require the mobile phone to have special hardware such as a LiDAR sensor or an RGB-D camera. The 3D modeling process is completed in the cloud and does not require the mobile phone to be equipped with a high-performance dedicated graphics card. This 3D modeling method can significantly lower the threshold for 3D modeling and has higher universality for terminal devices. Moreover, during the process of users collecting the data required for 3D modeling using their mobile phones, the mobile phone can enhance user interactivity through dynamic UI guidance, allowing users to intuitively perceive the data collection process.
[0257] Furthermore, in this 3D modeling method, when the mobile phone takes a picture of the target object at a certain location, it performs blur detection on each frame of the captured image to obtain images with sufficient clarity. This allows for the selection of keyframes, resulting in keyframes that are beneficial for modeling. The mobile phone extracts the matching information of each keyframe and sends a first file containing the matching information of each keyframe, along with a frame sequence file composed of the keyframes, to the cloud for cloud-based modeling (without needing to send all captured images). This significantly reduces the complexity of 3D modeling on the cloud side, minimizes the consumption of cloud hardware resources during the 3D modeling process, effectively reduces the computational load on cloud-based modeling, and improves the speed and quality of 3D modeling.
[0258] The following is an example illustrating the process of downloading a 3D model from the cloud onto a mobile phone for users to preview.
[0259] For example, after the cloud completes the 3D modeling of the target object and obtains the 3D model of the target object, it can send an instruction message to the mobile phone to indicate that the cloud has completed the 3D modeling.
[0260] In some embodiments, after receiving the aforementioned instruction message from the cloud, the mobile phone can automatically download a 3D model of the target object from the cloud for the user to preview.
[0261] For example, Figure 12 This is a schematic diagram of a 3D model preview interface provided in an embodiment of this application. In the second part described above, after the mobile phone sends the preprocessed data to the cloud, the display interface can be changed from... Figure 11 The 3D modeling data acquisition interface shown has been switched to the following: Figure 12 The 3D model preview interface shown. Figure 12As shown, the mobile phone can display the prompt message "Modeling" 1201 in the 3D model preview interface, which is used to inform the user that 3D modeling of the target object is being performed. "Modeling" 1201 can be referred to as the third prompt message.
[0262] In some embodiments, the mobile phone may display a third prompt message after detecting that the user has clicked on the above-mentioned function control "Upload Cloud Modeling" 1101.
[0263] After completing the 3D modeling of the target object in the cloud, the cloud can send an indication message to the mobile phone to indicate that the 3D modeling is complete. Upon receiving this indication message, the mobile phone can automatically download the 3D model of the target object from the cloud. For example, the mobile phone can send a download request message to the cloud, and the cloud can send the 3D model of the target object (i.e., the basic 3D model of the target object and the texture of the target object's surface) to the mobile phone based on the download request message.
[0264] Figure 13 Another schematic diagram of the 3D model preview interface provided in an embodiment of this application. For example... Figure 13 As shown, after receiving the instruction message, the mobile phone can change the prompt message from "Modeling in progress" 1201 to "Modeling completed" 1301 to inform the user that the 3D model of the target object has been completed. The 3D model preview interface can also include a view button 1302. The user can click the view button 1302, and the mobile phone can respond to this action by displaying the 3D model of the target object downloaded from the cloud in the 3D model preview interface. "Modeling completed" 1301 can be considered the fourth prompt message.
[0265] Taking the toy car described in the foregoing embodiments as an example, Figure 14 This is another schematic diagram of a 3D model preview interface provided in an embodiment of this application. For example... Figure 14 As shown, the mobile phone can respond to the user's click on the view button 1302, displaying the 3D model 1401 of the toy car in the 3D model preview interface. The user can... Figure 14 The 3D model preview interface shown displays the 3D model 1401 of the toy car. Clicking the view button 1302 allows the user to preview the 3D model corresponding to the target object.
[0266] Optionally, when a user views the 3D model of a toy car in the 3D model preview interface, the user can rotate the 3D model of the toy car counterclockwise or clockwise in any direction (such as horizontal or vertical). The mobile phone can respond to the user's operation and display the presentation effect of the 3D model of the toy car at different angles (360 degrees) in the 3D model preview interface.
[0267] For example, Figure 15 This is a schematic diagram illustrating a user performing a counter-clockwise rotation operation along the horizontal direction on a 3D model of a toy car, as provided in an embodiment of this application. Figure 15 As shown, users can use their fingers to drag the 3D model of the toy car in the 3D model preview interface and rotate it counterclockwise along the horizontal direction.
[0268] Figure 16 This is another schematic diagram of a 3D model preview interface provided in an embodiment of this application. For example... Figure 16 As shown, when a user drags the 3D model of a toy car counter-clockwise along the horizontal direction using their finger in the 3D model preview interface, the phone can respond to this operation by displaying the following on the 3D model preview interface: Figure 16 The presentation effect of angles shown in (a) and (b) in the text.
[0269] Understandable. Figure 16 The angles shown in (a) and (b) are for illustrative purposes only. The angles displayed by the 3D model of the toy car are related to the direction, distance, and number of times the user drags the 3D model of the toy car, which will not be shown in detail here.
[0270] Optionally, when a user views the 3D model of a toy car in the 3D model preview interface, they can also zoom in or out on the 3D model of the toy car. The mobile phone can respond to the user's zoom in or out operation on the 3D model of the toy car and display the zoomed-in or zoomed-out effect of the 3D model of the toy car to the user in the 3D model preview interface.
[0271] For example, Figure 17 This is a schematic diagram illustrating a user performing a scaling-down operation on a 3D model of a toy car, as provided in an embodiment of this application. Figure 17 As shown, users can use two fingers to slide inwards (in opposite directions) on the 3D model preview interface; this sliding operation is the zoom-out operation.
[0272] Figure 18 This is another schematic diagram of a 3D model preview interface provided in an embodiment of this application. For example... Figure 18As shown, when a user performs a zoom-out operation on the 3D model of a toy car, the mobile phone can respond to the zoom-out operation by displaying the zoomed-out appearance of the 3D model of the toy car on the 3D model preview interface.
[0273] Similarly, users can zoom in on a 3D model of a toy car by swiping outwards (in opposite directions) with two fingers on the 3D model preview screen. When a user zooms in on the 3D model of the toy car, the phone can respond by displaying the zoomed-in 3D model's appearance on the preview screen; details will not be elaborated further.
[0274] It should be noted that the zoom-in or zoom-out operations performed by the user on the 3D model of the toy car described above are merely illustrative examples. In other implementations, the zoom-in or zoom-out operations performed by the user on the 3D model of the toy car can also be double-click operations, long-press operations, or the 3D model preview interface can also include a functional control for zooming in or out, etc., without limitation.
[0275] In other embodiments, after receiving the aforementioned instruction message from the cloud, the mobile phone may also simply display the message as described above. Figure 13 The 3D model preview interface is shown. When the user clicks the view button 1302, the mobile phone responds to the user's click operation by downloading the 3D model of the target object from the cloud and displaying the 3D model of the target object in the 3D model preview interface for the user to preview. This application does not restrict the triggering conditions for the mobile phone to download the 3D model of the target object from the cloud.
[0276] As described above, in this embodiment, the user only needs to perform data-related operations required for 3D modeling on the terminal device side, and then view or preview the final 3D model on the terminal device. For the user, all operations are completed on the terminal device side, making the operation simpler and providing a better user experience.
[0277] To make the technical solutions provided in the embodiments of this application more concise and clear, the following will be combined with... Figure 19 and Figure 20 The implementation logic of the 3D modeling method provided in the embodiments of this application will be described by way of example.
[0278] For example, Figure 19 This is a flowchart illustrating the 3D modeling method provided in an embodiment of this application. Figure 19 As shown, the 3D modeling method may include S1901-S1913.
[0279] S1901, The mobile phone receives the first operation, which is to launch the first application.
[0280] For example, the first action could be the click or touch mentioned above. Figure 3 The operation of the first application icon 302 on the main interface of the phone shown. Alternatively, the first operation can be clicking or touching the function control to launch the first application on the pull-down interface or other display interfaces such as the negative one screen. Alternatively, the first operation can also be the operation described above, where the phone is controlled to launch and run the first application via a voice assistant.
[0281] S1902, The mobile phone responds to the first operation, starts and runs the first application, and displays the main interface of the first application.
[0282] The main interface of the first application can be referenced as described above. Figure 4 As shown. The main interface of the first application can be called the second interface.
[0283] S1903, The mobile phone receives a second operation, which is to start the 3D modeling function of the first application.
[0284] For example, the second operation could be the aforementioned click or touch. Figure 4 The operation of the "Start Modeling" function control 402 in the main interface 401 of the first application shown.
[0285] S1904. The mobile phone responds to the second operation, displays the 3D modeling data acquisition interface, and starts the camera's shooting function, displaying the image captured by the camera on the 3D modeling data acquisition interface.
[0286] The 3D modeling data acquisition interface displayed when the phone responds to the second operation can be referenced as described above. Figure 5 As shown, the 3D modeling data acquisition interface can be referred to as the first interface.
[0287] S1905, The mobile phone receives a third operation, which is to control the mobile phone to collect the data required for 3D modeling of the target object.
[0288] For example, the third operation may include the user clicking or touching the above. Figure 5 The operation of the scan button 502 in the 3D modeling data acquisition interface shown, as well as the operation of the user holding a mobile phone and taking pictures around the target object. The third operation can also be called the acquisition operation.
[0289] S1906, The mobile phone responds to the third operation by acquiring a frame sequence file composed of key frame images corresponding to the target object.
[0290] S1907. The mobile phone obtains the matching information of each key frame in the frame sequence file to obtain the first file.
[0291] The first file can refer to the description in the foregoing embodiments. For each keyframe, the matching information of that keyframe included in the first file may include: the identification information of the nearest keyframe corresponding to the four directions of the keyframe (up, down, left, right), for example, the identification information may be the number of the aforementioned image.
[0292] S1908, the mobile phone sends a frame sequence file and a first file to the cloud.
[0293] Accordingly, the cloud receives the frame sequence file and the first file.
[0294] S1909: The cloud performs 3D modeling based on the frame sequence file and the first file to obtain the 3D model of the target object.
[0295] The specific process of 3D modeling in the cloud based on the frame sequence file and the first file can be referred to in Part III of the aforementioned embodiments, and will not be repeated here. The 3D model of the target object may include the basic 3D model of the target object and the texture of the target object's surface. Applying the texture (texture map) of the target object's surface onto the basic 3D model of the target object constitutes the 3D model of the target object.
[0296] S1910: The cloud sends an instruction message to the mobile phone to indicate that the cloud has completed 3D modeling.
[0297] Accordingly, the mobile phone receives an instruction message.
[0298] S1911. The mobile phone sends a download request message to the cloud to request the download of the 3D model of the target object.
[0299] Accordingly, the cloud receives the download request message.
[0300] S1912, The cloud sends a 3D model of the target object to the mobile phone.
[0301] Accordingly, the mobile phone receives the 3D model of the target object.
[0302] S1913, The mobile phone displays a 3D model of the target object.
[0303] The phone displays a 3D model of the target object, allowing users to preview the 3D model.
[0304] For example, the effect of displaying a 3D model of a target object on a mobile phone can be referenced above. Figure 12 , Figure 13 , Figure 14 , Figure 16 , Figure 18As shown, when users preview the 3D model of a target object, they can rotate the 3D model of the target object displayed on the mobile phone, zoom in or out, etc.
[0305] The above Figure 19 For the specific implementation of the process shown and its beneficial effects, please refer to the foregoing embodiments, which will not be repeated here.
[0306] For example, Figure 20 This is a logical diagram illustrating the 3D modeling method implemented in the end-to-cloud collaborative system provided in this application embodiment.
[0307] like Figure 20 As shown in the embodiments of this application, the mobile phone may include at least an RGB camera (such as a webcam) and a first application. The first application is the aforementioned 3D modeling application.
[0308] An RGB camera can be used to achieve the same shooting function as a mobile phone, capturing images of the target object to be modeled and obtaining corresponding images. The RGB camera can then transmit the captured images to the primary application.
[0309] The first application may include a data acquisition and dynamic guidance module, a data processing module, a 3D model preview module, and a 3D model export module.
[0310] The data acquisition and dynamic guidance module enables functions such as blur detection, keyframe selection, guidance information calculation, and guidance interface updates. For images captured by an RGB camera, the blur detection function performs blur detection on each captured frame (referred to as the input frame), selecting images with acceptable sharpness as keyframes; if the current frame's sharpness does not meet the requirements, the next frame is acquired. The keyframe selection function determines whether an image is already stored in the frame sequence file; if not, the frame is added. The guidance information calculation function uses the camera pose information corresponding to the image to determine the relationship between the image and each facet in the facet model, obtaining the facet corresponding to the image. This correspondence between images and faces constitutes the guidance information. The guidance interface update function updates (i.e., changes) the display effect of faces in the facet model based on the calculated guidance information, such as illuminating faces.
[0311] The data processing module performs functions such as matching relationship calculation, matching list calculation, and data packaging. The matching relationship calculation function calculates the matching relationships between keyframes in the frame sequence file, such as whether they are adjacent. Specifically, the data processing module can calculate the matching relationships between keyframes in the frame sequence file based on the association relationships between patches in the patch model. The matching list calculation function generates a matching list for each keyframe based on the calculation results of the matching relationship calculation function. Each keyframe's matching list includes the matching information of each keyframe. For example, the matching information of a certain keyframe includes the identification information of other keyframes associated with it. The data packaging function packages the first file containing the matching information of each keyframe and the frame sequence file. After packaging the first file and the frame sequence file, the mobile phone can send the packaged data packet (including the first file and the frame sequence file) to the cloud.
[0312] The cloud-based system can include a data parsing module, a 3D modeling module, and a data storage module. The data parsing module parses received data packets to obtain a frame sequence file and a first file. The 3D modeling module performs 3D modeling based on the frame sequence file and the first file to obtain a 3D model.
[0313] For example, the 3D modeling module can perform functions such as key target extraction, feature detection and matching, global optimization and fusion, sparse point cloud computing, dense point cloud computing, surface reconstruction, and texture generation. The key target extraction function can separate the target object of interest from the background in a keyframe image, identifying and interpreting meaningful object entities from the image to extract different image features. The feature detection and matching function can detect unique pixels in a keyframe image as feature points; it describes significant feature points in different keyframe images and compares the similarity between two descriptions to determine if the feature points in different keyframe images are the same feature. The global optimization and fusion, sparse point cloud computing, and dense point cloud computing functions can generate 3D point cloud data corresponding to the target object based on the results of feature detection and matching. The surface reconstruction function can accurately restore the 3D surface shape of the object using 3D point cloud data, obtaining a basic 3D model of the target object. The texture generation function can generate textures (also called texture maps) on the surface of the target object based on the keyframe image or its features. After obtaining the texture of the target object's surface, the texture is mapped onto the surface of the target object's basic 3D model in a specific way to obtain the target object's 3D model.
[0314] After obtaining the 3D model of the target object, the 3D modeling module can store the 3D model of the target object in the data storage module.
[0315] The mobile phone's primary application can download a 3D model of the target object from the cloud's data storage module. After downloading the 3D model, the primary application can provide the user with a 3D model preview function through the 3D model preview module, or provide the user with a 3D model export function through the 3D model export module. For the specific process of the primary application providing a 3D model preview function through the 3D model preview module, please refer to the foregoing embodiments.
[0316] Optionally, in this embodiment of the application, when the mobile phone collects the data required for 3D modeling of the target object in the first part, the scanning progress can also be displayed on the 3D modeling data acquisition interface. For example, refer to the above. Figures 7A to 7E As shown, the mobile phone can display the scanning progress through the circular black fill effect of the scan button in the 3D modeling data acquisition interface. Understandably, different UI renderings of the scan button can lead to different ways the mobile phone displays the scanning progress in the 3D modeling data acquisition interface; therefore, no restrictions are imposed here.
[0317] In other embodiments, the mobile phone may not need to display the scanning progress; the user can understand the scanning progress by observing the lit-up areas in the area model.
[0318] The above embodiments illustrate the implementation of the 3D modeling method provided in this application within an end-to-cloud collaborative system composed of a terminal device and the cloud. Optionally, in some embodiments, all steps of the 3D modeling method provided in this application can also be implemented on the terminal device side. For example, for some terminal devices with strong processing capabilities and abundant computing resources, the functions implemented on the cloud side as described in the foregoing embodiments can also be implemented entirely on the terminal device. That is, after obtaining the aforementioned frame sequence file and first file, the terminal device can directly generate a 3D model of the target object locally based on the frame sequence file and the first file, and provide functions such as previewing and exporting the 3D model. The specific principle of the terminal device generating a 3D model of the target object locally based on the frame sequence file and the first file is the same as the principle of the cloud generating a 3D model of the target object based on the frame sequence file and the first file as described in the foregoing embodiments, and will not be repeated here.
[0319] It should be understood that the above embodiments are merely illustrative examples of the 3D modeling method provided in this application. In other possible implementations, some execution steps may be omitted or added to the above embodiments, or the order of some steps in the above embodiments may be adjusted, and this application does not impose any limitations on these aspects.
[0320] Corresponding to the 3D modeling method described in the foregoing embodiments, this application provides a modeling apparatus. This apparatus can be applied to a terminal device to implement the steps that the terminal device can perform in the 3D modeling method described in the foregoing embodiments. The function of this apparatus can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions.
[0321] For example, Figure 21 This is a schematic diagram of the modeling apparatus provided in an embodiment of this application. Figure 21 As shown, the device may include a display unit 2101 and a processing unit 2102. The display unit 2101 and the processing unit 2102 can be used to cooperate in implementing the functions of the terminal device in the modeling method described in the foregoing method embodiments.
[0322] For example, display unit 2101 is used to display a first interface, which includes the captured image from the terminal device.
[0323] The processing unit 2102 is configured to, in response to an acquisition operation, acquire multiple frames of images corresponding to the target object to be modeled, and obtain the correlation relationships between the multiple frames of images. Based on the multiple frames of images and the correlation relationships between them, a 3D model corresponding to the target object is obtained.
[0324] The display unit 2101 is also used to display the three-dimensional model corresponding to the target object.
[0325] During the acquisition of multiple frames of images corresponding to the target object, the display unit 2101 is also used to display a first virtual bounding volume; the first virtual bounding volume includes multiple facets. The processing unit 2102 is specifically used to acquire a first image and change the display effect of the facets corresponding to the first image when the terminal device is in a first pose; acquire a second image and change the display effect of the facets corresponding to the second image when the terminal device is in a second pose; and after changing the display effect of the multiple facets of the first virtual bounding volume, obtain the correlation relationship between the multiple frames of images based on the multiple facets.
[0326] Optionally, the display unit 2101 and the processing unit 2102 are also used to implement other display functions and processing functions of the terminal device in the modeling method described in the foregoing method embodiments, which will not be elaborated here.
[0327] Optionally, Figure 22 Another schematic diagram of the modeling apparatus provided in an embodiment of this application. (See attached diagram.) Figure 22As shown, in the modeling method described in the foregoing method embodiments, the terminal device sends the multi-frame images and the association between the multi-frame images to the server, and the server generates a three-dimensional model corresponding to the target object based on the multi-frame images and the association between the multi-frame images. The modeling device may also include a sending unit 2103 and a receiving unit 2104. The sending unit 2103 is used to send the multi-frame images and the association between the multi-frame images to the server, and the receiving unit 2104 is used to receive the three-dimensional model corresponding to the target object sent by the server.
[0328] Optionally, the sending unit 2103 is also used to implement other sending functions that the terminal device can implement in the method described in the foregoing method embodiments, such as sending a download request message. The receiving unit 2104 is also used to implement other receiving functions that the terminal device can implement in the method described in the foregoing method embodiments, such as receiving an indication message. These will not be described in detail here.
[0329] It should be understood that the device may also include other modules or units for implementing the functions of the terminal device described in the foregoing embodiments, which are not shown here one by one.
[0330] Optionally, embodiments of this application also provide a modeling apparatus, which can be applied to a server to implement the server functions in the 3D modeling method described in the foregoing embodiments. The functions of this apparatus can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the aforementioned functions.
[0331] For example, Figure 23 This is yet another schematic diagram of the modeling apparatus provided in an embodiment of this application. (See attached diagram.) Figure 23 As shown, the device may include a receiving unit 2301, a processing unit 2302, and a sending unit 2303. The receiving unit 2301, processing unit 2302, and sending unit 2303 can be used to cooperate in implementing the server function in the modeling method described in the foregoing method embodiments.
[0332] For example, receiving unit 2301 can be used to receive multiple frames of images corresponding to the target object and the correlation relationships between the multiple frames of images sent from the terminal device. Processing unit 2302 can be used to generate a three-dimensional model corresponding to the target object based on the multiple frames of images and the correlation relationships between the multiple frames of images. Sending unit 2303 can be used to send the three-dimensional model corresponding to the target object to the terminal device.
[0333] Optionally, the receiving unit 2301, processing unit 2302, and sending unit 2303 can be used to implement all the functions that the server in the modeling method described in the foregoing method embodiments can achieve, and will not be described in detail here.
[0334] It should be understood that the division of units (or modules) in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, all units in the device can be implemented entirely in software through processing element calls; all units can be implemented entirely in hardware; or some units can be implemented in software through processing element calls, while others can be implemented in hardware.
[0335] For example, each unit can be a separate processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as a program in memory, invoked and executed by a processing element within the device. Furthermore, these units can be integrated in whole or in part, or implemented independently. The processing element described here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In implementation, each step of the above method or each of the above units can be implemented through integrated logic circuits in the processor element or through software invoked by the processing element.
[0336] In one example, the unit in the above device may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.
[0337] For example, when the units in the device can be implemented through a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these units can be integrated together to form a system-on-a-chip (SOC).
[0338] In one implementation, the units that implement the corresponding steps in the above methods can be implemented in the form of a processing element scheduler. For example, the device may include a processing element and a storage element, wherein the processing element calls a program stored in the storage element to execute the method described in the above method embodiments. The storage element may be a storage element located on the same chip as the processing element, i.e., an on-chip storage element.
[0339] In another implementation, the program used to perform the above methods can be located on a storage element on a different chip than the processing element, i.e., an off-chip storage element. In this case, the processing element calls or loads the program from the off-chip storage element onto the on-chip storage element to call and execute the steps performed by the terminal device or server in the above method embodiments.
[0340] For example, embodiments of this application may also provide an apparatus, such as an electronic device. The electronic device may include: a processor; a memory; and a computer program; wherein the computer program is stored in the memory, and when the computer program is executed by the processor, it causes the electronic device to perform the steps executed by the terminal device or server in the 3D modeling method described in the foregoing embodiments. The memory may be located within or outside the electronic device. The processor may include one or more processors.
[0341] For example, the electronic device can be a mobile phone, a large screen (such as a smart screen), a tablet computer, a wearable device (such as a smartwatch, a smart bracelet, etc.), a television, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and other terminal devices.
[0342] In another implementation, the unit that implements the steps of the above method can be configured as one or more processing elements, which can be integrated circuits, such as one or more ASICs, or one or more DSPs, or one or more FPGAs, or combinations of these types of integrated circuits. These integrated circuits can be integrated together to form a chip.
[0343] For example, this application also provides a chip that can be applied to the aforementioned electronic device. The chip includes one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the processor receives and executes computer instructions from the memory of the electronic device through the interface circuits to implement the steps performed by the terminal device or server in the 3D modeling method described in the foregoing embodiments.
[0344] This application also provides a computer program product, including computer-readable code, which, when run in an electronic device, causes the electronic device to perform the steps executed by the terminal device or server in the 3D modeling method described in the foregoing embodiments.
[0345] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0346] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0347] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium.
[0348] Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product, such as a program. This software product is stored in a program product, such as a computer-readable storage medium, and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0349] For example, embodiments of this application may also provide a computer-readable storage medium, which includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps executed by the terminal device or server in the 3D modeling method described in the foregoing embodiments.
[0350] Optionally, embodiments of this application also provide an edge-cloud collaborative system, the composition of which can refer to the above. Figure 1 or Figure 20As shown, the system includes a terminal device and a server, with the terminal device connected to the server. The terminal device displays a first interface, which includes the image captured by the terminal device. In response to a capture operation, the terminal device captures multiple frames of images corresponding to the target object to be modeled and obtains the correlation between the multiple frames. During the capture of the multiple frames of images corresponding to the target object, the terminal device displays a first virtual bounding volume. The first virtual bounding volume includes multiple facets. The process of the terminal device capturing multiple frames of images corresponding to the target object to be modeled and obtaining the correlation between the multiple frames includes: when the terminal device is in a first pose, the terminal device captures the first image and changes the... The process involves: 1) Describing the display effect of the facets corresponding to the first image; 2) When the terminal device is in a second pose, acquiring a second image and changing the display effect of the facets corresponding to the second image; 3) After changing the display effect of the multiple facets of the first virtual bounding body, obtaining the association relationship between the multiple frames of images based on the multiple facets; 4) Sending the multiple frames of images and the association relationship between the multiple frames of images to the server; 5) Obtaining the 3D model corresponding to the target object based on the multiple frames of images and the association relationship between the multiple frames of images; 6) Sending the 3D model corresponding to the target object to the terminal device; 7) Displaying the 3D model corresponding to the target object.
[0351] Similarly, in this end-to-cloud collaborative system, the terminal device can realize all the functions that the terminal device can realize in the 3D modeling method described in the aforementioned method embodiments, and the server can realize all the functions that the server can realize in the 3D modeling method described in the aforementioned method embodiments, which will not be elaborated here.
[0352] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A modeling method, characterized in that, The method is applied to a terminal device, and the method includes: The terminal device displays a first interface, which includes the camera image captured by the terminal device. The terminal device responds to the acquisition operation by acquiring multiple frames of images corresponding to the target object to be modeled, and obtains the correlation between the multiple frames of images; wherein, during the acquisition of the multiple frames of images corresponding to the target object, the terminal device displays a first model; the first model includes multiple patches; The terminal device responds to the acquisition operation by acquiring multiple frames of images corresponding to the target object to be modeled, and obtains the correlation between the multiple frames of images, including: When the terminal device is in the first pose, the terminal device acquires the first image and changes the display effect of the patch corresponding to the first image; When the terminal device is in the second pose, the terminal device acquires the second image and changes the display effect of the corresponding patch of the second image; When the display effect of the multiple patches of the first model is changed, the terminal device obtains the correlation relationship between the multiple frames of images based on the multiple patches; The terminal device obtains the three-dimensional model corresponding to the target object based on the multi-frame images and the correlation between the multi-frame images; The terminal device displays a three-dimensional model corresponding to the target object.
2. The method according to claim 1, characterized in that, The terminal device includes a first application, and before the terminal device displays a first interface, the method further includes: The terminal device displays a second interface in response to the operation of opening the first application; The terminal device displays a first interface, including: The terminal device displays the first interface in response to the operation of activating the 3D modeling function of the first application on the second interface.
3. The method according to claim 1 or 2, characterized in that, The first model includes one or more layers, and the plurality of facets are distributed in the one or more layers.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: The terminal device displays a first prompt message, which is used to remind the user to place the target object in the center of the captured image.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: The terminal device displays a second prompt message; the second prompt message is used to remind the user to adjust one or more of the following: the shooting environment of the target object, the shooting method of the target object, and the screen ratio of the target object.
6. The method according to any one of claims 1-5, characterized in that, Before the terminal device obtains the 3D model corresponding to the target object based on the multi-frame images and the correlation between the multi-frame images, the method further includes: The terminal device detects the operation of generating a three-dimensional model; In response to the operation of generating a 3D model, the terminal device displays a third prompt message, which is used to prompt the user that the target object is being modeled.
7. The method according to any one of claims 1-6, characterized in that, After the terminal device obtains the 3D model corresponding to the target object based on the multi-frame images and the correlation between the multi-frame images, the method further includes: The terminal device displays a fourth prompt message, which is used to prompt the user that the modeling of the target object has been completed.
8. The method according to any one of claims 1-7, characterized in that, The terminal device displays the 3D model corresponding to the target object, which also includes: The terminal device responds to the operation of changing the display angle of the three-dimensional model corresponding to the target object by changing the display angle of the three-dimensional model corresponding to the target object; the operation of changing the display angle of the three-dimensional model corresponding to the target object includes dragging the three-dimensional model corresponding to the target object to rotate clockwise or counterclockwise along a first direction.
9. The method according to any one of claims 1-8, characterized in that, The terminal device displays the 3D model corresponding to the target object, which also includes: The terminal device responds to an operation that changes the display size of the 3D model corresponding to the target object by changing the display size of the 3D model corresponding to the target object; the operation of changing the display size of the 3D model corresponding to the target object includes an operation of enlarging or shrinking the 3D model corresponding to the target object.
10. The method according to any one of claims 1-9, characterized in that, The association between the multiple frames of images includes the matching information of each frame in the multiple frames of images; The matching information for each frame of the image includes the identification information of other images associated with the image in the multi-frame image set; The matching information for each frame of the image is obtained based on the association between each frame of the image and the corresponding patch, as well as the association between the multiple patches.
11. The method according to any one of claims 1-10, characterized in that, The terminal device, in response to the acquisition operation, acquires multiple frames of images corresponding to the target object to be modeled, and the acquisition of the correlation between the multiple frames of images also includes: The terminal device determines the target object based on the captured image; When the terminal device captures the multi-frame images, the target object is positioned in the center of the captured image.
12. The method according to any one of claims 1-11, characterized in that, The terminal device acquires multiple frames of images corresponding to the target object to be modeled, including: During the process of capturing images of the target object, the terminal device performs blur detection on each captured frame and collects images with a clarity greater than a first threshold as the images corresponding to the target object.
13. The method according to any one of claims 1-12, characterized in that, The terminal device displays a 3D model corresponding to the target object, including: The terminal device responds to the operation of previewing the three-dimensional model corresponding to the target object by displaying the three-dimensional model corresponding to the target object.
14. The method according to any one of claims 1-13, characterized in that, The three-dimensional model corresponding to the target object includes the basic three-dimensional model of the target object and the texture of the surface of the target object.
15. The method according to any one of claims 1-14, characterized in that, The terminal device is connected to the server; The terminal device obtains a 3D model of the target object based on the multiple frames of images and the correlation between the multiple frames of images, including: The terminal device sends the multi-frame images and the association relationships between the multi-frame images to the server; The terminal device receives a 3D model of the target object sent from the server.
16. The method according to claim 15, characterized in that, The method further includes: The terminal device sends to the server the camera intrinsic parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to the multiple frames of images.
17. The method according to claim 15 or 16, characterized in that, The method further includes: The terminal device receives an instruction message from the server, which indicates to the terminal device that the server has completed modeling the target object.
18. The method according to any one of claims 15-17, characterized in that, Before the terminal device receives the 3D model corresponding to the target object sent by the server, the method further includes: The terminal device sends a download request message to the server, the download request message being used to request the server to download the 3D model corresponding to the target object.
19. A modeling method, characterized in that, The method is applied to a server, which is connected to a terminal device; the method includes: The server receives multiple frames of images corresponding to the target object and the association between the multiple frames of images sent by the terminal device; The server generates a 3D model of the target object based on the multiple frames of images and the relationships between them. The server sends the 3D model of the target object to the terminal device.
20. The method according to claim 19, characterized in that, The method further includes: The server receives camera intrinsic parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to the multiple frames of images sent by the terminal device. The server generates a 3D model of the target object based on the multiple frames of images and the relationships between them, including: The server generates a 3D model of the target object based on the multiple frames of images, the relationships between the multiple frames of images, the camera intrinsic parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to each of the multiple frames of images.
21. The method according to claim 19 or 20, characterized in that, The association between the multiple frames of images includes the matching information of each frame in the multiple frames of images; The matching information for each frame of the image includes the identification information of other images associated with the image in the multi-frame image set; The matching information for each frame of the image is obtained based on the association between each frame of the image and the corresponding patch, as well as the association between the multiple patches.
22. An edge-cloud collaborative system, characterized in that, include: A terminal device and a server, wherein the terminal device is connected to the server; The terminal device displays a first interface, which includes the camera image captured by the terminal device. The terminal device responds to the acquisition operation by acquiring multiple frames of images corresponding to the target object to be modeled, and obtains the correlation between the multiple frames of images; wherein, during the acquisition of the multiple frames of images corresponding to the target object, the terminal device displays a first model; the first model includes multiple patches; The terminal device responds to the acquisition operation by acquiring multiple frames of images corresponding to the target object to be modeled, and obtains the correlation between the multiple frames of images, including: When the terminal device is in the first pose, the terminal device acquires the first image and changes the display effect of the patch corresponding to the first image; When the terminal device is in the second pose, the terminal device acquires the second image and changes the display effect of the corresponding patch of the second image; When the display effect of the multiple patches of the first model is changed, the terminal device obtains the correlation relationship between the multiple frames of images based on the multiple patches; The terminal device sends the multi-frame images and the association relationships between the multi-frame images to the server; The server obtains the 3D model corresponding to the target object based on the multi-frame images and the correlation between the multi-frame images; The server sends the 3D model corresponding to the target object to the terminal device; The terminal device displays a three-dimensional model corresponding to the target object.
23. An electronic device, characterized in that, include: processor; Memory; And a computer program; wherein the computer program is stored on the memory, and when the computer program is executed by the processor, causes the electronic device to perform the method as described in any one of claims 1-18, or the method as described in any one of claims 19-21.
24. A computer-readable storage medium comprising a computer program, characterized in that, When the computer program is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-18, or the method as described in any one of claims 19-21.
Citation Information
Cited By
Information terminal and guidance method using information terminal
JPWO2025027753A1