Modeling method, related electronic device and storage medium

By displaying virtual enclosures on terminal devices to acquire images and using RGB cameras, combined with cloud computing, the problems of high hardware requirements and complex operations in the existing 3D modeling methods are solved, and a low hardware requirements and efficient 3D modeling process is realized.

CN115526925BActive Publication Date: 2025-07-11HUAWEI TECH CO LTD

Patent Information

Application Number
CN202110715044.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-26
Publication Date
2025-07-11
Estimated Expiration
2041-06-26

AI Technical Summary

Technical Problem

In the existing 3D modeling methods, the data acquisition and reconstruction process is complex, the hardware requirements are high, special equipment such as LIDAR sensors or RGB-D cameras are required, and the user operation is complicated.

Method used

The virtual enclosure is displayed through the terminal device, multi-frame images are collected and related relationships are obtained, ordinary RGB cameras are used for 3D modeling, and model reconstruction is carried out in combination with cloud computing resources, reducing hardware requirements and simplifying operation processes.

Benefits of technology

It realizes 3D modeling with low hardware requirements, simplifies the data acquisition and reconstruction process, and improves modeling efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115526925B_ABST
    Figure CN115526925B_ABST
Patent Text Reader

Abstract

The present application provides a modeling method, related electronic devices, and storage media, which relate to the field of three-dimensional reconstruction. In this method, a terminal device can display a first interface including a captured image; in response to a collection operation, collect multiple frames of images corresponding to a target object to be modeled, and obtain the association relationship between the multiple frames of images; according to the multiple frames of images and the association relationship between the multiple frames of images, obtain a three-dimensional model corresponding to the target object and display it. When collecting multiple frames of images, a first virtual bounding volume including multiple patches is displayed. When the terminal device is in a first pose, collect a first image and change the display effect of the patch corresponding to the first image; when the terminal device is in a second pose, collect a second image and change the display effect of the patch corresponding to the second image; after changing the display effects of the multiple patches, obtain the association relationship between the multiple frames of images according to the multiple patches. This method simplifies the 3D modeling process and has low requirements for device hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of 3D reconstruction, and in particular, to a modeling method, related electronic devices, and storage media. Background Art

[0002] 3D reconstruction applications / software can be used to perform 3D modeling on objects. Currently, when implementing 3D modeling, users need to first use mobile tools (such as mobile phones, cameras, etc.) to collect data (such as pictures, depth information, etc.) required for 3D modeling. Then, the 3D reconstruction application can perform 3D reconstruction on the object based on the collected data required for 3D modeling to obtain a 3D model corresponding to the object.

[0003] However, in the current 3D modeling methods, the process of collecting data required for 3D modeling and the process of performing 3D reconstruction on the object based on the collected data required for 3D modeling are relatively complex, and the requirements for device hardware are relatively high. For example, the process of collecting data required for 3D modeling requires the collection device (such as the above-mentioned mobile tool) to be equipped with special hardware such as a lidar (light detection and ranging, LIDAR) sensor or an RGB depth (RGB-D) camera. The process of performing 3D reconstruction on the object based on the collected data required for 3D modeling requires the processing device running the 3D reconstruction application to be equipped with a high-performance independent graphics card. Summary of the Invention

[0004] Embodiments of this application provide a modeling method, related electronic devices, and storage media, which simplify the process of collecting data required for 3D modeling and the process of performing 3D reconstruction on the object based on the collected data required for 3D modeling, and have relatively low requirements for device hardware.

[0005] In a first aspect, embodiments of this application provide a modeling method, which is applied to a terminal device. The method includes:

[0006] The terminal device displays a first interface, and the first interface includes the shooting screen of the terminal device. The terminal device responds to a collection operation, collects multiple frames of images corresponding to a target object to be modeled, and obtains the association relationship between the multiple frames of images. The terminal device obtains a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images. The terminal device displays the three-dimensional model corresponding to the target object.

[0007] Wherein, during the process of collecting multiple frames of images corresponding to the target object, the terminal device displays a first virtual bounding volume; the first virtual bounding volume includes multiple patches. The terminal device responds to a collection operation, collects multiple frames of images corresponding to a target object to be modeled, and obtains the association relationship between the multiple frames of images, including:

[0008] When the terminal device is in the first pose, the terminal device captures a first image and changes the display effect of the patch corresponding to the first image; when the terminal device is in the second pose, the terminal device captures a second image and changes the display effect of the patch corresponding to the second image; after changing the display effects of the multiple patches of the first virtual enclosure, the terminal device obtains the correlation relationship between the multiple frames of images according to the multiple patches.

[0009] Exemplarily, for each key frame, the terminal device can determine the matching information of the key frame according to the correlation relationship between the patch corresponding to the key frame and other patches. Among them, the correlation relationship between the patch corresponding to the key frame and other patches may include: in the patch model, which patches correspond to the upper, lower, left, and right directions of the patch corresponding to the key frame.

[0010] For example, taking the patch model including two layers and each layer including 20 patches as an example, assuming that the patch corresponding to a certain key frame is patch 1 in the first layer, the correlation relationship between patch 1 and other patches may include: the patch below patch 1 is patch 21, the patch to the left of patch 1 is patch 20, and the patch to the right of patch 1 is patch 2. According to the foregoing correlation relationship between patch 1 and other patches, the mobile phone can determine that the other key frames associated with this key frame include the key frame corresponding to patch 21, the key frame corresponding to patch 20, and the key frame corresponding to patch 2. Thus, the mobile phone can obtain the matching information of this key frame including the identification information of the key frame corresponding to patch 21, the identification information of the key frame corresponding to patch 20, and the identification information of the key frame corresponding to patch 2.

[0011] It can be understood that for the mobile phone, the relationship between different patches in the patch model is a known quantity.

[0012] In this modeling method (or called 3D modeling method), the terminal device can realize 3D modeling only by relying on an ordinary RGB camera to collect the data required for 3D modeling. The process of collecting the data required for 3D modeling does not need to rely on special hardware such as a LIDAR sensor or an RGB-D camera on the terminal device. The terminal device obtains the three-dimensional model corresponding to the target object according to multiple frames of images and the correlation relationship between multiple frames of images, which can effectively reduce the computational load in the 3D modeling process and improve the efficiency of 3D modeling.

[0013] In addition, in this method, the user only needs to perform operations related to collecting the data required for 3D modeling on the terminal device side, and then view or preview the final 3D model on the terminal device. For the user, all operations are completed on the terminal device side, the operations are simpler, and the user experience can be better.

[0014] In a possible design, the terminal device includes a first application. Before the terminal device displays a first interface, the method further includes: the terminal device responds to an operation of opening the first application and displays a second interface.

[0015] The terminal device displays the first interface, including: the terminal device responds to an operation of starting the 3D modeling function of the first application on the second interface and displays the first interface.

[0016] For example, the second interface may include a function control for starting the 3D modeling function. The user can click or touch the function control on the second interface, and the mobile phone can respond to the operation of the user clicking or touching the function control on the second interface and start the 3D modeling function of the first application. That is, the operation of the user clicking or touching the function control on the second interface is the operation of starting the 3D modeling function of the first application on the second interface.

[0017] In some embodiments, the first virtual enclosure includes one or more layers, and the multiple patches are distributed on the one or more layers.

[0018] For example, in one implementation, the structure of the patch model may include upper and lower two layers, and each layer may include multiple patches. In another implementation, the structure of the patch model may include upper, middle, and lower three layers, and each layer may include multiple patches. In yet another implementation, the structure of the patch model may be a one-layer structure composed of multiple patches. There is no limitation here.

[0019] Optionally, the method further includes: the terminal device displays a first prompt message, and the first prompt message is used to remind the user to place the position of the target object in the center of the captured image.

[0020] For example, the first prompt message may be "Please place the target object in the center of the screen".

[0021] Optionally, the method further includes: the terminal device displays a second prompt message; the second prompt message is used to remind the user to adjust one or more of the shooting environment where the target object is located, the shooting method for the target object, and the screen occupation ratio of the target object.

[0022] For example, the second prompt message may be "Place the object statically on a solid - color plane, with soft lighting, shoot around the object for one week, and the screen occupation ratio of the object should be as large and complete as possible".

[0023] In this embodiment, after the user adjusts the shooting environment where the target object is located, the screen occupation ratio of the object, etc. according to the content of the second prompt message, the subsequent data acquisition process can be faster and the quality of the acquired data can be better.

[0024] Optionally, before the terminal device obtains the three-dimensional model corresponding to the target object according to the multi-frame images and the correlation relationship between the multi-frame images, the method further includes: the terminal device detecting an operation of generating a three-dimensional model; the terminal device, in response to the operation of generating a three-dimensional model, displaying a third prompt message for prompting the user that a model is being built for the target object.

[0025] For example, the third prompt message may be "Modeling".

[0026] Optionally, after the terminal device obtains the three-dimensional model corresponding to the target object according to the multi-frame images and the correlation relationship between the multi-frame images, the method further includes: the terminal device displaying a fourth prompt message for prompting the user that the modeling of the target object has been completed.

[0027] For example, the fourth prompt message may be "Modeling completed".

[0028] Optionally, the terminal device displaying the three-dimensional model corresponding to the target object further includes: the terminal device, in response to an operation of changing the display angle of the three-dimensional model corresponding to the target object, changing the display angle of the three-dimensional model corresponding to the target object; the operation of changing the display angle of the three-dimensional model corresponding to the target object includes an operation of dragging the three-dimensional model corresponding to the target object to rotate clockwise or counterclockwise along a first direction.

[0029] Wherein, the first direction may be any direction, such as a horizontal direction, a vertical direction, etc. The terminal device, in response to an operation of changing the display angle of the three-dimensional model corresponding to the target object, changing the display angle of the three-dimensional model corresponding to the target object, can achieve the effect of presenting the 3D model to the user at different angles.

[0030] Optionally, the terminal device displaying the three-dimensional model corresponding to the target object further includes: the terminal device, in response to an operation of changing the display size of the three-dimensional model corresponding to the target object, changing the display size of the three-dimensional model corresponding to the target object; the operation of changing the display size of the three-dimensional model corresponding to the target object includes an operation of enlarging or reducing the three-dimensional model corresponding to the target object.

[0031] For example, the reduction operation may be an operation in which the user slides two fingers inward (opposite directions) on the 3D model preview interface, and the enlargement operation may be an operation in which the user slides two fingers outward (opposite directions) on the 3D model preview interface. The 3D model preview interface is also the interface on which the terminal device displays the 3D model corresponding to the target object.

[0032] In some other implementation manners, the zoom-in operation or zoom-out operation performed by the user on the three-dimensional model corresponding to the target object may also be a double-click operation, a long-press operation, or alternatively, the 3D model preview interface may also include a function control for performing a zoom-in operation or a zoom-out operation, etc., which is not limited herein.

[0033] In some embodiments, the association relationship between the multiple frames of images includes the matching information of each frame of image in the multiple frames of images; the matching information of each frame of image includes the identification information of other images associated with the image in the multiple frames of images; the matching information of each frame of image is obtained according to the association relationship between each frame of image and the patch corresponding to each frame of image, and the association relationship between the multiple patches.

[0034] For example, for the key frame with the picture number 18, the identification information of other key frames associated with the key frame with the picture number 18 is the picture numbers of other key frames associated with the key frame with the picture number 18, such as 26, 45, 59, 78, 89, 100, 449, etc.

[0035] Optionally, when the terminal device responds to the acquisition operation, acquires multiple frames of images corresponding to the target object to be modeled, and obtains the association relationship between the multiple frames of images, it further includes: the terminal device determines the target object according to the captured picture; when the terminal device acquires multiple frames of images, the position of the target object in the captured picture is the central position of the captured picture.

[0036] Optionally, when the terminal device acquires multiple frames of images corresponding to the target object to be modeled, it includes: during the process of shooting the target object, the terminal device performs blur detection on each frame of image captured, and acquires the images with clarity greater than the first threshold as the images corresponding to the target object.

[0037] For each shooting position (one shooting position may correspond to one patch), the terminal device can obtain some key frame pictures with better quality by performing blur detection on the picture captured at this shooting position. The key frame pictures are the images corresponding to the target object, and the number of key frame pictures corresponding to each patch can be one or more.

[0038] Optionally, when the terminal device displays the three-dimensional model corresponding to the target object, it includes: the terminal device responds to the operation of previewing the three-dimensional model corresponding to the target object, and displays the three-dimensional model corresponding to the target object.

[0039] For example, the terminal device may display a view button, and the user can click this view button. The terminal device can respond to the operation of the user clicking this view button and display the three-dimensional model corresponding to the target object. The operation of the user clicking this view button is the operation of previewing the three-dimensional model corresponding to the target object.

[0040] Optionally, the three-dimensional model corresponding to the target object includes the basic three-dimensional model of the target object and the texture on the surface of the target object.

[0041] The texture on the surface of the target object may be a texture map of the surface of the target object. According to the basic 3D model of the target object and the texture on the surface of the target object, the 3D model of the target object can be generated. Mapping the texture on the surface of the target object onto the surface of the basic 3D model of the target object in a specific manner can more realistically restore the surface of the target object and make the target object look more real.

[0042] In a possible design, the terminal device is connected to the server; the terminal device obtains the three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images, including: the terminal device sends the multiple frames of images and the association relationship between the multiple frames of images to the server; the terminal device receives the three-dimensional model corresponding to the target object sent by the server.

[0043] In this design, the process of generating the three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images can be completed on the server side. That is, this design can use the computing resources of the server to implement 3D modeling. This design can be applied to scenarios where the computing power of some terminal devices is weak, improving the universality of this 3D modeling method.

[0044] Optionally, the method further includes: the terminal device sends the camera internal parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to the multiple frames of images to the server.

[0045] Optionally, the method further includes: the terminal device receives an indication message from the server, and the indication message is used to indicate to the terminal device that the server has completed the modeling of the target object.

[0046] For example, the terminal device may display the above fourth prompt message after receiving the indication message.

[0047] Optionally, before the terminal device receives the three-dimensional model corresponding to the target object sent by the server, the method further includes: the terminal device sends a download request message to the server, and the download request message is used to request the server to download the three-dimensional model corresponding to the target object.

[0048] After receiving the download request message, the server can send the three-dimensional model corresponding to the target object to the terminal device.

[0049] Optionally, in some embodiments, when the terminal device collects data required for 3D modeling of the target object, it can also display the scanning progress on the first interface. For example, the first interface may include a scanning button, and the terminal device can display the scanning progress through the circular black filling effect in the scanning button on the first interface.

[0050] It can be understood that since the UI rendering effects of the scanning buttons are different, the ways for the mobile phone to display the scanning progress on the first interface can be different, which is not limited herein.

[0051] In some other embodiments, the terminal device may also not need to display the scanning progress. The user can understand the scanning progress according to the lighting conditions of the patches in the first virtual bounding volume.

[0052] In a second aspect, an embodiment of the present application provides a modeling device, which can be applied to a terminal device and is used to implement the modeling method described in the first aspect above. The functions of the device can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, the device may include a display unit and a processing unit. The display unit and the processing unit can be used to cooperate to implement the modeling method described in the first aspect above.

[0053] For example, the display unit is used to display the first interface, and the first interface includes the captured image of the terminal device.

[0054] The processing unit is used to collect multiple frames of images corresponding to the target object to be modeled in response to a collection operation, and obtain the correlation relationship between the multiple frames of images. According to the multiple frames of images and the correlation relationship between the multiple frames of images, a three-dimensional model corresponding to the target object is obtained.

[0055] The display unit is further used to display the three-dimensional model corresponding to the target object.

[0056] Among them, during the process of collecting multiple frames of images corresponding to the target object, the display unit is further used to display a first virtual bounding volume; the first virtual bounding volume includes multiple patches. The processing unit is specifically used to collect a first image when the terminal device is in a first pose and change the display effect of the patch corresponding to the first image; collect a second image when the terminal device is in a second pose and change the display effect of the patch corresponding to the second image; after changing the display effects of the multiple patches of the first virtual bounding volume, obtain the correlation relationship between the multiple frames of images according to the multiple patches.

[0057] Optionally, the display unit and the processing unit are further used to implement other display functions and processing functions in the method described in the first aspect above, which will not be elaborated herein one by one.

[0058] Optionally, for the implementation manner in which the terminal device described in the above first aspect sends the multiple frames of images and the association relationship between the multiple frames of images to the server, and the server generates a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images, the modeling device may further include a sending unit and a receiving unit. The sending unit is configured to send the multiple frames of images and the association relationship between the multiple frames of images to the server, and the receiving unit is configured to receive the three-dimensional model corresponding to the target object sent from the server.

[0059] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor; a memory; and a computer program; wherein, the computer program is stored on the memory, and when the computer program is executed by the processor, the electronic device implements the method described in the first aspect and any possible implementation manner of the first aspect.

[0060] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the electronic device implements the method described in the first aspect and any possible implementation manner of the first aspect.

[0061] In a fifth aspect, an embodiment of the present application further provides a computer program product, including computer-readable code, and when the computer-readable code runs in an electronic device, the electronic device implements the method described in the first aspect and any possible implementation manner of the first aspect.

[0062] The beneficial effects of the above second aspect to fifth aspect can be referred to those described in the first aspect, and will not be elaborated here.

[0063] In a sixth aspect, an embodiment of the present application further provides a modeling method, the method is applied to a server, and the server is connected to a terminal device; the method includes: the server receives multiple frames of images corresponding to a target object and the association relationship between the multiple frames of images sent from the terminal device; the server generates a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images; the server sends the three-dimensional model corresponding to the target object to the terminal device.

[0064] This method can utilize the computing resources of the server to implement 3D modeling, and the process of generating a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images can be completed on the server side. The server combines the association relationship between the multiple frames of images for 3D modeling, which can effectively reduce the computing load of the server and improve the modeling efficiency.

[0065] For example, when the server performs 3D modeling, it can combine the correlation relationships between the multiple frames of images, and perform feature detection and matching on each frame of image and other images associated with that image, without the need to perform feature detection and matching between that image and all other images. In this way, two adjacent frames of images can be quickly compared, effectively reducing the computational load of the server and improving the efficiency of 3D modeling.

[0066] For another example, after the server determines the mapping relationship between the texture of the first frame of image and the surface of the basic 3D model of the target object, it can quickly and accurately determine the mapping relationship between the textures of other images associated with the first frame of image and the surface of the basic 3D model of the target object in combination with the matching information of the first frame of image. Similarly, for each subsequent frame of image, the server can quickly and accurately determine the mapping relationship between the textures of other images associated with that frame of image and the surface of the basic 3D model of the target object in combination with the matching information of that frame of image.

[0067] This method can be applicable to scenarios where the computing power of some terminal devices is weak, improving the universality of this 3D modeling method.

[0068] In addition, this method also has other beneficial effects described in the first aspect above, such as: in the process of collecting data required for 3D modeling, there is no need to rely on special hardware such as LIDAR sensors or RGB-D cameras on the terminal device. The server can obtain the three-dimensional model corresponding to the target object according to the multiple frames of images and the correlation relationships between the multiple frames of images, which can effectively reduce the computational load in the 3D modeling process and improve the efficiency of 3D modeling, etc., and will not be elaborated here one by one.

[0069] Optionally, the method further includes: the server receives the camera internal parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to the multiple frames of images sent from the terminal device.

[0070] The server generates the three-dimensional model corresponding to the target object according to the multiple frames of images and the correlation relationships between the multiple frames of images, including: the server generates the three-dimensional model corresponding to the target object according to the multiple frames of images, the correlation relationships between the multiple frames of images, the camera internal parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to the multiple frames of images.

[0071] Optionally, the association relationship between the multiple-frame images includes the matching information of each frame image in the multiple-frame images; the matching information of each frame image includes the identification information of other images associated with the image in the multiple-frame images; the matching information of each frame image is obtained according to the association relationship between each frame image and the patch corresponding to each frame image, and the association relationship between the multiple patches.

[0072] In a seventh aspect, an embodiment of the present application provides a modeling device, which can be applied to a server and is used to implement the modeling method described in the sixth aspect above. The functions of the device can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, the device may include: a receiving unit, a processing unit, and a sending unit. The receiving unit, the processing unit, and the sending unit can be used to cooperate to implement the modeling method described in the sixth aspect above.

[0073] For example, the receiving unit can be used to receive multiple-frame images corresponding to a target object and the association relationship between the multiple-frame images sent from a terminal device. The processing unit can be used to generate a three-dimensional model corresponding to the target object according to the multiple-frame images and the association relationship between the multiple-frame images. The sending unit can be used to send the three-dimensional model corresponding to the target object to the terminal device.

[0074] Optionally, the receiving unit, the processing unit, and the sending unit can be used to implement all the functions that the server can implement in the method described in the sixth aspect above, which will not be elaborated here one by one.

[0075] In an eighth aspect, an embodiment of the present application provides an electronic device, including: a processor; a memory; and a computer program; wherein, the computer program is stored on the memory, and when the computer program is executed by the processor, the electronic device is enabled to implement the method described in the sixth aspect and any possible implementation manner of the sixth aspect.

[0076] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the electronic device is enabled to implement the method described in the sixth aspect and any possible implementation manner of the sixth aspect.

[0077] In a tenth aspect, an embodiment of the present application further provides a computer program product, including computer-readable code. When the computer-readable code runs in an electronic device, the electronic device is enabled to implement the method described in the sixth aspect and any possible implementation manner of the sixth aspect.

[0078] For the beneficial effects of the seventh to tenth aspects described above, reference may be made to those described in the sixth aspect, and details are not repeated here.

[0079] Eleventh aspect, the embodiment of the present application further provides a terminal-cloud collaboration system, including: a terminal device and a server, the terminal device is connected to the server; the terminal device displays a first interface, and the first interface includes a captured image of the terminal device; the terminal device responds to a collection operation, collects multiple frames of images corresponding to a target object to be modeled, and obtains the association relationship between the multiple frames of images; wherein, during the process of collecting the multiple frames of images corresponding to the target object, the terminal device displays a first virtual bounding volume; the first virtual bounding volume includes multiple patches; the terminal device responds to a collection operation, collects multiple frames of images corresponding to a target object to be modeled, and obtains the association relationship between the multiple frames of images, including: when the terminal device is in a first pose, the terminal device collects a first image and changes the display effect of the patch corresponding to the first image; when the terminal device is in a second pose, the terminal device collects a second image and changes the display effect of the patch corresponding to the second image; after changing the display effects of the multiple patches of the first virtual bounding volume, the terminal device obtains the association relationship between the multiple frames of images according to the multiple patches; the terminal device sends the multiple frames of images and the association relationship between the multiple frames of images to the server; the server obtains a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images; the server sends the three-dimensional model corresponding to the target object to the terminal device; the terminal device displays the three-dimensional model corresponding to the target object.

[0080] For the beneficial effects of the eleventh aspect described above, reference may be made to those described in the first and sixth aspects, and details are not repeated here.

[0081] It should be understood that the description of technical features, technical solutions, beneficial effects or similar languages in the present application does not imply that all features and advantages can be achieved in any single embodiment. On the contrary, it can be understood that the description of features or beneficial effects means that at least one embodiment includes specific technical features, technical solutions or beneficial effects. Therefore, the description of technical features, technical solutions or beneficial effects in this specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and beneficial effects described in this embodiment can be combined in any appropriate manner. Those skilled in the art will understand that an embodiment can be implemented without one or more specific technical features, technical solutions or beneficial effects of a specific embodiment. In other embodiments, additional technical features and beneficial effects can also be identified in specific embodiments that do not embody all embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 Schematic diagram of the composition of the end-cloud collaboration system provided by the embodiments of the present application;

[0083] Figure 2 Schematic diagram of the structure of the terminal device provided by the embodiments of the present application;

[0084] Figure 3 Schematic diagram of the main interface of the mobile phone provided by the embodiments of the present application;

[0085] Figure 4 Schematic diagram of the main interface of the first application provided by the embodiments of the present application;

[0086] Figure 5 Schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0087] Figure 6 Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0088] Figure 7A Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0089] Figure 7B Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0090] Figure 7C Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0091] Figure 7D Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0092] Figure 7E Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0093] Figure 7F Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0094] Figure 8 Schematic diagram of the structure of the patch model provided by the embodiments of the present application;

[0095] Figure 9 Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0096] Figure 10 Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application;

[0097] Figure 11Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiment of the present application;

[0098] Figure 12 Schematic diagram of the 3D model preview interface provided by the embodiment of the present application;

[0099] Figure 13 Another schematic diagram of the 3D model preview interface provided by the embodiment of the present application;

[0100] Figure 14 Another schematic diagram of the 3D model preview interface provided by the embodiment of the present application;

[0101] Figure 15 Schematic diagram of the user performing a counterclockwise rotation operation along the horizontal direction on the 3D model of the toy car provided by the embodiment of the present application;

[0102] Figure 16 Another schematic diagram of the 3D model preview interface provided by the embodiment of the present application;

[0103] Figure 17 Schematic of the user performing a reduction operation on the 3D model of the toy car provided by the embodiment of the present application;

[0104] Figure 18 Another schematic diagram of the 3D model preview interface provided by the embodiment of the present application;

[0105] Figure 19 Flow schematic diagram of the 3D modeling method provided by the embodiment of the present application;

[0106] Figure 20 Logical schematic diagram of the end-cloud collaboration system implementing the 3D modeling method provided by the embodiment of the present application;

[0107] Figure 21 Structural schematic diagram of the modeling device provided by the embodiment of the present application;

[0108] Figure 22 Another structural schematic diagram of the modeling device provided by the embodiment of the present application;

[0109] Figure 23 Another structural schematic diagram of the modeling device provided by the embodiment of the present application. Detailed implementation manners

[0110] The terms used in the following embodiments are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The character " / " generally indicates an "or" relationship between the related objects before and after.

[0111] Reference to "one embodiment" or "some embodiments" or the like described in this specification means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way. The term "connection" includes direct connection and indirect connection, unless otherwise stated.

[0112] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0113] In the embodiments of the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0114] Three-dimensional (3D) reconstruction technology is widely used in fields such as virtual reality (VR), augmented reality (AR), extended reality (XR), mixed reality (MR), gaming, film and television, education, and healthcare. For example, 3D reconstruction technology can be used to model characters, props, vegetation, etc. in games, or to model human figures in film and television. It can also be used to achieve modeling related to chemical analysis structures in the field of education and modeling related to human structures in the field of healthcare, etc.

[0115] Currently, most 3D reconstruction applications / software that can be used to achieve 3D modeling need to be implemented on the personal computer (PC) side (such as a computer). A small number of 3D reconstruction applications can achieve 3D modeling on the mobile device side (such as a mobile phone). When a 3D reconstruction application on the PC side achieves 3D modeling, the user needs to first use a mobile device tool (such as a mobile phone, camera, etc.) to collect the data required for 3D modeling (such as pictures, depth information, etc.), and upload the collected data to the PC side. Then, the 3D reconstruction application on the PC side can perform 3D modeling processing based on the uploaded data. When a 3D reconstruction application on the mobile device side achieves 3D modeling, the user can directly use the mobile device to collect the data required for 3D modeling, and the 3D reconstruction application on the mobile device side can directly perform 3D modeling processing based on the data collected by the mobile device.

[0116] However, in the above two 3D modeling methods, when the user uses the mobile device to collect the data required for 3D modeling, they both rely on special hardware such as a light detection and ranging (LIDAR) sensor or an RGB depth (RGB-D) camera on the mobile device. The data collection process for 3D modeling requires relatively high hardware requirements. The 3D reconstruction applications on the PC side / mobile device side also have relatively high hardware requirements for achieving 3D modeling. For example, it may be necessary to configure a high-performance independent graphics card on the PC side / mobile device side.

[0117] In addition, the above method of achieving 3D modeling on the PC side is also relatively cumbersome. For example, after the user performs relevant data collection operations on the mobile device, not only does the user need to copy or transmit the collected data to the PC side through the network, but also relevant modeling operations need to be performed on the 3D reconstruction application on the PC side.

[0118] Under this background art, an embodiment of the present application provides a 3D modeling method, which can be applied to an end-cloud collaborative system composed of a terminal device and a cloud. Among them, the "end" in end-cloud collaboration refers to the terminal device, and the "cloud" refers to the cloud, which can also be referred to as a cloud server or a cloud platform. In this method, the terminal device can collect the data required for 3D modeling, preprocess the data required for 3D modeling, and then upload the preprocessed data required for 3D modeling to the cloud; the cloud can perform 3D modeling based on the received preprocessed data required for 3D modeling; the terminal device can download the 3D model obtained by the cloud for 3D modeling from the cloud and provide a preview function for the 3D model.

[0119] In this 3D modeling method, the terminal device can realize 3D modeling only by relying on an ordinary RGB camera to collect the data required for 3D modeling. In the process of collecting the data required for 3D modeling, it does not need to rely on the terminal device having special hardware such as a LIDAR sensor or an RGB-D camera; the process of performing 3D modeling is completed in the cloud, and it does not need to rely on the terminal device being configured with a high-performance independent graphics card. That is to say, this 3D modeling method has low hardware requirements for the terminal device.

[0120] In addition, compared with the above-mentioned method of realizing 3D modeling on the PC side, in this method, the user only needs to perform operations related to collecting the data required for 3D modeling on the terminal device side, and then view or preview the final 3D model on the terminal device. For the user, all operations are completed on the terminal device side, the operation is simpler, and the user experience can be better.

[0121] Exemplarily, Figure 1 is a schematic diagram of the composition of the end-cloud collaborative system provided by the embodiment of the present application. As Figure 1 shown, the end-cloud collaborative system provided by the embodiment of the present application can include: a cloud 100 and a terminal device 200, and the terminal device 200 can be connected to the cloud 100 through a wireless network.

[0122] Among them, the cloud 100 is also the server. For example, in some embodiments, the cloud 100 can be a single server or a server cluster composed of multiple servers, and the present application does not limit the implementation architecture of the cloud 100.

[0123] Optionally, in the embodiments of the present application, the terminal device 200 may be an interactive electronic whiteboard with a shooting function, a mobile phone, a wearable device (such as a smart watch, a smart bracelet, etc.), a tablet computer, a notebook computer, a desktop computer, a portable electronic device (such as a laptop), an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart TV (such as a smart screen), an in-vehicle computer, a smart speaker, an augmented reality (AR) device, a virtual reality (VR) device, and other intelligent devices with a display screen. Alternatively, it may also be a professional shooting device such as a digital camera, a single-lens reflex camera / mirrorless camera, an action camera, a gimbal camera, a drone, etc. The embodiments of the present application do not limit the specific type of the terminal device.

[0124] It should be understood that when the terminal device is a shooting device such as a gimbal camera or a drone, it will also include a display device that can provide a shooting interface for displaying a data acquisition interface for collecting data required for 3D modeling, a preview interface of the 3D model, etc. For example, the display device of a gimbal camera may be a mobile phone, and the display device of an aerial drone may be a remote control device, etc.

[0125] It should be noted that Figure 1 an exemplary terminal device 200 is given. However, it should be understood that the terminal device 200 in the terminal-cloud collaboration system may include one or more. The multiple terminal devices 200 may be the same, or different or partially the same, and no limitation is made here. The 3D modeling method provided by the embodiments of the present application is a process for realizing 3D modeling through the interaction between each terminal device 200 and the cloud 100.

[0126] Exemplarily, taking the terminal device 200 as a mobile phone as an example, Figure 2 is the structural schematic diagram of the terminal device provided by the embodiments of the present application. As Figure 2As shown in the figure, the mobile phone may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0127] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0128] Among them, the controller may be the nerve center and command center of the mobile phone. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0129] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0130] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM interface, and / or a USB interface, etc.

[0131] The external memory interface 220 may be used to connect to an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the mobile phone. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0132] The internal memory 221 may be used to store computer-executable program codes, and the executable program codes include instructions. The processor 210 executes various functional applications and data processing of the mobile phone by running the instructions stored in the internal memory 221.

[0133] The internal memory 221 may also include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function (such as the first application described in the embodiments of the present application), etc. The storage data area may store data created during the use of the mobile phone (such as image data, phone book), etc. In addition, the internal memory 221 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0134] The charging management module 240 is used to receive a charging input from a charger. While the charging management module 240 charges the battery 242, it can also supply power to the mobile phone through the power management module 241. The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 can also receive the input from the battery 242 to supply power to the mobile phone.

[0135] The wireless communication function of the mobile phone can be realized through Antenna 1, Antenna 2, Mobile Communication Module 250, Wireless Communication Module 260, Modulation and Demodulation Processor, Baseband Processor, etc. Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the mobile phone can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, Antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0136] The mobile phone can realize audio functions through Audio Module 270, Speaker 270A, Receiver 270B, Microphone 270C, Headphone Jack 270D, and Application Processor, etc. Such as music playback, recording, etc.

[0137] Sensor Module 280 may include Pressure Sensor 280A, Gyroscope Sensor 280B, Barometric Pressure Sensor 280C, Magnetic Sensor 280D, Acceleration Sensor 280E, Distance Sensor 280F, Proximity Light Sensor 280G, Fingerprint Sensor 280H, Temperature Sensor 280J, Touch Sensor 280K, Ambient Light Sensor 280L, Bone Conduction Sensor 280M, etc.

[0138] Camera 293 can include various types. For example, Camera 293 can include a telephoto camera, a wide-angle camera, or an ultra-wide-angle camera with different focal lengths, etc. Among them, the telephoto camera has a small field of view and is suitable for shooting scenes in a small range in the distance; the wide-angle camera has a relatively large field of view; the ultra-wide-angle camera has a field of view larger than that of the wide-angle camera and can be used to shoot panoramic and other large-range pictures. In some embodiments, the telephoto camera with a smaller field of view can be rotated, so that scenes in different ranges can be shot.

[0139] The mobile phone can capture the original image (also known as RAW image or digital negative) through the camera 293. For example, the camera 293 includes at least a lens and a sensor. When taking a photo or shooting a video, the shutter is opened, and the light can be transmitted to the sensor through the lens of the camera 293. The sensor can convert the optical signal passing through the lens into an electrical signal, and then perform analogue-to-digital (A / D) conversion on the electrical signal to output the corresponding digital signal. This digital signal is the RAW image. Subsequently, the mobile phone can perform subsequent ISP processing and YUV domain processing on the RAW image through a processor (such as ISP, DSP, etc.) to convert the RAW image into an image that can be used for display, such as a JPEG image or a high efficiency image file format (HEIF) image. The JPEG image or HEIF image can be transmitted to the display screen of the mobile phone for display, and / or transmitted to the memory of the mobile phone for storage. Thus, the mobile phone can achieve the function of taking pictures.

[0140] In a possible design, the photosensitive element of the sensor can be a charge coupled device (CCD), and the sensor also includes an A / D converter. In another possible design, the photosensitive element of the sensor can be a complementary metal-oxide-semiconductor (CMOS).

[0141] Exemplarily, ISP processing may include: bad pixel correction (DPC), RAW domain noise reduction, black level correction (BLC), lens shading correction (LSC), auto white balance (AWB), demosaicing color interpolation, color correction matrix (CCM), dynamic range compression (DRC), gamma, 3D look up table (LUT), YUV domain noise reduction, sharpening, detail enhance, etc. YUV domain processing may include: multi-frame registration, fusion, noise reduction of high-dynamic range images (HDR), and super-resolution (SR) algorithms, beauty filtering algorithms, distortion correction algorithms, defocusing algorithms, etc. for enhancing clarity.

[0142] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel may adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone may include one or N display screens 294, where N is a positive integer greater than 1. For example, the display screen 294 may be used to display the application program interface.

[0143] The mobile phone realizes the display function through the GPU, the display screen 294, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change the display information.

[0144] It can be understood that Figure 2The structure shown does not specifically limit the mobile phone. In some embodiments, the mobile phone may also include more or fewer components than Figure 2 shown, or combine certain components, or split certain components, or have different component arrangements, etc. Alternatively, Figure 2 some of the components shown may be implemented in hardware, software, or a combination of software and hardware.

[0145] In addition, when the terminal device 200 is an interactive electronic whiteboard, a wearable device, a tablet computer, a laptop computer, a desktop computer, a portable electronic device, a UMPC, a netbook, a PDA, a smart TV, a vehicle-mounted computer, a smart speaker, an AR device, a VR device, and other intelligent devices with a display screen, or other forms of terminal devices such as a digital camera, a single-lens reflex camera / mirrorless camera, an action camera, a gimbal camera, a drone, etc., the specific structures of these other forms of terminal devices may also refer to Figure 2 shown. Exemplarily, other forms of terminal devices may be those with components added or reduced based on the structure given in Figure 2 which will not be elaborated here one by one.

[0146] It should also be understood that one or more application programs that can collect data required for 3D modeling, preprocess the data required for 3D modeling, and support functions such as previewing 3D models can run in the terminal device 200 (such as a mobile phone), such as: a 3D modeling application or a 3D reconstruction application can be called. When the terminal device 200 runs the aforementioned 3D modeling application or 3D reconstruction application, the application program can call the camera of the terminal device 200 to take pictures according to the user's operation, collect the data required for 3D modeling, and preprocess the data required for 3D modeling. In addition, the application program can also display a preview interface of the 3D model through the display screen of the terminal device 200 for the user to view and preview the 3D model.

[0147] Next, taking the terminal device 200 in the end-cloud collaboration system shown above Figure 1 as a mobile phone as an example, in combination with the scenario where the user uses the mobile phone for 3D modeling, an exemplary description of the 3D modeling method provided in the embodiments of the present application will be given.

[0148] It should be noted that although the embodiments of the present application are described by taking the terminal device 200 as a mobile phone as an example, it should be understood that the 3D modeling method provided in the embodiments of the present application is equally applicable to the above-mentioned other terminal devices with a shooting function, and the present application does not limit the specific type of the terminal device.

[0149] Taking the 3D modeling of a certain target object as an example, the 3D modeling method provided in the embodiments of the present application may include the following three parts:

[0150] In the first part, the user uses the mobile phone to collect the data required for 3D modeling of the target object.

[0151] In the second part, the mobile phone preprocesses the data collected for 3D modeling of the target object and uploads the preprocessed data to the cloud.

[0152] In the third part, the cloud performs 3D modeling based on the data uploaded by the mobile phone to obtain the 3D model of the target object.

[0153] Through the above first part to the third part, 3D modeling of the target object can be achieved. After obtaining the 3D model of the target object, the mobile phone can download the 3D model from the cloud for the user to preview.

[0154] The following specifically describes the above first part to the third part respectively.

[0155] For the first part:

[0156] In the embodiments of the present application, a first application can be installed in the mobile phone, and this first application is the 3D modeling application or 3D reconstruction application described in the foregoing embodiments. For example, the name of the first application can be "3D Magic Cube", and the name of the first application is not limited herein. When the user wants to perform 3D modeling on the target object, the first application can be started and run on the mobile phone. After the mobile phone starts and runs the first application, the main interface of the first application can include a function control for starting the 3D modeling function, and the user can click or touch this function control on the main interface of the first application. The mobile phone can respond to the operation of the user clicking or touching this function control and start the 3D modeling function of the first application. After the 3D modeling function of the first application is started, the mobile phone can switch the display interface from the main interface of the first application to the 3D modeling data collection interface and start the shooting function of the camera, and display the picture captured by the camera on the 3D modeling data collection interface. During the process of the mobile phone displaying the 3D modeling data collection interface, the user can hold the mobile phone to collect the data required for 3D modeling of the target object.

[0157] In some embodiments, the mobile phone can display the function control for starting the first application on the main interface (or called the desktop), such as: the application icon (or called the button) of the first application. When the user wants to use the first application to perform 3D modeling on a certain target object, the application icon of the first application can be clicked or touched. After the mobile phone receives the operation of the user clicking or touching the application icon of the first application, it can respond to the operation of the user clicking or touching the application icon of the first application, start and run the first application, and display the main interface of the first application.

[0158] For example, Figure 3 is a schematic diagram of the main interface of the mobile phone provided by the embodiments of the present application. As Figure 3As shown, the application icon 302 of the first application may be included in the main interface 301 of the mobile phone. The main interface 301 of the mobile phone may also include application icons of other applications such as Application A, Application B, and Application C. The user can click or touch the application icon 302 on the main interface 301 of the mobile phone to trigger the mobile phone to start running the first application and display the main interface of the first application.

[0159] In some other embodiments, the mobile phone can also display a function control for starting the first application on a pull-down interface or other display interfaces such as the negative first screen. On the pull-down interface or the negative first screen, the function control of the first application can be presented in the form of an application icon or in the form of other function buttons, which is not limited herein.

[0160] Among them, the pull-down interface refers to the display interface that appears after sliding the top of the main interface of the mobile phone downward. Function buttons for common functions used by the user, such as WLAN and Bluetooth, can be displayed on the pull-down interface, facilitating the user to quickly use relevant functions. For example, when the current display interface of the mobile phone is the desktop, the user can perform a downward sliding operation on the top of the mobile phone screen to trigger the mobile phone to switch the display interface from the desktop to the pull-down interface (or superimpose and display the pull-down interface on the desktop). The negative first screen refers to the display interface that appears after sliding the main interface (or called the desktop) of the mobile phone to the right. Frequently used applications, functions, and subscribed services and information of the user can be displayed on the negative first screen, facilitating the user to quickly browse and use. For example, when the current display interface of the mobile phone is the desktop, the user can perform a rightward sliding operation on the mobile phone screen to trigger the mobile phone to switch the display interface from the desktop to the negative first screen.

[0161] It can be understood that the "negative first screen" is just a term used in the embodiments of the present application, and its meaning has been recorded in the embodiments of the present application, but its name does not constitute any limitation to the embodiments of the present application; in addition, in some other embodiments, the "negative first screen" can also be called other names such as "desktop assistant", "quick menu", "Widget collection interface", etc., which is not limited herein.

[0162] In still some other embodiments, when the user wants to use the first application to perform 3D modeling on a certain target object, the mobile phone can also be controlled to start running the first application through the voice assistant. The present application does not limit the startup method of the first application herein.

[0163] Exemplarily, Figure 4 is a schematic diagram of the main interface of the first application provided by the embodiments of the present application. As Figure 4As shown, the main interface 401 of the first application may include a function control: "Start Modeling" 402, and "Start Modeling" 402 is the function control for starting the 3D modeling function as described above. The user can click or touch "Start Modeling" 402 on the main interface 401 of the first application. The mobile phone can respond to the operation of the user clicking or touching "Start Modeling" 402, start the 3D modeling function of the first application, switch the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface, and start the shooting function of the camera, and display the picture captured by the camera on the 3D modeling data acquisition interface.

[0164] Exemplarily, Figure 5 is a schematic diagram of the 3D modeling data acquisition interface provided by the embodiment of the present application. As Figure 5 shown, the 3D modeling data acquisition interface 501 displayed on the mobile phone may include a function control: a scan button 502, and the picture captured by the camera of the mobile phone. For example, please continue to refer to Figure 5 shown. Assume that the user wants to perform 3D modeling on a toy car placed on a table. Then the user can aim the mobile phone camera at the toy car. At this time, the picture captured by the camera displayed in the 3D modeling data acquisition interface 501 may include the toy car 503 and the table 504. The user can move the shooting angle of the mobile phone to adjust the position of the toy car 503 in the picture to the central position of the mobile phone screen (i.e., the 3D modeling data acquisition interface 501), and click or touch the scan button 502 in the 3D modeling data acquisition interface 501. The mobile phone can respond to the operation of the user clicking or touching the scan button 502 and start collecting the data required for 3D modeling of the target object (i.e., the toy car 503) located at the central position of the mobile phone screen.

[0165] That is, in the embodiment of the present application, when the mobile phone responds to the operation of the user clicking or touching the scan button 502 and starts collecting the data required for 3D modeling of the target object, it can use the object whose position in the picture is at the central position of the mobile phone screen as the target object.

[0166] In the embodiment of the present application, the 3D modeling data acquisition interface may be referred to as the first interface, and the main interface of the first application may be referred to as the second interface.

[0167] Optionally, Figure 6 is another schematic diagram of the 3D modeling data acquisition interface provided by the embodiment of the present application. As Figure 6 shown, after the mobile phone receives the operation of the user clicking or touching the scan button 502, it may also display a prompt message in the modeling data acquisition interface 501: "Please place the target object in the center of the screen" 505. Among them, the target object is also the target object, and this prompt message can be used to remind the user to place the position of the target object in the picture at the central position of the mobile phone screen.

[0168] Exemplarily, the display position of "Please place the target object in the center of the screen" 505 in the modeling data acquisition interface 501 can be above the scan button 502, or at a position slightly below the center of the screen, etc. There is no limitation on the display position of "Please place the target object in the center of the screen" 505 in the modeling data acquisition interface 501.

[0169] In addition, it should be noted that the prompt message: "Please place the target object in the center of the screen" 505 is only for exemplary illustration. In some other embodiments, other text labels can also be used to remind the user to place the target object in the center of the mobile phone screen in the picture. The present application also does not limit the content of this prompt message. In the embodiments of the present application, the prompt message for reminding the user to place the target object in the center of the mobile phone screen in the picture can be referred to as the first prompt message.

[0170] In the embodiments of the present application, when the mobile phone responds to the operation of the user clicking or touching the scan button 502 and starts to collect the data required for 3D modeling of the target object, the user can hold the mobile phone and take circumferential photos around the target object. During the circumferential process, the mobile phone can collect 360-degree panoramic data of the target object. Among them, the data collected by the mobile phone for 3D modeling of the target object can include: the pictures / images of the target object taken by the mobile phone during the circumferential shooting around the target object. The pictures can be in JPG / JPEG format. For example, the mobile phone can collect the RAW image corresponding to the target object through the camera, and then the processor of the mobile phone can perform ISP processing and JPEG encoding on the RAW image to obtain the picture in JPG / JPEG format corresponding to the target object.

[0171] Figure 7A This is another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application. As Figure 7A shown, in a possible design, when the mobile phone responds to the operation of the user clicking or touching the scan button 502 and starts to collect the data required for 3D modeling of the target object (taking a toy car as an example), a patch model 701 (or called an enclosing body or virtual enclosing body) can also be displayed around the target object in the picture of the 3D modeling data acquisition interface 501. The patch model 701 can use the center of the target object as the central axis and cover the periphery of the target object. The patch model 701 can include two layers, and each layer can include multiple patches. The layer located above can be called the first layer, and the layer located below can be called the second layer. Each patch in each layer can correspond to an angular range within 360 degrees around the target object. For example, assuming that the number of patches in each layer is 20, each patch in each layer corresponds to an angular range of 18 degrees.

[0172] When the user holds the mobile phone to take pictures of the target object, it is necessary to take two circles around the target object. The first circle is to take pictures of the target object with the mobile phone camera looking down at the target object (such as looking down at a 30-degree angle, without limitation), and the second circle is to take pictures of the target object with the mobile phone camera facing the target object. When the user takes the first circle of pictures around the target object with the mobile phone camera looking down at the target object, the mobile phone can sequentially light up the patches on the first layer as it moves around the target object. When the user takes the second circle of pictures around the target object with the mobile phone camera facing the target object, the mobile phone can sequentially light up the patches on the second layer as it moves around the target object.

[0173] For example, assume that the user looks down at the target object with the mobile phone camera and takes the first circle of pictures around the target object in a clockwise or counterclockwise direction. Taking the initial shooting angle as 0 degrees and the number of patches in each layer as 20 as an example, when the mobile phone takes a picture of the target object within the angle range of 0 degrees to 18 degrees, the mobile phone can light up the first patch on the first layer. When the mobile phone takes a picture of the target object within the angle range of 18 degrees to 36 degrees, the mobile phone can light up the second patch on the first layer. And so on, when the mobile phone takes a picture of the target object within the angle range of 342 degrees to 360 degrees, the mobile phone can light up the 20th patch on the first layer. That is, after the user takes the first circle of pictures around the target object with the mobile phone camera looking down at the target object in a clockwise or counterclockwise direction, the 20 patches on the first layer can be all lit up. Similarly, after the user takes the second circle of pictures around the target object with the mobile phone camera facing the target object in a clockwise or counterclockwise direction, the 20 patches on the second layer can be all lit up.

[0174] In the embodiments of the present application, the picture corresponding to the first patch can be called the first image, and the picture corresponding to the second patch can be called the second image. The posture of the mobile phone when taking the first image can be called the first pose, and the posture of the mobile phone when taking the second image can be called the second pose.

[0175] Exemplarily, please continue to refer to Figure 7A As shown, when the user looks down at the target object with the mobile phone camera and takes the first circle of pictures around the target object in a counterclockwise direction, the effect that the mobile phone can light up the first patch on the first layer can be as shown in Figure 7A 702 in. The lit patch can present a different pattern or color (i.e., change the display effect) compared to other unlit patches.

[0176] Figure 7B This is another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application. As shown in Figure 7B As shown, in the above Figure 7AOn the basis shown, when the user points the mobile phone camera downward at the target object and takes the first lap of shots around the target object in the counterclockwise direction, when the user holds the mobile phone and rotates it counterclockwise around the target object by a certain angle (moves a certain distance), the mobile phone can continue to light up the second patch, the third patch, etc. in the first layer.

[0177] Figure 7C Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiment of the present application. As Figure 7C shown, on the basis shown above Figure 7B when the user points the mobile phone camera downward at the target object and takes the first lap of shots around the target object in the counterclockwise direction, when the user holds the mobile phone and moves it counterclockwise around the target object to the other side of the target object, the number of patches in the first layer that the mobile phone can light up can reach half or more than half.

[0178] Figure 7D Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiment of the present application. As Figure 7D shown, on the basis shown above Figure 7C when the user points the mobile phone camera downward at the target object and takes the first lap of shots around the target object in the counterclockwise direction, when the user holds the mobile phone and moves it counterclockwise around the target object to Figure 7A the initial shooting position described in

[0179] After the mobile phone lights up all the patches in the first layer, the user can adjust the shooting position of the mobile phone relative to the target object, lower the mobile phone by a certain distance, make the mobile phone camera face the target object directly, and take the second lap of shots around the target object in the counterclockwise direction.

[0180] For example, Figure 7E Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiment of the present application. As Figure 7E shown, when the user points the mobile phone camera directly at the target object and takes the second lap of shots around the target object in the counterclockwise direction, the mobile phone can light up the first patch in the second layer.

[0181] Similarly, when the user points the mobile phone camera directly at the target object and takes the second lap of shots around the target object in the counterclockwise direction, as the user holds the mobile phone and moves it counterclockwise around the target object, the mobile phone can gradually light up the patches in the second layer. For example, Figure 7F Another schematic diagram of the 3D modeling data acquisition interface provided by the embodiment of the present application. As Figure 7FAs shown, when the user faces the camera of the mobile phone towards the target object and takes the second circle of pictures around the target object in the counterclockwise direction, when the user holds the mobile phone and moves counterclockwise around the target object to the Figure 7E initial shooting position described in (or when the user holds the mobile phone and rotates counterclockwise around the target object for one circle), the mobile phone can light up all the patches in the second layer. The process of lighting up the first patch in the second layer to all the patches being lit is similar to the process of lighting up the patches in the first layer, and reference can be made to the Figures 7A to 7D process shown above, and no further details will be elaborated.

[0182] Optionally, in the embodiments of the present application, the rule for the mobile phone to light up each patch can be as follows:

[0183] 1) When the mobile phone takes pictures of the target object at a certain position, perform blur detection and key frame selection on each frame of the captured picture (which can be called the input frame) to obtain pictures with clarity and picture features meeting the requirements.

[0184] For example, in the embodiments of the present application, when the mobile phone takes pictures of the target object, it can collect a preview stream corresponding to the target object, and the preview stream includes multiple frames of pictures. For example: the mobile phone can take pictures of the target object at a frame rate of 24 frames per second, 30 frames per second, etc., and there is no limit here. The mobile phone can perform blur detection on each frame of the captured picture to obtain pictures with clarity greater than the first threshold. The first threshold can be determined according to requirements and the blur detection algorithm, and there is no limit on the size. If the clarity of the current frame picture does not meet the requirements (such as being less than or equal to the first threshold), then continue to obtain the next frame of picture. Then, the mobile phone can perform key frame selection (or called key frame screening) on the pictures with clarity greater than the first threshold to obtain pictures with picture features meeting the requirements. The picture features meeting the requirements can include: the features included in the picture are relatively clear and rich, the features included in the picture are easy to extract, the redundant information of the features included in the picture is less, etc. There is no limit on the algorithm and specific requirements for key frame selection here.

[0185] For each shooting position (one shooting position can correspond to one patch), the mobile phone can obtain some key frame pictures with better quality by performing blur detection and key frame selection on the pictures taken at this shooting position. The number of key frame pictures corresponding to each patch can be one or more.

[0186] 2) The mobile phone calculates the camera pose information corresponding to the pictures obtained in 1) (that is, the pose information of the mobile phone camera).

[0187] For example, when the mobile phone supports capabilities such as AR engine ability, or AR core ability, or AR KIT ability, etc., the mobile phone can directly obtain the camera pose information corresponding to the pictures by invoking the foregoing capabilities.

[0188] Exemplarily, the camera pose information may include qw, qx, qy, qz, tx, ty, and tz. Among them, qw, qx, qy, and qz represent the rotation matrix composed of unit quaternions, and tx, ty, and tz can form a translation matrix. This rotation matrix and translation matrix can represent the relative position relationship and angle between the camera (mobile phone camera) and the target object. The mobile phone can convert the coordinates of the target object from the world coordinate system to the camera coordinate system through the aforementioned rotation matrix and translation matrix, and obtain the coordinates of the target object in the camera coordinate system. Among them, the world coordinate system may refer to the coordinate system with the center of the target object as the origin, and the camera coordinate system may refer to the coordinate system with the camera center as the origin.

[0189] 3) The mobile phone determines the relationship between the picture obtained in 2) and each patch in the patch model according to the camera pose information corresponding to the picture, and obtains the patch corresponding to the picture.

[0190] For example, as shown in 2) above, the mobile phone can convert the coordinates of the target object from the world coordinate system to the camera coordinate system according to the camera pose information (rotation matrix and translation matrix) corresponding to the picture, and obtain the coordinates of the target object in the camera coordinate system. Then, the mobile phone can determine the connection line between the camera coordinate and the coordinates of the target object according to the coordinates of the target object in the camera coordinate system and the camera coordinate. The patch that intersects with the connection line between the camera coordinate and the coordinates of the target object in the patch model is the patch corresponding to this frame of picture. Among them, the camera coordinate is a known parameter for the mobile phone.

[0191] 4) Store the picture in the frame sequence file and light up the patch corresponding to the picture. The frame sequence file includes the pictures corresponding to each lit patch, and these pictures can be used as the data required for 3D modeling of the target object.

[0192] The format of the pictures included in the frame sequence file can be JPG format. For example, the pictures saved in the frame sequence file can be numbered 001.jpg, 002.jpg, 003.jpg... etc. in sequence.

[0193] It can be understood that after a certain patch is lit, the user can continue to move the position of the mobile phone to take pictures at the next angle, and the mobile phone can continue to light up the next patch according to the above rules.

[0194] Each frame of picture included in the above frame sequence file can be called a key frame, and these key frames can be used as the data required for 3D modeling of the target object collected by the mobile phone in the first part. The frame sequence file can also be called a key frame sequence file.

[0195] Optionally, in an embodiment of the present application, a user can hold a mobile phone and take a photo of the target object within 1.5 meters of the target object. When the shooting distance is too close (such as when the 3D modeling data acquisition interface cannot present the full picture of the target object), the mobile phone can turn on the wide-angle camera to take pictures.

[0196] In the first part above, when the user holds the mobile phone and takes surround shots around the target object, the mobile phone lights up the facets in the facet model one by one as it moves around the target object, thereby guiding the user to collect the data required for 3D modeling of the target object. The dynamic UI guidance using the 3D guidance interface (i.e. the 3D modeling data collection interface that displays the facet model) enhances user interactivity and allows the user to intuitively perceive the data collection process.

[0197] It should be noted that the description of the patch model (or bounding volume) in the first part above is only for exemplary purposes. In some other embodiments, the number of layers of the patch model may include more or fewer layers, and the number of patches in each layer may be greater than 20 or less than 20. The present application does not limit the number of layers of the patch model and the number of patches in each layer.

[0198] For example, Figure 8 This is a schematic diagram of the structure of the facet model provided in the embodiment of the present application. Please refer to Figure 8 As shown, in one implementation, the structure of the patch model can be as follows Figure 8 As shown in (a) of FIG. 1 , the structure of the patch model may include two layers, each of which may include multiple facets (i.e., the structure described in the above embodiment). Figure 8 As shown in (b) of FIG. 1 , the structure of the patch model may include three layers: upper, middle and lower. Each layer may include multiple facets. In another implementation, the structure of the facet model may be as follows: Figure 8 As shown in (c) of FIG. 1 , it is a layer structure composed of multiple facets. In another implementation, the structure of the facet model can also be as follows: Figure 8 As shown in (d) in the figure, it includes two layers, an upper layer and a lower layer, and each layer may include multiple face patches.

[0199] Figure 8 The structures of the face model shown in the figure are all exemplary. The present application does not limit the structure of the face model and the inclination angle of each layer in the face model (the inclination angle relative to the central axis).

[0200] It should be understood that the patch model described in the embodiment of the present application is a virtual model, and the patch model can be preset in the mobile phone, such as being configured in the file directory of the first application in the form of a configuration file.

[0201] In some embodiments, multiple patch models may be pre - installed in the mobile phone. When the mobile phone collects data required for 3D modeling of the target object, it may recommend a target patch model that matches the target object according to the shape of the target object, or select a target patch model from multiple patch models according to the user's selection operation, and use the target patch model to implement the guiding function described in the foregoing embodiments. The target patch model may also be referred to as the first virtual enclosure.

[0202] Optionally, Figure 9 is another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application. As Figure 9 shown, after the mobile phone switches the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface 501, when it detects that there is a target object (target object) in the screen, it may also display a prompt message in the modeling data acquisition interface 501: "Target object detected, click the button to start scanning" 506. Among them, the button is also the scanning button 502, and this prompt message can be used to remind the user to click the scanning button 502 so that the mobile phone starts to collect data required for 3D modeling of the target object.

[0203] Optionally, in the embodiments of the present application, the mobile phone may respond to the operation of the user clicking or touching "Start Modeling" 402, start the 3D modeling function of the first application, and after switching the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface, it may first display in the 3D modeling data acquisition interface prompt information related to reminding the user to adjust the shooting environment where the target object is located, the way of shooting the target object, and the object screen occupation ratio, etc.

[0204] For example, Figure 10 is another schematic diagram of the 3D modeling data acquisition interface provided by the embodiments of the present application. As Figure 10 shown, the mobile phone may respond to the operation of the user clicking or touching "Start Modeling" 402, start the 3D modeling function of the first application, and after switching the display interface from the main interface 401 of the first application to the 3D modeling data acquisition interface, it may first display prompt information 1001 in the 3D modeling data acquisition interface. The content of the prompt information 1001 may be "Place the object statically on a pure - color plane, with soft lighting, take pictures around the object for one week, and the object screen occupation ratio should be as large and complete as possible", which can be used to remind the user to adjust the shooting environment where the target object is located to be that the object is statically placed on a pure - color plane and the lighting is soft; the way of shooting the target object is to take pictures around the object for one week; at the same time, the object screen occupation ratio should be as large and complete as possible. Continuing to refer to Figure 10As shown, the mobile phone can first display the function control 1002 on the 3D modeling data collection interface. For example, the function space can be "Got it", "Confirm", etc. After the user clicks the function control 1002, the mobile phone no longer displays the prompt message 1001 and the function control 1002, and presents the 3D modeling data collection interface as described above Figure 5 shown.

[0205] In this embodiment, after the user adjusts the shooting environment where the target object is located, the object screen occupation ratio, etc. according to the prompt content of the prompt message 1001, the subsequent data collection process can be faster and the quality of the collected data can be better. The prompt message 1001 can also be referred to as the second prompt message.

[0206] In some embodiments, after the mobile phone switches the display interface from the main interface 401 of the first application to the 3D modeling data collection interface, it can also display the prompt message 1001 on the 3D modeling data collection interface for a preset duration. When the preset duration is reached, the mobile phone can automatically stop displaying the prompt message 1001 and present the 3D modeling data collection interface as described above Figure 5 shown. Among them, the size of the preset duration can be 20 seconds, 30 seconds, etc., which is not limited here.

[0207] In some embodiments, after the mobile phone collects the data (each frame picture included in the frame sequence file) required for 3D modeling of the target object in the manner described in the first part above, it can automatically execute the second part.

[0208] In some other embodiments, after the mobile phone collects the data (each frame picture included in the frame sequence file) required for 3D modeling of the target object in the manner described in the first part above, it can display a function control for uploading to the cloud for 3D modeling on the 3D modeling data collection interface. The user can click the function control for uploading to the cloud for 3D modeling. The mobile phone can execute the second part in response to the operation of the user clicking the function control for uploading to the cloud for 3D modeling.

[0209] For example, Figure 11 is another schematic diagram of the 3D modeling data collection interface provided by the embodiment of the present application. As Figure 11As shown, after the mobile phone collects the data required for 3D modeling of the target object (each frame of picture included in the frame sequence file) in the manner described in the above first part, function controls can be displayed on the 3D modeling data collection interface: "Upload to Cloud for Modeling" 1101. "Upload to Cloud for Modeling" 1101 is a function control for uploading to the cloud for 3D modeling. The user can click "Upload to Cloud for Modeling" 1101, and the mobile phone can execute the second part in response to the user's click on "Upload to Cloud for Modeling" 1101. In the embodiments of the present application, the operation of the user clicking "Upload to Cloud for Modeling" 1101 can be referred to as an operation for generating a 3D model.

[0210] Optionally, please continue to refer to Figure 11 As shown, after the mobile phone collects the data required for 3D modeling of the target object (each frame of picture included in the frame sequence file) in the manner described in the above first part, a prompt message can also be displayed on the 3D modeling data collection interface to prompt the user that the mobile phone has collected the data required for 3D modeling of the target object, such as: "Scanning Completed" 1102.

[0211] Optionally, please refer to the foregoing Figures 5 to 7E , and from 9 to Figure 11 As shown, the mobile phone can also display an exit button 1103 in the 3D modeling data collection interface (only marked in Figure 11 ). During the execution of the first part by the mobile phone, the user can click the exit button 1103 at any time, and the mobile phone can respond to the user's click on the exit button 1103 and exit the execution process of the first part. After exiting the execution process of the first part, the mobile phone can switch the display interface from the 3D modeling data collection interface to the main interface of the first application as shown in Figure 4 .

[0212] For the second part:

[0213] As can be seen from the above first part, the data required for 3D modeling of the target object collected by the mobile phone in the first part is each frame of picture (i.e., key frame) included in the frame sequence file mentioned in the first part. In the second part, the preprocessing of the data required for 3D modeling of the target object collected by the mobile phone refers to: the mobile phone preprocesses the key frames included in the frame sequence file collected in the first part, specifically as follows:

[0214] 1) The mobile phone calculates the matching information of each frame of key frame and saves the matching information of each frame of key frame in a first file. For example, the first file can be a file in the format of JavaScript Object Notation (JSON).

[0215] Among them, for each key frame, the matching information of this key frame may include: the identification information of other key frames associated with this key frame. For example, the matching information of a certain key frame may include the identification information of the key frames corresponding to the upper, lower, left, and right directions of this key frame (such as the nearest key frames), and the key frames corresponding to the upper, lower, left, and right directions of this key frame are the other key frames associated with this key frame. Among them, the identification information of the key frame may be the picture number of the key frame. The matching information of this key frame can be used to indicate which pictures in the frame sequence file are the other frames associated with this key frame.

[0216] Exemplarily, the matching information of each key frame is obtained according to the association relationship between each key frame and the patch corresponding to each key frame, and the association relationship between the multiple patches. That is, for each key frame, the mobile phone can determine the matching information of this key frame according to the association relationship between the patch corresponding to this key frame and other patches. Among them, the association relationship between the patch corresponding to this key frame and other patches may include: in the patch model, which patches correspond to the upper, lower, left, and right directions of the patch corresponding to this key frame.

[0217] For example, taking the patch model shown above Figure 7A as an example, assume that the patch corresponding to a certain key frame is patch 1 in the first layer. Then the association relationship between patch 1 and other patches may include: the patch below patch 1 is patch 21, the patch to the left of patch 1 is patch 20, and the patch to the right of patch 1 is patch 2. Based on the aforementioned association relationship between patch 1 and other patches, the mobile phone can determine that the other key frames associated with this key frame include the key frame corresponding to patch 21, the key frame corresponding to patch 20, and the key frame corresponding to patch 2. Thus, the mobile phone can obtain that the matching information of this key frame includes the identification information of the key frame corresponding to patch 21, the identification information of the key frame corresponding to patch 20, and the identification information of the key frame corresponding to patch 2.

[0218] It can be understood that for the mobile phone, the relationship between different patches in the patch model is a known quantity.

[0219] Optionally, the first file further includes: the camera intrinsics, gravity direction information, picture name, picture number, camera pose information, timestamp, etc. corresponding to each key frame.

[0220] Exemplarily, the first file includes three parts: "intrinsics", "keyframes", and "matching_list". Among them, the "intrinsics" part is the camera intrinsics; "keyframes" is information such as the gravity direction information, image name, image number, camera pose information, and timestamp corresponding to each frame keyframe; "matching_list" is the matching information of each frame keyframe.

[0221] For example, the content of the first file can be as follows:

[0222]

[0223]

[0224] In the content of the first file given above by way of example, cx, cy, fx, fy, height, k1, k2, k3, p1, p2, and width are all camera intrinsics. Among them, cx and cy represent the offsets of the optical axis from the coordinate center of the projection plane; fx and fy represent the focal lengths in the x - direction and y - direction respectively when the camera takes pictures; k1, k2, and k3 represent the radial distortion coefficients; p1 and p2 represent the tangential distortion coefficients; height and width represent the resolution when the camera takes pictures.

[0225] x, y, and z represent the gravity direction information. The gravity direction information can be obtained by the mobile phone according to the built - in gyroscope and can represent the offset angle when the mobile phone takes pictures.

[0226] 18.jpg represents the image name, and 18 is the image number (only taking 18.jpg as an example here). That is, the above example is for the camera intrinsics (intrinsics), gravity direction information (gravity), image name, image number (index), camera pose information (slam pose), timestamp, and matching information corresponding to 18.jpg.

[0227] qw, qx, qy, qz, tx, ty, and tz are all camera pose information. qw, qx, qy, and qz represent the rotation matrix composed of unit quaternions, and tx, ty, and tz can form a translation matrix. This rotation matrix and translation matrix can represent the relative position relationship and angle between the camera (mobile phone camera) and the target object. The mobile phone can convert the coordinates of the target object from the world coordinate system to the camera coordinate system through the aforementioned rotation matrix and translation matrix, and obtain the coordinates of the target object in the camera coordinate system. Among them, the world coordinate system can refer to the coordinate system with the center of the target object as the origin, and the camera coordinate system can refer to the coordinate system with the camera center as the origin.

[0228] timestamp represents the timestamp, which means the time when the camera captures this key frame.

[0229] src_id represents the picture number of each key frame. For example, in the content of the first file given above as an example, the picture number is 18, and the "matching_list" part is the matching information of the key frame with the picture number 18. tgt_id represents the picture numbers of other key frames associated with the key frame with the picture number 18 (that is, the identification information of other key frames associated with the key frame with the picture number 18). For example, in the content of the first file given above as an example, the picture numbers of other key frames associated with the key frame with the picture number 18 include: 26, 45, 59, 78, 89, 100, 449, etc. That is, the other key frames associated with the key frame with the picture number 18 include: the key frame with the picture number 26, the key frame with the picture number 45, the key frame with the picture number 59, the key frame with the picture number 78, the key frame with the picture number 89, the key frame with the picture number 100, the key frame with the picture number 449, etc.

[0230] It should be noted that the above only takes the key frame with the picture number 18 as an example to give part of the content of the first file, and is not used to limit the content of the first file.

[0231] 2) The mobile phone packs the first file with all the key frames in the frame sequence file (that is, all the frame pictures included in the above frame sequence file).

[0232] The result obtained after packing the first file with all the key frames in the frame sequence file (such as called the packed file or data packet) is the preprocessed data obtained by the mobile phone in the second part for preprocessing the data required for 3D modeling of the target object collected in the first part.

[0233] That is, in the embodiments of the present application, the data obtained after preprocessing the data required for 3D modeling of the target object collected in the first part by the mobile phone may include: each key frame picture saved in the frame sequence file during the process of the mobile phone taking circumferential pictures around the target object, and the first file including the matching information of each frame of key frame.

[0234] After obtaining the above preprocessed data, the mobile phone can send (i.e., upload) the preprocessed data to the cloud, and the cloud can execute the third part to perform 3D modeling based on the data uploaded by the mobile phone to obtain a 3D model of the target object.

[0235] For the third part:

[0236] The process of the cloud performing 3D modeling based on the data uploaded by the mobile phone can be as follows:

[0237] 1) The cloud decompresses the data packet received from the mobile phone (the data packet includes the frame sequence file and the first file), and extracts the frame sequence file and the above first file.

[0238] 2) The cloud performs 3D modeling processing based on the key frame pictures included in the frame sequence file and the above first file to obtain a 3D model of the target object.

[0239] Exemplarily, the steps of the cloud performing 3D modeling processing based on the key frame pictures included in the frame sequence file and the above first file may at least include: key target extraction, feature detection and matching, global optimization and fusion, sparse point cloud computing, dense point cloud computing, surface reconstruction, texture generation.

[0240] Among them, key target extraction refers to the operation of separating the target object of interest in the key frame picture from the background and identifying and interpreting meaningful object entities in the image to extract different image features.

[0241] Feature detection and matching means: detecting unique pixel points existing in the key frame picture as the feature points of the key frame picture; describing the feature points with significant features in different key frame pictures, and comparing the similarity of the two descriptions to determine whether the feature points in different key frame pictures are the same feature.

[0242] In the embodiments of the present application, when the cloud performs feature detection and matching, for each key frame, the cloud can determine the other key frames associated with the key frame according to the matching information of the key frame included in the first file (i.e., the identification information of the other key frames associated with the key frame), and perform feature detection and matching on the key frame and the other key frames associated with the key frame.

[0243] For example, taking the content of the first file exemplarily given with the key frame numbered 18 as an example, the cloud can determine other key frames associated with the key frame numbered 18 from the first file, including: the key frame numbered 26, the key frame numbered 45, the key frame numbered 59, the key frame numbered 78, the key frame numbered 89, the key frame numbered 100, the key frame numbered 449, etc. Then, the cloud can perform feature detection and matching on the key frame numbered 18 with the key frame numbered 26, the key frame numbered 45, the key frame numbered 59, the key frame numbered 78, the key frame numbered 89, the key frame numbered 100, the key frame numbered 449, etc., without performing feature detection and matching on the key frame numbered 18 with all other key frames in the frame sequence file.

[0244] It can be seen that in the embodiment of the present application, for each key frame, the cloud can combine the matching information of the key frame included in the first file, and perform feature detection and matching on the key frame and other key frames associated with the key frame, without the need to perform feature detection and matching on the key frame with all other key frames in the frame sequence file. In this way, the computing load of the cloud can be effectively reduced, and the efficiency of 3D modeling can be improved.

[0245] Global optimization and fusion refer to using the global optimization and fusion algorithm to optimize and fuse the matching results of feature detection and matching. The results of global optimization and fusion can be used to generate a basic 3D model.

[0246] Sparse point cloud computing and dense point cloud computing refer to: generating three-dimensional point cloud data corresponding to the target object according to the results of global optimization and fusion. Compared with images, point clouds have an irreplaceable advantage - depth. Three-dimensional point cloud data directly provides data in three-dimensional space, while images need to infer three-dimensional data through perspective geometry.

[0247] Surface reconstruction refers to: accurately restoring the three-dimensional surface shape of the object using the three-dimensional point cloud data to obtain the basic 3D model of the target object.

[0248] Texture generation refers to: generating the texture (also called texture mapping) on the surface of the target object according to the key frame picture or the features of the key frame picture. After obtaining the texture on the surface of the target object, mapping the texture to the surface of the basic 3D model of the target object in a specific manner can more realistically restore the surface of the target object and make the target object look more real.

[0249] In the embodiments of the present application, the cloud can also quickly and accurately determine the mapping relationship between the texture and the surface of the basic 3D model of the target object according to the matching information of each key frame included in the first file, which can further improve the modeling efficiency and effect.

[0250] For example, after the cloud determines the mapping relationship between the texture of the first key frame and the surface of the basic 3D model of the target object, it can quickly and accurately determine the mapping relationship between the textures of other key frames associated with the first key frame and the surface of the basic 3D model of the target object in combination with the matching information of the first key frame. Similarly, for each subsequent key frame, the cloud can quickly and accurately determine the mapping relationship between the textures of other key frames associated with the key frame and the surface of the basic 3D model of the target object in combination with the matching information of the key frame.

[0251] After the cloud obtains the basic 3D model of the target object and the texture on the surface of the target object, it can generate the 3D model of the target object according to the basic 3D model of the target object and the texture on the surface of the target object. The cloud can save the basic 3D model of the target object and the texture on the surface of the target object for the mobile phone to download.

[0252] It can be seen that in the process of 3D modeling by the cloud according to the data uploaded by the mobile phone, the matching information of each key frame included in the first file can effectively improve the processing speed of 3D modeling, reduce the computing load of the cloud, and improve the efficiency of 3D modeling by the cloud.

[0253] Exemplarily, the basic 3D model of the target object can be stored in the OBJ format, and the texture on the surface of the target object can be stored in the JPG format (such as a texture map). For example: the basic 3D model of the target object can be an OBJ file, and the texture on the surface of the target object can be a JPG file.

[0254] Optionally, the cloud can save the basic 3D model of the target object and the texture on the surface of the target object for a certain period of time (such as 7 days), and after reaching this period, the cloud can automatically delete the basic 3D model of the target object and the texture on the surface of the target object. Or, the cloud can also permanently retain the basic 3D model of the target object and the texture on the surface of the target object, which is not limited here.

[0255] Through the above first part to the third part, 3D modeling of the target object can be achieved, and a 3D model of the target object can be obtained. As can be seen from the above first part to the third part, in the 3D modeling method provided by the embodiments of the present application, the mobile phone can achieve 3D modeling only by relying on an ordinary RGB camera (camera) to collect the data required for 3D modeling. In the process of collecting the data required for 3D modeling, it does not need to rely on the mobile phone having special hardware such as a LIDAR sensor or an RGB-D camera. The process of performing 3D modeling is completed in the cloud, and it does not need to rely on the mobile phone being equipped with a high-performance independent graphics card. This 3D modeling method can significantly lower the threshold of 3D modeling and has higher universality for terminal devices. Moreover, during the process of the user using the mobile phone to collect the data required for 3D modeling, the mobile phone can enhance the user interaction through the dynamic UI guidance, enabling the user to intuitively perceive the data collection process.

[0256] In addition, in this 3D modeling method, when the mobile phone takes pictures of the target object at a certain position, it performs blurring detection on each captured picture, obtains pictures with clarity meeting the requirements, can achieve the screening of key frames, and obtains key frames beneficial to modeling. The mobile phone extracts the matching information of each frame of key frames and sends the first file including the matching information of each frame of key frames and the frame sequence file composed of key frames to the cloud for the cloud to perform modeling (without sending all the captured pictures), which can greatly reduce the complexity of 3D modeling on the cloud side, reduce the consumption of the cloud's hardware resources during the 3D modeling process, effectively reduce the computing load of the cloud for modeling, and improve the speed and effect of 3D modeling.

[0257] The following is an exemplary description of the process of the mobile phone downloading the 3D model from the cloud for the user to preview the 3D model.

[0258] Exemplarily, after the cloud completes the 3D modeling of the target object and obtains the 3D model of the target object, it can send an indication message to the mobile phone to indicate that the cloud has completed 3D modeling.

[0259] In some embodiments, after receiving the above indication message from the cloud, the mobile phone can automatically download the 3D model of the target object from the cloud for the user to preview the 3D model.

[0260] For example, Figure 12 is a schematic diagram of the 3D model preview interface provided by the embodiments of the present application. In the above second part, after the mobile phone sends the preprocessed data to the cloud, it can switch the display interface from Figure 11 the 3D modeling data collection interface shown to Figure 12 the 3D model preview interface shown in Figure 12As shown, the mobile phone can display a prompt message: "Modeling in progress" 1201 in the 3D model preview interface to prompt the user that 3D modeling of the target object is in progress. "Modeling in progress" 1201 can be referred to as the third prompt message.

[0261] In some embodiments, after detecting the user's operation of clicking on the above-mentioned function control "Upload to Cloud for Modeling" 1101, the mobile phone can display the third prompt message.

[0262] After the cloud completes the 3D modeling of the target object and obtains the 3D model of the target object, it can send an indication message to the mobile phone to indicate that the 3D modeling has been completed. After receiving the indication message, the mobile phone can automatically download the 3D model of the target object from the cloud. For example: the mobile phone can send a download request message to the cloud, and the cloud can send the 3D model of the target object (i.e., the basic 3D model of the target object and the texture on the surface of the target object) to the mobile phone according to the download request message.

[0263] Figure 13 This is another schematic diagram of the 3D model preview interface provided by the embodiments of the present application. As Figure 13 shown, after receiving the indication message, the mobile phone can also change the prompt message from "Modeling in progress" 1201 to "Modeling completed" 1301 to prompt the user that the 3D model of the target object has been completed. The 3D model preview interface can also include a view button 1302. The user can click on the view button 1302, and the mobile phone can, in response to the user's operation of clicking on the view button 1302, display the 3D model of the target object downloaded from the cloud in the 3D model preview interface. "Modeling completed" 1301 can be referred to as the fourth prompt message.

[0264] Taking the target object as the toy car described in the foregoing embodiments as an example, Figure 14 This is yet another schematic diagram of the 3D model preview interface provided by the embodiments of the present application. As Figure 14 shown, the mobile phone can, in response to the user's operation of clicking on the view button 1302, display the 3D model 1401 of the toy car in the 3D model preview interface. The user can view the 3D model 1401 of the toy car in the Figure 14 shown 3D model preview interface. The user's operation of clicking on the view button 1302 is an operation to preview the three-dimensional model corresponding to the target object.

[0265] Optionally, when the user views the 3D model of the toy car in the 3D model preview interface, the user can perform a counterclockwise rotation operation or a clockwise rotation operation on the 3D model of the toy car in any direction (such as the horizontal direction, the vertical direction, etc.). The mobile phone can respond to the foregoing operation of the user and display for the user the rendering effects of the 3D model of the toy car at different angles (360 degrees) in the 3D model preview interface.

[0266] Exemplarily, Figure 15 is a schematic diagram of the user performing a counterclockwise rotation operation on the 3D model of the toy car along the horizontal direction provided by the embodiment of the present application. As Figure 15 shown, the user can use a finger to drag the 3D model of the toy car in the 3D model preview interface to perform a counterclockwise rotation along the horizontal direction.

[0267] Figure 16 is another schematic diagram of the 3D model preview interface provided by the embodiment of the present application. As Figure 16 shown, when the user can use a finger to drag the 3D model of the toy car in the 3D model preview interface to perform a counterclockwise rotation along the horizontal direction, the mobile phone can respond to the operation of the user dragging the 3D model of the toy car to perform a counterclockwise rotation along the horizontal direction, and display for the user in the 3D model preview interface the rendering effects at angles such as Figure 16 (a), (b), etc. shown in

[0268] It can be understood that Figure 16 the rendering effects at angles such as (a), (b), etc. shown in

[0269] are only exemplary illustrations. The angles at which the 3D model of the toy car is rendered are related to the direction, distance, number of times, etc. of the user dragging the 3D model of the toy car, and will not be represented one by one here.

[0270] Exemplarily, Figure 17 is a schematic diagram of the user performing a reduction operation on the 3D model of the toy car provided by the embodiment of the present application. As Figure 17 shown, the user can use two fingers to slide inward (opposite directions) in the 3D model preview interface, and this sliding operation is the reduction operation.

[0271] Figure 18 is another schematic diagram of the 3D model preview interface provided by the embodiment of the present application. As Figure 18As shown, when the user performs a reduction operation on the 3D model of the toy car, the mobile phone can respond to the reduction operation performed by the user on the 3D model of the toy car and display the rendering effect after the reduction of the 3D model of the toy car for the user on the 3D model preview interface.

[0272] Similarly, the enlargement operation performed by the user on the 3D model of the toy car can be achieved by sliding two fingers outward (in opposite directions) on the 3D model preview interface. When the user performs an enlargement operation on the 3D model of the toy car, the mobile phone can respond to the enlargement operation performed by the user on the 3D model of the toy car and display the rendering effect after the enlargement of the 3D model of the toy car for the user on the 3D model preview interface, which will not be elaborated in detail.

[0273] It should be noted that the above enlargement operation or reduction operation performed by the user on the 3D model of the toy car are all exemplary illustrations. In some other implementation manners, the enlargement operation or reduction operation performed by the user on the 3D model of the toy car can also be a double-click operation, a long-press operation, or alternatively, the 3D model preview interface can also include a function control for performing an enlargement operation or a reduction operation, etc., which is not limited herein.

[0274] In some other embodiments, after receiving the above instruction message from the cloud, the mobile phone can also only display the 3D model preview interface as shown above. Figure 13 When the user clicks the view button 1302, the mobile phone then responds to the operation of the user clicking the view button 1302, downloads the 3D model of the target object from the cloud, and displays the 3D model of the target object in the 3D model preview interface for the user to preview. The present application also does not limit the trigger condition for the mobile phone to download the 3D model of the target object from the cloud.

[0275] As described above, in the embodiments of the present application, the user only needs to perform operations related to collecting data required for 3D modeling on the terminal device side, and then view or preview the final 3D model on the terminal device. For the user, all operations are completed on the terminal device side, the operations are simpler, and the user experience can be better.

[0276] To make the technical solutions provided in the embodiments of the present application more concise and clear, the implementation logics of the 3D modeling methods provided in the embodiments of the present application are respectively exemplified below in combination with Figure 19 and Figure 20 to exemplarily illustrate the implementation logics of the 3D modeling methods provided in the embodiments of the present application.

[0277] Exemplarily, Figure 19 is a flow diagram of the 3D modeling method provided in the embodiments of the present application. As Figure 19 shown, the 3D modeling method can include S1901 - S1913.

[0278] S1901. The mobile phone receives a first operation, which is an operation to start a first application.

[0279] For example, the first operation can be the above-mentioned click or touch Figure 3 operation on the application icon 302 of the first application in the main interface of the mobile phone shown. Or, the first operation can be an operation of clicking or touching a function control for starting the first application in other display interfaces such as the drop-down interface or the negative first screen. Or, the first operation can also be the above-mentioned operation of controlling the mobile phone to start and run the first application through a voice assistant.

[0280] S1902. The mobile phone responds to the first operation, starts and runs the first application, and displays the main interface of the first application.

[0281] The main interface of the first application can be referred to the above Figure 4 shown. The main interface of the first application can be called the second interface.

[0282] S1903. The mobile phone receives a second operation, which is an operation to start the 3D modeling function of the first application.

[0283] For example, the second operation can be the above-mentioned click or touch Figure 4 operation on the function control "Start Modeling" 402 in the main interface 401 of the first application shown.

[0284] S1904. The mobile phone responds to the second operation, displays a 3D modeling data acquisition interface, and starts the shooting function of the camera, and displays the picture captured by the camera on the 3D modeling data acquisition interface.

[0285] The 3D modeling data acquisition interface displayed by the mobile phone in response to the second operation can be referred to the above Figure 5 shown. The 3D modeling data acquisition interface can be called the first interface.

[0286] S1905. The mobile phone receives a third operation, which is an operation to control the mobile phone to collect data required for 3D modeling of a target object.

[0287] For example, the third operation can include the operation of the user clicking or touching the scan button 502 in the 3D modeling data acquisition interface shown above Figure 5 and the operation of the user holding the mobile phone and making a circumferential shooting around the target object. The third operation can also be called the acquisition operation.

[0288] S1906. The mobile phone responds to the third operation and obtains a frame sequence file composed of key frame pictures corresponding to the target object.

[0289] S1907. The mobile phone obtains the matching information of each frame key frame in the frame sequence file and gets a first file.

[0290] The first file can be referred to as described in the foregoing embodiments. For each key frame, the matching information of the key frame included in the first file may include: the identification information of the nearest key frames corresponding to the four directions of up, down, left, and right of the key frame. For example, the identification information may be the number of the foregoing picture.

[0291] S1908. The mobile phone sends the frame sequence file and the first file to the cloud.

[0292] Correspondingly, the cloud receives the frame sequence file and the first file.

[0293] S1909. The cloud performs 3D modeling based on the frame sequence file and the first file to obtain a 3D model of the target object.

[0294] The specific process of the cloud performing 3D modeling based on the frame sequence file and the first file can be referred to as described in the third part of the foregoing embodiments and will not be elaborated here. The 3D model of the target object may include the basic 3D model of the target object and the texture on the surface of the target object. Attaching the texture (texture map) of the surface of the target object to the basic 3D model of the target object is the 3D model of the target object.

[0295] S1910. The cloud sends an indication message to the mobile phone to indicate that the cloud has completed 3D modeling.

[0296] Correspondingly, the mobile phone receives the indication message.

[0297] S1911. The mobile phone sends a download request message to the cloud to request to download the 3D model of the target object.

[0298] Correspondingly, the cloud receives the download request message.

[0299] S1912. The cloud sends the 3D model of the target object to the mobile phone.

[0300] Correspondingly, the mobile phone receives the 3D model of the target object.

[0301] S1913. The mobile phone displays the 3D model of the target object.

[0302] The mobile phone displaying the 3D model of the target object allows the user to preview the 3D model of the target object.

[0303] For example, the effect of the mobile phone displaying the 3D model of the target object can be referred to the above Figure 12 、 Figure 13 、 Figure 14 、 Figure 16 、 Figure 18As shown in the figure, when the user previews the 3D model of the target object, the user can rotate the angle, zoom in or out, etc. of the 3D model of the target object displayed on the mobile phone.

[0304] The above Figure 19 For the specific implementation of the process shown above and the beneficial effects achieved, please refer to the foregoing embodiments and will not be elaborated herein.

[0305] Exemplarily, Figure 20 FIG. is a logical schematic diagram of a 3D modeling method implemented by the end-cloud collaboration system provided in the embodiment of the present application.

[0306] As Figure 20 shown, in the embodiment of the present application, the mobile phone may at least include an RGB camera (such as a camera) and a first application. The first application is the above-mentioned 3D modeling application.

[0307] The RGB camera can be used to implement the shooting function of the mobile phone, shoot the target object to be modeled, and obtain a picture corresponding to the target object. The RGB camera can transmit the captured picture to the first application.

[0308] The first application may include a data acquisition and dynamic guidance module, a data processing module, a 3D model preview module, and a 3D model export module.

[0309] The data acquisition and dynamic guidance module can implement functions such as blur detection, key frame selection, guidance information calculation, and guidance interface update. For the pictures captured by the RGB camera, through the blur detection function, each captured picture (which can be called an input frame) can be subjected to blur detection, and a picture with a clarity meeting the requirements is obtained as a key frame; if the clarity of the current frame picture does not meet the requirements, the next frame picture is continuously obtained. Through the key frame selection function, it can be determined whether the picture is stored in the frame sequence file. If the picture is not stored in the frame sequence file, the frame picture is added to the frame sequence file. Through the guidance information calculation function, according to the camera pose information corresponding to the picture, the relationship between the picture and each patch in the patch model can be determined, and the patch corresponding to the picture can be obtained. The correspondence between the picture and the patch is the guidance information. Through the guidance interface update function, according to the foregoing calculated guidance information, the display effect of the patches in the patch model can be updated (i.e., changed), such as lighting up the patches.

[0310] The data processing module can implement functions such as matching relationship calculation, matching list calculation, and data packaging. Through the matching relationship calculation function, the matching relationship between key frames in the frame sequence file can be calculated, such as whether they are adjacent. Specifically, the data processing module can calculate the matching relationship between key frames in the frame sequence file according to the association relationship between patches in the patch model through the matching relationship calculation function. Through the matching list calculation function, a matching list for each key frame can be generated based on the calculation result of the matching relationship calculation function. The matching list for each key frame includes the matching information of each key frame. For example, the matching information of a certain key frame includes the identification information of other key frames associated with this key frame. Through the data packaging function, the first file including the matching information of each key frame and the frame sequence file can be packaged. After packaging the first file and the frame sequence file, the mobile phone can send the packaged data packet (including the first file and the frame sequence file) to the cloud.

[0311] The cloud can include a data parsing module, a 3D modeling module, and a data storage module. The data parsing module can parse the received data packet to obtain the frame sequence file and the first file. The 3D modeling module can perform 3D modeling based on the frame sequence file and the first file to obtain a 3D model.

[0312] For example, the 3D modeling module can implement functions such as key target extraction, feature detection and matching, global optimization and fusion, sparse point cloud calculation, dense point cloud calculation, surface reconstruction, and texture generation. Through the key target extraction function, the target object of interest in the key frame image can be separated from the background, and meaningful object entities can be recognized and interpreted from the image to extract different image features. Through the feature detection and matching function, unique pixel points existing in the key frame image can be detected as the feature points of the key frame image; the feature points with significant features in different key frame images are described, and the similarity of the two descriptions is compared to determine whether the feature points in different key frame images are the same feature. Through functions such as global optimization and fusion function, sparse point cloud calculation function, and dense point cloud calculation function, three-dimensional point cloud data corresponding to the target object can be generated according to the results of feature detection and matching. Through the surface reconstruction function, the three-dimensional surface shape of the object can be accurately restored using the three-dimensional point cloud data to obtain the basic 3D model of the target object. Through the texture generation function, the texture (also called texture map) of the surface of the target object can be generated according to the key frame image or the features of the key frame image. After obtaining the texture of the surface of the target object, the texture is mapped onto the surface of the above-mentioned basic 3D model of the target object in a specific manner to obtain the 3D model of the target object.

[0313] After the 3D modeling module obtains the 3D model of the target object, it can store the 3D model of the target object in the data storage module.

[0314] The first application on the mobile phone can download the 3D model of the target object from the data storage module in the cloud. After downloading the 3D model of the target object, the first application can provide the 3D model preview function for the user through the 3D model preview module, or provide the function of exporting the 3D model for the user through the 3D model export module. For the specific process of the first application providing the 3D model preview function for the user through the 3D model preview module, please refer to the foregoing embodiments.

[0315] Optionally, in the embodiments of the present application, when the mobile phone collects the data required for 3D modeling of the target object in the first part, it can also display the scanning progress on the 3D modeling data collection interface. For example, as described above Figures 7A to 7E As shown, the mobile phone can display the scanning progress through the circular black filling effect in the scan button on the 3D modeling data collection interface. It can be understood that since the UI rendering effect of the scan button is different, the way for the mobile phone to display the scanning progress on the 3D modeling data collection interface can be different, and this is not limited herein.

[0316] In some other embodiments, the mobile phone may not need to display the scanning progress, and the user can understand the scanning progress according to the lighting situation of the patches in the patch model.

[0317] The above embodiments are described by taking the implementation of the 3D modeling method provided in the embodiments of the present application in the end-cloud collaborative system composed of the terminal device and the cloud as an example. Optionally, in some other embodiments, the steps of the 3D modeling method provided in the embodiments of the present application can also be all implemented on the terminal device side. For example, for some terminal devices with strong processing capabilities and abundant computing resources, the functions implemented on the cloud side in the foregoing embodiments can also be all implemented in the terminal device. That is, after the terminal device obtains the above frame sequence file and the first file, it can directly generate the 3D model of the target object locally according to the frame sequence file and the first file, and provide functions such as preview and export of the 3D model. The specific principle of the terminal device generating the 3D model of the target object locally according to the frame sequence file and the first file is the same as the principle of the cloud generating the 3D model of the target object according to the frame sequence file and the first file in the foregoing embodiments, and will not be elaborated here.

[0318] It should be understood that the above descriptions in each embodiment are only exemplary descriptions of the 3D modeling method provided in the embodiments of the present application. In some other possible implementation manners, some execution steps in the above embodiments may be deleted or added, or the order of some steps described in the above embodiments may also be adjusted, and the present application does not limit this.

[0319] Corresponding to the 3D modeling method described in the foregoing embodiments, an embodiment of the present application provides a modeling device. This device can be applied to a terminal device and is used to implement the steps that the terminal device can implement in the 3D modeling method described in the foregoing embodiments. The functions of this device can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions.

[0320] For example, Figure 21 is a schematic structural diagram of the modeling device provided by an embodiment of the present application. As Figure 21 shown, the device may include: a display unit 2101 and a processing unit 2102. The display unit 2101 and the processing unit 2102 can be used to cooperate to implement the functions of the terminal device in the modeling method described in the foregoing method embodiments.

[0321] For example, the display unit 2101 is used to display a first interface, and the first interface includes a captured image of the terminal device.

[0322] The processing unit 2102 is used to, in response to a collection operation, collect multiple frames of images corresponding to a target object to be modeled and obtain the association relationship between the multiple frames of images. According to the multiple frames of images and the association relationship between the multiple frames of images, obtain a three-dimensional model corresponding to the target object.

[0323] The display unit 2101 is further used to display the three-dimensional model corresponding to the target object.

[0324] Among them, during the process of collecting multiple frames of images corresponding to the target object, the display unit 2101 is further used to display a first virtual enclosure; the first virtual enclosure includes multiple patches. The processing unit 2102 is specifically used to, when the terminal device is in a first pose, collect a first image and change the display effect of the patch corresponding to the first image; when the terminal device is in a second pose, collect a second image and change the display effect of the patch corresponding to the second image; after changing the display effects of the multiple patches of the first virtual enclosure, obtain the association relationship between the multiple frames of images according to the multiple patches.

[0325] Optionally, the display unit 2101 and the processing unit 2102 are further used to implement other display functions and processing functions of the terminal device in the modeling method described in the foregoing method embodiments, which will not be elaborated herein one by one.

[0326] Optionally, Figure 22 is another schematic structural diagram of the modeling device provided by an embodiment of the present application. As Figure 22As shown, for the modeling method described in the foregoing method embodiments, in the implementation where the terminal device sends the multiple frames of images and the association relationships between the multiple frames of images to the server, and the server generates a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationships between the multiple frames of images, the modeling apparatus may further include a sending unit 2103 and a receiving unit 2104. The sending unit 2103 is configured to send the multiple frames of images and the association relationships between the multiple frames of images to the server, and the receiving unit 2104 is configured to receive the three-dimensional model corresponding to the target object sent from the server.

[0327] Optionally, the sending unit 2103 is further configured to implement other sending functions that the terminal device can implement in the method described in the foregoing method embodiments, such as: sending a download request message, and the receiving unit 2104 is further configured to implement other receiving functions that the terminal device can implement in the method described in the foregoing method embodiments, such as: receiving an indication message, which will not be elaborated one by one here.

[0328] It should be understood that the apparatus may further include other modules or units for implementing the functions of the terminal device that can be implemented in the method described in the foregoing embodiments, which are not shown one by one here.

[0329] Optionally, an embodiment of the present application further provides a modeling apparatus, which can be applied to a server and is used to implement the functions of the server in the 3D modeling method described in the foregoing embodiments. The functions of the apparatus can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions.

[0330] For example, Figure 23 is another structural schematic diagram of the modeling apparatus provided by an embodiment of the present application. As Figure 23 shown, the apparatus may include: a receiving unit 2301, a processing unit 2302, and a sending unit 2303. The receiving unit 2301, the processing unit 2302, and the sending unit 2303 may be used to cooperate to implement the functions of the server in the modeling method described in the foregoing method embodiments.

[0331] For example, the receiving unit 2301 may be used to receive multiple frames of images corresponding to the target object sent from the terminal device and the association relationships between the multiple frames of images. The processing unit 2302 may be used to generate a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationships between the multiple frames of images. The sending unit 2303 may be used to send the three-dimensional model corresponding to the target object to the terminal device.

[0332] Optionally, the receiving unit 2301, the processing unit 2302, and the sending unit 2303 may be used to implement all the functions that the server in the modeling method described in the foregoing method embodiments can implement, which will not be elaborated one by one here.

[0333] It should be understood that the division of units (or modules) in the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity, or physically separated. And the units in the device can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some units can be implemented in the form of software called by a processing element, and some units can be implemented in the form of hardware.

[0334] For example, each unit can be a separately established processing element, or can be integrated in a certain chip of the device. In addition, it can also be stored in the memory in the form of a program and called and executed by a certain processing element of the device to perform the function of the unit. In addition, all or part of these units can be integrated together or can be independently implemented. The processing element mentioned here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above units can be implemented through the integrated logic circuit of the hardware in the processor element or in the form of software called by the processing element.

[0335] In an example, the units in the above device can be one or more integrated circuits configured to implement the above method. For example: one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.

[0336] Again, when the units in the device can be implemented in the form of a processing element scheduling program, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call programs. Again, these units can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0337] In one implementation, the units in the above device that implement the corresponding steps in the above method can be implemented in the form of a processing element scheduling program. For example, the device can include a processing element and a storage element. The processing element calls the program stored in the storage element to execute the method described in the above method embodiments. The storage element can be a storage element on the same chip as the processing element, that is, an on-chip storage element.

[0338] In another implementation, the program for executing the above method may be stored in a storage element on a different chip from the processing element, i.e., an off-chip storage element. At this time, the processing element calls or loads the program from the off-chip storage element onto the on-chip storage element to call and execute the steps performed by the terminal device or the server in the method described in the above method embodiments.

[0339] For example, an embodiment of the present application may further provide a device, such as an electronic device. The electronic device may include: a processor; a memory; and a computer program; wherein, the computer program is stored on the memory, and when the computer program is executed by the processor, the electronic device implements the steps performed by the terminal device or the server in the 3D modeling method described in the foregoing embodiments. The memory may be located inside the electronic device or outside the electronic device. And the processor includes one or more.

[0340] Exemplarily, the electronic device may be a mobile phone, a large screen (such as a smart screen), a tablet computer, a wearable device (such as a smart watch, a smart bracelet, etc.), a television, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), or other terminal devices.

[0341] In yet another implementation, the units for implementing the above steps in the device may be configured as one or more processing elements, and the processing elements here may be integrated circuits. For example: one or more ASICs, or one or more DSPs, or one or more FPGAs, or a combination of these types of integrated circuits. These integrated circuits may be integrated together to form a chip.

[0342] For example, an embodiment of the present application further provides a chip, which may be applied to the above electronic device. The chip includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the processors receive and execute computer instructions from the memory of the electronic device through the interface circuits to implement the steps performed by the terminal device or the server in the 3D modeling method described in the foregoing embodiments.

[0343] An embodiment of the present application further provides a computer program product, including computer-readable code, which, when running on an electronic device, enables the electronic device to implement the steps performed by the terminal device or the server in the 3D modeling method described in the foregoing embodiments.

[0344] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above function modules is used as an example. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above.

[0345] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0346] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium.

[0347] Based on such an understanding, the technical solution of the embodiment of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, such as a program. This software product is stored in a program product, such as a computer-readable storage medium, and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0348] For example, the embodiment of the present application can also provide a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the electronic device is enabled to implement the steps executed by the terminal device or the server in the 3D modeling method described in the foregoing embodiment.

[0349] Optionally, the embodiment of the present application also provides an edge-cloud collaboration system, and the composition of the edge-cloud collaboration system can refer to the above Figure 1 or Figure 20As shown, it includes a terminal device and a server, and the terminal device is connected to the server; the terminal device displays a first interface, and the first interface includes the captured image of the terminal device; the terminal device responds to a collection operation, collects multiple frames of images corresponding to a target object to be modeled, and obtains the association relationship between the multiple frames of images; wherein, during the process of collecting the multiple frames of images corresponding to the target object, the terminal device displays a first virtual bounding volume; the first virtual bounding volume includes multiple patches; the terminal device responds to a collection operation, collects multiple frames of images corresponding to a target object to be modeled, and obtains the association relationship between the multiple frames of images, including: when the terminal device is in a first pose, the terminal device collects a first image and changes the display effect of the patch corresponding to the first image; when the terminal device is in a second pose, the terminal device collects a second image and changes the display effect of the patch corresponding to the second image; after changing the display effects of the multiple patches of the first virtual bounding volume, the terminal device obtains the association relationship between the multiple frames of images according to the multiple patches; the terminal device sends the multiple frames of images and the association relationship between the multiple frames of images to the server; the server obtains a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images; the server sends the three-dimensional model corresponding to the target object to the terminal device; the terminal device displays the three-dimensional model corresponding to the target object.

[0350] Similarly, in this terminal-cloud collaboration system, the terminal device can implement all the functions that the terminal device can implement in the 3D modeling method described in the foregoing method embodiments, and the server can implement all the functions that the server can implement in the 3D modeling method described in the foregoing method embodiments, which will not be elaborated herein one by one.

[0351] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A modeling method, characterized in that, The method is applied to a terminal device, and the method includes: The terminal device displays a first interface, and the first interface includes a captured image of the terminal device; The terminal device responds to a collection operation, collects multiple frames of images corresponding to a target object to be modeled, and obtains the association relationship between the multiple frames of images; wherein, during the process of collecting the multiple frames of images corresponding to the target object, the terminal device displays a first virtual bounding volume; the first virtual bounding volume includes multiple patches; The terminal device responds to a collection operation, collects multiple frames of images corresponding to a target object to be modeled, and obtaining the association relationship between the multiple frames of images includes: When the terminal device is in a first pose, the terminal device collects a first image and changes the display effect of the patch corresponding to the first image; When the terminal device is in a second pose, the terminal device collects a second image and changes the display effect of the patch corresponding to the second image; After changing the display effects of the multiple patches of the first virtual bounding volume, the terminal device obtains the association relationship between the multiple frames of images according to the multiple patches; The terminal device obtains a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images; The terminal device displays the three-dimensional model corresponding to the target object.

2. The method according to claim 1, wherein The terminal device includes a first application. Before the terminal device displays the first interface, the method further includes: The terminal device responds to an operation of opening the first application and displays a second interface; The terminal device displays the first interface, including: The terminal device responds to an operation of starting the 3D modeling function of the first application on the second interface and displays the first interface.

3. The method according to claim 1 or 2, characterized in that, The first virtual bounding volume includes one or more layers, and the multiple patches are distributed on the one or more layers.

4. The method according to claim 1 or 2, characterized in that, The method further includes: The terminal device displays a first prompt message, and the first prompt message is used to remind the user to place the position of the target object in the captured image at the central position.

5. The method according to claim 1 or 2, characterized in that, The method further includes: The terminal device displays a second prompt message; the second prompt message is used to remind the user to adjust one or more of the shooting environment where the target object is located, the shooting method for the target object, and the screen occupation ratio of the target object.

6. The method according to claim 1 or 2, characterized in that, Before the terminal device obtains a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images, the method further includes: The terminal device detects an operation of generating a three-dimensional model; The terminal device responds to the operation of generating the three-dimensional model and displays a third prompt message, and the third prompt message is used to prompt the user that the target object is being modeled.

7. The method according to claim 1 or 2, characterized in that, After the terminal device obtains a three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images, the method further includes: The terminal device displays a fourth prompt message, and the fourth prompt message is used to prompt the user that the modeling of the target object has been completed.

8. The method according to claim 1 or 2, characterized in that, The terminal device displaying the three-dimensional model corresponding to the target object further includes: The terminal device changes the display angle of the three-dimensional model corresponding to the target object in response to an operation of changing the display angle of the three-dimensional model corresponding to the target object; the operation of changing the display angle of the three-dimensional model corresponding to the target object includes an operation of dragging the three-dimensional model corresponding to the target object to rotate clockwise or counterclockwise along a first direction.

9. The method according to claim 1 or 2, characterized in that, The terminal device displaying the three-dimensional model corresponding to the target object further includes: The terminal device changes the display size of the three-dimensional model corresponding to the target object in response to an operation of changing the display size of the three-dimensional model corresponding to the target object; the operation of changing the display size of the three-dimensional model corresponding to the target object includes an operation of enlarging or reducing the three-dimensional model corresponding to the target object.

10. The method according to claim 1 or 2, characterized in that, The association relationship between the multiple frames of images includes the matching information of each frame of image in the multiple frames of images; The matching information of each frame of the image includes the identification information of other images associated with the image in the multiple frames of images; The matching information of each frame of the image is obtained according to the association relationship between each frame of the image and the patch corresponding to each frame of the image, and the association relationship between the multiple patches.

11. The method according to claim 1 or 2, characterized in that, The terminal device collecting multiple frames of images corresponding to the target object to be modeled in response to a collection operation, and obtaining the association relationship between the multiple frames of images further includes: The terminal device determines the target object according to the captured picture; When the terminal device collects the multiple frames of images, the position of the target object in the captured picture is the central position of the captured picture.

12. The method according to claim 1 or 2, characterized in that, The terminal device collecting multiple frames of images corresponding to the target object to be modeled includes: During the process of the terminal device photographing the target object, the terminal device performs a blur detection on each frame of the captured image, and collects the image with a clarity greater than a first threshold as the image corresponding to the target object.

13. The method according to claim 1 or 2, characterized in that, The terminal device displaying the three-dimensional model corresponding to the target object includes: The terminal device displays the three-dimensional model corresponding to the target object in response to an operation of previewing the three-dimensional model corresponding to the target object.

14. The method according to claim 1 or 2, characterized in that, The three-dimensional model corresponding to the target object includes the basic three-dimensional model of the target object and the texture on the surface of the target object.

15. The method according to claim 1 or 2, characterized in that, The terminal device is connected to the server; The terminal device obtaining the three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images includes: The terminal device sends the multiple frames of images and the association relationship between the multiple frames of images to the server; The terminal device receives the three-dimensional model corresponding to the target object sent from the server.

16. The method according to claim 15, wherein The method further includes: The terminal device sends the camera internal parameters, gravity direction information, image name, image number, camera pose information, and time stamp respectively corresponding to the multiple frames of images to the server.

17. The method according to claim 15, wherein The method further includes: The terminal device receives an indication message from the server, and the indication message is used to indicate to the terminal device that the server has completed the modeling of the target object.

18. The method according to claim 15, wherein Before the terminal device receives the three-dimensional model corresponding to the target object sent by the server, the method further includes: The terminal device sends a download request message to the server, where the download request message is used to request the server to download the three-dimensional model corresponding to the target object.

19. A modeling method, characterized in that, The method is applied to a server, and the server is connected to a terminal device; the method includes: The server receives multiple frames of images corresponding to the target object and the association relationship between the multiple frames of images sent by the terminal device; among them, the multiple frames of images corresponding to the target object include: a first image and a second image; the first image is collected when the terminal device is in a first pose; the second image is collected when the terminal device is in a second pose; the association relationship between the multiple frames of images is obtained by the terminal device according to multiple patches after changing the display effect; the multiple patches are included in a first virtual bounding volume, and the first virtual bounding volume is displayed during the process of the terminal device collecting the multiple frames of images corresponding to the target object; the server generates the three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images; The server sends the three-dimensional model corresponding to the target object to the terminal device.

20. The method according to claim 19, wherein The method further includes: The server receives the camera internal parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to each of the multiple frames of images sent by the terminal device; The server generates the three-dimensional model corresponding to the target object according to the multiple frames of images and the association relationship between the multiple frames of images, including: The server generates the three-dimensional model corresponding to the target object according to the multiple frames of images, the association relationship between the multiple frames of images, the camera internal parameters, gravity direction information, image name, image number, camera pose information, and timestamp corresponding to each of the multiple frames of images.

21. The method according to claim 19 or 20, characterized in that, The association relationship between the multiple frames of images includes the matching information of each frame of the multiple frames of images; The matching information of each frame of the image includes the identification information of other images associated with the image in the multiple frames of images; The matching information of each frame of the image is obtained according to the association relationship between each frame of the image and the patch corresponding to each frame of the image, and the association relationship between the multiple patches.

22. A terminal-cloud collaboration system, characterized in that, Includes: A terminal device and a server, where the terminal device is connected to the server; The terminal device displays a first interface, and the first interface includes the shooting screen of the terminal device; The terminal device responds to a collection operation, collects multiple frames of images corresponding to the target object to be modeled, and obtains the association relationship between the multiple frames of images; among them, during the process of collecting the multiple frames of images corresponding to the target object, the terminal device displays a first virtual bounding volume; the first virtual bounding volume includes multiple patches; The terminal device responds to a collection operation, collects multiple frames of images corresponding to the target object to be modeled, and obtains the association relationship between the multiple frames of images, including: When the terminal device is in the first pose, the terminal device acquires a first image and changes the display effect of the patch corresponding to the first image; When the terminal device is in the second pose, the terminal device acquires a second image and changes the display effect of the patch corresponding to the second image; After changing the display effects of the multiple patches of the first virtual enclosure, the terminal device obtains the correlation relationship between the multiple frames of images according to the multiple patches; The terminal device sends the multiple frames of images and the correlation relationship between the multiple frames of images to the server; The server obtains the three-dimensional model corresponding to the target object according to the multiple frames of images and the correlation relationship between the multiple frames of images; The server sends the three-dimensional model corresponding to the target object to the terminal device; The terminal device displays the three-dimensional model corresponding to the target object.

23. An electronic device, characterized in that, Including: A processor; A memory; And a computer program; wherein, the computer program is stored on the memory, and when the computer program is executed by the processor, the electronic device implements the method according to any one of claims 1-18, or the method according to any one of claims 19-21.

24. A computer-readable storage medium, the computer-readable storage medium comprising a computer program, characterized in that, When the computer program runs on the electronic device, the electronic device implements the method according to any one of claims 1-18, or the method according to any one of claims 19-21.

Citation Information

Patent Citations

  • Information processing method and device and electronic equipment

    CN109658507A

Cited By

  • Modeling method, related electronic device, and storage medium

    EP4779595A2