Method, electronic device, program product for three-dimensional modeling
By segmenting the image into areas of interest and non-interest, using the neural network of servers and terminal devices to generate a three-dimensional model, the problem of inefficient resource utilization in the prior art is solved, and faster and better three-dimensional modeling effect is achieved.
Patent Information
- Application Number
- CN202410107438.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
现有三维建模方法在资源利用上不够高效,导致处理速度慢且资源消耗过多,难以处理复杂任务或大数据集。
The image is segmented into an area of interest and a non-interested area, and the first 3D model is generated using a high-performance GPU on the server side, and the second 3D model is generated through a neural network on the terminal device, combining the two to generate the final model.
It improves the processing speed and quality of three-dimensional modeling, reduces resource consumption, and can model faster and better.
Smart Images

Figure CN120374829A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computers, and more particularly, to methods, electronic devices, and products for three-dimensional modeling. Background Art
[0002] Three-dimensional modeling is the process of using software to create a three-dimensional digital representation of an object. In three-dimensional space, objects can be created and manipulated to build models ranging from simple geometric shapes to complex visualizations and animations. This process is widely applied in multiple fields, including video game design, film production, architectural design, engineering, and medicine, among others.
[0003] For example, three-dimensional modeling can be used in technologies such as the metaverse, which will lead humanity into a new digital era and allow people to obtain real and ultimate experiences. People can use metaverse applications anywhere on Earth without worrying about delays in computing and data access. Learning three-dimensional (3D) models from two-dimensional (2D) images / videos is an efficient method that can easily elevate image content to the 3D space. Currently, high-performance GPUs have made great progress in acceleration. 2D to 3D modeling and rendering can be achieved. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, an electronic device, and a computer program product for three-dimensional modeling.
[0005] According to a first aspect of the present disclosure, there is provided a method for three-dimensional modeling. The method includes segmenting an image into a region of interest and a non-region of interest, and sending data associated with the region of interest to a server. The method further includes receiving from the server a first three-dimensional (3D) model corresponding to the region of interest, the first 3D model being generated by a first neural network at the server, and generating a second 3D model corresponding to the non-region of interest by a second neural network, the first neural network having more network parameters than the second neural network, and generating a three-dimensional model corresponding to the image based on the first 3D model and the second 3D model.
[0006] According to a second aspect of the present disclosure, there is provided an electronic device for 3D modeling. The device includes at least one processor, and a memory coupled to the at least one processor and having instructions stored thereon that, when executed by the at least one processor, cause the electronic device to perform actions including segmenting an image into a region of interest and a non-region of interest, and sending data associated with the region of interest to a server. The method further includes receiving from the server a first three-dimensional (3D) model corresponding to the region of interest, the first 3D model being generated by a first neural network at the server, and generating a second 3D model corresponding to the non-region of interest by a second neural network, the first neural network having more network parameters than the second neural network, and generating a 3D model corresponding to the image based on the first 3D model and the second 3D model.
[0007] According to a third aspect of the present disclosure, there is provided a computer program product tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions that, when executed, cause the machine to perform the steps of the method implemented in the first aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, in which like reference numerals generally represent like elements in the exemplary embodiments of the present disclosure.
[0009] Figure 1 FIG. illustrates a schematic diagram of a system for 3D modeling in which the device and / or method according to an embodiment of the present disclosure can be implemented;
[0010] Figure 2 FIG. illustrates a flowchart of a method for 3D modeling according to an embodiment of the present disclosure;
[0011] Figure 3 FIG. illustrates a schematic diagram of an architecture of a multi-level neural network according to an embodiment of the present disclosure;
[0012] Figure 4 FIG. illustrates a schematic diagram of a multi-stage training process based on ROI for scene segmentation according to an embodiment of the present disclosure;
[0013] Figure 5 FIG. illustrates a schematic diagram of a multi-stage process for scene modeling and rendering according to an embodiment of the present disclosure;
[0014] Figure 6 FIG. illustrates graphs of multiple experimental results according to an embodiment of the present disclosure; and
[0015] Figure 7A schematic block diagram of an example device that can be used to implement embodiments of the present disclosure is shown. Detailed implementation manners
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0017] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0018] Currently, processing data associated with images based on 2D to 3D neural networks is extremely time-consuming because they do not maximize the use of computing capabilities (such as resources like CPUs and GPUs) for acceleration. Many methods do not perform well in real-time rendering. Since these methods mostly train neural networks offline and then render free-view videos or 3D meshes through offline calculations, it is more time-consuming and resource-intensive locally. In addition, in some methods, the collected image regions are not distinguished, and a large amount of resources are spent on rendering and drawing in each region. Excessive resource consumption may limit the system's ability to process more complex tasks or larger datasets.
[0019] To at least solve the above and other potential problems, embodiments of the present disclosure provide a method for three-dimensional modeling. The method includes segmenting an image into regions of interest and regions of non-interest, and sending data associated with the regions of interest to a server. The method also includes receiving from the server a first three-dimensional (3D) model corresponding to the regions of interest, the first 3D model being generated by a first neural network at the server, and generating a second 3D model corresponding to the regions of non-interest by a second neural network, the first neural network having more network parameters than the second neural network, and generating a three-dimensional model corresponding to the image based on the first 3D model and the second 3D model. By using this method, the processing speed of three-dimensional modeling can be accelerated, the modeling quality and accuracy can be improved, thereby achieving better and faster modeling.
[0020] The basic principles and several example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Figure 1FIG. shows a schematic diagram of a system 100 for 3D modeling in which an apparatus and / or method according to an embodiment of the present disclosure may be implemented. It should be understood that Figure 1 The number and arrangement of the objects, components, and elements shown are merely examples, and the schematic diagram may include different numbers and arrangements of components, elements, nodes, objects, and various additional elements.
[0021] As Figure 1 shown, a user may use a terminal device 102 such as a mobile phone or a camera to capture a relevant scene 104 including a person and a background, and generate an image and / or video 106 corresponding to these scenes. In some embodiments, the generated image and / or video 106 may be preprocessed. For example, image or video data conversion, video sampling, and upsampling or downsampling of an image or a video, adjusting the resolution of these data, and so on. For example, the terminal device 102 may convert image or video data into a format more suitable for processing, storing, or displaying through image or video data conversion. As an example, the terminal device 102 may convert a RAW format image into a JPEG or PNG format, or convert an AVI format video into an MP4 format, etc. Thus, the file size can be reduced, or the color space can be converted to meet the requirements of different display devices.
[0022] In some embodiments, the terminal device 102 may extract important frames or segments from the original video through video sampling to reduce the data volume and highlight key content. For example, sampling is performed according to a time interval (e.g., extracting one frame per second), or according to the importance of content changes (e.g., extracting frames when a scene transition or an important event occurs), so as to select specific frames.
[0023] Additionally or alternatively, in some embodiments, the terminal device 102 may increase the resolution of an image or a video through upsampling of the image or the video to enhance details and clarity. For example, interpolation methods such as nearest neighbor interpolation, bilinear interpolation, and cubic interpolation may be used to increase the number of pixels. It works by estimating and filling in the color values of the newly added pixels. Thus, the image details can be magnified or the video playback quality can be improved. In some embodiments, the terminal device 102 may also reduce the resolution of an image or a video through downsampling of the image or the video to reduce the file size and facilitate storage and transmission between different devices or networks. For example, the terminal device 102 may remove pixels by selectively removing certain pixels or averaging neighboring pixels, thereby optimizing the image and video to adapt to bandwidth limitations and storage limitations.
[0024] The generated image or video frame 106, intermediate data, etc. can be uploaded to the server 108, where there are (multiple) neural networks 110 that can be used for 3D modeling or rendering. According to some embodiments of the present disclosure, these neural networks 110 can be arranged on (multiple) GPUs 112 for training, thereby providing more robust and efficient audio-visual processing capabilities. The (multiple) GPUs 112 contain several cores and are designed to simultaneously and parallelly process multiple computing tasks, especially graphics and image-related computations, thus significantly accelerating the processing speed and enhancing the training of the (multiple) neural networks 110.
[0025] In some embodiments, the user can also select different (multiple) neural networks 110 based on different GPU array configurations. For example, the neural network 110 with more GPU array configurations can be used for more complex and delicate 3D modeling. For example, the neural network 110 above a specific GPU number threshold can be used for 3D modeling of people, plants, clothes, textures, animal hairs, etc. While the neural network 110 with fewer GPU array configurations or less than the specific GPU number threshold can be used for relatively simple 3D modeling, such as roads, skies, etc.
[0026] In some examples, the neural network 110 can be a deep learning model for 3D scene reconstruction and rendering. For example, the neural network 110 reconstructs a 3D representation of a continuous scene by observing a set of images captured from multiple different angles. The neural network 110 can first use a series of 2D pictures taken from different angles as inputs. In some embodiments, the neural network 110 can also utilize volume rendering techniques to generate a 3D scene and consider the interaction of light with objects when passing through the scene, such as factors like scattering and absorption.
[0027] In some embodiments, the neural network 110 can represent the scene based on a fully connected neural network. During the rendering process, the neural network 110 can sample along the rays emitted from the virtual camera position to simulate ray casting and estimate the color and volume density of these points. The neural network 110 can then integrate the sampled color and density. For example, the neural network 110 can determine the final pixel value by integrating the color and density along each ray and consider the cumulative effect of light when passing through space. Finally, the neural network 110 can output a 3D model of the scene that can be observed from any new angle.
[0028] In some embodiments, the neural network 110 can be trained based on a set of images taken from different angles. During the training process, the neural network 110 can adjust the parameters of the neural network in such a way that the (multiple) neural networks 110 can better predict the 3D scene corresponding to each input image.
[0029] In some embodiments, the neural network 110 may apply background removal and object detection to the generated images and / or videos 106. For example, in some embodiments, the neural network 110 may select a Region of Interest (ROI) 114 by generating a binary mask. Specifically, in some embodiments, the neural network 110 may separate the object or region of interest 114 from the image or video by removing or segmenting the background. In some embodiments, the neural network 110 may identify the object or region of interest 114 by comparing features such as color, texture, edges, motion features in multiple images of different parts of the object or region of interest compared to the ordinary background, so as to distinguish the foreground object and the background. In some embodiments, the neural network 110 may also separate the foreground and background in the image by edge detection, threshold processing, etc.
[0030] According to some embodiments of the present disclosure, the neural network 110 may also determine the object 114 in the scene based on template matching, feature point matching, etc. Additionally or alternatively, in some embodiments, the neural network 110 may generate a mask by creating a mask layer on the image or video frame 106. For example, the region of interest 114 of the object is marked with one value (such as the numerical value 1), while other regions are marked with another value (such as the numerical value 0), and the regions marked with the other value are removed to achieve background removal, object detection, image segmentation, etc.
[0031] Additionally or alternatively, in some embodiments, the neural network 110 may also convert the 2D image into a sparse point cloud 116. For example, the neural network 110 may also identify and extract key feature points in the uploaded image 118. The neural network 110 may also generate a disparity map to estimate depth information. The disparity map reflects the disparity change of the same scene under different perspectives. For multi-view images, the neural network 110 may improve the accuracy of depth estimation by fusing depth information from multiple perspectives.
[0032] Finally, the neural network 110 may convert the pixels of the image into points in 3D space according to the estimated depth information to form a point cloud. The position of each point can be determined according to its position in the image and the depth value. The neural network 110 may finally selectively extract key points from the generated point cloud to form a sparse point cloud. Additionally or alternatively, the neural network 110 may clean the generated point cloud to remove noise and irrelevant points, and may also smooth and optimize the generated point cloud model to improve the quality. In some embodiments, the neural network 110 may merge the point cloud with other data sources such as the point cloud generated by a laser scanner to improve accuracy and integrity.
[0033] According to embodiments of the present disclosure, some of the trained neural networks 110, and the generated 3D models can be sent or downloaded to the same or different terminal devices 120 as the terminal device 102. In some embodiments, the terminal device 120 can communicate with the server 108 through sockets, bindings, etc., so as to receive and send associated data. At the terminal device 120, the 3D model for the ROI can also be fine-tuned, for example, by applying an image mask to the 3D model to achieve 3D model conversion.
[0034] According to embodiments of the present disclosure, the terminal device 102 or the terminal device 120 can be any computing device with processing computing resources or storage resources. For example, the computing device can have common capabilities such as receiving and sending data requests, real-time data analysis, local data storage, and real-time network connection. Computing devices generally can include various types of devices. Examples of computing devices can include, but are not limited to: database servers, rack servers, server clusters, desktop computers, laptop computers, smart phones, wearable devices, security devices, intelligent manufacturing devices, smart home devices, Internet of Things devices, smart cars, drones, etc., and the present disclosure does not make any restrictions on this.
[0035] Additionally or alternatively, in some embodiments, online rendering 122 can be achieved by setting network IP and ports, broadcasting web links. In some embodiments, video rendering can be achieved by setting a camera path, offline rendering for customized videos. In some embodiments, point cloud rendering 124 can be achieved by mesh reconstruction. The finally rendered 3D model (such as point cloud or mesh) can be shared for various 3D applications. Thus, according to embodiments of the present disclosure, a complete automatic multi-stage framework can be realized, and end-to-end processing of 2D photos for 3D modeling can be achieved without manual adjustment or cumbersome parameter settings. In some embodiments, a user-friendly ROI detection module can be provided, allowing users to select key components from the image for view rendering from coarse to fine. It results in multi-stage optimization to enhance key 3D object modeling and rendering.
[0036] The above combines Figure 1 a block diagram of an environment in which some embodiments of the present disclosure can be implemented is described. The following combines Figure 2 a flowchart of a method 200 for three-dimensional modeling according to embodiments of the present disclosure is described. The method 200 can be executed at Figure 1 the terminal device 102 and / or 102.
[0037] At block 202, the image is segmented into a region of interest and a non-region of interest. According to embodiments of the present disclosure, Figure 1The terminal device 102 described in [description] can segment the received image data or video data based on user selection or predetermined settings. For example, in some embodiments, the terminal device 102 can divide the image data or video data into foreground data or background data based on data such as the people or objects selected or marked by the user on the screen and by generating a binary mask. The foreground data is the area that the user is concerned about or interested in, while the background data can be the data that the user is less concerned about.
[0038] Additionally or alternatively, the terminal device 102 can determine the region of interest by comparing one or more of color, texture, and edges in multiple different regions of the image. For example, by identifying a specific color or color pattern to mark the region. For example, in a natural image, the terminal device 102 can identify a green region as a tree or grassland, etc. In some embodiments, the terminal device 102 can locate the region by identifying the texture pattern in the image. The texture can be a repeating pattern or a region with a specific directionality. In other examples, the terminal device 102 can identify an object by recognizing the boundary of a region with significant color or brightness changes.
[0039] At block 204, data associated with the region of interest is sent to the server. In some embodiments, the terminal device 102 can send the previously identified region of interest to the server 108. As an example, the region of interest can be an object such as a person, animal, clothing, hair, texture, etc. that the user is concerned about or focused on. In some embodiments, methods such as object detection methods (e.g., methods based on deep learning for object detection) can be used to identify objects such as people and animals in the image. For a video or a sequence of consecutive images, object tracking techniques can be applied to continuously track these objects. In some embodiments, a machine learning model can also be used to learn and predict the user's points of interest. This can be based on the user's historical interactions, image content analysis, or other relevant data. These regions that the user is more interested in can be processed with higher priority.
[0040] At block 206, a first three-dimensional 3D model corresponding to the region of interest is received from the server, and the first 3D model is generated by a first neural network at the server. The data associated with the region of interest can be processed by the first neural network. According to an embodiment of the present disclosure, the first neural network is deployed in one or more GPUs in the server. Multiple GPUs can effectively assist the neural network in processing parallel tasks, thereby enabling faster and more efficient processing of large datasets or complex image data.
[0041] According to embodiments of the present disclosure, these neural networks can generate 3D models by simulating the propagation of light using volume rendering based on a set of images captured from multiple perspectives by utilizing a GPU. For example, the neural network can first simulate the propagation of light inside a substance (such as scattering and absorption, etc.) to generate multiple 2D images, and then the neural network can reconstruct a detailed and accurate 3D scene model from the multiple 2D images.
[0042] At block 208, a second 3D model corresponding to the non-interested region is generated by a second neural network, and the first neural network has more network parameters than the second neural network. According to embodiments of the present disclosure, the neural network deployed on the GPU in the terminal device 102 can process the segmented non-interested region to generate a corresponding 3D model, such as models of the background, road, sky, etc. In some embodiments, these neural networks are also trained and adjusted based on the image regions that can be selected by the user and the image regions not selected by the user.
[0043] According to embodiments of the present disclosure, compared with the neural network deployed on the GPU in the terminal device 102, the neural network deployed on the GPU in the server 108 can have more adjustable parameters, a faster modeling speed, and more available modeling resources, etc. These neural networks or the associated GPUs can be dynamically scheduled and configured by the task management based on factors such as the required data volume, task complexity, and hardware configuration for processing the interested region and the non-interested region.
[0044] At block 210, a 3D model corresponding to the image is generated based on the first 3D model and the second 3D model. According to embodiments of the present disclosure, the terminal device 102 can combine the 3D models generated by different devices to generate a final 3D model. For example, it is determined according to different factors such as the required computing amount and hardware configuration. In some embodiments, the terminal device 102 can determine the final 3D model only using local computing and the generated 3D model. In this case, the terminal device is fully responsible for the 3D model generation process. All data processing is performed locally, thereby reducing the processing time caused by network latency.
[0045] In some embodiments, the terminal device 102 can be combined with the 3D model generated on the server side to generate a final 3D model, which can be used to process more complex tasks and dynamically adjust the computing burden between the local and the server as needed. In other embodiments, the terminal device 102 can also only call the neural network deployed on the server side to generate the final 3D model, which reduces various requirements for the terminal device 102.
[0046] Additionally or alternatively, in some embodiments, the method implemented according to the present disclosure can also be used to process video data. For example, the terminal device 102 can split the video into a region of interest video and a non-region of interest video. The terminal device 102 can receive the 3D video corresponding to the region of interest video generated by the neural network deployed on the server, and generate the 3D video corresponding to the non-region of interest video, and finally can combine and splice these 3D videos.
[0047] The above has been described in conjunction with Figure 2 the flowchart of the method 200 for 3D modeling according to an embodiment of the present disclosure. The following will be described in conjunction with Figure 3 the schematic diagram of the architecture diagram 300 of the multi-level neural network according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, it may include a training stage and a visualization rendering stage for the multi-level neural network.
[0048] According to an embodiment of the present disclosure, multiple neural networks 302 can be first stored in the database 306a. Subsequently, the multiple neural networks 302 can be copied into N copies and stored in the sub-databases 308-1 to 308-n respectively. Each copy neural network has the same neural network structure. Similarly, the photos and videos 304 can also be stored in the database 306b. Subsequently, the multiple neural networks 302 can be copied into N copies and sampled into N batches of data by a random sampler and stored in the sub-databases 310-1 to 310-n.
[0049] In some embodiments, the stored copy neural networks and photos and / or videos can be managed and scheduled by the task manager 312. For example, the task manager 312 can deploy the copy neural networks to multiple GPUs located on the server 314, such as GPUs 316-1 to 316-n. In some embodiments, the task manager 312 can split the neural network and photo-video data into small tasks, package them for parallel training. For example, in the training stage, the task manager 312 can use multiple GPUs on multiple clouds to accelerate the training of (multiple) copy neural networks.
[0050] In some embodiments, the task manager 312 can determine the available resources and task complexity, including the number and performance of GPUs on the server. The task manager 312 can then schedule and allocate the rendering or modeling tasks to the appropriate GPUs 316-1 to 316-n based on the resource availability and task requirements.
[0051] Additionally or alternatively, in some embodiments, the task manager 312 may define the number of available GPUs for neural network initialization. For example, one neural network may correspond to one or more GPUs, or multiple neural networks may correspond to one GPU. The task manager 312 may need to dynamically schedule or determine the optimal correspondence based on the computational requirements of the replicated neural network and the computational capabilities of the GPUs.
[0052] For example, in some embodiments, the task manager 312 may dynamically allocate the number of GPUs based on the requirements of rendering or modeling tasks and the usage of current resources. For example, for tasks that require high computational resources, the task manager 312 may allocate more GPUs. Additionally or alternatively, in some implementations, the task manager 312 may monitor the usage and performance metrics of the GPUs in real time, thereby enabling performance monitoring and adjustment.
[0053] When deploying the replicated neural network, the task manager 312 may need to consider load balancing to ensure balanced utilization of each GPU and avoid overloading some GPUs while others are idle. In some embodiments, the task manager 312 may limit the number of GPUs used for neural network initialization, thereby helping to optimize resource usage, especially in resource-constrained scenarios.
[0054] In some embodiments, the task manager 312 may further include a user interface, which may provide an interface for the user to view the task status, resource usage, and perform manual adjustments when necessary. In some embodiments, the task manager 312 may have a user feedback program, which may dynamically adjust the resource management and task scheduling policies according to the specific needs of the user. Thus, the task manager 312 can ensure the effective management and scheduling of neural networks and their related data in a multi-GPU environment, thereby improving resource utilization, accelerating task processing speed, and providing flexibility to meet various computational requirements.
[0055] In some embodiments, one or more trained neural networks deployed on the cloud server 314 may be transmitted to the terminal device 318 via a network connection. The terminal device 318 may then deploy the one or more trained neural networks to the local GPU 320. According to an embodiment of the present disclosure, the GPUs 316-1 to 316-n deployed in the cloud server 314 and the local GPUs 320 deployed in the terminal device 318 may form a heterogeneous GPU.
[0056] These GPUs can have multiple graphics processing units of different types or performance specifications. For example, these GPUs can be from different manufacturers and have different architectures, different numbers of cores, computing capabilities, processing speeds, memory capacities, or energy efficiency ratios. For different application scenarios, heterogeneous GPUs can optimize the allocation of computing resources and assign specific tasks to the most suitable GPU.
[0057] For example, according to some embodiments of the present disclosure, 3D modeling tasks with intensive details such as those for characters, animals, hair, textures, clothes, etc. can be processed by GPUs 316-1 to 316-n deployed in server 314. These GPUs have faster processing speeds and higher memory bandwidths, and can continuously adjust the resource allocation strategy according to the actual running situation and performance feedback. In some embodiments of the present disclosure, the 3D modeling tasks with intensive details can be the regions of interest in an image or video specified by the user.
[0058] According to some embodiments of the present disclosure, 3D modeling tasks with coarse granularity such as those for backgrounds, roads, sky environments, etc. can be processed by GPU 320 deployed in terminal device 318. These tasks generally do not require highly complex computations and are therefore suitable for being performed on hardware with lower performance. In some embodiments of the present disclosure, the 3D modeling tasks with coarse granularity can be the non-regions of interest in an image or video specified by the user. In some embodiments, a visual website can also be provided to the user for directly observing the 360-degree view of the scene. In some embodiments, GPUs 316-1 to 316-n and GPU 320 can be in the same virtual environment for the convenience of training and visualization. Finally, the 3D models rendered at GPUs 316-1 to 316-n can be combined with the 3D models rendered at GPU 320 to form.
[0059] Figure 4 FIG. shows a schematic diagram of an ROI multi-stage training process 400 for scene segmentation according to an embodiment of the present disclosure. As Figure 4 shown, according to some embodiments of the present disclosure, the scene segmentation model 402 stored in the cloud server can segment the input image or video 404. In some embodiments, the obtained segmentation map can divide the scene content in the image or video 404 into F classes 408 based on pixel-level segmentation.
[0060] In some examples, the user 410 may divide the segmented scene content into regions of interest (ROIs) as foreground images as needed, and use the remaining scene content as different layers 412 of the background image. A binary mask can be generated based on the selection of the user 401 to cover the background area. In some embodiments, the scene segmentation model 402 may be any known or unknown segmentation model that can well detect cars, bicycles, people, ground, grass, plantations, etc. In some embodiments, multiple GPUs deployed on the server 416 may be used to train the neural network based on the background area and the ROI.
[0061] In this way, the neural network can be made more focused on reconstructing the details 420 associated with the ROI region. After the neural network training is completed, it can be downloaded and deployed to an edge device 422 such as a client, and users can use their computers to refine the entire scene in a very short time. During the multi-stage training process based on the ROI, the ROI region selected by the user can be trained and rendered on the server, and the remaining background will be refined in the edge computer, and finally a combined 3D rendering 422 including the rendered characters and scenes will be generated. The neural network can be selected according to the speed, quality, function of different neural networks and different scenes.
[0062] Figure 5 FIG. shows a schematic diagram of a multi-stage process 500 for scene modeling and rendering according to an embodiment of the present disclosure. As Figure 5 shown, the external image or video data 502 can be determined as the ROI 506 and the background area based on the area 504 selected by the user. The ROI 506 that needs to be refined or processed with fine granularity can be modeled and drawn by the GPU deployed in the server, while the background area can be processed or fine-tuned at the user terminal, and the ROI 506 and the background area are combined to finally generate a complete 3D model 510.
[0063] It can be seen that the human model is well reconstructed from the ROI 506 stage. In the combination stage, the floating noise around the human body can be removed, and the background can be refined to obtain good visual quality. FIGS. 512 to 518 respectively show schematic diagrams of 3D models modeled and drawn based on different neural networks. These images show the quantitative results of different methods on different GPUs.
[0064] Figure 6 FIG. shows multiple experimental result diagrams 600 according to an embodiment of the present disclosure. As Figure 6As shown, FIGS. 602 to 608 respectively show schematic diagrams of 4K resolution scene reconstruction 3D models modeled and drawn based on different neural networks, demonstrating depth estimation between different methods. According to embodiments of the present disclosure, quantitative comparisons can be made among multiple different neural networks and can be provided to users as a benchmark for use. Both users or administrators can make corresponding improvements for specific scene modeling and rendering.
[0065] Figure 7 FIG. shows a schematic block diagram of an example device 700 that can be used to implement embodiments of the present disclosure. Figure 1 The terminal device in can be implemented using device 700. As shown, device 700 includes a central processing unit (CPU) 701, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or computer program instructions loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of device 700 can also be stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0066] Multiple components in device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage page 708, such as a disk, optical disc, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0067] The various processes and treatments described above, such as method 200, can be executed by the processing unit 701. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the CPU 701, one or more actions of method 200 described above can be executed.
[0068] The present disclosure can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present disclosure.
[0069] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0070] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0071] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0072] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer - readable program instructions.
[0073] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium that causes a computer, a programmable data - processing apparatus, and / or other devices to function in a particular manner, so that the computer - readable medium storing the instructions comprises a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0074] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0075] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0076] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.
Claims
1. A method for three-dimensional modeling, comprising: Segmenting an image into a region of interest and a non-region of interest; Sending data associated with the region of interest to a server; Receiving from the server a first three-dimensional 3D model corresponding to the region of interest, the first 3D model being generated by a first neural network at the server; Generating, by a second neural network, a second 3D model corresponding to the non-region of interest, the first neural network having more network parameters than the second neural network; And Generating a 3D model corresponding to the image based on the first 3D model and the second 3D model.
2. The method according to claim 1, wherein the first neural network and the second neural network generate a 3D model by simulating the propagation of light using volume rendering based on an image set captured from multiple angles.
3. The method according to claim 1, wherein segmenting the image into the region of interest and the non-region of interest comprises: Segmenting the region of interest and the non-region of interest by generating a binary mask based on user selection and identification.
4. The method according to claim 1, wherein segmenting the image into the region of interest and the non-region of interest comprises: Determining the region of interest by comparing one or more of color, texture, and edges in multiple different regions of the image.
5. The method according to claim 1, wherein the first neural network is deployed in one or more GPUs in the server; The second neural network is deployed in one or more GPUs in a terminal device; and The one or more GPUs are dynamically scheduled and configured by a task manager based on one or more of the required data volume, task complexity, and hardware configuration factors for processing the region of interest and the non-region of interest.
6. The method according to claim 1, wherein the first neural network is used for rendering and reconstruction of the region of interest, and the second neural network is used for rendering and reconstruction of the non-region of interest.
7. The method according to claim 3, wherein the first neural network is further trained and adjusted based on the image region selected by the user, and the second neural network is trained and adjusted based on the image region not selected by the user.
8. The method according to claim 1, further comprising: Segmenting a video into a region-of-interest video and a non-region-of-interest video; And Receiving a first 3D video corresponding to the region-of-interest video generated by the first neural network.
9. The method according to claim 8, further comprising: Generating, by a second neural network, a second 3D video corresponding to the non-region-of-interest video; And Generating a 3D video corresponding to the video based on the first 3D video and the second 3D video.
10. An electronic device, comprising: At least one processor; And A memory coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the electronic device to perform actions, the actions including: Segment the image into a region of interest and a non - region of interest; Send data associated with the region of interest to the server; Receive a first three - dimensional (3D) model corresponding to the region of interest from the server, the first 3D model being generated by a first neural network at the server; Generate a second 3D model corresponding to the non - region of interest by a second neural network, the first neural network having more network parameters than the second neural network; and Generate a 3D model corresponding to the image based on the first 3D model and the second 3D model.
11. The electronic device according to claim 10, wherein the first neural network and the second neural network generate 3D models by simulating the propagation of light using volume rendering based on an image set captured from multiple angles.
12. The electronic device according to claim 10, wherein segmenting the image into the region of interest and the non - region of interest includes: Segmenting the region of interest and the non - region of interest by generating a binary mask based on user selection and identification.
13. The electronic device according to claim 10, wherein segmenting the image into the region of interest and the non - region of interest includes: Determining the region of interest by comparing one or more of color, texture, and edges in multiple different regions of the image.
14. The electronic device according to claim 10, wherein the first neural network is deployed in one or more GPUs in the server; The second neural network is deployed in one or more GPUs in the terminal device; and The one or more GPUs are dynamically scheduled and configured by a task manager based on one or more of the required data volume, task complexity, and hardware configuration factors for processing the region of interest and the non - region of interest.
15. The electronic device according to claim 10, wherein the first neural network is used for rendering and reconstruction of the region of interest, and the second neural network is used for rendering and reconstruction of the non - region of interest.
16. The electronic device according to claim 12, wherein the first neural network is also trained and adjusted based on the image region selected by the user, and the second neural network is trained and adjusted based on the image regions not selected by the user.
17. The electronic device according to claim 10, further comprising: Segment the video into a region - of - interest video and a non - region - of - interest video; And Receive a first 3D video corresponding to the region - of - interest video generated by the first neural network.
18. The electronic device according to claim 10, further comprising: Generate a second 3D video corresponding to the non - region - of - interest video by the second neural network; And Generate a 3D video corresponding to the video based on the first 3D video and the second 3D video.
19. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform a method, the method comprising: Segment the image into a region of interest and a non - region of interest; Send data associated with the region of interest to the server; Receiving, from the server, a first three-dimensional (3D) model corresponding to the region of interest, the first 3D model being generated by a first neural network at the server; Generating, by a second neural network, a second 3D model corresponding to the region of non-interest, the first neural network having more network parameters than the second neural network; And Generating, based on the first 3D model and the second 3D model, a 3D model corresponding to the image.
20. The computer program product according to claim 19, wherein the first neural network and the second neural network generate the 3D model by simulating the propagation of light using volume rendering based on an image set captured from multiple angles.