Scene construction method and electronic device

By combining multi-scale feature extraction and 3D convolutional networks, the problem of low accuracy in 3D scene reconstruction under sparse perspectives is solved, and high-precision scene reconstruction results are achieved.

CN122115703APending Publication Date: 2026-05-29HONOR DEVICE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-11-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, the accuracy and reliability of 3D scene construction under sparse observation perspective are low, resulting in inaccurate scene reconstruction results.

Method used

Feature extraction is performed on multi-view images using a multi-scale feature extraction network. The feature volume is aggregated using a 3D convolutional network with differentiable homography transformation and multi-scale structure to generate an aggregated feature volume based on the reference image view. Scene is then constructed by combining scene geometry and texture features.

Benefits of technology

It improves the reconstruction accuracy and robustness of new perspective images of scenes under sparse perspective, and realizes the gradual combination of geometric information from coarse to fine and from high dimension to low dimension, thereby improving the accuracy and robustness of scene geometric information reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115703A_ABST
    Figure CN122115703A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a scene construction method. In the method, a multi-scale feature extraction network is used to perform feature extraction on N input images of different perspectives in the same scene to obtain multi-scale feature maps of the input images. An arbitrary input image in the N input images is selected as a reference image, and the other images are taken as source images. The multi-scale feature maps of the input images are mapped into a perspective cone in which the reference image is located by using differentiable homography transformation to obtain multi-scale feature volumes based on the perspective of the reference image. A three-dimensional convolution network is used to aggregate the multi-scale feature volumes to obtain aggregated feature volumes based on the perspective of the reference image. Based on the aggregated feature volumes based on the perspective of the reference image, Gaussian sphere geometric properties corresponding to any scene position in the perspective of the reference image are determined as scene geometric features, and current scene construction is performed according to the scene geometric features matched with each perspective of the reference image to obtain a scene construction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to a scene construction method and an electronic device. Background Technology

[0002] High-precision 3D scene construction plays a crucial role in many applications, such as 3D video, augmented reality, 3D mapping, and autonomous driving. 3D scene construction based on multi-view images is an important research problem in the field of computer vision.

[0003] In related technologies, the accuracy of scene construction largely depends on the density of known viewpoint images. For 3D scene construction under sparse observation viewpoints, there are problems such as low scene construction accuracy and poor reliability of construction results. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a scene construction method. In this method, features are extracted from N input images from multiple perspectives within the same scene to obtain multi-scale feature maps for each input image. One of the N input images is randomly selected as a reference image, and the other input images are used as source images. Through differentiable homography transformation, the multi-scale feature maps of each input image are mapped onto the view frustum of the reference image to obtain a multi-scale feature volume based on the reference image's perspective. A 3D convolutional network with a multi-scale structure is used to aggregate the multi-scale feature volumes to obtain an aggregated feature volume based on the reference image's perspective. Based on the aggregated feature volume based on the reference image's perspective, scene geometric features matching the reference image's perspective are determined. Based on the scene geometric features matching each of the N reference image perspectives, the current scene is constructed to obtain the scene construction result.

[0005] By extracting multi-scale feature maps from N input images and aggregating them, it is beneficial to achieve a gradual combination of geometric information from coarse to fine manner and from high-dimensional to low-dimensional. It is also beneficial to optimize from global to local methods, which can gradually improve the robustness and accuracy of scene geometric information reconstruction and effectively improve the reconstruction accuracy of new perspective images of scenes under sparse perspectives.

[0006] In a first aspect, embodiments of this application provide a scene construction method, comprising: using a multi-scale feature extraction network to extract features from N input images from multiple perspectives within the same scene, obtaining multi-scale feature maps for each input image, wherein any input image among the N input images is selected as a reference image, and other input images besides the reference image are used as source images; mapping the multi-scale feature maps of each input image to the view frustum where the reference image is located through differentiable homography transformation, to obtain a multi-scale feature volume based on the reference image's perspective; aggregating the multi-scale feature volume using a three-dimensional convolutional network with a multi-scale structure, to obtain an aggregated feature volume based on the reference image's perspective; determining scene geometric features matching the reference image's perspective based on the aggregated feature volume based on the reference image's perspective; and constructing the current scene based on the scene geometric features matching each of the N reference image perspectives, to obtain a scene construction result.

[0007] According to the first aspect, the step of mapping the multi-scale feature maps of each input image to the frustum of the reference image through differentiable homography transformation to obtain a multi-scale feature body based on the viewpoint of the reference image includes: sampling the frustum of the reference image at equal intervals based on a preset depth interval to obtain multiple discretized sampling depth planes; and mapping the multi-scale feature maps of each input image to each sampling depth plane through differentiable homography transformation to obtain the multi-scale feature body based on the viewpoint of the reference image.

[0008] According to the first aspect, or any implementation thereof, the step of using a multi-scale three-dimensional convolutional network to aggregate the multi-scale feature volumes to obtain an aggregated feature volume based on the reference image viewpoint includes: using the multi-scale three-dimensional convolutional network to aggregate N smallest-scale feature volumes based on the reference image viewpoint to obtain an aggregated sub-feature volume at the current scale; and repeating the following operations in ascending order of feature volume scale until the aggregation of N largest-scale feature volumes based on the reference image viewpoint is completed: using the multi-scale three-dimensional convolutional network to aggregate the aggregated sub-feature volume at the current scale obtained from the previous aggregation operation and the N next-level scale feature volumes based on the reference image viewpoint; wherein, the aggregated sub-feature volume at the current scale obtained from the last aggregation operation constitutes the aggregated feature volume based on the reference image viewpoint.

[0009] According to the first aspect, or any implementation of the first aspect above, any scene position under the reference image view is mapped to a Gaussian sphere. The step of determining the scene geometric features matching the reference image view based on the aggregated feature volume based on the reference image view includes: inputting the aggregated feature volume into a multilayer perceptron to integrate global scene information and local scene information based on the reference image view to obtain the Gaussian sphere geometric attributes corresponding to any scene position under the reference image view, as the scene geometric features. The scene geometric features include at least one of the following parameters: opacity parameter, scale parameter, rotation parameter, and position parameter.

[0010] According to the first aspect, or any implementation of the first aspect above, the method further includes: fusing the multi-scale feature maps of each input image to obtain fused image features of the corresponding input image; determining global feature information matching the corresponding scene location based on the fused image features of each input image based on arbitrary scene location; and determining scene texture features matching the reference image viewpoint based on the global feature information matching the arbitrary scene location and the aggregated feature body matching the corresponding scene location based on the viewpoint of the reference image.

[0011] According to the first aspect, or any implementation of the first aspect above, the step of constructing the current scene based on the scene geometric features matched with each of the N reference image viewpoints to obtain the scene construction result includes: constructing the current scene based on the scene geometric features and the scene texture features matched with each of the reference image viewpoints to obtain the scene construction result.

[0012] According to the first aspect, or any implementation of the first aspect above, determining the feature global information matching the corresponding scene position based on the fused image features of each of the input images based on arbitrary scene positions includes: calculating the mean and variance of the image features of the fused image features of the N input images based on arbitrary scene positions, wherein the mean and variance of the image features constitute the feature global information matching the corresponding scene position.

[0013] According to the first aspect, or any implementation of the first aspect above, any scene position under the reference image view is mapped to a Gaussian sphere. The step of determining the scene texture features matching the reference image view based on the global feature information matching the arbitrary scene position and the aggregated feature body matching the corresponding scene position based on the reference image view includes: concatenating the mean of the image features matching the target scene position, the variance of the image features, and the aggregated feature body to obtain a concatenated feature based on the target scene position, where the target scene position includes any scene position of the current scene; inputting the concatenated feature into a multi-head attention mechanism to obtain a weight vector matching the target scene position; performing a weighted summation of the fused image features based on the target scene position of the N input images according to the weight vector to obtain a multi-view perception feature matching the target scene position; and inputting the multi-view perception feature based on the arbitrary scene position into a multilayer perceptron to obtain the Gaussian sphere texture attribute corresponding to any scene position under the reference image view, as the scene texture feature.

[0014] According to the first aspect, or any implementation of the first aspect above, the step of constructing the current scene based on the scene geometric features and scene texture features matched with the viewpoints of each of the reference images, and obtaining the scene construction result, includes: constructing N reference image viewpoint Gaussian spheres based on the scene geometric features and scene texture features based on the viewpoints of each of the reference images, the N reference image viewpoint Gaussian spheres constituting the current scene Gaussian sphere set; and performing Gaussian splashing differentiable rendering on the current scene Gaussian sphere set to obtain the scene construction result based on a preset viewpoint direction.

[0015] According to the first aspect, or any implementation of the first aspect above, the multi-scale feature maps of each input image are fused to obtain the fused image features of the corresponding input image, including: performing an upsampling operation on the feature maps of each scale of any input image to obtain multiple resampled feature maps of the same scale that match the corresponding input image; and concatenating the multiple resampled feature maps of the same scale to obtain the fused image features of the corresponding input image.

[0016] Secondly, embodiments of this application provide an electronic device, including: one or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and when the computer programs are executed by the one or more processors, the electronic device performs the following steps: using a multi-scale feature extraction network to extract features from N input images from multiple perspectives in the same scene, obtaining multi-scale feature maps of each input image, wherein any input image among the N input images is selected as a reference image, and other input images besides the reference image are used as source images; mapping the multi-scale feature maps of each input image to the view frustum where the reference image is located through differentiable homography transformation, to obtain a multi-scale feature volume based on the viewpoint of the reference image; using a three-dimensional convolutional network with a multi-scale structure to aggregate the multi-scale feature volume, obtaining an aggregated feature volume based on the viewpoint of the reference image; determining scene geometric features matching the viewpoint of the reference image based on the aggregated feature volume based on the viewpoint of the reference image; and constructing the current scene based on the scene geometric features matching each of the N reference image views, to obtain a scene construction result.

[0017] Thirdly, embodiments of this application provide a computer-readable storage medium including a computer program that, when run on an electronic device, causes the electronic device to execute instructions for the method as described in any possible implementation of the first aspect. Attached Figure Description

[0018] Figure 1 A schematic diagram illustrating the principles of scene construction is provided.

[0019] Figure 2 A schematic diagram of the structure of the electronic device 100 is shown.

[0020] Figure 3 A schematic diagram of the software structure of electronic device 100 is shown.

[0021] Figure 4 The schematic diagram illustrates a flowchart of a scene construction method according to an embodiment of this application;

[0022] Figure 5 This illustration schematically shows a process for determining scene geometric features according to an embodiment of this application;

[0023] Figure 6 The illustration shows a schematic diagram of a scene construction process according to an embodiment of this application;

[0024] Figure 7 The diagram illustrates another scenario construction process according to an embodiment of this application.

[0025] Figure 8 The diagram illustrates another scenario construction process according to an embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0028] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.

[0029] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0030] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.

[0031] This application provides a scene construction method, apparatus, and computer-readable storage medium. The scene construction apparatus can be integrated into an electronic device, which can be a terminal device or a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal device can be a mobile phone, computer, intelligent voice interaction device, smart home appliance, vehicle terminal, aircraft, etc.

[0032] The embodiments of this application can be applied to various scenarios, such as 3D camera movement, 3D video, artificial intelligence, augmented reality, 3D map construction, and assisted driving.

[0033] 3D scene construction based on multi-view images is an important research problem in the field of computer vision. Figure 1 This diagram illustrates the principles of scene construction. (For example...) Figure 1 As shown, during scene construction, sparse viewpoint images are input into the electronic device. These sparse viewpoint images include, for example, the i-th viewpoint image and the j-th viewpoint image. Based on these sparse viewpoint images, the electronic device can perform 3D reconstruction, new viewpoint generation, and target attribute editing on the current large-scale, boundless scene. Figure 1 As shown, electronic devices can construct 3D scenes based on sparse viewpoint images to obtain scene reconstruction videos with preset viewpoint directions.

[0034] The accuracy of scene reconstruction largely depends on the density of known viewpoint images. For 3D scene reconstruction under sparse observation viewpoints, the incompleteness of observation data can easily lead to the loss of information in reconstructing new viewpoints, which limits the reliability of scene reconstruction in practical applications.

[0035] This application proposes a scene construction scheme for use in electronic devices. In this scheme, the electronic device can utilize a multi-scale feature extraction network to extract features from N input images from multiple perspectives within the same scene, obtaining multi-scale feature maps for each input image. Any one of the N input images can be selected as a reference image, and the other input images besides the reference image can be used as source images. The electronic device can use a differentiable homography transformation to map the multi-scale feature maps of each input image onto the view frustum of the reference image, obtaining a multi-scale feature volume based on the reference image's perspective. Next, the electronic device can use a multi-scale 3D convolutional network to aggregate the multi-scale feature volumes, obtaining an aggregated feature volume based on the reference image's perspective. The electronic device can then determine the scene geometric features matching the reference image's perspective based on the aggregated feature volume, and construct the current scene based on the scene geometric features matching each of the N reference image perspectives, obtaining the scene construction result.

[0036] By extracting multi-scale feature maps from N input images and aggregating them, it is beneficial to achieve a gradual combination of geometric information from coarse to fine manner and from high-dimensional to low-dimensional. It is also beneficial to optimize from global to local methods, which can gradually improve the robustness and accuracy of scene geometric information reconstruction and effectively improve the reconstruction accuracy of new perspective images of scenes under sparse perspectives.

[0037] To better understand the embodiments of this application, the structure of the electronic device 100 of the embodiments of this application will be described below.

[0038] like Figure 2 This is a schematic diagram illustrating the structure of an electronic device 100 as an example. Figure 2 The diagram shows the structure of electronic device 100. Optionally, electronic device 100 can be referred to as a terminal or terminal device, and its specific product form can be a smart terminal, such as a mobile phone, tablet, digital video camera, smartwatch, smart wearable device, laptop computer, smart speaker, etc. Specifically, the functional modules involved in this application can be deployed on the DSP chip of the relevant device, specifically as application programs or software. A frequency response calibration function can be provided through software installation or upgrades, and through hardware calls and coordination.

[0039] It should be understood that, Figure 2 The electronic device 100 shown is only one example of an electronic device, and the electronic device 100 may have more or fewer components than shown in the figure, may combine two or more components, or may have different component configurations. Figure 2The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0040] Electronic device 100 may include: processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include pressure sensors, gyroscope sensors, accelerometers, temperature sensors, motion sensors, barometric pressure sensors, magnetic sensors, distance sensors, proximity sensors, fingerprint sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0041] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, memory, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0042] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0043] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory.

[0044] USB interface 130 is an interface that conforms to the USB standard specification, specifically it can be a Mini USB interface, Micro USB interface, USB Type C interface, etc.

[0045] The charging management module 140 receives charging input from a charger, which can be a wireless charger or a wired charger. While charging the battery 142, the charging management module 140 can also power the electronic device via the power management module 141. The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and powers the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc.

[0046] The wireless communication function of electronic device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.

[0047] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.

[0048] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.

[0049] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.

[0050] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.

[0051] The electronic device 100 can implement audio functions, such as voice communication functions and audio playback functions, through the speaker 171, earpiece (i.e. receiver) 172, microphone 173 and application processor in the audio module 170.

[0052] Electronic device 100 implements display functions through a GPU, display screen 194, and application processor. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0053] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, electronic device 100 may include one or N displays screens 194, where N is a positive integer greater than 1.

[0054] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0055] The ISP is used to process data fed back from the camera. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's image sensor. The light signal is converted into an electrical signal, and the image sensor transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye.

[0056] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0057] The camera 193 can be located at the edge of the electronic device, and can be an under-display camera or a pop-up camera. The camera 193 may include a rear-facing camera, or a rear-facing camera. This application embodiment does not limit the specific location and shape of the camera 193. The electronic device 100 may include one or more cameras with different focal lengths, such as telephoto cameras, wide-angle cameras, ultra-wide-angle cameras, or panoramic cameras.

[0058] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage functions.

[0059] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to illustrate the software structure of electronic device 100.

[0060] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to illustrate the software structure of electronic device 100.

[0061] like Figure 3 The software architecture diagram of the illustrative electronic device 100 illustrates a layered architecture that divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime, the system layer, and the kernel layer.

[0062] The application layer can include a series of application packages, such as Figure 3 As shown, the application package may include a gallery, maps, augmented reality (AR) applications, 3D videos, etc.

[0063] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer, including various components and services to support Android development. The application framework layer includes some predefined functions. For example... Figure 3 As shown, the application framework layer may include a window manager, a resource manager, a window management service, and a scene building service.

[0064] Window managers are used to manage windowed applications, for example. A window manager can obtain the screen size, determine if a status bar is present, lock the screen, capture screenshots, and perform other actions.

[0065] The file explorer can provide applications with various resources, such as localized strings, icons, images, layout files, video files, etc.

[0066] Window management services are used, for example, to assign and manage window attributes (such as hierarchy, size, display order, etc.) for applications. Window management services are also used to manage the display and switching of lock screen windows and wallpaper windows.

[0067] Scene building services, for example, are used to build the current scene based on multiple perspective input images of the same scene, and obtain the scene building result.

[0068] The system layer includes system libraries and the Android Runtime. System libraries can include multiple functional modules, such as image rendering libraries, image compositing libraries, function libraries, and media libraries. The image rendering library provides image processing functions to help meet various image processing needs.

[0069] The Android runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for scheduling and managing the Android system. The core libraries comprise two parts: one part contains the functionalities that Java calls, and the other part consists of the Android core libraries. The application layer and application framework layer run in the virtual machine, which executes the Java files of the application layer and application framework layer into binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0070] Understandable Figure 3 The components included in the system framework layer, system library, and runtime layer shown do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than shown, or combine some components, or split some components, or have different component arrangements.

[0071] The kernel layer is the layer between the hardware and the aforementioned software layers. The kernel layer includes at least display drivers, audio drivers, and sensor drivers. Hardware may include devices such as cameras, displays, microphones, processors, and memory.

[0072] Figure 4 The schematic diagram illustrates a flowchart of a scene construction method according to an embodiment of this application. For example... Figure 4 As shown, the scene construction method includes, for example, operations S101 to S105.

[0073] In operation S101, the electronic device uses a multi-scale feature extraction network to extract features from N input images from multiple perspectives within the same scene, obtaining multi-scale feature maps for each input image. Specifically, any one of the N input images is selected as a reference image, and all other input images are used as source images.

[0074] In operation S102, the electronic device maps the multi-scale feature maps of each input image to the view frustum where the reference image is located through differentiable homography transformation, so as to obtain a multi-scale feature volume based on the viewpoint of the reference image.

[0075] In operation S103, the electronic device uses a three-dimensional convolutional network with a multi-scale structure to aggregate multi-scale feature volumes, thereby obtaining aggregated feature volumes based on the reference image perspective.

[0076] In operation S104, the electronic device determines scene geometric features that match the reference image viewpoint based on the aggregated feature volume.

[0077] In operation S105, the electronic device constructs the current scene based on the scene geometric features matched with the viewpoints of each of the N reference images, and obtains the scene construction result.

[0078] The following is an illustrative explanation of the implementation process of each operation in the scene construction method.

[0079] In operation S101, the electronic device uses a multi-scale feature extraction network to extract features from N input images from multiple perspectives within the same scene, obtaining multi-scale feature maps for each input image. Specifically, any one of the N input images is selected as a reference image, and all other input images are used as source images.

[0080] For example, in response to acquiring N input images from multiple perspectives within the same scene, an electronic device utilizes a multi-scale feature extraction network to extract features from the N input images, obtaining multi-scale feature maps for each input image. For instance, a multi-scale feature extraction network formed by multiple concatenated convolutional layers can be used to extract features from each input image, resulting in a feature pyramid for each input image composed of feature maps at different scales.

[0081] For example, a multi-scale feature extraction network consisting of m concatenated convolutional layers can be used to perform convolution operations on each input image to extract key features such as color, texture, and shape, resulting in multi-scale feature maps for each input image. Specifically, when the multi-scale feature extraction network consists of m concatenated convolutional layers, the feature maps at different scales output by the m1, m2, m3, and m4th convolutional layers can, for example, constitute the multi-scale feature maps of the corresponding input images.

[0082] In operation S102, the electronic device maps the multi-scale feature maps of each input image to the view frustum where the reference image is located through differentiable homography transformation, so as to obtain a multi-scale feature volume based on the viewpoint of the reference image.

[0083] For example, the electronic device samples the view frustum of the reference image at equal intervals based on a preset depth interval, obtaining multiple discretized sampling depth planes. The electronic device then maps the multi-scale feature maps of each input image to each sampling depth plane through a differentiable homography transformation, obtaining a multi-scale feature body based on the viewpoint of the reference image.

[0084] The electronic device can move the reference image from the minimum depth d at preset depth intervals. min Mapped to maximum depth d max This yields a camera frustum containing multiple depth intervals. The electronic device then projects the multi-scale feature maps of each source image onto each sampled depth plane of the reference image frustum using a differentiable homography transformation, resulting in a multi-scale feature volume based on the reference image's viewpoint.

[0085] In one alternative approach, the electronic device maps the feature maps of the input image at various scales from their respective view coordinate systems to the view coordinate system of the reference image. During the coordinate transformation, given the camera's calibration matrix C = [K, R, t], where K represents the camera's intrinsic parameter matrix, and R and t represent the camera's rotation and translation matrices, respectively, the differentiable homography transformation process can be represented using equation (1):

[0086]

[0087] Among them, H j (d) represents the torsion matrix from the j-th view feature to the reference view at depth d, K j R represents the intrinsic parameter matrix of the camera relative to the j-th view. j and t j Let I and T represent the rotation and translation matrices of the camera relative to the j-th view feature, respectively, where I represents the identity matrix, T represents the transpose of the matrix, and n represents the translation matrix. i Indicates the principal axis direction of the reference view camera, t i K represents the translation matrix of the reference view. i This represents the camera intrinsic parameter matrix of the reference view.

[0088] Equation (2) can be used to represent the process of deforming each feature map in the reference view:

[0089] f jd (a,b)=f j (X j (d)[u,v,1] T Equation (2)

[0090] Among them, f jd (a,b) represents the twist feature map at depth d, and (u,v,1) represents the pixel position in the reference view.

[0091] In operation S103, the electronic device uses a three-dimensional convolutional network with a multi-scale structure to aggregate multi-scale feature volumes, thereby obtaining aggregated feature volumes based on the reference image perspective.

[0092] For example, an electronic device uses a multi-scale three-dimensional convolutional network to aggregate N smallest-scale feature volumes based on a reference image viewpoint, obtaining an aggregated sub-feature volume at the current scale. The electronic device repeats the following operation in ascending order of feature volume scale until the aggregation of the N largest-scale feature volumes based on the reference image viewpoint is completed: using a multi-scale three-dimensional convolutional network, the aggregated sub-feature volume at the current scale obtained from the previous aggregation operation is aggregated with the N next-level feature volumes based on the reference image viewpoint. The aggregated sub-feature volume at the current scale obtained from the final aggregation operation constitutes the aggregated feature volume based on the reference image viewpoint.

[0093] Optionally, the electronic device performs convolution processing on feature volumes at different scales based on the reference image viewpoint, resulting in a feature volume pyramid with scales decreasing from large to small. Then, the electronic device uses a top-down approach, employing a multi-scale 3D convolutional network, to aggregate the N smallest-scale feature volumes based on the reference image viewpoint, obtaining the smallest-scale aggregated sub-feature volume. The electronic device further utilizes a multi-scale 3D convolutional network to fuse the N second-smallest-scale feature volumes and the smallest-scale aggregated sub-feature volume, obtaining the second-smallest-scale aggregated sub-feature volume.

[0094] The electronic device repeats the following operations in ascending order of feature volume scale until the aggregation of the N largest-scale feature volumes based on the reference image viewpoint is completed, resulting in an aggregated feature volume based on the reference image viewpoint: A multi-scale 3D convolutional network is used to aggregate the current-scale aggregated sub-feature volume obtained from the previous aggregation operation and the N next-level scale feature volumes based on the reference image viewpoint. After aggregating the N largest-scale feature volumes and the second-largest-scale aggregated sub-feature volumes, the current-scale aggregated sub-feature volume obtained from this aggregation operation (i.e., the final aggregation operation) is used as the aggregated feature volume based on the reference image viewpoint. This aggregated feature volume constitutes the cost volume based on the reference image viewpoint.

[0095] In operation S104, the electronic device determines the scene geometric features that match the reference image viewpoint based on the aggregated feature volume based on the reference image viewpoint.

[0096] For example, any pixel in a multi-view image can be mapped to a corresponding Gaussian sphere. An electronic device can input an aggregated feature volume into a multilayer perceptron to integrate global and local scene information based on the reference image viewpoint, obtaining the Gaussian sphere geometric properties corresponding to any scene location under the reference image viewpoint, as scene geometric features. Scene geometric features may include at least one of the following parameters: opacity parameter, scale parameter, rotation parameter, and position parameter.

[0097] In operation S105, the electronic device constructs the current scene based on the scene geometric features matched with the viewpoints of each of the N reference images, and obtains the scene construction result.

[0098] This application embodiment extracts features from the input image at multiple scales and optimizes them step by step from coarse to fine and from high-dimensional to low-dimensional geometric information, i.e. from global to local. This can gradually improve the robustness and accuracy of the extracted scene geometric information.

[0099] Since different perspectives can reflect the geometric information of the current scene to a certain extent, differentiable homography transformation and 3D convolution operations can effectively fuse and diffuse multi-view feature information to generate a continuous geometric representation of the current scene. By making full use of the shared information between multiple perspectives, computational ambiguities in areas such as texture repetition, occlusion, lack of texture, and surface reflection in the scene reconstruction results can be effectively resolved.

[0100] Figure 5 The illustration shows a schematic diagram of the process for determining scene geometric features according to an embodiment of this application.

[0101] The following explanation uses multiple viewpoint images of the same scene, including the i-th viewpoint image and the j-th viewpoint image, as an example. The i-th viewpoint image is selected as the reference image, and the j-th viewpoint image constitutes the source image.

[0102] Electronic devices utilize a multi-scale feature extraction network to extract features from an image at the i-th viewpoint, obtaining a multi-scale feature map of that image. The low-dimensional feature map within the multi-scale feature map has high feature resolution and can better reflect the detailed features of the image at the i-th viewpoint. The high-dimensional feature map within the multi-scale feature map has lower feature resolution and can better reflect the global features of the image at the i-th viewpoint.

[0103] Similarly, electronic devices utilize multi-scale feature extraction networks to extract features from the j-th viewpoint image, obtaining a multi-scale feature map of the j-th viewpoint image. The low-dimensional feature map within the multi-scale feature map has high feature resolution and can better reflect the detailed features of the j-th viewpoint image. The high-dimensional feature map within the multi-scale feature map has lower feature resolution and can better reflect the global features of the j-th viewpoint image.

[0104] Electronic devices use differentiable homography transformation to map the multi-scale feature map of the j-th viewpoint image onto the view frustum containing the i-th viewpoint image, thereby obtaining a multi-scale feature volume based on the reference image viewpoint (e.g., ...). Figure 5 The warped twisted feature j is shown. Differentiable homography transformation refers to the process in plane geometry of mapping a set of points on one plane to a set of points on another plane. Assuming the depth interval of the camera's view frustum in the i-th viewpoint image is [d1, d2], and the depth resolution is Δd, we can obtain D = (d2-d1) / Δd depth planes, each depth plane d corresponding to a homography transformation matrix X.

[0105] The electronic device uses the homography transformation matrix X(i) corresponding to the depth plane d(i) to transform the multi-scale feature map of the j-th viewpoint image, obtaining the corresponding feature value of each pixel in the camera pose of the i-th viewpoint image. Assuming there are D depth planes, D transformed feature maps corresponding to each scale feature map can be obtained. Each transformed feature map represents the corresponding feature value of the pixel when its true depth is the current depth. The D transformed feature maps constitute a feature body matching the corresponding scale feature map. By mapping the multi-scale feature maps of the source image from their respective view coordinate systems to the camera coordinate system of the reference image, it can be ensured that all feature maps are aligned to the same coordinate system.

[0106] Electronic devices can also use differentiable homography transformation to convert the multi-scale feature map of the i-th viewpoint image into transformed feature maps corresponding to D depth planes. The D transformed feature maps constitute a feature body matching the corresponding scale feature map (e.g., ...). Figure 5 The feature body i is shown. Since the i-th viewpoint image is selected as the reference image, when performing homography transformation on the multi-scale feature map of the i-th viewpoint image, the homography transformation matrix X corresponding to the D depth planes is the identity matrix.

[0107] Electronic devices utilize a multi-scale 3D convolutional network to aggregate multi-scale feature volumes corresponding to the i-th and j-th viewpoint images, obtaining aggregated feature volumes based on the reference image viewpoint. For example, the electronic device uses a 3D convolutional network to aggregate the smallest-scale feature volume i and the warped feature volume j, obtaining a smallest-scale aggregated sub-feature volume. Then, using the 3D convolutional network, it aggregates the smallest-scale aggregated sub-feature volume, the second-smallest-scale feature volume i, and the second-smallest-scale warped feature volume j, obtaining a second-smallest-scale aggregated sub-feature volume.

[0108] The electronic device repeats the following operation in ascending order of feature volume scale until the aggregation of the N largest-scale feature volumes based on the reference image viewpoint is completed: A 3D convolutional network is used to aggregate the current-scale aggregated sub-feature volume obtained from the previous aggregation operation and the N next-level-scale feature volumes based on the reference image viewpoint. The current-scale aggregated sub-feature volume obtained from the final aggregation operation constitutes the i-th viewpoint aggregated feature volume, which is the i-th viewpoint cost volume.

[0109] The electronic device inputs the aggregated feature volume of the i-th viewpoint into the multilayer perceptron (MLP) to integrate the global and local scene information of the i-th viewpoint and obtain the scene geometric features of the i-th viewpoint. The scene geometric features include, for example, the opacity parameter α, the scale parameter S, the rotation parameter R, and the position parameter X.

[0110] By utilizing multi-head attention mechanisms to learn the dependencies between different features, we can make full use of the correlation between images from different perspectives. This is beneficial for improving the geometric shape and appearance features of the scene construction results, enhancing the detail and quality of the reconstructed scene, effectively improving the reconstruction accuracy of new perspective images under sparse perspectives, and effectively mitigating the impact of observation noise on the reconstruction effect.

[0111] Figure 6 The illustration shows a schematic diagram of a scene construction process according to an embodiment of this application.

[0112] like Figure 6 As shown, N input images based on different scenes within the same scenario are input into the scene construction unit of an electronic device. These N input images include, for example, an image from the i-th viewpoint and an image from the j-th viewpoint. Any one of the N input images is selected as a reference image, and all other images are used as source images. The scene construction unit is located, for example, in the application framework layer of the electronic device's software architecture layer. The scene construction unit may include a multi-scale feature extraction module, a scene geometric feature determination module, a scene texture feature determination module, and a scene rendering module.

[0113] A multi-scale feature extraction module is used to extract features from N input images, resulting in multi-scale feature maps for each input image. A scene geometric feature determination module is then used to determine the scene geometric features based on the viewpoint of a reference image, using the multi-scale feature maps of each input image.

[0114] As an optional approach, a scene geometry feature determination module is used to perform differentiable homography transformation on the multi-scale feature maps of each input image. This maps the multi-scale feature maps of each input image onto the camera's view frustum containing the reference image, resulting in a multi-scale feature volume based on the reference image's perspective. A 3D convolutional network within the scene geometry feature determination module is then used to aggregate the multi-scale feature volumes, resulting in an aggregated feature volume based on the reference image's perspective—the cost volume under the reference image's perspective. Finally, the scene geometry feature determination module inputs the cost volume under the reference image's perspective into a multilayer perceptron to integrate global and local scene information based on the reference image's perspective, obtaining scene geometry features that match the reference image's perspective.

[0115] The scene texture feature determination module uses multi-scale feature maps of each input image to determine scene texture features based on the viewpoint of the reference image. For example, based on the multi-scale feature maps of each input image, it predicts the color attribute information of the Gaussian sphere in the current scene, that is, determines the spherical harmonics of the Gaussian sphere in the current scene.

[0116] As an optional approach, a scene texture feature determination module is used to fuse multi-scale feature maps of each input image to obtain fused image features for the corresponding input image. The scene texture feature determination module then uses these fused image features based on arbitrary scene locations to determine global feature information matching the corresponding scene location. Furthermore, the scene texture feature determination module uses the global feature information matching arbitrary scene locations and the aggregated feature body matching the corresponding scene location based on the reference image's viewpoint to determine scene texture features matching the reference image's viewpoint.

[0117] For example, the scene texture feature determination module is used to upsample the feature maps of any input image at various scales to obtain multiple resampled feature maps of the same scale that match the corresponding input image. The multiple resampled feature maps of the same scale are then concatenated to obtain the fused image features of the corresponding input image.

[0118] For example, for the input image The feature maps at three different scales are upsampled to obtain resampled feature maps at scales of h×w×C1, h×w×C2, and h×w×C3, respectively. The resampled feature maps at scales of h×w×C1, h×w×C2, and h×w×C3 are then concatenated to obtain a fused image feature map at scale h×w×(C1+C2+C3).

[0119] Using the scene texture feature determination module, the fused image features of N input images are calculated based on the mean and variance of image features at any scene location. The mean and variance of image features constitute global feature information that matches the corresponding scene location.

[0120] For example, for any scene location (u,v) in the current scene, the scene texture feature determination module calculates the mean image feature of the fused image features of N input images at scene location (u,v). and image feature variance

[0121] The scene texture feature determination module concatenates the mean, variance, and aggregated feature volume of the image features matching the target scene location to obtain concatenated features based on the target scene location, which includes any scene location within the current scene. This concatenated feature is then input into a multi-head attention mechanism to obtain a weight vector matching the target scene location. Next, based on the determined weight vector, the fused image features of N input images based on the target scene location are weighted and summed to obtain multi-view perceptual features matching the target scene location. Finally, these multi-view perceptual features based on arbitrary scene locations are input into a multilayer perceptron to obtain the Gaussian sphere texture attribute corresponding to any scene location from the reference image's viewpoint, which serves as the scene texture feature.

[0122] For example, for the target scene position (u,v) in the current scene, the scene texture feature determination module inputs the concatenated features into a multi-head attention mechanism to obtain a weight vector {ω1,…,ω...} that matches the target scene position (u,v). N Using the weight vector {ω1,…,ω} N}, fuse image features of N input images at target scene location (u,v) We perform weighted summation to obtain multi-view perception features that match the target scene location.

[0123] By inputting multi-view perceptual features matched to arbitrary scene locations into a multilayer perceptron, scene texture features based on the reference image's viewpoint are obtained. For example, by inputting multi-view perceptual features matched to arbitrary scene locations into a multilayer perceptron, the spherical harmonic coefficients c of the Gaussian sphere of the current scene based on the reference image's viewpoint are obtained.

[0124] As an optional approach, the scene rendering module is used to construct the current scene based on the scene's geometric and texture features matched with the viewpoints of each reference image, resulting in a scene construction result. For example, based on the scene's geometric and texture features from each reference image viewpoint, N reference image viewpoint Gaussian spheres are constructed, and these N reference image viewpoint Gaussian spheres constitute the current scene Gaussian sphere set. The scene rendering module then performs Gaussian splash differentiable rendering on the current scene Gaussian sphere set to obtain the scene construction result based on a preset viewpoint direction.

[0125] Because color is viewpoint-dependent, the color representation of the same scene location may differ depending on changes in viewpoint and lighting conditions. Weighted fusion of multi-view feature information can improve scene construction quality and enhance the stability and robustness of scene representation. It can also effectively solve the problem of flickering between frames when generating new viewpoint video sequences.

[0126] Figure 7 The diagram illustrates another scenario construction process according to an embodiment of this application.

[0127] This explanation uses multiple viewpoint images of the same scene, including the i-th viewpoint image and the j-th viewpoint image, as an example. The i-th viewpoint image is selected as the reference image, and the j-th viewpoint image constitutes the source image. For details on determining the scene's geometric features, please refer to [link to documentation / documentation]. Figure 5 The textual descriptions are omitted here.

[0128] The electronic device fuses the multi-scale feature maps of the i-th viewpoint image and the j-th viewpoint image respectively to obtain the fused image features of the corresponding viewpoint image. For example, the electronic device upsamples the feature maps of each scale of the i-th viewpoint image and the j-th viewpoint image respectively to obtain multiple resampled feature maps of the same scale that match the corresponding viewpoint image. Then, the multiple resampled feature maps of the same scale are concatenated to obtain the fused image features of the corresponding viewpoint image.

[0129] Taking the image from the i-th viewpoint as an example, for the image from the i-th viewpoint... The feature maps at three different scales are upsampled to obtain resampled feature maps at scales of h×w×C1, h×w×C2, and h×w×C3, respectively. The resampled feature maps at scales of h×w×C1, h×w×C2, and h×w×C3 are then concatenated to obtain a fused image feature map at scale h×w×(C1+C2+C3).

[0130] Electronic devices calculate the fused image features of the i-th and j-th viewpoint images based on the mean image features at any scene location (u,v). and image feature variance Image feature mean and image feature variance This constitutes global feature information that matches the corresponding scene position (u,v).

[0131] It is worth noting that the same location in the current scene may correspond to different pixel locations in images from different viewpoints. Taking the i-th viewpoint image and the j-th viewpoint image as an example, the top corner of the door frame in the current scene corresponds to different pixel locations in the i-th viewpoint image and the j-th viewpoint image.

[0132] Mean value of image features based on target scene location (u,v) by electronic device Image feature variance Aggregate feature The features are concatenated to obtain concatenated features based on the target scene location (u,v). These concatenated features are then input into a multi-head attention mechanism (MHA), and the MHA output is input again into a multilayer perceptron (MLP) to obtain a weight vector [ω] that matches the target scene location (u,v). i ω j ].

[0133] The electronic device is based on the determined weight vector [ω] i ω j ], fused image features based on target scene location (u,v) for the i-th view image and the j-th view image. Weighted summation is performed to obtain multi-view perceptual features matching the target scene position (u,v). The multi-view perceptual features based on arbitrary scene positions are input into a multilayer perceptron (MLP) to obtain the spherical harmonic coefficient component c of the Gaussian sphere of the scene in the i-th view, which is the scene texture feature in the i-th view.

[0134] Figure 8 The diagram illustrates another scenario construction process according to an embodiment of this application.

[0135] like Figure 8 As shown, during the construction of the current scene, a sparse viewpoint RGB image / video sequence is input to the electronic device. The sparse viewpoint RGB images include, for example, the i-th viewpoint image and the j-th viewpoint image.

[0136] For example, the electronic device inputs the i-th viewpoint image into the FPN structural feature extraction network to obtain a multi-scale feature map of the i-th viewpoint image. Similarly, the electronic device inputs the j-th viewpoint image into the FPN structural feature extraction network to obtain a multi-scale feature map of the j-th viewpoint image.

[0137] The electronic device upsamples the feature maps at each scale of the i-th viewpoint image to obtain multiple resampled feature maps of the same scale that match the i-th viewpoint image. Then, these multiple resampled feature maps of the same scale are stitched together to obtain the 2D fused image features of the i-th viewpoint image. The electronic device can use a similar method to determine the 2D fused image features of the j-th viewpoint image, which will not be elaborated upon here.

[0138] Electronic devices can use the identity matrix as a homography transformation matrix to transform the feature maps at each scale of the i-th viewpoint image, obtaining the i-th viewpoint feature volume corresponding to each scale feature map. Similarly, electronic devices can also use the identity matrix as a homography transformation matrix to transform the feature maps at each scale of the i-th viewpoint image, obtaining the j-th viewpoint feature volume corresponding to each scale feature map.

[0139] When the i-th viewpoint image is used as the reference image, homography transformation can be performed on the j-th viewpoint feature volume to obtain the transformed j-th viewpoint feature volume. Electronic devices can utilize a multi-level 3D convolutional network to aggregate the i-th viewpoint feature volumes at various scales and the transformed j-th viewpoint feature volumes to obtain the i-th viewpoint aggregated feature volume, i.e., the i-th viewpoint cost volume.

[0140] Electronic devices can input the cost volume of the i-th viewpoint into a multilayer perceptron (MLP) to obtain the scene geometric features under the i-th viewpoint, such as the opacity component, scale component, rotation component, and depth component under the i-th viewpoint.

[0141] The electronic device can determine the 2D fused image features of the i-th viewpoint image based on the mean and variance of the image features at the target scene location. Then, the electronic device can concatenate the mean and variance of the image features matched with the target scene location and the cost volume of the i-th viewpoint to obtain the concatenated features based on the target scene location, which can include any scene location in the current scene.

[0142] The electronic device can input the stitched features into a multi-head attention mechanism (MHA), and then input the MHA output into a multilayer perceptron (MLP) to obtain a weight vector matching the target scene location. Based on the determined weight vector, the electronic device performs a weighted summation of the 2D fused image features from each viewpoint, based on the target scene location, to obtain multi-view perceptual features matching the target scene location. The electronic device can also input multi-view perceptual features based on arbitrary scene locations into the MLP to obtain scene texture features at the i-th viewpoint, such as the spherical harmonic coefficient components at the i-th viewpoint.

[0143] The electronic device constructs a Gaussian sphere from the i-th viewpoint based on the scene's geometric and texture features.

[0144] Similarly, when the j-th viewpoint image is used as the reference image, homography transformation can be performed on the i-th viewpoint feature volume to obtain the transformed i-th viewpoint feature volume. Electronic devices can utilize multi-level 3D convolutional networks to aggregate the j-th viewpoint feature volumes at various scales and the transformed i-th viewpoint feature volumes to obtain the j-th viewpoint aggregated feature volume, i.e., the j-th viewpoint cost volume.

[0145] Electronic devices can input the cost volume of the j-th viewpoint into a multilayer perceptron (MLP) to obtain the scene geometric features under the j-th viewpoint, such as the opacity component, scale component, rotation component, and depth component under the j-th viewpoint.

[0146] The electronic device can determine the 2D fused image features of the j-th viewpoint image based on the mean and variance of image features at the target scene location. Then, the electronic device can concatenate the mean and variance of image features matched with the target scene location and the j-th viewpoint cost volume to obtain the concatenated features based on the target scene location, which can include any scene location within the current scene.

[0147] The electronic device can input the stitched features into a multi-head attention mechanism (MHA), and then input the MHA output into a multilayer perceptron (MLP) to obtain a weight vector matching the target scene location. Based on the determined weight vector, the electronic device performs a weighted summation of the 2D fused image features from each viewpoint, based on the target scene location, to obtain multi-view perceptual features matching the target scene location. The electronic device can also input multi-view perceptual features based on arbitrary scene locations into the MLP to obtain scene texture features at the j-th viewpoint, such as the spherical harmonic coefficient components at the j-th viewpoint.

[0148] The electronic device constructs a Gaussian sphere from the j-th viewpoint based on the scene's geometric and textural features.

[0149] A set of Gaussian spheres from multiple perspectives can be constructed to represent the current scene. Electronic devices can perform Gaussian splashing differentiable rendering on this set of Gaussian spheres based on a preset viewpoint direction to obtain a new perspective RGB image. In the Gaussian splashing algorithm, the current scene can be represented as a 3D point set, where each point is described by a corresponding 3D Gaussian distribution. This 3D Gaussian distribution includes parameters such as position (x, y, z), covariance, opacity, and color. Electronic devices can fit the 3D Gaussian distribution of each point to obtain matching current scene data. Furthermore, electronic devices can perform image rendering based on the current scene data, for example, integrating the geometric properties and spherical harmonic coefficients of each 3D Gaussian distribution into the image rendering result to obtain the scene construction result for the current scene.

[0150] In the Gaussian splashing algorithm, a visible sparse point cloud can be estimated from multi-view images. For each point in the point cloud, a Gaussian ellipsoidal probabilistic prediction model, similar to a scattering field, can be constructed. Learning is performed through a neural network model to obtain the probability parameters corresponding to each ellipsoid (the probability parameters indicate whether an object exists on the corresponding ellipsoid), thus obtaining a discrete representation similar to volume pixels to support multi-angle volume rendering and rasterization. 3D reconstruction based on 3D Gaussian splashing technology enables real-time scene rendering and explicit editing.

[0151] It is understood that, in order to achieve the above-mentioned functions, electronic devices include hardware and / or software modules that perform the respective functions. Based on the algorithmic steps of the examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of this application.

[0152] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0153] This embodiment also provides an electronic device, including: one or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and when the computer programs are executed by one or more processors, the electronic device performs the following steps: using a multi-scale feature extraction network to extract features from N input images from multiple perspectives in the same scene, obtaining multi-scale feature maps of each input image, wherein any input image among the N input images is selected as a reference image, and the other input images besides the reference image are used as source images; mapping the multi-scale feature maps of each input image to the view frustum where the reference image is located through differentiable homography transformation, to obtain a multi-scale feature volume based on the viewpoint of the reference image; using a three-dimensional convolutional network with a multi-scale structure to aggregate the multi-scale feature volume, obtaining an aggregated feature volume based on the viewpoint of the reference image; determining scene geometric features that match the viewpoint of the reference image based on the aggregated feature volume based on the viewpoint of the reference image; and constructing the current scene based on the scene geometric features that match the viewpoints of each of the N reference image views, to obtain a scene construction result.

[0154] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement the scene construction method in the above embodiment.

[0155] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory. The memory is used to store computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory to cause the chip to execute the scene construction method in the above method embodiments.

[0156] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.

[0157] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0158] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0159] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0160] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0161] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.

[0162] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0163] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0164] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium well known in the art. One exemplary embodiment couples a storage medium to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0165] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0166] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A scene construction method, characterized in that, include: A multi-scale feature extraction network is used to extract features from N input images from multiple perspectives in the same scene to obtain multi-scale feature maps of each input image. In this process, any input image among the N input images is selected as a reference image, and the other input images besides the reference image are used as source images. By using differentiable homography transformation, the multi-scale feature maps of each input image are mapped to the frustum where the reference image is located, so as to obtain a multi-scale feature body based on the viewpoint of the reference image; The multi-scale feature volume is aggregated using a three-dimensional convolutional network with a multi-scale structure to obtain an aggregated feature volume based on the viewpoint of the reference image. Based on the aggregated feature body based on the reference image viewpoint, determine the scene geometric features matching the reference image viewpoint; and Based on the scene geometric features matched with each of the N reference image viewpoints, the current scene is constructed to obtain the scene construction result.

2. The method according to claim 1, characterized in that, The step of mapping the multi-scale feature maps of each input image to the view frustum where the reference image is located through differentiable homography transformation to obtain a multi-scale feature volume based on the viewpoint of the reference image includes: Based on a preset depth interval, the view frustum of the reference image is sampled at equal intervals to obtain multiple discretized sampling depth planes; and By using differentiable homography transformation, the multi-scale feature maps of each input image are mapped to each sampling depth plane to obtain the multi-scale feature volume based on the viewpoint of the reference image.

3. The method according to claim 1, characterized in that, The method of using a multi-scale three-dimensional convolutional network to aggregate the multi-scale feature volumes to obtain an aggregated feature volume based on the reference image viewpoint includes: Using the aforementioned multi-scale 3D convolutional network, N minimum-scale feature volumes based on the reference image viewpoint are aggregated to obtain an aggregated sub-feature volume at the current scale; and The following operations are repeated in ascending order of feature volume scale until the aggregation of the N largest-scale feature volumes based on the reference image viewpoint is completed: using the multi-scale structured 3D convolutional network, the aggregated sub-feature volumes at the current scale obtained from the previous aggregation operation and the N next-level feature volumes based on the reference image viewpoint are aggregated. The aggregated sub-features obtained from the last aggregation operation at the current scale constitute the aggregated feature based on the reference image perspective.

4. The method according to claim 1, characterized in that, Mapping any scene location from the reference image perspective to a Gaussian sphere, and determining scene geometric features matching the reference image perspective based on the aggregated feature volume, includes: The aggregated feature volume is input into a multilayer perceptron to integrate global and local scene information based on the reference image viewpoint, thereby obtaining the Gaussian sphere geometric properties corresponding to any scene location under the reference image viewpoint, which are then used as the scene geometric features. The scene geometric features include at least one of the following parameters: opacity parameter, scale parameter, rotation parameter, and position parameter.

5. The method according to claim 1, characterized in that, The method further includes: The multi-scale feature maps of each input image are fused to obtain the fused image features of the corresponding input image; Based on the fused image features of each input image at an arbitrary scene location, global feature information matching the corresponding scene location is determined; and Based on the global feature information matching any scene location and the aggregated feature body matching the corresponding scene location based on the reference image viewpoint, the scene texture features matching the reference image viewpoint are determined.

6. The method according to claim 5, characterized in that, The step of constructing the current scene based on the scene geometric features matched with each of the N reference image viewpoints, and obtaining the scene construction result, includes: Based on the scene geometric features and scene texture features that match the viewpoints of each of the reference images, the current scene is constructed to obtain the scene construction result.

7. The method according to claim 5, characterized in that, The step of determining global feature information matching the corresponding scene location based on the fused image features of each input image at an arbitrary scene location includes: The fused image features of the N input images are calculated based on the mean and variance of image features at any scene location. The mean and variance of image features constitute the global feature information that matches the corresponding scene location.

8. The method according to claim 7, characterized in that, Mapping any scene location from the reference image perspective to a Gaussian sphere, and determining the scene texture features matching the reference image perspective based on the global feature information matching the arbitrary scene location and the aggregated feature body matching the corresponding scene location based on the reference image perspective, includes: The mean of the image features, the variance of the image features, and the aggregated feature body that match the target scene location are concatenated to obtain the concatenated features based on the target scene location, where the target scene location includes any scene location of the current scene; The concatenated features are input into a multi-head attention mechanism to obtain a weight vector that matches the target scene location; Based on the weight vector, the fused image features of the N input images based on the target scene location are weighted and summed to obtain multi-view perception features matching the target scene location; and The multi-view perception features based on arbitrary scene locations are input into a multilayer perceptron to obtain the Gaussian sphere texture attributes corresponding to arbitrary scene locations under the viewpoint of the reference image, which are then used as the scene texture features.

9. The method according to claim 6, characterized in that, The step of constructing the current scene based on the scene geometric features and scene texture features matched with the viewpoints of each of the reference images, and obtaining the scene construction result, includes: Based on the scene geometric features and scene texture features according to the viewpoints of each of the reference images, N reference image viewpoint Gaussian spheres are constructed, and the N reference image viewpoint Gaussian spheres constitute the current scene Gaussian sphere set; and Gaussian splashing differentiable rendering is performed on the current scene Gaussian sphere set to obtain the scene construction result based on the preset view direction.

10. The method according to claim 5, characterized in that, The multi-scale feature maps of each input image are fused to obtain the fused image features of the corresponding input image, including: Upsampling is performed on feature maps of any input image at various scales to obtain multiple resampled feature maps of the same scale that match the corresponding input image; and The multiple resampled feature maps of the same scale are stitched together to obtain the fused image features of the corresponding input image.

11. An electronic device, characterized in that, include: One or more processors, memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and when the computer programs are executed by the one or more processors, cause the electronic device to perform the following steps: A multi-scale feature extraction network is used to extract features from N input images from multiple perspectives in the same scene to obtain multi-scale feature maps of each input image. In this process, any input image among the N input images is selected as a reference image, and the other input images besides the reference image are used as source images. By using differentiable homography transformation, the multi-scale feature maps of each input image are mapped to the frustum where the reference image is located, so as to obtain a multi-scale feature body based on the viewpoint of the reference image; The multi-scale feature volume is aggregated using a three-dimensional convolutional network with a multi-scale structure to obtain an aggregated feature volume based on the viewpoint of the reference image. Based on the aggregated feature body based on the reference image viewpoint, determine the scene geometric features matching the reference image viewpoint; and Based on the scene geometric features matched with each of the N reference image viewpoints, the current scene is constructed to obtain the scene construction result.

12. A computer-readable storage medium, characterized in that, The method includes a computer program that, when run on an electronic device, causes the electronic device to perform the scene construction method as described in any one of claims 1 to 10.