Methods, apparatus, electronic devices and storage media for generating depth maps
By combining the N-step phase-shifting method with the target model and fusing multi-view image information, the problem of increased depth map acquisition and generation time was solved, achieving efficient and accurate depth map generation, and improving the inspection efficiency and 3D reconstruction quality of the SMT production line.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JABIL INC
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, the multi-step phase-shifting method in SMT 3D AOI systems leads to increased depth map acquisition and generation time, reducing detection efficiency.
An N-step phase-shifting method is used to acquire initial depth maps, average brightness maps, and surface brightness maps from multiple viewpoints. These image information are then fused using a target model to generate a high-quality depth map identical to that obtained using the M-step phase-shifting method, simplifying the acquisition and processing workflow.
It improves the efficiency and accuracy of depth map generation, shortens acquisition time, and enhances the inspection efficiency and 3D reconstruction quality of SMT production lines.
Smart Images

Figure CN122089955A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of optical detection technology, specifically relating to a method, apparatus, electronic device, and storage medium for generating depth maps. Background Technology
[0002] With the miniaturization and multi-functionality of electronic products, surface mount technology (SMT) has been widely used in the electronics manufacturing industry. To ensure manufacturing quality and improve production efficiency on SMT production lines, automated optical inspection (AOI) technology has become a crucial element. Compared to traditional 2D AOI, 3D AOI technology, by combining automated optical inspection with 3D imaging technology, can acquire the 3D contours and surface features of components, achieving high-precision inspection at the micron level, and is gradually becoming an industry trend. Currently, one typical optical imaging scheme in SMT 3D AOI systems uses a single camera with dual telecentric lenses, combined with multiple digital light processing (DLP) optomechanical systems at different angles for structured light 3D measurement.
[0003] To obtain high signal-to-noise ratio and high-precision single-view phase measurement results, multi-step phase-shifting methods are commonly used for 3D reconstruction. The accuracy and anti-interference capability of the multi-step phase-shifting method are positively correlated with the number of projection steps; to improve the signal-to-noise ratio and measurement accuracy, the number of projection steps needs to be increased. However, increasing the number of projection steps directly leads to each DLP optical engine needing to project more images, and correspondingly, the camera needs to simultaneously acquire more images. This results in a significant increase in the time for depth map acquisition and generation, reducing the efficiency of depth map acquisition and generation, and consequently reducing the inspection efficiency of the SMT production line. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and storage medium for generating depth maps, which can solve the problem of reduced efficiency in depth map acquisition and generation.
[0005] In a first aspect, embodiments of this application provide a method for generating a depth map. The method includes: obtaining a first initial depth map of a first target object from multiple viewpoints using an N-step phase-shifting method; obtaining a first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple viewpoints; obtaining a first surface brightness map of the first target object from multiple viewpoints; inputting the first initial depth map, the first average brightness map, and the first surface brightness map from multiple viewpoints into a target model, and outputting a corresponding first target depth map through the target model. The first depth information of the first target depth map is the same as the second depth information of the second target depth map of the first target object obtained by an M-step phase-shifting method, wherein M is greater than N.
[0006] Secondly, embodiments of this application provide a depth map generation apparatus, comprising: a first acquisition module, configured to acquire a first initial depth map of a first target object from multiple viewpoints using an N-step phase-shifting method; a second acquisition module, configured to acquire a first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple viewpoints; a third acquisition module, configured to acquire a first surface brightness map of the first target object from multiple viewpoints; and a generation module, configured to input the first initial depth map, the first average brightness map, and the first surface brightness map from multiple viewpoints into a target model, and output a corresponding first target depth map through the target model, wherein the first depth information of the target depth map is the same as the second depth information of the second target depth map of the first target object acquired by an M-step phase-shifting method, and M is greater than N.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0009] In this embodiment, an initial depth map of a first target object from multiple viewpoints is obtained using an N-step phase-shifting method; a first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple viewpoints is obtained; a first surface brightness map of the first target object from multiple viewpoints is obtained; the first initial depth map, the first average brightness map, and the first surface brightness map from multiple viewpoints are input into a target model, and the target model outputs a corresponding first target depth map. The first depth information of the first target depth map is the same as the second depth information of the second target depth map of the first target object obtained using an M-step phase-shifting method, where M is greater than N. This improves the generation efficiency and accuracy of the depth map and avoids potential loss of accuracy and efficiency due to reduced acquisition and processing. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic flowchart of a method for generating a depth map provided in an embodiment of this application; Figure 2 This is a schematic diagram of the architecture of a target model provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a depth map generation device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0014] The following description, in conjunction with the accompanying drawings, details the method, apparatus, electronic device, and storage medium for generating depth maps provided in this application, through specific embodiments and application scenarios.
[0015] Figure 1 This application illustrates a method for generating a depth map according to an embodiment of the present application. This method can be performed by an electronic device. In other words, the method can be installed by an electronic device... The method, executed by the device's software or hardware, includes the following steps: Step S101: Obtain the first initial depth map of the first target object from multiple perspectives using the N-step phase shift method.
[0016] A typical optical imaging scheme for SMT 3D AOI is to use a single camera with dual telecentric lenses and multiple DLP optical engines at different angles to perform fringe projection 3D reconstruction. The first initial depth map of the first target object is obtained from multiple viewpoints by using multiple DLP optical engines at different angles.
[0017] Digital Light Processing (DLP) optical engine: This is a device that uses digital micromirror devices (DMDs) to project patterns (such as stripes). In 3D AOI, it is used to project structured light patterns onto the surface of the target object to generate depth maps for 3D reconstruction.
[0018] Structured Light 3D Reconstruction: An optical 3D measurement technique that calculates the 3D information of an object by projecting a known light pattern (usually stripes) onto the object's surface and using a camera to observe the deformation of the pattern caused by the object's surface shape.
[0019] Phase Shifting / Phase Measuring Method: A method in fringe projection techniques that projects multiple (usually N steps) fringe patterns with phase shifts and analyzes the resulting images to calculate the phase value of each pixel, thereby obtaining depth information of the object. This process includes calculating the wrapped phase and the unwrapped phase.
[0020] Wrapped Phase: A phase value directly calculated using the phase method, typically limited to a single period (e.g., 0 to 2π). It contains object height information but is ambiguous and requires further processing.
[0021] Phase Unwrapping: The process of unfolding the wrapped phase to recover the continuous, unambiguous absolute phase, which is a key step in calculating the true depth (height) of an object from phase information.
[0022] Phase to Depth Conversion: The process of converting the unwrapped phase value into the actual physical depth or height value based on the geometry of the optical system.
[0023] Depth Map: A two-dimensional image representing the distance (depth) from each pixel in the scene to the camera.
[0024] In this embodiment, the depth map generation process for each DLP optical engine typically involves calculating the wrapping phase and unwrapping phase using an N-step phase method, followed by the phase-to-depth conversion. This allows multiple DLP cameras to acquire initial depth maps of the first target object from multiple viewpoints. Specifically, multiple (e.g., four) DLP optical engines can be used to project structured light patterns (such as stripes) onto the surface of the target object (e.g., a PCB board) from different preset angles. This multi-angle projection design aims to overcome problems such as occlusion, shadows, and secondary reflection interference caused by surface reflection that may exist in single-view measurement, thereby obtaining more comprehensive information about the object's surface.
[0025] Using one or more industrial cameras and a specific optical system (such as a dual telecentric lens), deformable pattern images projected onto the surface of an object by different DLP optical engines are captured simultaneously. Subsequently, for the image sequence acquired from each DLP optical engine's viewpoint, an N-step phase-shifting method is used to process the images, calculating the wrapping phase, performing unwrapping operations, and finally converting the phase information into a first initial depth map under that viewpoint. This allows for the acquisition of first initial depth maps of the first target object from multiple viewpoints. In this embodiment, N is a positive integer, such as 3, 4, or 5, and the corresponding N-step phase-shifting method can be a 3-step, 4-step, or 5-step phase-shifting method, etc.
[0026] Step S102: Obtain the first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple viewpoints.
[0027] In this embodiment, the reflection intensity of the N-step phase-shifting projection pattern on the first target object can also be obtained from multiple viewing angles. The reflection intensity can be obtained from the first average brightness map synchronously acquired by the camera when the projection pattern of multiple DLP optical engine projection stripes is projected onto the first target object. This first average brightness map can reflect the reflection intensity of the projection pattern on the surface of the first target object.
[0028] Step S103: Obtain the first surface brightness map of the first target object from multiple perspectives.
[0029] In this embodiment, a first surface brightness map of a first target object can be captured from multiple angles using a conventional light source (non-structured light) of an AOI system. This first surface brightness map reflects the inherent reflective properties of the object's surface.
[0030] Step S104: Input the first initial depth map, the first average brightness map, and the first surface brightness map from multiple perspectives into the target model, and output the corresponding first target depth map through the target model.
[0031] In this embodiment, the first initial depth map, first average brightness map, and first surface brightness map of the first target object from multiple viewpoints can be used as inputs to the target model. In the target model, a first target depth map corresponding to the first initial depth map, first average brightness map, and first surface brightness map of the first target object from multiple viewpoints can be generated. The first depth information of the first target depth map is the same as the second depth information of the second target depth map of the first target object obtained through the M-step phase-shifting method, where M is a positive integer and M is greater than N. For example, the N-step phase-shifting method can be a 3-step phase-shifting method, and the M-step phase-shifting method can be a 12-step phase-shifting method, or even a 3-step phase-shifting method and a 15-step phase-shifting method, etc. In this embodiment, the target model can be a deep learning neural network model. The goal of this model is to directly predict the high-quality first target depth map generated by the more complex 12-step phase-shifting + HDR + traditional fusion process based on the simplified data (first initial depth map, first average brightness map, first surface brightness map) obtained from N-step phase shifting + single exposure.
[0032] The depth map generation method provided in this application embodiment obtains a first initial depth map of a first target object from multiple viewpoints using an N-step phase-shifting method; obtains a first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple viewpoints; obtains a first surface brightness map of the first target object from multiple viewpoints; inputs the first initial depth map, the first average brightness map, and the first surface brightness map from multiple viewpoints into a target model, and outputs a corresponding first target depth map through the target model. The first depth information of the first target depth map is the same as the second depth information of the second target depth map of the first target object obtained by an M-step phase-shifting method, where M is greater than N. This method can significantly reduce the number of phase-shifting steps of each DLP optical engine from M steps to N steps, eliminates the need for multiple HDR exposures, greatly shortens the time required for image acquisition, simplifies the depth map generation steps for a single viewpoint, and replaces the computationally complex traditional depth fusion algorithm with an efficient target model, significantly reducing the overall algorithm processing time. The target model can still output a high-quality, high-precision first target depth map even with simplified input data, avoiding the accuracy and efficiency loss that may result from reducing the acquisition and processing processes, and improving the generation efficiency and accuracy of the depth map.
[0033] In one embodiment, the step of outputting the corresponding first target depth map through the target model includes: fusing and encoding the first initial depth map, the first average brightness map, and the first surface brightness map under each viewpoint through the target model to generate multiple first feature maps representing the first initial depth map, the first average brightness map, and the first surface brightness map; performing feature fusion on the multiple first feature maps to generate a second feature map; and converting the second feature map into the first target depth map through the target model.
[0034] Figure 2 This application provides an embodiment of a target model with a schematic diagram of its architecture. Figure 2 As shown, the architecture of the target model includes multiple single DLP feature fusion modules. This module takes three types of initial depth maps, first average brightness maps, and first surface brightness maps with different sources and properties as inputs. Through a network structure composed of a series of typical convolutional layers (Conv), batch normalization layers (BN), and ReLU activation functions, it extracts, fuses, and encodes these input features to generate a first feature map that can represent the comprehensive information of the first initial depth map, first average brightness map, and first surface brightness map under this specific DLP perspective.
[0035] Since there are typically multiple DLP optical engines at different angles (e.g., four), this module will run in parallel multiple times (e.g., four times), processing the first initial depth map, first average brightness map, and first surface brightness map acquired from one DLP each time. To ensure the generalization ability and parameter efficiency of the target model, instances of this module processing different DLP inputs share the same set of network weights (Weight Sharing). This means that the feature extraction and fusion rules learned by the model are universal and applicable to processing similar input patterns from different perspectives.
[0036] like Figure 2 As shown, the target model architecture also includes a multi-DLP feature fusion module. This module is responsible for integrating the first feature maps from all DLP optical engines at multiple perspectives to generate a global second feature map that incorporates multi-view information. The second feature map can then be converted into a first target depth map using the target model.
[0037] In one embodiment, the step of fusing features from multiple first feature maps to generate a second feature map includes: merging multiple first feature maps along the channel dimension to generate a third feature map; and fusing features from the third feature map using a preset encoder and decoder to generate the second feature map.
[0038] In this embodiment, the multi-DLP feature fusion module receives a first feature map output from the previous single-DLP feature fusion module. First, through a channel concatenation operation, these multiple first feature maps are merged along the channel dimension to form a third feature map containing information from all viewpoints. Next, to effectively fuse this information and obtain a larger receptive field for understanding spatial relationships and contextual information over a wider range, this module employs a typical U-Net architecture encoder and decoder to perform feature fusion on the third feature map, generating the second feature map. The U-Net encoder (downsampling path) captures high-level semantic information and global context of the image, while the decoder (upsampling path), combined with skip connections, helps recover spatial details. Through this structure, the target model can learn how to effectively integrate, denoise, and complete features from different viewpoints.
[0039] The choice of the U-Net architecture allows the target model to utilize both global and local information during the fusion process, which is crucial for handling occlusion, inconsistencies, and noise in multi-view data. Skip connections help preserve and pass on fine structural information from earlier layers, which is very important for the boundary and detail accuracy of the final depth map.
[0040] like Figure 2 As shown, the architecture of the target model also includes a depth map prediction module, which is the final output layer of the target model. This module is responsible for directly decoding (or mapping) the second feature map after multi-view fusion into the final, high-quality first target depth map.
[0041] This depth map prediction module receives a second feature map output from the multi-DLP feature fusion module (the last layer of the U-Net decoder). Its structure typically consists of one or a few convolutional layers, and its core function is to convert the high-dimensional feature representation into a single-channel depth value matrix corresponding to the input image size, i.e., the final first target depth map.
[0042] In the embodiments of this application, by leveraging the powerful feature extraction and nonlinear mapping capabilities of deep learning, the target model can more intelligently handle problems such as multi-view information fusion, noise suppression, and invalid value repair, achieving faster processing speed and better robustness compared to traditional algorithms.
[0043] In one embodiment, before obtaining the initial depth map of the first target object from multiple viewpoints using the N-step phase-shifting method, the method further includes: using the second initial depth map, the second average brightness map, and the second surface brightness map of the second target object obtained by the N-step phase-shifting method as input data of the target model, and using the third target depth map of the second target object obtained by the M-step phase-shifting method as output data of the target model to train the target model, so that the target model outputs the first target depth map based on the first initial depth map, the first average brightness map, and the first surface brightness map.
[0044] In this embodiment, training the target model requires a large amount of input and output data. The input data consists of a second initial depth map, a second average brightness map, and a second surface brightness map, obtained through a simplified process (N-step phase-shifting method + single exposure) for each DLP viewpoint, and preliminarily processed. The output data is a second target depth map, considered realistic or of high quality, generated using a high-quality process (M-step phase-shifting method + HDR exposure + traditional depth fusion algorithm). During the training phase, a large amount of different PCB board image data needs to be collected as both input and output data.
[0045] Loss Function: The core objective of training the target model is to make the target depth map predicted by the target model using the N-step phase shift method as close as possible to the target depth map predicted by the M-step phase shift method. The loss function used is a combination of the L2 and L1 losses between the target depth maps predicted by the N-step and M-step phase shift methods. This loss function directly measures the numerical difference between the two target depth maps at each pixel, and drives the network learning by minimizing this difference.
[0046] Where H represents the height of the target depth map, and W represents the width of the target depth map. Represents the pixels of the target model in the target depth map The predicted depth value, This indicates the pixel point The corresponding actual depth value.
[0047] The combined loss function is defined as a weighted sum of the L2 loss and the L1 loss: Ltotal = α × LL² + (1 α)×LL1 Here, α is a hyperparameter used to balance the contributions of L2 loss and L1 loss (e.g., α = 0.5).
[0048] The depth map generation method provided in this application introduces a carefully designed target model and cleverly combines multimodal feature fusion (depth, brightness) with multi-view information integration (U-Net architecture). It successfully optimizes the traditionally time-consuming "M-step phase shift + HDR + traditional fusion" process into a fast "N-step phase shift + single exposure + deep learning fusion" process. This significantly shortens the SMT 3D AOI detection time while improving the generation efficiency of depth maps, thereby enhancing the accuracy, quality, and efficiency of 3D reconstruction.
[0049] It should be noted that the depth map generation method provided in this application embodiment can be executed by a depth map generation device or a control module within that device for executing the depth map generation method. This application embodiment uses the execution of the depth map generation method by a depth map generation device as an example to illustrate the depth map generation device provided in this application embodiment.
[0050] Figure 3 This is a schematic diagram of the structure of a depth map generation apparatus according to an embodiment of this application. Figure 3 As shown, the depth map generation device 300 includes: a first acquisition module 310, a second acquisition module 320, a third acquisition module 330, and a generation module 340.
[0051] The first acquisition module 310 is used to acquire a first initial depth map of a first target object from multiple viewpoints using an N-step phase-shifting method; the second acquisition module 320 is used to acquire a first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple viewpoints; the third acquisition module 330 is used to acquire a first surface brightness map of the first target object from multiple viewpoints; and the generation module 340 is used to input the first initial depth map, the first average brightness map, and the first surface brightness map from multiple viewpoints into a target model, and output a corresponding first target depth map through the target model. The first depth information of the target depth map is the same as the second depth information of the second target depth map of the first target object acquired by an M-step phase-shifting method, where M is greater than N.
[0052] In one embodiment, the generation module 340 is configured to: fuse and encode the first initial depth map, the first average brightness map, and the first surface brightness map under each viewpoint using the target model to generate a plurality of first feature maps representing the first initial depth map, the first average brightness map, and the first surface brightness map; perform feature fusion on the plurality of feature maps to generate a second feature map; and convert the second feature map into the first target depth map using the target model.
[0053] In one embodiment, the generation module 340 is configured to: merge multiple first feature maps in the channel dimension to generate a third feature map; and perform feature fusion on the third feature map using a preset encoder and decoder to generate a second feature map.
[0054] In one embodiment, the generation module 340 is further configured to: use the second initial depth map, the second average brightness map, and the second surface brightness map of the second target object obtained by the N-step phase shift method as input data of the target model, and use the third target depth map of the second target object obtained by the M-step phase shift method as output data of the target model to train the target model, so that the target model outputs the first target depth map based on the first initial depth map, the first average brightness map, and the first surface brightness map.
[0055] The depth map generation device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0056] The depth map generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0057] The depth map generation apparatus provided in this application embodiment can achieve Figures 1 to 2 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0058] Optionally, such as Figure 4As shown in the illustration, this application embodiment also provides an electronic device 400, including a processor 401 and a memory 402. The memory 402 stores a program or instructions that can run on the processor 401. When the program or instructions are executed by the processor 401, they perform the following: obtaining a first initial depth map of a first target object from multiple viewing angles using an N-step phase-shifting method; obtaining a first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple viewing angles; obtaining a first surface brightness map of the first target object from multiple viewing angles; inputting the first initial depth map, the first average brightness map, and the first surface brightness map from multiple viewing angles into a target model; and outputting a corresponding first target depth map from the target model. The first depth information of the first target depth map is the same as the second depth information of the second target depth map of the first target object obtained by an M-step phase-shifting method, where M is greater than N.
[0059] In one embodiment, the first initial depth map, the first average brightness map, and the first surface brightness map under each viewpoint are fused and encoded by the target model to generate a plurality of first feature maps representing the first initial depth map, the first average brightness map, and the first surface brightness map; the plurality of feature maps are fused to generate a second feature map; and the second feature map is converted into the first target depth map by the target model.
[0060] In one embodiment, multiple first feature maps are merged along the channel dimension to generate a third feature map; the third feature map is then fused using a preset encoder and decoder to generate a second feature map.
[0061] In one embodiment, before acquiring the first initial depth map of the first target object from multiple viewpoints using the N-step phase-shifting method, the second initial depth map, the second average brightness map, and the second surface brightness map of the second target object acquired using the N-step phase-shifting method are used as input data for the target model, and the third target depth map of the second target object acquired using the M-step phase-shifting method is used as output data to train the target model, so that the target model outputs the first target depth map based on the first initial depth map, the first average brightness map, and the first surface brightness map.
[0062] The specific execution steps can be found in the various steps of the above-described method embodiment for generating depth maps, and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0063] It should be noted that the electronic devices in the embodiments of this application include: servers, terminals, or other devices besides terminals.
[0064] The above electronic device structure does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or arrange them differently. For example, an input unit may include a Graphics Processing Unit (GPU) and a microphone, and a display unit may use a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar display panels. User input units include at least one of a touch panel and other input devices. A touch panel is also called a touchscreen. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be elaborated further here.
[0065] Memory can be used to store software programs and various data. Memory can primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area can store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, memory can include volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0066] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly handles operations related to the operating system, user interface, and applications, while the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor.
[0067] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described depth map generation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0068] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk.
[0069] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0070] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0071] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for generating a depth map, characterized in that, include: The first initial depth map of the first target object from multiple perspectives is obtained by the N-step phase shift method; Obtain the first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple perspectives; Obtain the first surface brightness map of the first target object from multiple perspectives; The first initial depth map, the first average brightness map, and the first surface brightness map from multiple perspectives are input into the target model, and the target model outputs the corresponding first target depth map. The first depth information of the first target depth map is the same as the second depth information of the second target depth map of the first target object obtained by the M-step phase shift method, where M is greater than N.
2. The method according to claim 1, characterized in that, The step of outputting the corresponding first target depth map through the target model includes: The target model is used to fuse and encode the first initial depth map, the first average brightness map, and the first surface brightness map under each viewpoint to generate multiple first feature maps representing the first initial depth map, the first average brightness map, and the first surface brightness map. The multiple feature maps are fused to generate a second feature map; The second feature map is converted into the first target depth map using the target model.
3. The method according to claim 2, characterized in that, The step of fusing features from multiple first feature maps to generate a second feature map includes: Multiple first feature maps are merged along the channel dimension to generate a third feature map; The second feature map is generated by fusing features from the third feature map using a preset encoder and decoder.
4. The method according to claim 1, characterized in that, Before obtaining the first initial depth map of the first target object from multiple viewpoints using the N-step phase-shifting method, the method further includes: The second initial depth map, second average brightness map, and second surface brightness map of the second target object obtained by the N-step phase shift method are used as input data for the target model. The third target depth map of the second target object obtained by the M-step phase shift method is used as output data for training the target model, so that the target model outputs the first target depth map based on the first initial depth map, the first average brightness map, and the first surface brightness map.
5. A depth map generation apparatus, characterized in that, include: The first acquisition module is used to acquire the first initial depth map of the first target object from multiple perspectives using the N-step phase-shifting method; The second acquisition module is used to acquire the first average brightness map of the projection pattern of the N-step phase-shifting method on the first target object from multiple perspectives. The third acquisition module is used to acquire the first surface brightness map of the first target object from multiple perspectives; The generation module is used to input the first initial depth map, the first average brightness map, and the first surface brightness map from multiple perspectives into the target model, and output the corresponding first target depth map through the target model. The first depth information of the target depth map is the same as the second depth information of the second target depth map of the first target object obtained by the M-step phase shift method, and M is greater than N.
6. The apparatus according to claim 5, characterized in that, The generation module is used for: The target model is used to fuse and encode the first initial depth map, the first average brightness map, and the first surface brightness map under each viewpoint to generate multiple first feature maps representing the first initial depth map, the first average brightness map, and the first surface brightness map. The multiple feature maps are fused to generate a second feature map; The second feature map is converted into the first target depth map using the target model.
7. The apparatus according to claim 6, characterized in that, The generation module is used for: Multiple first feature maps are merged along the channel dimension to generate a third feature map; The second feature map is generated by fusing features from the third feature map using a preset encoder and decoder.
8. The apparatus according to claim 5, characterized in that, The generation module is further configured to: The second initial depth map, second average brightness map, and second surface brightness map of the second target object obtained by the N-step phase shift method are used as input data for the target model. The third target depth map of the second target object obtained by the M-step phase shift method is used as output data for training the target model, so that the target model outputs the first target depth map based on the first initial depth map, the first average brightness map, and the first surface brightness map.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the depth map generation method as described in any one of claims 1-4.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the depth map generation method as described in any one of claims 1-4.