Method and apparatus for asymmetric normalized correlation layers for deep neural network feature matching

By using two image sensors and an asymmetric normalized correlation layer on a mobile electronic device to generate a depth map, the problem that the camera on a mobile device cannot blur a specific depth of field is solved, achieving high-quality defocus effects and image detail enhancement.

CN114830176BActive Publication Date: 2025-09-12SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080079531.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-28
Filing Date
2020-11-04
Publication Date
2025-09-12
Estimated Expiration
2040-11-04

AI Technical Summary

Technical Problem

Cameras on mobile electronic devices are typically unable to selectively blur parts of an image outside a specific depth of field, resulting in degraded image quality and making it difficult for existing machine learning algorithms to effectively create a bokeh effect.

Method used

An image of a scene is captured using two image sensors of an electronic device to generate a feature map, and the spatial resolution is restored through an asymmetric normalized correlation layer and a reversible wavelet layer to generate a depth map, thereby achieving feature matching and Boke effect of the image.

Benefits of technology

Improves the quality of images on mobile electronic devices, can create a soft bokeh effect, enhance the depth of field effect of images, and improve the quality and details of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114830176B_ABST
    Figure CN114830176B_ABST
Patent Text Reader

Abstract

A method includes obtaining a first image of a scene using a first image sensor of an electronic device, and obtaining a second image of the scene using a second image sensor of the electronic device. The method also includes generating a first feature map from the first image and generating a second feature map from the second image. The method also includes generating a third feature map based on the first feature map, the second feature map, and an asymmetric search window. The method also includes generating a depth map by restoring spatial resolution of the third feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to image capture systems and more particularly to asymmetric normalized correlation layers for deep neural network feature matching. Background Art

[0002] Many mobile electronic devices, such as smartphones and tablets, include cameras that can be used to capture both still and video images. While convenient, cameras on mobile electronic devices often have a number of drawbacks that reduce their image quality. Various machine learning algorithms can be used in many image processing-related applications to improve the quality of images captured using mobile electronic devices or other devices. For example, different neural networks can be trained and then used to perform different image processing tasks to improve the quality of captured images. As a specific example, a neural network can be trained and used to blur specific portions of a captured image. Summary of the Invention

[0003] Solution

[0004] A method includes obtaining a first image of a scene using a first image sensor of an electronic device, and obtaining a second image of the scene using a second image sensor of the electronic device. The method also includes generating a first feature map from the first image and generating a second feature map from the second image. The method also includes generating a third feature map based on the first feature map, the second feature map, and an asymmetric search window. The method additionally includes generating a depth map by restoring spatial resolution to the third feature map. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein like reference numerals represent like parts:

[0006] Figure 1 illustrates an example network configuration including electronic devices according to the present disclosure;

[0007] Figure 2A 、 Figure 2B and Figure 2C illustrates an example input image and example processing results that may be obtained using an asymmetric normalized correlation layer in a neural network according to the present disclosure;

[0008] Figure 3 illustrates an example neural network architecture according to the present disclosure;

[0009] Figure 4 illustrates a detailed example of a neural network including an asymmetric normalized correlation layer according to the present disclosure;

[0010] Figure 5illustrates an example application of a reversible wavelet layer of a neural network according to the present disclosure;

[0011] Figure 6A and Figure 6B illustrates an example asymmetric search window used in an asymmetric normalized correlation layer and an example application of an asymmetric normalized correlation layer according to the present disclosure; and

[0012] Figure 7 Illustrated is an example method for deep neural network feature matching using an asymmetric normalized correlation layer according to the present disclosure. DETAILED DESCRIPTION

[0013] Best Mode

[0014] This disclosure provides an asymmetric normalized correlation layer for deep neural network feature matching.

[0015] In a first embodiment, a method includes obtaining a first image of a scene using a first image sensor of an electronic device, and obtaining a second image of the scene using a second image sensor of the electronic device. The method also includes generating a first feature map from the first image and a second feature map from the second image. The method also includes generating a third feature map based on the first feature map, the second feature map, and an asymmetric search window. Furthermore, the method includes generating a depth map by restoring spatial resolution to the third feature map.

[0016] In a second embodiment, an electronic device includes a first image sensor, a second image sensor, and at least one processor operably coupled to the first image sensor and the second image sensor. The at least one processor is configured to obtain a first image of a scene using the first image sensor and to obtain a second image of the scene using the second image sensor. The at least one processor is also configured to generate a first feature map from the first image and a second feature map from the second image. The at least one processor is further configured to generate a third feature map based on the first feature map, the second feature map, and an asymmetric search window. Furthermore, the at least one processor is configured to generate a depth map by restoring spatial resolution to the third feature map.

[0017] In a third embodiment, a non-transitory machine-readable medium includes instructions that, when executed, cause at least one processor of an electronic device to obtain a first image of a scene using a first image sensor of the electronic device and to obtain a second image of the scene using a second image sensor of the electronic device. The medium also includes instructions that, when executed, cause the at least one processor to generate a first feature map from the first image and a second feature map from the second image. The medium further includes instructions that, when executed, cause the at least one processor to generate a third feature map based on the first feature map, the second feature map, and an asymmetric search window. Furthermore, the medium includes instructions that, when executed, cause the at least one processor to generate a depth map by restoring spatial resolution to the third feature map.

[0018] Other technical features will be apparent to those skilled in the art from the following drawings, description, and claims.

[0019] Before proceeding with the detailed description below, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms "send," "receive," and "communicate," and their derivatives, include direct and indirect communications. The terms "include" and "comprises," and their derivatives, mean including, but not limited to, such. The term "or" is inclusive, meaning and / or. The phrase "associated with," and its derivatives, means including, included within, interconnected with, contains, contained within, connected to or connected with, coupled to or coupled with, communicating with, cooperating with, interleaved with, juxtaposed with, proximate to, incorporated into or combined with, having the nature of, relating to, and the like.

[0020] In addition, the various functions described below can be implemented or supported by one or more computer programs, each of which is formed of a computer-readable program code and contained in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or parts thereof that are suitable for implementation in a suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read-only memory (ROM), random access memory (RAM), a hard drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. "Non-transitory" computer-readable media does not include wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Non-transitory computer-readable media include media that can permanently store data and media that can store data and rewrite it later, such as rewritable optical discs or erasable storage devices.

[0021] As used herein, terms and phrases such as "having", "can have", "include", or "can include" a feature (such as a number, function, operation, or component, such as a part) indicate the presence of the feature and do not exclude the presence of other features. In addition, as used herein, the phrases "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", and "at least one of A or B" may mean all (1) include at least one A, (2) include at least one B, or (3) include at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components, regardless of their importance, and do not limit these components. These terms are only used to distinguish different components. For example, a first user device and a second user device may indicate user devices that are different from each other, regardless of the order or importance of the devices. Without departing from the scope of the present disclosure, a first component may be represented as a second component, and vice versa.

[0022] It should be understood that when an element (such as a first element) is referred to as being “coupled to” another element (such as a second element), “coupled to” another element (such as a second element), or “connected to” another element (such as a second element), it can be coupled or connected to the other element directly or through a third element. Conversely, it should be understood that when an element (such as a first element) is referred to as being “directly coupled to” or “directly connected to” another element (such as a second element), or “directly connected to” another element (such as a second element), there are no other elements (such as a third element) between the element and the other element.

[0023] As used herein, the phrase “configured (or set) to” may be used interchangeably with the phrases “suitable for,” “capable of,” “designed to,” “applicable to,” “manufactured to,” or “capable of,” as the case may be. The phrase “configured (or set) to” does not inherently mean “specially designed in hardware to.” On the contrary, the phrase “configured to” may mean that one device can perform an operation together with another device or component. For example, the phrase “a processor configured (or set) to perform A, B, and C” may refer to a general-purpose processor (such as a CPU or application processor) that can perform operations by executing one or more software programs stored in a storage device, or may refer to a dedicated processor (such as an embedded processor) for performing operations.

[0024] The terms and phrases used herein are intended only to describe some embodiments of the present disclosure and are not intended to limit the scope of other embodiments of the present disclosure. It should be understood that the singular forms "a", "an" and "the" include plural references unless the context clearly dictates otherwise. All terms and phrases used herein, including technical and scientific terms and phrases, have the same meanings as those generally understood by those of ordinary skill in the art to which the embodiments of the present disclosure belong. It should also be understood that terms and phrases, such as those defined in commonly used dictionaries, should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and unless clearly defined herein, will not be interpreted in an idealized or overly formal sense. In some cases, the terms and phrases defined herein may be interpreted to exclude embodiments of the present disclosure.

[0025] Examples of “electronic devices” according to embodiments of the present disclosure may include at least one of a smartphone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, etc.) Other examples of electronic devices include smart home appliances. Examples of smart home appliances may include a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave oven, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), a smart speaker or a speaker with an integrated digital assistant (such as SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a game console (such as XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camcorder, or an electronic photo frame. Other examples of electronic devices include at least one of various medical devices (such as various portable medical measuring devices (e.g., blood glucose measuring devices, heart rate measuring devices, or body temperature measuring devices), magnetic resonance angiography (MRA) devices, magnetic resonance imaging (MRI) devices, computed tomography (CT) devices, imaging devices, or ultrasound devices), navigation devices, global positioning system (GPS) receivers, event data recorders (EDRs), flight data recorders (FDRs), automotive infotainment devices, marine electronic devices (such as marine navigation devices or gyrocompasses), avionics equipment, security devices, vehicle head units, industrial or domestic robots, automated teller machines (ATMs), point-of-sale (POS) devices, or Internet of Things (IoT) devices (such as light bulbs, various sensors, electricity or gas meters, sprinklers, fire alarms, thermostats, streetlights, toasters, fitness equipment, hot water tanks, heaters, or boilers). Other examples of electronic devices include at least a portion of a piece of furniture or a building / structure, an electronic board, an electronic signature receiving device, a projector, or various measuring devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). Note that, according to various embodiments of the present disclosure, the electronic device may be one or a combination of the above-mentioned devices. According to some embodiments of the present disclosure, the electronic device may be a flexible electronic device. The electronic devices disclosed herein are not limited to the devices listed above and may include new electronic devices as technology develops.

[0026] In the following description, an electronic device is described with reference to the accompanying drawings according to various embodiments of the present disclosure. As used herein, the term "user" may refer to a person using the electronic device or another device (such as an artificial intelligence electronic device).

[0027] Definitions for other specific words and phrases may be provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most instances, such definitions apply to prior and future uses of such defined words and phrases.

[0028] Nothing in this application should be read as implying that any particular element, step, or function is essential to the scope of the claims. The scope of the patented subject matter is limited solely by the claims.

[0029] Invention Mode

[0030] Discussed below Figures 1 to 7 Various embodiments of the present disclosure are described with reference to the accompanying drawings. However, it should be understood that the present disclosure is not limited to these embodiments, and all variations and / or equivalents or alternatives thereof also fall within the scope of the present disclosure. Throughout the specification and drawings, the same or similar reference symbols may be used to refer to the same or similar elements.

[0031] As mentioned above, many mobile electronic devices, such as smartphones and tablets, include cameras that can be used to capture both still and video images. However, cameras on mobile electronic devices typically have a number of disadvantages compared to digital single-lens reflex (DSLR) cameras. For example, DSLR cameras can create a soft focus effect (also known as a bokeh effect) due to changes in the depth of field (DoF) of the captured image. The bokeh effect can be created by using a lens with a wide aperture in a DSLR camera, which causes a soft or blurry appearance outside of a specific depth of field at which objects in the image are in focus. Cameras on mobile electronic devices are typically not able to selectively blur portions of an image outside of a specific depth of field because most cameras on mobile electronic devices generate images in which the entire image is in focus.

[0032] Various machine learning algorithms can be used in many applications related to image processing, including applications that generate the Boke effect computationally (rather than optically) in images captured using mobile electronic devices or other devices. For example, different neural networks can be trained and used to perform different image processing tasks to improve the quality of captured images. Each neural network is typically trained to perform a specific task. For example, in the field of image processing, different neural networks can be trained to identify the type of scene or objects in the scene, identify the depth of objects in the scene, segment an image based on objects in the scene, or generate high dynamic range (HDR) images, Boke images, or super-resolution images.

[0033] Embodiments of the present disclosure describe various techniques for creating a boke effect and other image processing effects in images captured using a mobile electronic device or other device. As described in more detail below, a synthetic graphics engine can be used to generate training data with specific characteristics. The synthetic graphics engine is used to generate training data customized for a specific mobile electronic device or other device. Evaluation methods can be used to test the quality of depth maps (or disparity maps), which can be generated by a neural network trained using the training data. The depth map or disparity map can be used to identify depth in a scene, which (in some cases) allows more distant portions of the scene image to be computationally blurred to provide a boke effect. In some embodiments, a wavelet synthesis neural network (WSN) architecture can be used to generate high-definition depth maps. To generate high-definition depth maps, the WSN architecture includes a reversible wavelet layer and a normalized correlation layer. The reversible wavelet layer is applied to iteratively decompose and synthesize feature maps, and the normalized correlation layer is used for robust dense feature matching that is coupled to camera specifications, including baseline distances between multiple cameras and calibration accuracy when calibrating images from multiple cameras.

[0034] Additional details about the neural network architecture including the asymmetric normalization layer are provided below. It should be noted here that although the feature maps generated based on the reversible wavelet layer and the asymmetric normalization layer are generally described as being used to perform specific image processing tasks, the neural network architecture provided in this disclosure is not limited to use for these specific image processing tasks or for general image processing. Instead, the asymmetric normalization layer of the neural network can be used in any suitable system to perform feature matching.

[0035] Figure 1 An example network configuration 100 including electronic devices according to the present disclosure is illustrated. Figure 1 The embodiment of the network configuration 100 shown is for illustration only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.

[0036] According to an embodiment of the present disclosure, an electronic device 101 is included in a network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, or one or more sensors 180. In some embodiments, the electronic device 101 may exclude at least one of these components, or may add at least one other component. The bus 110 includes circuits for connecting the components 120-180 to each other and for transmitting communications (such as control messages and / or data) between the components.

[0037] The processor 120 includes one or more of a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), or a communication processor (CP). The processor 120 is capable of controlling at least one other component of the electronic device 101 and / or performing communication-related operations or data processing. In some embodiments, the processor 120 processes image data using a neural network architecture to perform feature matching using a reversible wavelet layer and an asymmetric normalized correlation layer to generate a single feature map from multiple images of a scene. This can support various image processing functions, such as creating a Bokeh effect in an image.

[0038] The memory 130 can include volatile and / or non-volatile memory. For example, the memory 130 can store commands or data related to at least one other component of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 can store software and / or programs 140. The programs 140 include, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or "application") 147. At least a portion of the kernel 141, middleware 143, or application programming interface 145 can be represented as an operating system (OS).

[0039] The kernel 141 is capable of controlling or managing system resources (such as the bus 110, processor 120, or memory 130) used to perform operations or functions implemented in other programs (such as middleware 143, application programming interface 145, or application 147). The kernel 141 provides an interface that allows the middleware 143, application programming interface 145, or application 147 to access various components of the electronic device 101 to control or manage system resources. The application 147 includes one or more applications for image capture and image processing using the neural network architecture described below. These functions can be performed by a single application or by multiple applications, each performing one or more of these functions. For example, the middleware 143 can act as a relay to allow the API 145 or application 147 to communicate data with the kernel 141. Multiple applications 147 can be provided. The middleware 143 is capable of controlling work requests received from the application 147, such as by assigning priority to at least one of the multiple applications 147 for use of the electronic device 101's system resources (such as the bus 110, processor 120, or memory 130). The API 145 is an interface that allows the application 147 to control functions provided from the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (such as a command) for file control, window control, image processing, or text control.

[0040] The I / O interface 150 serves as an interface that can, for example, transmit commands or data input from a user or other external devices to other components of the electronic device 101. The I / O interface 150 can also output commands or data received from other components of the electronic device 101 to the user or other external devices.

[0041] The display 160 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. The display 160 can also be a depth perception display, such as a multi-focal display. The display 160 can display, for example, various contents (such as text, images, videos, icons, or symbols) to the user. The display 160 can include a touch screen and can receive, for example, touch, gesture, proximity, or hover input using an electronic pen or a body part of the user.

[0042] The communication interface 170 can, for example, establish communication between the electronic device 101 and an external electronic device (such as the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 can connect to the network 162 or 164 via wireless or wired communication to communicate with the external electronic device. The communication interface 170 can be a wired or wireless transceiver or any other component for transmitting and receiving signals such as images.

[0043] Wireless communication can use, for example, at least one of Long Term Evolution (LTE), Long Term Evolution Advanced (LTE-A), fifth generation wireless systems (5G), millimeter wave or 60 GHz wireless communication, wireless USB, code division multiple access (CDMA), wideband code division multiple access (WCDMA), universal mobile telecommunications system (UMTS), wireless broadband (WiBro), or global system for mobile communications (GSM) as a cellular communication protocol. Wired connections can include, for example, at least one of universal serial bus (USB), high-definition multimedia interface (HDMI), recommended standard 232 (RS-232), or plain old telephone service (POTS). Network 162 or 164 includes at least one communication network, such as a computer network (e.g., a local area network (LAN) or a wide area network (WAN)), the Internet, or a telephone network.

[0044] The electronic device 101 also includes one or more sensors 180, which can measure and detect the physical quantity or activation state of the electronic device 101 and convert the measured or detected information into an electrical signal. For example, one or more sensors 180 can include one or more cameras or other imaging sensors for capturing scene images. The sensor 180 can also include one or more buttons for touch input, gesture sensors, gyroscopes or gyro sensors, air pressure sensors, magnetic sensors or magnetometers, acceleration sensors or accelerometers, grip sensors, proximity sensors, color sensors (such as red, green, and blue (RGB) sensors), biophysical sensors, temperature sensors, humidity sensors, illumination sensors, ultraviolet (UV) sensors, electromyography (EMG) sensors, electroencephalography (EEG) sensors, electrocardiography (ECG) sensors, infrared (IR) sensors, ultrasonic sensors, iris sensors, or fingerprint sensors. The sensor 180 can also include an inertial measurement unit (IMU), which can include one or more accelerometers, gyroscopes, and other components. In addition, the sensor 180 can include a control circuit for controlling at least one sensor included herein. Any of these sensors 180 can be located within the electronic device 101 .

[0045] The first external electronic device 102 or the second external electronic device 104 can be a wearable device or an electronic device that can be mounted on a wearable device (such as an HMD). When the electronic device 101 is mounted on the electronic device 102 (such as an HMD), the electronic device 101 can communicate with the electronic device 102 through the communication interface 170. The electronic device 101 can be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network. The electronic device 101 can also be an augmented reality wearable device including one or more cameras, such as glasses.

[0046] Each of the first external electronic device 102 and the second external electronic device 104 and the server 106 can be a device of the same or different type as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. In addition, according to certain embodiments of the present disclosure, all or some operations performed on the electronic device 101 can be performed on another one or more other electronic devices (such as the electronic devices 102 and 104 or the server 106). In addition, according to certain embodiments of the present disclosure, when the electronic device 101 should automatically or upon request perform certain functions or services, the electronic device 101 can request another device (such as the electronic devices 102 and 104 or the server 106) to perform at least some functions associated therewith, rather than performing the function or service itself, or the electronic device 101 can additionally request another device (such as the electronic devices 102 and 104 or the server 106) to perform at least some functions associated therewith. The other electronic devices (such as the electronic devices 102 and 104 or the server 106) can perform the requested functions or additional functions and transmit the execution results to the electronic device 101. The electronic device 101 can provide the requested function or service by processing the received result as it is or additionally. To this end, for example, cloud computing, distributed computing or client-server computing technology can be used. Figure 1 The electronic device 101 is shown to include a communication interface 170 to communicate with the external electronic device 104 or the server 106 via the network 162 or 164 , but according to some embodiments of the present disclosure, the electronic device 101 may operate independently without a separate communication function.

[0047] The server 106 can include components 110-180 (or a suitable subset thereof) that are the same or similar to the electronic device 101. The server 106 can support driving the electronic device 101 by performing at least one of the operations (or functions) implemented on the electronic device 101. For example, the server 106 can include a processing module or processor that can support the processor 120 implemented in the electronic device 101. In some embodiments, the server 106 uses a neural network architecture to process image data to perform feature matching using a reversible wavelet layer and an asymmetric normalized correlation layer to generate a single feature map from multiple images of a scene. This can support various image processing functions, such as creating a Bokeh effect in an image.

[0048] although Figure 1 An example of a network configuration 100 including an electronic device 101 is shown, but the network configuration 100 may be configured as follows: Figure 1 Various changes may be made. For example, network configuration 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems come in a variety of configurations, and Figure 1The scope of the present disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document can be used, but these features can be used in any other suitable system.

[0049] Figure 2A 、 Figure 2B and Figure 2C 1 illustrates an example input image and example processing results that can be obtained using an asymmetric normalized correlation layer in a neural network according to the present disclosure. In this particular example, a neural network (such as a wavelet synthesis neural network) is used to generate a depth map, which is then used to create a Boke effect from the original image. However, a neural network such as a wavelet synthesis neural network can be used to perform any other suitable task, whether related to image processing or not. For ease of explanation, Figure 2A 、 Figure 2B and Figure 2C The input image and processing results shown in are about Figure 1 The network configuration 100 is described with reference to the electronic device 101 or the server 106. However, the neural network with the asymmetric normalized correlation layer may be used by any other suitable device and any other suitable system.

[0050] like Figure 2A As shown, an image 202 to be processed by a neural network is received, such as when image 202 is received from at least one camera (sensor 180) of electronic device 101. In this example, image 202 represents an image of a person in the foreground next to a chain-link fence, and the background includes both a field and a building. Although the person's face is obscured for privacy reasons, both the foreground and the background are in focus, as is common on devices such as smartphones and tablets. In some embodiments, image 202 can be generated using two images captured by two different cameras of electronic device 101. In these embodiments, the two images are calibrated to account for any differences between the two cameras, such as the use of different lenses, different fields of view, different focus points, etc.

[0051] like Figure 2BAs shown, depth map 204 is generated by a neural network. Depth map 204 generally identifies different depths in different parts of the scene captured in image 202 (or the image pair used to generate image 202). In this example, lighter colors represent shallower or smaller depths, and darker colors represent deeper or larger depths. In some embodiments, depth map 204 is generated using two input images. For example, two cameras separated by a known distance can each capture an image of the same scene. The neural network can then compare the positions of the same points in the scene in the different images to determine the disparity of those points in the images. There is an inverse relationship between the disparity of each point in the image and the depth of that point in the scene. For example, a larger disparity indicates that the point is closer to electronic device 101, and a smaller disparity indicates that the point is farther away from electronic device 101. Therefore, the disparity of each point in the scene can be calculated and used to generate depth map 204 (or the disparity can be used to generate a disparity map).

[0052] Figure 2B The depth map 204 in the image identifies the distances between the electronic device 101 and different regions or portions of the imaged scene on a pixel-by-pixel basis. As shown herein, the background is generally dark, indicating that the background is sufficiently far away from the camera (in some cases, it can be referred to as infinite). That is, the parallax between common points in the background captured in multiple images is negligible. The brighter portions of the depth map 204 include people and chain-link fences, indicating that there are larger or more measurable parallaxes between common points in the foreground captured in multiple images.

[0053] like Figure 2C As shown, image 206 is generated based on image 202 and depth map 204. As shown in image 206, the background of the scene has been computationally blurred to produce a blur effect in image 206, while objects in the foreground of the scene (e.g., a person and a chain-link fence) are in focus. Electronic device 101 or server 106 can generate image 206 by applying a variable amount of blur to image 202, where the amount of blur applied to each portion (or each pixel) of image 202 is based on depth map 204. Thus, for example, maximum blur can be applied to pixels of image 202 associated with the darkest colors in depth map 204, and minimum blur or no blur can be applied to pixels of image 202 associated with the lightest colors in depth map 204.

[0054] As described in more detail below, a neural network (such as a wavelet synthesis neural network) is used to generate the depth map 204, and the resulting depth map 204 is then used to perform some image processing functions (such as Bock generation). The neural network includes a reversible wavelet layer and a normalized correlation layer, which will be described in more detail below.

[0055] although Figure 2A 、 Figure 2B and Figure 2C The diagrams illustrate an example of an input image and an example of a processing result that can be obtained using an asymmetric normalized correlation layer in a neural network, but various changes can be made to these diagrams. For example, these diagrams are merely meant to illustrate one example of the type of results that can be obtained using the methods described in this disclosure. Images of scenes can vary greatly, and the results obtained using the methods described in this patent document can also vary greatly depending on the environment.

[0056] Figure 3 An example neural network architecture 300 according to the present disclosure is illustrated. For ease of explanation, the neural network architecture 300 is described using Figure 1 The neural network architecture 300 is implemented as an electronic device 101 or server 106 in the network configuration 100. However, the neural network architecture 300 can be used by any other suitable device and any other suitable system. In addition, the neural network architecture 300 is described as being used to perform specific image processing related tasks, such as creating a bok effect in an image. However, the neural network architecture 300 can be used to perform any other suitable tasks, including non-image processing tasks.

[0057] like Figure 3 As shown, neural network architecture 300 is configured to receive and process input data, which in this example includes input image 302 and input image 304. Input images 302 and 304 can be received from any suitable source (or sources), such as from two cameras (or one or more sensors 180) of electronic device 101. Neural network architecture 300 generally operates here to process input images 302 and 304 and generate various outputs. In this example, the outputs include depth map 312 and Boke image 316.

[0058] The depth map 312 may be similar to Figure 2B The depth map 204 is used as a proxy for the depth of the scene being imaged because it can identify the depth in the scene being imaged (possibly on a pixel-by-pixel basis). Thus, the depth map 312 represents the apparent pixel difference between the input images 302 and 304 (for disparity) or the apparent depth of one or more pixels in the images 302 and 304 (for depth). In the absence of motion, the disparity between the same points in the input images 302 and 304 is inversely proportional to the depth, so the disparity map can be used when calculating the depth map (and vice versa). The Boke image 316 can be similar to Figure 2C 206 because it can include a computationally blurred background. Thus, Boke image 316 generally represents an image in which the background of the image has been digitally blurred, where the image is based on input image 302 and / or input image 304.

[0059] In this example, neural network architecture 300 includes a calibration engine 308 that accounts for differences between input images 302 and 304 (such as differences based on the cameras that captured images 302 and 304). For example, if the camera that captured input image 302 used a wide-angle lens, while the camera that captured input image 304 used a telephoto lens, then input images 302 and 304 have captured different portions of the same scene. For example, input image 304 may represent a greater magnification of the scene than input image 302. Calibration engine 308 modifies one or both of input images 302 and 304 so that the images depict similar views of the scene. Calibration engine 308 can also calibrate input images 302 and 304 based on other differences associated with the cameras, such as different focused objects, different fields of view, etc.

[0060] Neural network 310 receives input images 302 and 304 (modified by calibration engine 308) and processes the calibrated images to generate depth map 312. In this example, the two inputs to neural network 310 correspond to the two input images 302 and 304 calibrated by calibration engine 308. As described in more detail below, neural network 310 generally includes a feature extractor (encoder), a normalized correlation layer, and a refinement layer (decoder) for generating depth map 312 from two or more images. In some embodiments, neural network 310 also includes a reversible wavelet layer. Note that although neural network 310 herein receives two input images, it can also receive and process more than two input images of a scene. It should be noted that as the number of input images received by neural network 310 increases, the fidelity of depth map 312 also increases.

[0061] The feature extractor of neural network 310 generally operates to extract high-level features from the calibrated input images 302 and 304 to generate two or more feature maps. Neural network 310 can use a feature extractor including convolutional layers and pooling layers to reduce the spatial resolution of the input images while increasing the depth of the feature maps. In some embodiments, neural network 310 uses the same number of feature extractors as the number of input images, such that each feature extractor branch corresponds to one input image. For example, if two images (e.g., input images 302 and 304) are input to neural network 310, a first feature extractor can generate a first feature map corresponding to input image 302, and a second feature extractor can generate a second feature map corresponding to input image 304. In those embodiments, the input to each feature extractor is an RGB image (such as input image 302 or 304) or other image data. In some embodiments, the feature extractor can feed intermediate feature maps to a refinement layer. In some cases, the feature maps generated by the feature extractors of neural network 310 can include three-dimensional (3D) feature maps, where the dimensions include height (H), width (W), and channels (C).

[0062] After generating the feature maps, the normalized correlation layer of the neural network 310 performs matching in the feature map space to generate a new feature map. For example, the normalized correlation layer can calculate the mutual correlation between two or more feature maps. In some embodiments, the asymmetric normalized correlation layer performs a normalized comparison between feature maps. In each search direction w, the asymmetric normalized correlation layer identifies the similarity d between the two feature maps. In a specific embodiment, the following equation (1) describes how the asymmetric normalized correlation layer identifies the similarity between multiple feature maps.

[0063]

[0064] The new feature map generated by the normalized correlation layer can have dimensions (H, W, C'), where C' is determined based on the size of the asymmetric search window used by the normalized correlation layer. The asymmetric search window (and the corresponding size C') is based on physical parameters between the cameras that captured the input images 302 and 304 being processed. In some cases, the parameters are based on the distance between the cameras. The asymmetric search window (and the corresponding size C') is also based on the accuracy of the calibration engine 308, so the value of C' decreases as the accuracy of the calibration engine 308 increases or as the distance between the two cameras decreases.

[0065] Pooling layers can be used in neural network 310 to increase the receptive field of the feature extractor, enabling neural network 310 to gain a global context or understanding of input images 302 and 304. Convolutional layers can be used to stack and increase the receptive field, while pooling layers can multiply the receptive field. Note that pooling layers can introduce information loss. For example, in a 2×2 max pooling layer, 75% of the information may be discarded. Generally, in classification applications, five 2×2 pooling operations can be used to achieve an output stride of 32, which effectively discards a significant amount of information. However, in pixel-to-pixel applications such as semantic segmentation, disparity, or optical flow estimation, the output resolution is often the same as the input resolution. This requires more information to pass through neural network 310. Therefore, wavelet and inverse wavelet transforms can be used to provide both spatial resolution reduction and information preservation. Wavelet transforms are reversible and can achieve the same spatial resolution reduction effect as pooling layers without information loss, so they can be used in neural network 310. Further details on wavelet and inverse wavelet transforms are provided below.

[0066] The refinement layers of the neural network 310 restore spatial resolution to the feature maps generated by the normalized correlation layers. This results in the generation of a depth map 312, which can be output by the neural network 310. More details of the neural network are provided below.

[0067] In some embodiments, the neural network 310 also generates a confidence map associated with the depth map 312. The confidence map can be obtained by applying a softmax operation on the channel dimension of the feature map. The confidence map can indicate the decrease in confidence of pixel matches in homogeneous and occluded areas of the input images 302 and 304. The confidence map can be used for filtering, blending, or other rendering purposes.

[0068] Renderer 314 is configured to generate a background image 316 based on depth map 312 and at least one of images 302 and 304. For example, renderer 314 may generate background image 316 based on focus point 306, input image 302, and depth map 312. In some embodiments, the cameras that captured input images 302 and 304 can be designated as a primary camera and a secondary camera. For example, if a user wishes to capture an image of a scene using a telephoto lens, the camera including the telephoto lens of electronic device 101 can be designated as the primary camera, while another camera of electronic device 101 can be designated as the secondary camera. Similarly, if a user wishes to capture an image of a scene using a wide-angle lens, the camera including the wide-angle lens of electronic device 101 can be designated as the primary camera, while another camera of electronic device 101 (such as a camera including an ultra-wide-angle lens) can be designated as the secondary camera. Regardless of the designation, focus point 306 can correspond to a focal position within the image captured by the primary camera. As a result, focus point 306, when combined with depth map 312, can identify a focal plane. The focal plane represents the distance (or depth) in the scene at which the primary camera is desired to be in focus.

[0069] Renderer 314 also generates a blur effect in blurry image 316 by applying an appropriate blur to image 302. For example, renderer 314 can generate a circle of confusion (CoC) map based on the primary camera's focus point 306 and depth map 312. In the CoC map, blur levels increase with increasing distance from the focal plane. That is, as content in image 302 moves further from the focal plane, it is assigned increasingly greater blur levels, as shown in depth map 312. If neural network 310 also generates and outputs a confidence map, renderer 314 can use this confidence map when generating the blur effect for blurry image 316. For example, renderer 314 can perform alpha blending, which uses the confidence map to blend focused image 302 with the CoC map. Because the confidence map indicates the accuracy of the pixel matching used in creating depth map 312, renderer 314 can increase or decrease the alpha blending accordingly.

[0070] In addition to generating a boke image 316, the renderer 314 can use the focus 306 and the depth map 312 to provide various other effects, such as variable focus, variable aperture, and art boke. The variable focus effect generates a new image that changes the position of the focus within the image relative to the primary camera. The variable aperture effect corresponds to an adjustable CoC map. The art boke effect enables an adjustable kernel shape of the light spot in the image relative to the primary camera, such as by changing the shape of the background light in the image.

[0071] To generate depth maps 312 for various scenes, neural network 310 is trained before being put into use. Training establishes the parameters of neural network 310 for performing various functions, such as generating and processing feature maps. In some embodiments, neural network 310 undergoes three training phases before being put into use. During the first phase of training, neural network 310 can be trained using synthetic data, and weights between feature extractors can be shared when processing features extracted from stereo images. During the second phase of training, neural network 310 learns the photometric mapping between the cameras that captured the calibration images. Because the cameras of electronic device 101 will typically have different lenses (such as telephoto lenses, wide-angle lenses, ultra-wide-angle lenses, etc.), different image signal processors, different settings, different tuning, etc., photometric differences may exist. During the third phase of training, neural network 310 does not share weights between feature extractors, allowing the feature extractors to be trained with independent weights.

[0072] The various operations performed in neural network architecture 300 can be implemented in any suitable manner. For example, each operation performed in neural network architecture 300 can be implemented or supported using one or more software applications or other software instructions executed by at least one processor 120 of electronic device 101 or server 106. In other embodiments, at least some of the operations performed in neural network architecture 300 can be implemented or supported using dedicated hardware components. In general, the operations of neural network architecture 300 can be performed using any suitable hardware or any suitable combination of hardware and software / firmware instructions.

[0073] although Figure 3 One example of a neural network architecture 300 is shown, but may be used for Figure 3 Various changes may be made. For example, neural network architecture 300 may be capable of receiving and processing more than two input images. Furthermore, the tasks performed using neural network architecture 300 may or may not involve image processing.

[0074] Figure 4 Detailed example of a neural network 410 including an asymmetric normalized correlation layer 420 according to the present disclosure is shown. For example, Figure 4 The neural network 410 shown represents Figure 3 A more detailed view of the neural network 310 is shown. For ease of explanation, the neural network 410 is described using Figure 1 10. However, neural network 410 may be implemented by electronic device 101 or server 106 in network configuration 100. However, neural network 410 may be used by any other suitable device and any other suitable system. Furthermore, neural network 410 is described as being used to perform specific image processing-related tasks, such as creating a bok effect in an image. However, neural network 410 may be used to perform any other suitable tasks, including non-image processing tasks.

[0075] like Figure 4 As shown, neural network 410 generally operates to receive a plurality of calibrated input images 402 and 404 and generate a depth map 428. Calibrated input images 402 and 404 may, for example, represent input images 302 and 304 after being processed by calibration engine 308. Note that neural network 410 as described herein can be used to process any suitable input data and is not limited to processing image data. Note also that neural network 410 can receive and process more than two calibration images. In other embodiments, additional calibration images can be input to neural network 410. For each additional calibration image, an additional feature extractor can be provided in neural network 410.

[0076] In this example, calibration image 402 is input to feature extractor 412 and calibration image 404 is input to feature extractor 416. Feature extractor 412 generates feature map 414, such as a feature map having dimensions (H, W, C). Similarly, feature extractor 416 generates feature map 418, such as a feature map having dimensions (H, W, C). In some embodiments, feature extractors 412 and 416 utilize convolutional layers and pooling layers to reduce the spatial resolution of calibration images 402 and 404 while increasing the depth of feature maps 414 and 418. In certain embodiments, a reversible wavelet layer performs the spatial resolution reduction.

[0077] Feature maps 414 and 418 are input to an asymmetric normalized correlation layer 420. In some embodiments, the asymmetric normalized correlation layer 420 applies independent random binary masks to the feature maps 414 and 418. The binary mask blocks random pixels along the channel dimension of each feature map 414 and 418. For example, at a specific (H, W) position in each feature map 414 and 418, the channel dimension can be blocked. The binary mask is random so that a random pixel in the feature map 414 and a random pixel in the feature map 418 are blocked. In some embodiments, a zero value is assigned to each blocked pixel in the feature maps 414 and 418 with a probability of 0.25. The binary mask can be applied to the feature maps 414 and 418 to force the asymmetric normalized correlation layer 420 to learn how to match features even if a small portion of the view is blocked. The binary mask can be used to determine the accuracy of the calibration engine 308.

[0078] The asymmetric normalized correlation layer 420 can use an asymmetric search window to perform matching between feature maps 414 and 418, which helps ensure that the search is asymmetric in order to maximize search efficiency. For example, the size of the asymmetric search window can be based on the distance between the cameras that capture the input images calibrated to form calibration images 402 and 404 and the accuracy of the calibration engine 308. The size of the asymmetric search window can also be based on various dimensions represented as dx+, dx-, dy-, and dy+. For cameras with larger baselines, a larger dx+ value can be assigned to the search window. For cameras with smaller baselines, a smaller dx+ value can be assigned to the search window. The accuracy of the calibration can also change the dimensions. For example, when the accuracy of the calibration engine 308 is high, the dimensions dx-, dy-, and dy+ can be set to smaller values. More details about the asymmetric search window are provided below.

[0079] The dx+ dimension is typically larger than the other dimensions because dx+ is based on the physical distance between the cameras, while dx-, dy-, and dy+ are based on the calibration accuracy. For example, for a feature map spatial resolution of 256×192 (H×W), dx+ can be 16, dx- can be 2, dy- can be 2, and dy+ can be 2. When dx+ is 16, dx- is 2, dy- is 2, and dy+ is 2, the size of the asymmetric search window is 72 because (16+2)×(2+2) equals 72. Note that the asymmetric search window is an improvement over the symmetric search window because the symmetric search window is based on the largest dimension, which results in a much larger size for the symmetric search window. In some embodiments, the asymmetric normalized correlation layer 420 sets the size of the asymmetric search window based on the identified calibration accuracy and the physical distance between the cameras that captured the image. In some cases, the physical distance between the cameras can vary from image to image because each camera can include an optical image stabilizer (OIS) that slightly moves the camera sensor to compensate for movement when capturing the image.

[0080] The size of the asymmetric search window indicates the number of search directions (u, v) for which the asymmetric normalized correlation layer 420 calculates the channel-normalized cross-correlations. Thus, the asymmetric normalized correlation layer 420 can calculate the channel-normalized correlation between the feature map 414 and the shifted version of the feature map 418 to generate one channel of the new feature map 422. The asymmetric normalized correlation layer 420 can repeat this process for all directions based on the size of the asymmetric search window. For example, if the size of the asymmetric search window is 72 (based on the previous example), the asymmetric normalized correlation layer 420 can calculate the channel-normalized correlation between the feature map 414 and the shifted feature map 418, where the feature map 418 is shifted 72 times to generate the new feature map 422. In this example, the new feature map 422 will have dimensions of 256×192×72.

[0081] The asymmetric normalization correlation layer 420 can also normalize the values ​​of the new feature map, such as by normalizing the values ​​to the range [0, 1]. In some embodiments, the feature map values ​​can be normalized by subtracting the mean (average) and dividing by the residual variance in the input feature map. Equations (2) and (3) below describe one possible implementation of normalization to ensure that the output feature map is constrained to the range [0, 1].

[0082]

[0083] here, Represents the two-dimensional (2D) output feature map, F L and represents the three-dimensional (3D) left input feature map and the right input feature map, and F L and Represents feature maps 414 and 418. In addition, var e represents the variance of the feature map in the channel dimension, and ∈ represents a specific value (such as 10 -5 ) to prevent the possibility of division by zero. Equations (2) and (3) can be applied to all directions (u, v) in the search window and stacked on the 2D feature map along the channel dimension. , to generate a 3D feature map 422.

[0084] Note that, although shown and described as processing two calibrated input images 402 and 404, asymmetric normalized correlation layer 420 is not limited to stereo matching applications. Rather, asymmetric normalized correlation layer 420 can be used by any neural network that performs feature map matching, regardless of whether the feature maps are associated with two or more inputs. Furthermore, asymmetric normalized correlation layer 420 can be used by any neural network to support other image processing functions or other functions. As a specific example, asymmetric normalized correlation layer 420 can be applied to face verification for matching high-level features of multiple faces.

[0085] The refinement layer 426 generates a depth map 428 by restoring the spatial resolution of the generated feature map 422. In this example, the feature extractor 412 feeds one or more intermediate feature maps 424 forward to the refinement layer 426 for restoring the spatial resolution of the generated feature map 422. In some embodiments, a reversible wavelet layer performs spatial resolution reduction in the feature extractor 412, and the reversible wavelet layer can provide the necessary information to the refinement layer 426 to restore the spatial resolution of the generated feature map 422.

[0086] The various operations performed in neural network 410 can be implemented in any suitable manner. For example, each operation performed in neural network 410 can be implemented or supported using one or more software applications or other software instructions executed by at least one processor 120 of electronic device 101 or server 106. In other embodiments, at least some of the operations performed in neural network 410 can be implemented or supported using dedicated hardware components. In general, the operations of neural network 410 can be performed using any suitable hardware or any suitable combination of hardware and software / firmware instructions.

[0087] although Figure 4 A detailed example of a neural network 410 including an asymmetric normalized correlation layer 420 is shown, but may be used for Figure 4Various changes may be made. For example, neural network 410 may include any suitable number of convolutional layers, pooling layers, or other layers as needed or desired. Furthermore, neural network 410 may be capable of receiving and processing more than two input images. Furthermore, the tasks performed using neural network 410 may or may not include image processing.

[0088] Figure 5 FIGURE 5 illustrates an example application of a reversible wavelet layer 500 for a neural network according to the present disclosure. The reversible wavelet layer 500 may be used, for example, to Figure 3 Neural Network 310 or Figure 4 For ease of explanation, the reversible wavelet layer 500 is described as using Figure 1 The reversible wavelet layer 500 is implemented by the electronic device 101 or the server 106 in the network configuration 100. However, the reversible wavelet layer 500 can be used by any other suitable device(s) and any other suitable system(s). In addition, the reversible wavelet layer 500 is described as being used to perform specific image processing related tasks, such as creating a Boke effect in an image. However, the reversible wavelet layer 500 can be used to perform any other suitable tasks, including non-image processing tasks.

[0089] As described above, the reversible wavelet layer 500 can be applied to iteratively decompose and synthesize feature maps. Figure 4 In , the reversible wavelet layer 500 can be used in one or more feature extractors 412 and 416 to reduce the spatial resolution of the calibration images 402 and 404 while increasing the depth of the feature maps 414 and 418. Figure 5 In FIG, the reversible wavelet layer 500 receives the feature map 510 and decomposes it into four elements, namely a low-frequency component 520 (such as average information) and three high-frequency components 530 (such as detailed information). The high-frequency components 530 can be stacked in the channel dimension to form a new feature map.

[0090] The low frequency components 520 may represent a first feature map generated by the reversible wavelet layer 500. In some cases, the low frequency components 520 have dimensions of (H / 2, W / 2, C). The high frequency components 530 may collectively represent a second feature map generated by the reversible wavelet layer 500. In some cases, the high frequency components 530 collectively have dimensions of (H / 2, W / 2, 3C). The low frequency components 520 and the high frequency components 530 are represented by Figure 3 Neural Network 310 or Figure 4 The neural network 410 of FIG. 410 can be processed differently. For example, the neural network 310 or 410 can iteratively process the low-frequency component 520 to obtain a global contextual understanding of the image data without being distracted by local details. The high-frequency component 530 can be used to restore the spatial resolution of the output of the neural network 310 or 410, such as the new feature map 422.

[0091] In some embodiments, the feature maps 414 and 418 are Figure 4 Before being processed by the asymmetric normalized correlation layer 420, the reversible wavelet layer 500 reduces the low frequency components 520 by a factor of eight (although other reduction factors can also be used). In addition, in some embodiments, one or more convolution modules in the neural network 310 or 410 can have a stride of 1. In addition, in some embodiments, a convolution module in the neural network 310 or 410 can include more than one convolution block, where each convolution block performs a 1×1 convolution expansion step, a 3×3 depthwise convolution step, and a 1×1 convolution projection step. If the feature map generated (after projection) has the same number of channels as the input feature map, an additional recognition branch connects the input feature map and the output feature map.

[0092] although Figure 5 One example of application of a reversible wavelet layer 500 to a neural network is shown, but can be applied to Figure 5 Various changes may be made. For example, any other suitable layers may be used in neural network architecture 300 or neural network 410.

[0093] Figure 6A and Figure 6B An example asymmetric search window 600 and an example application of the asymmetric normalized correlation layer 420 are illustrated for use in the asymmetric normalized correlation layer 420 according to the present disclosure. For ease of explanation, the asymmetric search window 600 and the asymmetric normalized correlation layer 420 are described as being used. Figure 1 The asymmetric search window 600 and the asymmetric normalized correlation layer 420 may be implemented in the electronic device 101 or server 106 of the network configuration 100. However, the asymmetric search window 600 and the asymmetric normalized correlation layer 420 may be used by any other suitable device and any other suitable system. Furthermore, the asymmetric search window 600 and the asymmetric normalized correlation layer 420 are described as being used to perform specific image processing-related tasks, such as creating a bokeh effect in an image. However, the asymmetric search window 600 and the asymmetric normalized correlation layer 420 may be used to perform any other suitable tasks, including non-image processing tasks.

[0094] like Figure 6AAs shown and described above, asymmetric search window 600 is based on four dimensions, namely dimension 602 (dy+), dimension 604 (dy-), dimension 606 (dx-), and dimension 608 (dx+). Dimensions 602, 604, 606, and 608 are parametric measurements from pixel 610 to asymmetric search window 600. The size of dimensions 602, 604, 606, and 608 can be based on the camera baseline distance and the accuracy of calibration engine 308. For example, if dy+ is 2, dy- is 2, dx- is 2, and dx+ is 16, then the size of asymmetric search window 600 is 72. Given this, asymmetric normalized correlation layer 420 can shift feature map 418 a total of 72 times and perform a channel-normalized cross-correlation operation to generate feature map 422. In some embodiments, dimensions 602, 604, and 606 are the same size, and dimension 608 is larger than dimensions 602, 604, and 606.

[0095] like Figure 6B As shown, the asymmetric normalized correlation layer 420 receives the feature map 612 and the feature map 614. The feature map 612 can represent Figure 4 The feature map 414 and the feature map 614 can represent Figure 4 4. Asymmetric normalized correlation layer 420 randomly applies a binary mask to feature map 612 to create mask feature map 616, and asymmetric normalized correlation layer 420 randomly applies a binary mask to feature map 614 to create mask feature map 618. As described above, the binary mask blocks random channel values ​​in feature maps 612 and 614 to produce mask feature maps 616 and 618. Blocking random channel values ​​can force neural network 310 or 410 to learn matching even if a small portion of the view in the image is blocked.

[0096] The mask feature map 618 is subjected to a shift operation 620 that shifts the mask feature map 618 a number of times in one or more directions 622. The shifts here are based on Figure 6A 6. Asymmetric search window 600 is shown. For each shift of mask feature map 618 in a particular (u, v) direction 622, multiple feature maps 624, 626, and 628 are generated. The number of times mask feature map 618 is shifted can be based on the size of asymmetric search window 600. For example, when the dimensions of asymmetric search window 600 are dy+=2, dy-=2, dx-=2, and dx+=16, mask feature map 618 is shifted 72 times, resulting in 72 sets of feature maps 624, 626, and 628. The shifting of mask feature map 618 can occur in the (u, v) direction, where u is between -2 and 16, and v is between -2 and 2.

[0097] To generate each set of feature maps 624, 626, and 628, the asymmetric normalized correlation layer 420 can perform feature matching by calculating the inner product and the average of the mask feature map 616 and the shifted mask feature map 618. For example, the asymmetric normalized correlation layer 420 can calculate the inner product between the mask feature map 616 and the shifted mask feature map 618 while shifting along the channel dimension to generate the feature map 626. The asymmetric normalized correlation layer 420 can also calculate the average of the mask feature map 616 along the channel dimension to generate the feature map 628, and the asymmetric normalized correlation layer 420 can calculate the average of the mask feature map 618 while shifting along the channel dimension to generate the feature map 624. The set of feature maps 624, 626, and 628 represents a single-channel feature map.

[0098] Then, the asymmetric normalized correlation layer 420 normalizes the feature map 626 using the feature maps 624 and 628 to generate a normalized feature map 630. In some embodiments, the asymmetric normalized correlation layer 420 normalizes the feature map 626 using equation (4) below.

[0099]

[0100] The normalized feature map 630 is a 2D feature map because it corresponds to a single channel. However, by generating a normalized feature map 630 for each shift of the mask feature map 618, the asymmetric normalized correlation layer 420 generates new feature maps 624, 626, and 628, and generates a new normalized feature map 630 for that shift of the mask feature map 618. Each new normalized feature map 630 corresponds to a different channel, and multiple normalized feature maps 630 can be stacked. The stacking of the normalized feature maps 630 increases the depth and thereby forms a 3D feature map of dimensions (H, W, C'), where C' corresponds to the number of shifts of the mask feature map 618 (which is based on the size of the asymmetric search window 600).

[0101] The set of normalized feature maps 630 can represent the output to Figure 4 Note that when the reversible wavelet layer 500 is used to reduce the low frequency components 520 by the factor described above, the refinement layer 426 (using the high frequency components 530) operates to restore the spatial resolution to the normalized feature map 630 in order to generate the depth map 428.

[0102] although Figure 6A and Figure 6B An example of an asymmetric search window 600 used in an asymmetric normalized correlation layer 420 and an example application of the asymmetric normalized correlation layer 420 are illustrated, but may be used for Figure 6A and Figure 6BVarious changes can be made. For example, the size of the asymmetric search window 600 can be varied based on characteristics of the electronic device 101, such as the physical distance between the cameras and the accuracy of the calibration. In addition, the asymmetric normalized correlation layer 420 can process any other number of input feature maps, which can be based on the number of input images being processed.

[0103] Figure 7 7. An example method 700 for deep neural network feature matching using an asymmetric normalized correlation layer according to the present disclosure is shown. More specifically, Figure 7 An example method 700 is illustrated for generating a depth map using an asymmetric normalized correlation layer 420 in a neural network 310 or 410, wherein the generated depth map is used to perform an image processing task. For ease of explanation, Figure 7 The method 700 is described as comprising Figure 1 Used in network configuration 100 Figure 3 However, the method 700 may include the use of any suitable neural network architecture designed according to the present disclosure, and the asymmetric normalized correlation layer 420 may be used in any other suitable device or system.

[0104] In step 702, neural network architecture 300 obtains input data, such as a plurality of input images. The input images represent two or more images of a scene, such as images captured by different cameras or other image sensors of an electronic device. For example, a first image of the scene can be obtained using a first image sensor of the electronic device, and a second image of the scene can be obtained using a second image sensor of the electronic device. Note that neural network architecture 300 can be implemented in an end-user device (such as electronic device 101, 102, or 104) and process data collected or generated by the end-user device, or neural network architecture 300 can be implemented in one device (such as server 106) and process data collected or generated by another device (such as electronic device 101, 102, or 104).

[0105] In step 704, neural network architecture 300 generates a first feature map from the first image and a second feature map from the second image. For example, images 302 and 304 can be processed by calibration engine 308 to modify at least one of images 302 and 304 and produce calibration images 402 and 404. Calibration images 402 and 404 can then be processed by feature extractors 412 and 416 to produce feature maps 414 and 418. In some embodiments, neural network architecture 300 uses separate feature extractors to generate different feature maps. For example, feature map 414 can be generated by feature extractor 412, and feature map 418 can be generated by feature extractor 416. If additional input images are obtained in step 702, additional feature extractors can be utilized to generate additional feature maps for those images. In some embodiments, the feature extractors operate to generate feature maps in parallel, meaning simultaneously during the same or similar time periods.

[0106] In step 706, the neural network architecture 300 generates a third feature map based on the first and second feature maps using an asymmetric search window. The size of the asymmetric search window is based on the accuracy of the calibration algorithm used to calibrate the input image and the distance (or distances) between the cameras capturing the images. In some cases, the asymmetric search window can be longer horizontally than vertically. The size of the asymmetric search window corresponds to the number of times the second feature map is shifted when performing feature matching to generate the third feature map. In some embodiments, to generate the third feature map, the neural network architecture 300 applies a binary mask to random channels of the first and second feature maps. The binary mask can be used to identify errors in the calibration process or the accuracy level of the calibration process when generating the calibration image. After applying the mask to the second feature map, the second feature map is shifted a number of times based on the size of the asymmetric search window. For each shift of the second feature map, the neural network architecture 300 calculates the channel-normalized cross-correlation between the shifted versions of the first and second feature maps to identify the channel values ​​of the third feature map. This can occur as described above. This process is repeated for each shift of the second feature map, thereby generating multiple single-channel feature maps. Multiple single-channel feature maps can then be stacked to form a third feature map.

[0107] In step 708, the neural network architecture 300 generates a depth map by restoring spatial resolution to the third feature map. For example, the neural network architecture 300 can use the refinement layer 426 to restore spatial resolution to the third feature map. In some cases, the neural network architecture 300 can decompose the first feature map into multiple components, such as multiple high-frequency components 530 and low-frequency components 520. In these embodiments, the neural network architecture 300 can use a reversible wavelet layer to decompose the first feature map. The low-frequency components 520 of the first feature map provide global context of the image without being distracted by local details, while the high-frequency components 530 of the first feature map are used to restore spatial resolution to the third feature map when generating the depth map.

[0108] In step 710, an image processing task is performed using the depth map. For example, the neural network architecture 300 can identify a focal point within one of the captured images. Based on the location of the focal point, the neural network architecture 300 can identify a depth plane within the depth map that corresponds to the location of the focal point in the image. The neural network architecture 300 can then blur portions of the captured image based on their identified distance from the depth plane, such as by increasing the degree of blur at greater depths. This allows the neural network architecture 300 to produce a bokeh effect in the final image of the scene.

[0109] although Figure 7 An example of a method 700 for deep neural network feature matching using an asymmetric normalized correlation layer 420 is shown, but it can be used for Figure 7 For example, although shown as a series of steps, Figure 7 The steps in the method 700 may overlap, occur in parallel, or occur any number of times. In addition, the method 700 may process any suitable input data and is not limited to image processing tasks.

[0110] Although the present disclosure has been described with reference to various exemplary embodiments, various changes and modifications may occur to those skilled in the art. The present disclosure is intended to encompass such changes and modifications as fall within the scope of the appended claims.

Claims

1. A method for asymmetric normalized correlation layers for deep neural network feature matching, comprising: obtaining a first image of a scene using a first image sensor of the electronic device, and obtaining a second image of the scene using a second image sensor of the electronic device; calibrating the first image and the second image by resolving a difference between the first image and the second image based on the first image sensor and the second image sensor; generating a first feature map from the first image and a second feature map from the second image by reducing spatial resolution of the calibrated first image and the calibrated second image; generating a third feature map based on the first feature map and the second feature map by calculating a channel-normalized cross-correlation between the first feature map and the second feature map using an asymmetric search window, wherein a size of the asymmetric search window indicates a number of search directions, and wherein an asymmetric dimension of the asymmetric search window is based on a distance between the first image sensor and the second image sensor; and The depth map is generated by restoring the spatial resolution of the third feature map.

2. The method according to claim 1, wherein Generating the first feature map and the second feature map includes: modifying at least one of the first image and the second image to generate a calibration image pair; and The calibration image pair is used to generate a first feature map and a second feature map.

3. The method according to claim 1, further comprising: A high-frequency component and a low-frequency component of the first feature map are identified, wherein the high-frequency component is used to restore spatial resolution of the third feature map.

4. The method according to claim 1, wherein The asymmetric search window includes at least two different distances with respect to at least two different directions in the asymmetric search window.

5. The method according to claim 1, wherein The first feature map and the second feature map are generated in parallel using different feature extractors in the neural network.

6. The method according to claim 1, wherein Generating the third feature map includes: applying a random binary mask across the first feature map and the second feature map to generate a masked first feature map and a masked second feature map; and The third feature map is identified by computing a channel-normalized cross-correlation between the first feature map of the mask and a shifted version of the second feature map of the mask, wherein the second feature map of the mask is shifted a number of times based on a size of the asymmetric search window.

7. The method according to claim 1, further comprising: obtaining a focus within the first image; as well as Using the depth map, a defocusing effect is generated by blurring portions of the first image corresponding to a depth different from a depth associated with a focal point.

8. An electronic device comprising: a first image sensor; a second image sensor; as well as At least one processor operatively coupled to the first image sensor and the second image sensor and configured to: obtaining a first image of a scene using a first image sensor and obtaining a second image of the scene using a second image sensor; calibrating the first image and the second image by resolving a difference between the first image and the second image based on the first image sensor and the second image sensor; generating a first feature map from the first image and a second feature map from the second image by reducing spatial resolution of the calibrated first image and the calibrated second image; generating a third feature map based on the first feature map and the second feature map by calculating a channel-normalized cross-correlation between the first feature map and the second feature map using an asymmetric search window, wherein a size of the asymmetric search window indicates a number of search directions, and wherein an asymmetric dimension of the asymmetric search window is based on a distance between the first image sensor and the second image sensor; and The depth map is generated by restoring the spatial resolution of the third feature map.

9. The electronic device according to claim 8, wherein To generate the first feature map and the second feature map, the at least one processor is configured to: modifying at least one of the first image and the second image to generate a calibration image pair; and The calibration image pair is used to generate a first feature map and a second feature map.

10. The electronic device according to claim 8, wherein: The at least one processor is further configured to identify a high frequency component and a low frequency component of the first feature map; and The at least one processor is configured to restore spatial resolution to the third feature map using the high frequency component.

11. The electronic device according to claim 8, wherein The asymmetric search window includes at least two different distances with respect to at least two different directions in the asymmetric search window.

12. The electronic device according to claim 8, wherein The at least one processor is configured to generate the first feature map and the second feature map in parallel using different feature extractors in the neural network.

13. The electronic device according to claim 8, wherein To generate the depth map, the at least one processor is configured to: applying a random binary mask across the first feature map and the second feature map to generate a masked first feature map and a masked second feature map; as well as identifying a third feature map by computing a channel-normalized cross-correlation between the first feature map of the mask and a shifted version of the second feature map of the mask; and The at least one processor is configured to shift the second feature map multiple times based on the size of the asymmetric search window.

14. The electronic device according to claim 8, wherein The at least one processor is further configured to: obtaining a focus within the first image; as well as Using the depth map, a defocusing effect is generated by blurring portions of the first image corresponding to a depth different from a depth associated with a focal point.

15. A machine-readable medium comprising instructions that, when executed, cause at least one processor of an electronic device to perform operations corresponding to one of the methods of claims 1-7.

Citation Information

Patent Citations

  • Depth map prediction method based on visual angle fusion

    CN110443842A

  • Panoramic camera systems

    US20180300590A1