Augmented visual perception on computing devices
By merging image data from multiple sensors, wearable devices improve visual perception by making indiscernible objects visible, addressing the challenge of displaying transparent or obstructed objects.
Patent Information
- Application Number
- PCT/US2024/015987
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-15
- Publication Date
- 2025-08-21
AI Technical Summary
Wearable devices struggle with displaying portions of frames and images that are indiscernible by the human eye, such as glass, plexiglass, or car license plates in the dark, due to their transparency or obstruction, making it difficult for users to perceive these objects.
Enhance human visual perception by merging or fusing image data from multiple sensors, including RGB and thermal cameras, to generate an enhanced frame that makes indiscernible objects perceptible.
Enables users to see otherwise invisible objects, allowing for better obstacle detection and thermal leakage observation, enhancing the user's ability to distinguish and identify previously unobservable details.
Smart Images

Figure US2024015987_21082025_PF_FP_ABST
Abstract
Description
AUGMENTED VISUAL PERCEPTION ONCOMPUTING DEVICESBACKGROUND
[0001] Computer vision systems can use algorithms and machine learning models to identify an object(s) in an image. The algorithms and machine learning models can rely on characteristics and features of the image to identify the object(s). A computing device can be used to display the identified objects.SUMMARY
[0002] A feature map or saliency map can be generated using multi-modal images. The multi-modal images can include an image (e.g., RGB image) including a portion that is indiscernible by the human eye (e.g., determined to be indiscernible by a human eye based on a criterion or criteria). The multi-modal images and the feature map or saliency map can be used in an image merging operation that includes a first image (e.g., RGB image) with an overlay of a portion of a second image (e g., infrared image) that includes the portion that is indiscernible by the human eye in an image mode that can be perceived by the human eye.
[0003] In a general aspect, a device, a system, a non-transitory computer- readable medium (having stored thereon computer executable program code which can be executed on a computer system), and / or a method can perform a process with a method including receiving a first image and a second image representing a frame of a video, the first image including a first image portion that is determined based on a visual criterion or criteria, generating a map based on the second image, merging the first image and the second image using the map as an enhanced frame representing the frame of the video, the enhanced frame including the first image and a second portion from the second image, the second image portion corresponding to the first image portion, and rendering the enhanced frame.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Example implementations will become more fully understood from the detailed description given herein below and the accompanying drawings, wherein likeelements are represented by like reference numerals, which are given by way of illustration only and thus are not limiting of the example implementations.
[0005] FIG. 1A illustrates a block diagram of a wearable device including a transparent display according to an example implementation.
[0006] FIG. IB illustrates a block diagram of a video see-through system according to an example implementation.
[0007] FIG. 2A illustrates a block diagram of a video see-through system according to an example implementation.
[0008] FIG. 2B illustrates a block diagram of a machine learning model according to an example implementation.
[0009] FIG. 3A illustrates a block diagram of a video see-through system according to an example implementation.
[0010] FIG. 3B illustrates a block diagram of a machine learning model according to an example implementation.
[0011] FIG. 4 illustrates a block diagram of a video see-through system according to an example implementation.
[0012] FIG. 5 illustrates a block diagram of a method of generating an enhanced frame of a video according to an example implementation.
[0013] FIG. 6 illustrates a block diagram of a method of a machine learning model according to an example implementation.
[0014] FIG. 7 illustrates a pictorial and block diagram of a video see-through system according to an example implementation.
[0015] FIG. 8 is a front view of an example head mounted wearable device according to an example implementation.
[0016] It should be noted that these Figures are intended to illustrate the general characteristics of methods, and / or structures utilized in certain example implementations and to supplement the written description provided below. The structure drawings are not. however, to scale and may not precisely reflect the precise structural or performance characteristics of any given implementation and should not be interpreted as defining or limiting the range of values or properties encompassed by example implementations. For example, the positioning of modules and / or structural elements may be reduced or exaggerated for clarity. The use of similar or identical reference numbers in the various drawings is intended to indicate the presence of a similar or identical element or feature.DETAILED DESCRIPTION
[0017] At least one technical problem with projecting frames of a video and / or images on a display of a wearable device can be that some portions of the frames and / or images can be difficult to view and / or indiscernible by the human eye. Portions of frames and / or images that are indiscernible by the human eye can be difficult for the human eye to distinguish when displayed on the display of the wearable device. In other words, some portions of the frames and / or images can be invisible when displayed on the display of the wearable device. A portion of the frame and / or image may also be indiscernible (invisible or difficult to see) in the real-world as well. For example, glass, plexiglass, transparent objects, car license plates in the dark, and / or the like may be indiscernible, inconspicuous, indistinguishable, hidden, unidentifiable, invisible (or difficult to see), and / or the like to the human eye in the real-world. Therefore, mixed reality wearable devices can have at least one of the same technical problems viewing some objects as wearable devices that only present frames and / or images to a user via a display.
[0018] A wearable device can include vanous video sensors, including RGB video, infrared cameras, time-of-flight depth cameras, thermal cameras, event cameras, and / or the like. At least one technical solution to the aforementioned technical problem can include enhancing human visual perception by generating an image using the output of two or more video sensors. For example, a portion of a frame and / or image that is indiscernible (e.g., determined to be indiscernible by a human eye based on a visual criterion or criteria) by the human eye in a first image (e.g., RGB image captured by a first camera) can be modified with a corresponding portion from a second image (e.g., thermal image captured by a second camera, or the first camera at a later time). Modifying the first image can include removal and replacement of the portion of the first image with the corresponding portion from the second image. Modifying the first image can include overlaying the corresponding portion from the second image onto the first image. Modifying the first image can include rendering both the first image and the second image. Other modification techniques can be used.
[0019] Said differently, at least one technical solution to the aforementioned technical problem can include enhancing human visual perception by, for example, fusing (e.g., fusing, combining, overlaying) video signals, images, and / or image data from multiple sensors in real-time. For example, some implementations can be configured to merge or fuse image data (e.g., multi-modal images) captured by videosensors, including RGB video, infrared cameras, time-of-flight depth cameras, thermal cameras, event cameras, and / or the like to generate an enhanced frame and / or image.
[0020] At least one technical effect of this technical solution can be that the enhanced frame and / or image can provide transparent (e.g., see-through, pass-through and / or the like) capabilities for the wearable device where some portions of the frame and / or image can be perceptible that would otherwise be indiscernible when displayed on the display of the wearable device. In other words, some implementations can enable users of the wearable device to see glass, plexiglass, transparent objects, car license plates in the dark, and / or the like. At least one benefit of the technical solution can be that users of the wearable device may be able to distinguish obstacles and / or distances, detect thermal leakage in rooms, observe fast movements, and the like that would otherwise have been unobservable by the user.
[0021] FIG. 1A illustrates a block diagram of a wearable device including a transparent (e.g., see-through) display according to an example implementation. FIG. 1A illustrates a wearable device 100 configured to enhance human visual perception by generating an image using the output (e.g., multiple images) of two or more of video sensors. The wearable device 100 is configured to solve at least one technical problem including the projecting of frames of a video and / or images on a display of a wearable device that can include some portions of the frames and / or images can be difficult to view and / or indiscernible by the human eye.
[0022] As shown in FIG. 1 A, the wearable device 100 (e.g., AR headset or glasses) can include a portion of a frame 102. The portion of frame 102 can include optical portions in the form of lenses 104. The lenses 104 can include a transparent display 106 (e.g., a see-through display, a pass-through display). The transparent display 106 is illustrated as being included in one of the lenses 104, however, the transparent display 106 can be included in the other of lenses 104 or both lenses 104. Light rays 108 can pass-through both of the transparent display 106 and the lenses 104 to be seen by an eye(s) 114.
[0023] A portion 112 of an image captured by camera 1 (e.g., RGB video sensor) can include a first image portion that is determined based on a visual criterion or criteria (e.g., ahuman visual criterion or criteria). A portion 112 of an image captured by camera 1 (e.g., RGB video sensor) can be indiscernible by the human eye (e.g., determined to be indiscernible by a human eye based on a visual criterion or criteria) when displayed or rendered on transparent display 106. For example, portion 112 canbe (or include) an object in the real-world captured by camera 1. The object may not be seen and / or may be blurry enough to be visually muddled, cloaked, concealed, and / or the like. As an example, the object can be a piece of glass that may not be seen when displayed or the object can be a sign with blurry text. The portion 112 of the image can be determined to need additional processing (e.g., prior to being rendered) based on a visual criterion or criteria. For example, determining portion 112 may not be seen and / or may be blurry to a human eye can trigger additional processing to modify and / or replace portion 112.
[0024] The portion 112 of the image can be determined to be indiscernible by a human eye based on the visual criterion or criteria. For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be invisible, unclear, blurry, and / or difficult for the human eye to distinguish when displayed on the display of the wearable device can result in replacing a first image portion or portion 112 with a second image portion or portion 112’ where the second image portion can correspond to the first image portion. The portion 112 of the image can be determined to be imperceptible by a human eye based on the visual criterion or criteria. For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be invisible, barely visible, faint, and / or difficult for the human eye to distinguish when displayed on the display of the wearable device can result in replacing a first image portion or portion 112 with a second image portion or portion 1 12’ where the second image portion can correspond to the first image portion.
[0025] The portion 112 of the image can be determined to be unclear by a human eye based on the visual criterion or criteria. For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be hazy, blurry, undecipherable, and / or difficult for the human eye to distinguish when displayed on the display of the wearable device can result in replacing a first image portion or portion 112 with a second image portion or portion 112’ where the second image portion can correspond to the first image portion. The portion 112 of the image can be determined to be transparent to the human eye based on the visual criterion or criteria. For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be invisible, see-through, translucent, and / or difficult for the human eye to distinguish when displayed on the display of the wearable device can result in replacing a first image portion or portion 112 with a second image portion or portion 112’ where the second image portion can correspond to the first image portion.
[0026] The portion 112 of the image can be determined to be unseen by a human eye based on the visual criterion or criteria. For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be obscured and / or blend in with the surroundings causing portion 112 to be determined to be unseen when displayed on the display of the wearable device. As a result, portion 112 can be replaced with portion 112’ prior to displaying the image. The portion 112 of the image can be determined to be hidden from a human eye based on the visual criterion or criteria. For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be fully or partially blocked by another object, blend in with the surroundings, inside another object (e.g., inside a house), and / or under another object (e.g., buried underground) when displayed on the display of the wearable device can result in replacing a first image portion or portion 1 12 with a second image portion or portion 112’ where the second image portion can correspond to the first image portion.
[0027] The portion 112 of the image can be determined to be concealed from a human eye based on the visual criterion or criteria. For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be masked by another object or the surroundings, disguised, camouflaged, covered, and / or the like when displayed on the display of the wearable device can result in replacing a first image portion or portion 112 with a second image portion or portion 112’ where the second image portion can correspond to the first image portion. Below, indiscernible, indiscernible by a human eye, and / or indiscernible by a human eye based on the visual criterion or criteria may be used to refer to one or more of the aforementioned conditions and / or an equivalent condition. The visual criterion or criteria can be based on or satisfied based on a transparency or a level of transparency (e.g., an object is determined to be at least partially or completely see-through, or translucent). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be glass, plexiglass, water, and / or the like. The visual criterion or criteria can be based on or satisfied based on a level of light and / or a level of darkness associated with an ambient environment and / or a captured image (e.g., light obscures an object or it is too dark to see an object). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be in a low level of light environment such that it is difficult to determine the object is in a location, difficult to identify the object, and / or difficult to determine characteristics of the objects. As anexample, the object can be a license plate on a car at night (without any lighting) making it difficult to read the text on the license plate.
[0028] The visual criterion or criteria can be based on or satisfied based on object obstruction (e.g., another object is blocking or partially blocking the object). For example, portion 112 can be (or include) an object in the real -world captured by camera 1. The object can be fully or partially blocked by another object, inside another object (e.g., inside a house), and / or under another object (e.g., buried underground). As an example, the object can be behind a wall and completely blocked for capture by camera 1. Further camera 2 can be a thermal camera that senses the heat of the object through the wall. The visual criterion or criteria can be based on or satisfied based on object differentiation (e.g., similarity to adjacent or nearby objects). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be disguised and / or camouflaged by the surroundings. As an example, some animals have a coat that blends in with nature making the animal unseen in an image captured by camera 1. Further camera 2 can be a thermal camera that senses the heat of the heat of the animal.
[0029] The visual criterion or criteria can be based on or satisfied based on context (e.g., identification is unclear based on the surroundings). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be disguised and / or camouflaged by the surroundings. As an example, a hiker can be wearing a green coat that is causing the hiker to be obscured by the leaves of a tree making the hiker unseen in an image captured by camera 1. Further camera 2 can be a thermal camera that senses the heat of the hiker. The visual criterion or criteria can be based on or satisfied based on object features (e.g., shape, color, patterns, etc.). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be a street sign having a specific shape. As an example, a stop sign can be an octagon. The shape can identify the sign as having special interest should the sign be indiscernible for any reason (e.g., lack of light).
[0030] The visual criterion or criteria can be based on or satisfied based on pixel values (e.g., a color, a hue, etc. of one or more pixels). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can be a street sign having a specific color. As an example, a stop sign can be red. The color can identify the sign as having special interest should the sign be indiscernible for any reason (e.g., lack of light). The visual criterion or criteria can be based on or satisfiedbased on pixel borders (e.g., paterns, fuzzy, color difference, etc.). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can include leters and / or numbers that have standard paterns and borders defined by foreground and background colors. The visual criterion or criteria can be based on or satisfied based on pixel transitions (e.g., transitions between shapes, transitions between colors, transitions between shades, etc.). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. The object can include leters and / or numbers that have borders defined by color transitions from the color of the leter and / or number to a background color.
[0031] The visual criterion or criteria can be based on or satisfied based on pixel shading (e.g., shades and shadows can obscure objects). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. Pixel shading can be based on color, texture, lighting, and other properties where a pixel value(s) is calculated using mathematical operations. Factors like light sources, shadows, reflections, and other visual effects can be included in the calculation. The object can be in a low level of light environment causing the calculated pixel to include the reflections, shades and / or shadows such that it is difficult to determine the object is in a location, difficult to identify the object, and / or difficult to determine characteristics of the objects. As an example, the object can be a license plate on a car at night (without any lighting) making it difficult to read the text on the license plate.
[0032] The visual criterion or criteria can be based on pixel gradients (e.g., color gradients across a region). For example, portion 112 can be (or include) an object in the real-world captured by camera 1. A color gradient can be based on a change in pixel intensity in a direction. The object can be in a low level of pixel intensify or a high level of pixel intensify causing difficulty in determining the object is in a location, difficult to identify the object, and / or difficult to determine characteristics of the objects. The visual criterion or criteria can be based on or satisfied based on a discernable image characteristic(s). For example, portion 112 can be (or include) an object in the real- world captured by camera 1. The object can have a combination of a shape(s). size(s). color(s), border(s), and / or the like that are of special interest in addition to being indiscernible for any reason (e.g., lack of light).
[0033] In some implementations, the portion 112 can be captured by another camera of the wearable device 100. For example, portion 112’ of an image can be captured by camera 2 (e.g., an infrared camera(s)). The portion 112’ of the imagecaptured by camera 2 can be perceptible by the human eye (e.g. visible). For example, FIG. 1A illustrates portion 112’ of the image captured by camera 2 including a stop sign. Some implementations can include modifying portion 1 12 of the image captured by camera 1 with the corresponding portion 112’ of the image captured by camera 2. As show n in FIG. 1A camera 1 + camera 2 can be used to generate an image (sometimes called merging or fusing two or more images) such that portion 112” of the image can be perceptible by the human eye (e.g. visible). This generated image can then be displayed or rendered on transparent display 106. In some implementations, displaying or rendering the generated (e.g., modified) image on transparent display 106 enhance the human visual perception of a user of the wearable device 100.
[0034] FIG. IB illustrates a block diagram of a video see-through system according to an example implementation. The video see-through system can be implemented in, for example, the wearable device 100 of (e.g., a processor of the wearable device 100 of) FIG. 1 A. The see-through sy stem can be a system of a wearable device. The see-through system can be incorporated in the wearable device and / or a companion device of the wearable device. The wearable device can be a virtual reality (VR) device, an augmented reality (AR) device, an AR / VR device, a mixed reality device, a head mounted display, and the like. As shown in FIG. IB, the video see- through system includes camera 105-1, 105-2, 105-3, ..., 105-n, a feature map module 110, map 115-1, 115-2, 115-3. ..., 115-m, and a fusion module 120. The camera 105-1. 105-2, 105-3, ..., 105-n can be configured to generate image 5-1 , 5-2, 5-3, ..., 5-n. The camera 105-1, 105-2, 105-3, ..., 105-n can be a video sensor including RGB video, infrared camera, time-of-flight depth camera, thermal camera, event camera, and / or the like. In some implementations, the images 5-1, 5-2, 5-3, ..., 5-n can be multi-modal images. In some implementations, the images 5-1, 5-2, 5-3, ..., 5-n can be associated with one frame of a video to be displayed on a display of the wearable device. In other words, image 5-1, 5-2, 5-3, ..., 5-n can be frame and / or image data used to generate a frame.
[0035] The feature map module 110 can be configured to generate maps 115-1. 115-2, 115-3, ..., 115-m. The cameras 105-1, 105-2, 105-3, ..., 105-n can be a video sensor including RGB video, infrared camera, time-of-flight depth camera, thermal camera, event camera, and / or the like. Accordingly, the feature map module 110 can be configured to generate a map based on RGB video data, infrared image data, time-of- flight depth image data, thermal image data, event image data, and / or the like.Therefore, the feature map module 110 can be configured to generate a map based on multi-modal images. Further, the map 115-1, 115-2. 115-3, ..., 115-m can be a thermal map, a depth map, a map based on one sensor (or camera), a map based on two or more sensors (or cameras), a map based on all cameras, and / or the like. A map or feature map can be the output of a convolutional layer after the convolution operation has been performed on an image passed to the convolutional layer. A map or feature map can identify features of interest. A map or feature map can identify low-level features or patterns, but as the features or patterns propagate to successive layers, the map or feature map can identify new patterns and combine them to form high-level features or patterns. A map or feature map can be propagated to multiple layers in the neural network for image processing. Therefore, the features identified in previous layers can be propagated to each successive layer of the neural network. A map or feature map can be used to predict an object in the image. A map or feature map can divide the input image into different portions representing a part of the input image.
[0036] In some implementations, m and n can be equal, or the feature map module 110 can be configured to generate the same number of maps as there are cameras. In some implementations, m can be greater than n, or the feature map module 110 can be configured to generate more maps than there are cameras. In some implementations, m can be less than n, or the feature map module 110 can be configured to generate fewer maps than there are cameras. In some implementations, the map 115- 1 , 115-2, 1 15-3, ..., 115-m can be a saliency map or, for example, a saliency thermal map, a saliency depth map, and / or the like. In some implementations, the map 115-1, 115-2, 115-3, ..., 115-m can be a semantic segmentation map or, for example, a semantic segmentation thermal map, a semantic segmentation depth map. and / or the like. In some implementations, the map 115-1, 115-2, 115-3, ..., 115-m can be a distance map or, for example, a distance thermal map, a distance depth map, and / or the like.
[0037] The fusion module 120 is configured to generate an enhanced frame 10. The fusion module 120 can be configured to generate the enhanced frame 10 based on the image 5-1, 5-2. 5-3, .... 5-n. The fusion module 120 can be configured to generate the enhanced frame 10 based on the map 1 15-1, 115-2, 115-3, ..., 115-m. The fusion module 120 can be configured to generate the enhanced frame 10 based on the image 5-1, 5-2, 5-3, ..., 5-n and the map 115-1, 115-2, 115-3. ..., 115-m. Therefore, the fusion module 120 can be configured to generate the enhanced frame 10 based on multi-modal images (e.g., the image 5-1, 5-2, 5-3, ..., 5-n) and the map 115-1, 115-2, 1 15-3, ..., 115-m. In some implementations, camera 105-1 can be an RGB video camera and one of camera 105-2, 105-3, .... 105-n can be an infrared camera, and the fusion module 120 can be configured to generate the enhanced frame 10 based on the image 5-1, and the image 5-2, 5-3, ..., 5-n associated with the infrared camera. In some implementations, camera 105-1 can be an RGB video camera and one of camera 105-2, 105-3, ..., 105-n can be a thermal camera, and the fusion module 120 can be configured to generate the enhanced frame 10 based on the image 5-1. and the image 5-2. 5-3, .... 5-n associated with the thermal camera. In some implementations, camera 105-1 can be an RGB video camera and one of camera 105-2, 105-3, ..., 105-n can be a depth camera, and the fusion module 120 can be configured to generate the enhanced frame 10 based on the image 5-1, and the image 5-2, 5-3. .... 5-n associated with the depth camera. In some implementations, camera 105-1 can be an RGB video camera and camera 105-2, 105- 3, ..., 105-n can be two or more cameras, and the fusion module 120 can be configured to generate the enhanced frame 10 based on the image 5-1, and the image 5-2, 5-3, ..., 5-n associated with the two or more cameras.
[0038] The feature map module 110 can be configured to generate map 115-1. 115-2, 1 15-3, ..., 115-m using, for example, saliency analysis and / or semantic segmentation. In some implementations, the saliency analysis and / or semantic segmentation can be based on an algorithm (e g., a Fourier transformation, a Gaussian function, a Fourier feature mapping algorithm, and the like). For example, a Fourier transform can be (or can generate) a spectrum transformation. In other words, a Fourier transform can be used to generate map 115-1, 115-2, 115-3, ..., 115-m as a heat map where each point on the map represents a sine wave’s amplitude, and a position on the map represents a frequency. The sine waves’ amplitude and frequency can be used as a filter indicating portions of the image 5-1, image 5-2, 5-3, ..., 5-n that can be difficult for the human eye to perceive or be invisible when displayed on the display of the wearable device.
[0039] In some implementations, the saliency analysis and / or semantic segmentation can be based on a neural network or machine learning model. Saliency analysis can be an analysis of an image to determine a likelihood (e.g., a degree of importance) of being focused on by a human viewer. Saliency analysis can identify aSaliency can be based on image features, objects, motion (e.g., optical flow, frame-to- frame motion, etc.), statistics, and / or the like. In some implementations, saliencyanalysis can use an algorithm to generate bounding boxes indicating a location of an object witbin an image. In some implementations, a saliency analysis can be used to generate a saliency map. A saliency map can be an image where the brightness of each pixel in the image indicates the saliency of the pixel. A saliency map can be used to filter an image where the less salient pixels (e.g., a brightness below a threshold) are filtered from the image. Semantic segmentation can be an algorithm that associates a label or category to a pixel in an image. A plurality of pixels having a label, class, or category (e.g., predefined label, class, or category) can be identified as being (or predicted to be) of interest.
[0040] In some implementations, saliency analysis and / or semantic segmentation can be performed using a machine learning algorithm. In some implementations, the machine learning algorithm can be trained to identify portions of an image (e.g., objects or features) that can be indiscernible by the human eye (e.g., determined to be indiscernible by a human eye based on a visual criterion or criteria (e.g., the visual criterion or criteria discussed above)). For example, the saliency analysis can be based on a U-Net model, a deep convolutional neural network (CNN) or DCNN model, a fully convolutional network (FCN) model, a pyramid feature attention network model, and the like. For example, a U-Net model, a CNN, DCNN model, an FCN model, a pyramid feature attention network model, and the like can be trained to perform saliency analysis and / or semantic segmentation on image 5-1, 5-2. 5-3, ..., 5-n. In other words, a U-Net model, a CNN, DCNN model, an FCN model, a pyramid feature attention network model, and the like can be trained to generate map 115-1, 115-2. 115-3, ..., 115-m as a saliency map or a semantic segmentation map. The saliency map or a semantic segmentation map can be used as a filter indicating portions of the image 5-1, 5-2, 5-3, ..., 5-n that can be difficult for the human eye to perceive or be invisible when displayed on the display of the wearable device.
[0041] Some implementations can include enhancing human visual perception by merging or fusing video signals, images, and / or image data from multiple sensors in real-time. For example, some implementations can be configured to merge or fuse image data (e.g., multi-modal images) captured by video sensors, including RGB video, infrared cameras, time-of-flight depth cameras, thermal cameras, event cameras, and / or the like to generate an enhanced frame and / or image. The fusion module 120 can be configured to generate the enhanced frame 10. The fusion module 120 can be configured to generate the enhanced frame 10 based on the image 5-1, 5-2, 5-3, ..., 5-nand the map 115-1, 115-2, 115-3, ..., 115-m. In some implementations, camera 105-1 can be an RGB video camera, and the fusion module 120 can be configured to merge or fuse a portion of an image captured by one or more of camera 105-2, 105-3, ..., 105- n. The portion of an image captured by one or more of camera 105-2, 105-3, ..., 105-n can be based on filtering image captured by one or more of camera 105-2, 105-3, ..., 105-n using map 115-1, 115-2, 115-3, ..., 115-m. The portion of an image captured by one or more of cameras 105-2, 105-3. ..., 105-n can be difficult for the human eye to perceive or be invisible when rendered as an RGB frame or image. Therefore, the fusion module 120 can be configured to merge or fuse (e.g., overlay) a portion of, for example, a thermal image or an infrared image with the RGB image. The portion of. for example, the thermal image or the infrared image can represent an object that can be difficult for the human eye to perceive or be invisible when rendered as an RGB frame or image making the object visible as an infrared or thermal image overlaying the RGB frame.
[0042] In some implementations, the fusion module 120 can be configured to fuse two or more images using an image blending algorithm. In some implementations, the fusion module 120 can be configured to merge or fuse two or more images using a neural network or machine learning model (e g., a U-Net based neural network). Image fusion (e.g., merging or fusing two or more images) can be combining two or more into one image that can preserve details of each image. Image fusion can generate an image that would not otherwise exist (e.g., does not exist in the view of a camera). Image fusion can substitute (or overlay) a portion of a first image with a portion of a second image. For example, the first image can include a portion that is imperceivable, degraded, missing, and / or the like. Image fusion can substitute or overlay the imperceivable, degraded, missing, and / or the like portion of the first image with a portion of the second image such that the resultant image includes a perceivable, improved, and the like portion. In some implementations, a map, feature map, saliency map and the like can be used to filter the first and second images such that the imperceivable, degraded, missing, and / or the like portion of the first image is removed from the first image and included in the second image (e.g.. a background is removed) prior to merging or fusing (or overlaying) the second image with the first image.
[0043] In some implementations, the image blending algorithm can be configured to apply an additive depth image (d) onto camera images (s) to create a merged or fused video (r) as r = (d < 0.5) ? 2.0 * s * d : 1.0 - 2.0 * (1.0 - s) * (1.0 - d). Alternatively, the image blending algorithm can be configured to apply an additivedepth image (d) onto camera images (s) to create a merged fused video (r) as r = (s < 0.5) ? d - (1.0 - 2.0 * s) * d * (1.0 - d) : (d < 0.25) ? d + (2.0 * s - 1.0) * d * ((16.0 * d - 12.0) * d + 3.0) : d + (2.0 * s - 1.0) * (sqrt(d) - d).
[0044] In some implementations, the neural network or machine learning model can be a UNet-based neural network where a set of training data that contains source videos from various cameras and the desired selective fused video as the output can be used to train the U-Net based neural network to predict a fused video frame as an enhanced frame including portions of a frame that should be perceived by a user that may not have without the enhancement.
[0045] In some implementations, the feature map module 110 can be implemented as a machine learning model and the fusion module 120 can be implemented as an algorithm. In some implementations, the feature map module 110 can be implemented as an algorithm and the fusion module 120 can be implemented as a machine learning model. In some implementations, the feature map module 110 can be implemented as a machine learning model and the fusion module 120 can be implemented as a machine learning model.
[0046] In some implementations, a user of the w earable device can use hand gestures to augment the visible RGB video. For example, virtual magnified glasses can be rendered when the user holds up their hands in a holding gesture, and the fused video can be displayed on a display of the wearable device. Alternatively, virtual objects such as papers, phones, screens, windows, and / or the like can be used to enhance the visual experience.
[0047] In some implementations, the augmented frame 10 can be generated based on a distance or semantic implications of a scene. For example, the thermal camera image(s) can be merged or fused with the depth camera images if objects that are within a predetermined distance (e.g., Im - 1.5 meters) from the user, and / or the thermal camera images can be merged or fused within a window of the scene that the user is gazing at. In some implementations, the generating of the augmented frame 10 can be enabled and disabled by the user of the wearable device.
[0048] FIG. 2A illustrates a block diagram of a video see-through system according to an example implementation. In the implementation illustrated in FIG. 2A, the feature map module 110 can be implemented as a machine learning model. The see- through system can be a system of a wearable device. The see-through system can be incorporated in the w earable device and / or a companion device of the w earable device.The wearable device can be a virtual reality (VR) device, an augmented reality (AR) device, an AR / VR device, a mixed reality device, ahead mounted display, and the like. As shown in FIG. 2 A, the video see-through system includes camera 105-1, 105-2, 105- 3, ..., 105-n, a machine learning model 205, the map 1 15-1, 115-2, 115-3, ..., 115-m, and the fusion module 120. The camera 105-1, 105-2, 105-3, ..., 105-n can be configured to generate image 5-1, 5-2, 5-3. ..., 5-n. In some implementations, the image 5-1, 5-2, 5-3. ..., 5-n can be associated with one frame of a video to be displayed on a display of the wearable device. In other words, image 5-1, 5-2, 5-3, ..., 5-n can be frame and / or image date used to generate a frame.
[0049] The machine learning model 205 can be configured to generate map 115-1, 115-2, 115-3. ..., 115-m. The camera 105-1, 105-2. 105-3, ..., 105-n can be a video sensor including RGB video, infrared camera, time-of-flight depth camera, thermal camera, event camera, and / or the like. Accordingly, the machine learning model 205 can be configured to generate a map based on RGB video data, infrared image data, time-of-flight depth image data, thermal image data, event image data, and / or the like (together sometimes called multi-modal images). Further, the map 115- 1, 115-2, 115-3, ..., 115-m can be a thermal map, a depth map, a map based on one sensor (or camera), a map based on two or more sensors (or cameras), a map based on all cameras, and / or the like.
[0050] In some implementations, m and n can be equal, or the machine learning model 205 can be configured to generate the same number of maps as there are cameras. In some implementations, m can be greater than n, or the machine learning model 205 can be configured to generate more maps than there are cameras. In some implementations, m can be less than n, or the machine learning model 205 can be configured to generate fewer maps than there are cameras. In some implementations, the map 115-1, 115-2, 115-3, ..., 115-m can be a saliency map or, for example, a saliency thermal map, a saliency depth map, and / or the like. In some implementations, the map 115-1, 115-2, 115-3, ..., 115-m can be a semantic segmentation map or, for example, a semantic segmentation thermal map, a semantic segmentation depth map. and / or the like. In some implementations, the map 115-1, 115-2, 115-3, ..., 115-m can be a distance map or, for example, a distance thermal map, a distance depth map, and / or the like.
[0051] In some implementations, the machine learning model can be configured to generate map 115-1, 115-2, 1 15-3, ..., 115-m using, for example, saliency analysisand / or semantic segmentation. For example, the saliency analysis can be based on a U- Net model, a deep convolutional neural network (CNN) or DCNN model, a fully convolutional network (FCN) model, a pyramid feature attention network model, and the like. For example, aU-Net model, a CNN, DCNN model, an FCN model, a pyramid feature attention network model, and the like can be trained to perform saliency analysis and / or semantic segmentation on image 5-1, 5-2, 5-3. ..., 5-n. In other words, a U-Net model, a CNN. DCNN model, an FCN model, a pyramid feature attention network model, and the like can be trained to generate map 115-1, 115-2, 115-3, ..., 115-m as a saliency map or a semantic segmentation map. The saliency map or a semantic segmentation map can be used as a filter indicating portions of the image 5-1. 5-2, 5-3, ..., 5-n that can be difficult for the human eye to perceive or be invisible when displayed on the display of the wearable device.
[0052] FIG. 2B illustrates a block diagram of a machine learning model 205 according to an example implementation. FIG. 2B is a machine learning model 205 implemented as a saliency network based on a U-Net model. However, alternative implementations are within the scope of this disclosure. For example, the machine learning model 205 can be implemented as a semantic segmentation model. Further, the machine learning model 205 can be implemented as a neural network including, for example, a CNN, DCNN model, an FCN model, a pyramid feature attention network model, and / or the like.
[0053] As shown in FIG. 2B, the machine learning model 205 includes convolutional networks 220, feature map 235-1, 235-2, 235-3, 235-4, 235-5, 235-6, 235-7, convolution 245, 260, 265, a context-aware pyramid feature extraction (CPFE) module 240, a channel-wise attention (CA) module 250, a spatial attention (SA) module 255, an upsample module 270, and a loss function 275. In FIG. 2A, frame data 210 can be used as input to the machine learning module. The machine learning module 205 can be configured to generate a map 280.
[0054] In FIG. 2B a machine learning model 205 operation phase can include convolutional networks 220, feature map 235-1. 235-2, 235-3, 235-4, 235-5. 235-6. 235-7, convolution 245, 260, 265, the CPFE module 240, the CA module 250, and the SA module 255. In FIG. 2B a machine learning model training phase can include convolutional networks 220, feature map 235-1, 235-2, 235-3, 235-4, 235-5, 235-6, 235-7, convolution 245, 260, 265, the CPFE module 240. the CA module 250. the SA module 255, and the loss function 275. In the operation phase, frame data 210 can bereceived from wearable device sensors. For example, in the operation phase the frame data 210 can be (or correspond to) image 5-1, 5-2, 5-3, ..., 5-n. Therefore, the machine learning module 255 can be configured to generate a map 280 based on image 5-1, 5- 2, 5-3, ..., 5-n. Accordingly, in the operation phase the map 280 can be (or correspond to) map 115-1, 115-2, 115-3, ..., 115-m.
[0055] In the training phase, frame data 210 can be training data. In some implementations, the training data can be a large dataset that includes, for example. RGB images, infrared camera images, depth images, thermal camera images, and the like for the same scene. Therefore, the machine learning model 205 can be configured to generate a map 280 based on the training data. In some implementations, the training data can be preprocessed to resize and normalize the images in the dataset. In the training phase, a ground truth map 215 can be used as input to the machine learning module. The ground truth map 215 can be a target map that is desired to be generated by the machine learning model 205 during, for example, a supervised training operation. In the training phase, the loss function 275 can be configured to generate a loss indicating the result of a comparison of the ground truth map 215 and the map 280. In other words, the machine learning model 205 can be trained based on minimizing a loss function between the predicted map 280 generated by the machine learning model 205 and the ground truth map 280.
[0056] In some implementations, frame data 210 (as training data) can be a labelled dataset with the saliency maps. The saliency maps can indicate which objects in the scene are more important or relevant to human perception. In some implementations, the machine learning model 205 can have multiple hidden layers with an increasing number of filters and pooling layers to reduce the spatial dimension of the input images. A labelled dataset can be a dataset including data with a meaningful and informative label(s) that provides context for training a machine learning model An object can be relevant to human perception as a result of the training of the machine learning model. For example, an object can be labelled based on a prior decision. For example, a training image can include plexiglass that has been labelled. Labelling plexiglass indicates that plexiglass is an object relevant to human perception. Additional objects can be labelled. In some implementations, the labelled objects can be objects that can be indiscernible by the human eye. In some implementations, a portion of an image can be determined to be indiscernible by a human eye based on a visual criterion or criteria. The labelled objects can be labelled based on the visualcriterion or criteria such the machine learning model can be trained to determine (e.g., predict) that a portion of an image can be determined to be indiscernible by a human eye. Therefore, the visual criterion or criteria can be based on transparency, light / darkness, object obstruction, object differentiation (e.g., similarity to adjacent or nearby objects), context, object features, and / or the like.
[0057] In some implementations, the input to the feature map 235-1 can be low- level features 225. Accordingly, conv 1-2 and conv 2-2 together can be configured to generate low-level features 225. In some implementations, the input to the CPFE module 240 can be high-level features 230. Accordingly, conv 3-3, conv 4-3, and conv 5-3 together can be configured to generate high-level features 230. The convolutional networks 220 can be visual geometry group networks (VGGNet) sometimes called a VGG model. The VGGNet can be a deep CNN with multiple layers. Convolutional layers can learn to identify low-level features such as edges and comers, and then combine them to identify high-level features (e.g., more complex features) such as objects, shapes, and patterns.
[0058] The CPFE module 240 can be configured to identify context information. Context information (or contextual information) can be information corresponding to nearby pixels in the image. Context information (or contextual information) can be surrounding information, circumstances, and conditions that affect the predictions associated with a pixel(s). Context information (or contextual information) can be associated with patterns of pixels in a portion of the image around a pixel being evaluated. Context information (or contextual information) can be used to classify a pixel. Using context information (or contextual information) can minimize noise in predicting a classification. Context information (or contextual information) can be used in semantic analysis. For example, the CPFE module 240 can be configured to identify context information associated with the high-level features (of the images) generated by the conv 3-3, conv 4-3, and conv 5-3. In some implementations, the conv 3-3, conv 4-3, and conv 5-3 can assign weights (e.g.. during the training phase) to channels having a high response to objects that can be difficult for the human eye to perceive or be invisible when rendered as an RGB frame or image. Further, the CPFE module 240 can be configured to filter background details based on the high-level features which can focus more on the foreground regions. In some implementations, the CPFE module 240 can be configured to include a Laplacian of a Gaussian representation. The Laplacian of a Gaussian representation can fuse scale spacerepresentations and pyramid multi-resolution representations. The scale space representations can be processed by different Gaussian kernel functions with the same (or substantially the same) resolution. Further, the pyramid multi-resolution representations can be processed by down samples with different resolutions.
[0059] The CA module 250 can be configured to select a scale and receptive field for generating saliency regions. A receptive field can be a region in an input space that affects the features of a particular layer in a machine learning model (e.g., neural network). The receptive field of a feature (e.g., object) can correspond to the center location of a feature and the size of the feature. The receptive field can have a scale (or size) for a region in a layer. The scale can be fixed and / or variable and trained. A region of interest can be a portion of an input image that is being analyzed. Some machine learning models include layers. In a first layer the region of interest can be the entire image and in subsequent layers, there can be a plurality of regions of interest can be that can each be analyzed independently .
[0060] In the training phase, weights associated with the CA can be assigned to channels based on saliency detection (e.g.. a higher weight for higher saliency compared to a lower saliency). In order to refine the boundaries of saliency regions, low-level features can be merged or fused with edge information.
[0061] The SA module 255 can be configured to refine saliency boundaries based on low-level features. In some implementations, the SA module 255 can refine saliency boundaries by filtering (e g., removing) background details based on high-level features. During the training phase, the SA module 255 can be trained (e.g., weights modified) to generate low-level feature maps in order to refine salient object details. During the training phase, the SA module 255 can be trained (e.g.. weights modified) to preserve edge loss in order to learn information in boundary localization.
[0062] The upsample module 270 can be configured to increase the resolution of the data representing an image, a map, and / or the like. The loss function 275 can represent the price paid for inaccuracy of map predictions. In some implementations, the loss function 275 can be a cross-entropy loss between the map 280 and the ground truth map 215.
[0063] FIG. 3A illustrates a block diagram of a video see-through system according to an example implementation. In the implementation illustrated in FIG. 3A, the fusion module 120 can be implemented as a machine learning model. The see- through system can be a system of a wearable device. The see-through system can beincorporated in the wearable device and / or a companion device of the wearable device. The wearable device can be a virtual reality’ (VR) device, an augmented reality (AR) device, an AR / VR device, a mixed reality device, ahead mounted display, and the like. As shown in FIG. 3 A, the video see-through system includes camera 105-1, 105-2, 105- 3, ..., 105-n, a machine learning model 205, the map 115-1, 115-2, 115-3, ..., 115-m, and the fusion module 120. The camera 105-1, 105-2, 105-3, ..., 105-n can be configured to generate image 5-1, 5-2, 5-3. ..., 5-n. In some implementations, the image 5-1, 5-2, 5-3, ..., 5-n can be associated with one frame of a video to be displayed on a display of the wearable device. In other words, image 5-1, 5-2, 5-3, ..., 5-n can be frame and / or image date used to generate a frame.
[0064] The machine learning model 305 can be configured to generate the enhanced frame 10. The machine learning model 305 can be configured to generate the enhanced frame 10 based on the image 5-1, 5-2, 5-3, ..., 5-n and the map 115-1, 115-2, 115-3, ..., 115-m. In some implementations, camera 105-1 can be an RGB video camera, and the machine learning model 305 can be configured to merge or fuse a portion of an image captured by one or more of camera 105-2, 105-3. .... 105-n. The portion of an image captured by one or more of camera 105-2, 105-3, ..., 105-n can be based on filtering image captured by one or more of camera 105-2, 105-3, ..., 105-n using map 115-1, 115-2, 115-3, .... 115-m. The portion of an image captured by one or more of camera 105-2. 105-3, ..., 105-n can be difficult for the human eye to perceive or be invisible when rendered as an RGB frame or image. Therefore, the machine learning model 305 can be configured to merge or fuse (e.g., overlay) a portion of, for example, a thermal image or an infrared image with the RGB image. The portion of, for example, the thermal image or the infrared image can represent an object that can be difficult for the human eye to perceive or be invisible when rendered as an RGB frame or image making the object visible as an infrared or thermal image overlaying the RGB frame.
[0065] In some implementations, the machine learning model 305 can be a neural network (e.g.. a U-Net based neural network). For example, the input layer of the machine learning model 305 can take in the RGB image, infrared camera image, depth image, and thermal camera image, as well the 115-1, 115-2, 115-3, ..., 115-m. The machine learning model 305 can have multiple hidden layers with increasing number of filters and pooling layers to reduce the spatial dimension of the input images.The output layer of the machine learning model 305 can generate the enhanced image 10 that highlights objects that are invisible from human eyes.
[0066] FIG. 3B illustrates a block diagram of the machine learning model 305 according to an example implementation. FIG. 3B is the machine learning model 305 implemented as a neural network based on a U-Net model. However, alternative implementations are within the scope of this disclosure. For example, the machine learning model 305 can be implemented as a neural network including, for example, a CNN, DCNN model, an FCN model, and / or the like.
[0067] FIG. 3B illustrates the machine learning model 305 operation phase can include convolution layers 325, 330, 335, 340, 350, 355, 360, 365. In a training phase, the machine learning model 305 further includes a loss function 375. The machine learning model 305 receives input from camera 105-1, 105-2, 105-3, ..., 105-n (e.g., image 5-1, 5-2, 5-3, ..., 5-n) as input 310 and the map 115-1, 115-2, 115-3, ..., 115-m as input 315. In other words, input 310 can be frame data associated with RGB video, infrared cameras, time-of-flight depth cameras, thermal cameras, event cameras, and / or the hke and the input 315 can be a map(s) generated based on the frame data associated with RGB video, infrared cameras, time-of-flight depth cameras, thermal cameras, event cameras, and / or the like. The output 370 can be an enhanced frame that can correspond to enhanced frame 10 when the machine learning model 305 is used in an operation phase.
[0068] In some implementations, the convolution layers 325, 330, 335, 340 can form an inference model (e.g., an encoder). In some implementations, the convolution layers 325. 330, 335, 340 can include a plurality of convolutions (three (3) are illustrated). In some implementations, the convolution layers 350, 355, 360, 365 can form a segmentation model (e.g., a decoder). In some implementations, the convolution layers 350, 355, 360, 365 can include a plurality of convolutions (three (3) are illustrated). The inference-segmentation model (encoder-decoder) can include skip connection(s) (e.g., the output of a convolution of convolution layer 325 can be communicated as input to a convolution of convolution layer 365) and other computations needed to realize the segmentation model outputs.
[0069] A convolution layer (e.g., 325, 330, 335, 340, 345, 350, 355, 360, 365) or convolution can be configured to extract features from a frame, a map, and / or an image. Features can be based on color, frequency domain, edge detectors, and / or the like. A convolution can have a filter (sometimes called a kernel) and a stride. Forexample, a filter can be a 1x1 filter (or Ixlxn for a transformation to n output channels, a 1x1 filter is sometimes called a pointwise convolution) with a stride of 1 which results in an output of a cell generated based on a combination (e.g., addition, subtraction, multiplication, and / or the like) of the features of the cells of each channel at a position of the A xA grid. In other words, a feature map having more than one depth or channel is combined into a feature map having a single depth or channel. A filter can be a 3x3 filter with a stride of 1 which results in an output with fewer cells in / for each channel of the AfxAL grid or feature map. The output can have the same depth or number of channels (e.g., a 3x3xw filter, where n = depth or number of channels, sometimes called a depthwise filter) or a reduced depth or number of channels (e.g., a 3x3xk filter, where k<depth or number of channels). Each channel, depth, or feature map can have an associated filter. Each associated filter can be configured to emphasize different aspects of a channel. In other words, different features can be extracted from each channel based on the filter (this is sometimes called a depthwise separable filter). Other filters are within the scope of this disclosure.
[0070] Another type of convolution can be a combination of two or more convolutions. For example, a convolution can be a depthwise and pointwise separable convolution. This can include, for example, a convolution in two steps. The first step can be a depthwise convolution (e.g., a 3x3 convolution). The second step can be a pointwise convolution (e.g., a 1x1 convolution). The depthwise and pointwise convolution can be a separable convolution in that a different filter (e.g., filters to extract different features) can be used for each channel or each depth of a feature map. In an example implementation, the pointwise convolution can transform the feature map to include c channels based on the filter. For example, an 8x8x3 feature map (or image) can be transformed to an 8x8x256 feature map (or image) based on the filter. In some implementations, more than one filter can be used to transform the feature map (or image) to an MxMxc feature map (or image).
[0071] A convolution can be linear. A linear convolution describes the output, in terms of the input, as being linear time-invariant (LTI). Convolutions can also include a rectified linear unit (ReLU). A ReLU is an activation function that rectifies the LTI output of a convolution and limits the rectified output to a maximum. A ReLU can be used to accelerate convergence (e.g.. more efficient computation).
[0072] In some implementations, a combination of depthwise convolutions and depthwise and pointwise separable convolutions can be used. Each of the convolutionscan be configurable (e.g., configurable feature, stride and / or depth). For example, the convolution layers 325, 330, 335, 340 can transform the input 310 and the input 315 into a feature map 345. The convolution layers 325, 330, 335, 340 can incrementally transform the input 310 and the input 315 into feature map 345. For example, the convolution layers 350, 355, 360, 365 can transform the feature map 345 into output 370. The convolution layers 350, 355, 360, 365 can incrementally transform the feature map 345 into output 370.
[0073] In an operation phase, output 370 can correspond to enhanced frame 10. In a training phase, the input 310 and the input 315 can be training data. In some implementations, the training data can be a large dataset that includes, for example, RGB images, infrared camera images, depth images, thermal camera images, maps and the like for the same scene. Therefore, the machine learning model 305 can be configured to generate output 370 as a predicted frame based on the training data. In some implementations, the training data can be preprocessed to resize and normalize the images in the dataset. In the training phase, a ground truth frame 380 can be used as input to the machine learning module 305. The ground truth frame 380 can be a target frame (e.g., image) that is desired to be generated by the machine learning model 305 during, for example, a supervised training operation. In the training phase, the loss function 375 can be configured to generate a loss indicating the result of a comparison of the ground truth frame 380 and the output 370 as a predicted frame. The loss function 375 can represent the price paid for inaccuracy of image and / or frame predictions. In some implementations, the loss function 375 can be a cross-entropy loss between the output 370 (e.g., predicted frame) and the ground truth frame 380. In other words, the machine learning model 305 can be trained based on minimizing a loss function between the predicted frame (output 370) generated by the machine learning model 305 and the ground truth frame 380.
[0074] In some implementations, the machine learning model 205 and the machine model 305 can be trained together. The machine learning model 205 can be trained to predict the map(s) and the machine learning model 305 can be trained to predict the enhanced frame based on the map and the images. A suitable loss function can be used to evaluate the accuracy of the prediction and optimize the machine learning model 205 and the machine model 305 weights (e.g., weights associated with the convolutions) during the training. In some implementations, the performance of the trained machine learning models can be evaluated on a separate test dataset. In someimplementations, the results of the enhanced image produced by the machine learning model 305 can be compared to the original images to determine the effectiveness of the fusion. In some implementations, the architecture and hyperparameters of the machine learning models can be fine-tuned to improve the performance of the machine learning models. In some implementations, the training and evaluation steps can be repeated until the desired level of performance is achieved.
[0075] FIG. 4 illustrates a block diagram of a video see-through system according to an example implementation. In the implementation illustrated in FIG. 4, the feature map module 110 and the fusion module 120 can be implemented as a machine learning model. The see-through system can be a system of a wearable device. The see-through system can be incorporated in the wearable device and / or a companion device of the wearable device. The wearable device can be a virtual reality (VR) device, an augmented reality (AR) device, an AR / VR device, a mixed reality device, a head mounted display, and the like. As shown in FIG. 4, the video see-through system includes camera 105-1, 105-2. 105-3, ..., 105-n, the machine learning model 205, the map 115-1. 115-2, 115-3, .... 1 15-m. and the machine learning model 305. The camera 105-1, 105-2, 105-3, ..., 105-n can be configured to generate image 5-1, 5-2, 5-3, ..., 5- n. In some implementations, the image 5-1, 5-2, 5-3, ..., 5-n can be associated with one frame of a video to be displayed on a display of the wearable device. In other words, image 5-1, 5-2. 5-3, .... 5-n can be frame and / or image date used to generate a frame.
[0076] The machine learning model 205 can be configured to generate map 115-1, 115-2, 115-3, ..., 115-m. The camera 105-1, 105-2, 105-3, ..., 105-n can be a video sensor including RGB video, infrared camera, time-of-flight depth camera, thermal camera, event camera, and / or the like. Accordingly, the machine learning model 205 can be configured to generate a map based on RGB video data, infrared image data, time-of-flight depth image data, thermal image data, event image data, and / or the like (together sometimes called multi-modal images). Further, the map 115- 1, 115-2, 115-3, ..., 115-m can be a thermal map, a depth map, a map based on one sensor (or camera), a map based on two or more sensors (or cameras), a map based on all cameras, and / or the like.
[0077] In some implementations, m and n can be equal, or the machine learning model 205 can be configured to generate the same number of maps as there are cameras. In some implementations, m can be greater than n, or the machine learning model 205 can be configured to generate more maps than there are cameras. In someimplementations, m can be less than n, or the machine learning model 205 can be configured to generate fewer maps than there are cameras. In some implementations, the map 115-1, 115-2, 115-3, ..., 115-m can be a saliency map or, for example, a saliency thermal map, a saliency depth map, and / or the like. In some implementations, the map 115-1, 115-2, 115-3, ..., 115-m can be a semantic segmentation map or, for example, a semantic segmentation thermal map, a semantic segmentation depth map, and / or the like. In some implementations, the map 115-1. 115-2, 115-3, .... 115-m can be a distance map or, for example, a distance thermal map, a distance depth map, and / or the like.
[0078] In some implementations, the machine learning model can be configured to generate map 115-1, 115-2. 115-3, ..., 115-m using, for example, saliency analysis and / or semantic segmentation. For example, the saliency analysis can be based on a U- Net model, a deep convolutional neural network (CNN) or DCNN model, a fully convolutional network (FCN) model, a pyramid feature attention network model, and the like. For example, a U-Net model, a CNN, DCNN model, an FCN model, a pyramid feature attention network model, and the like can be trained to perform saliency analysis and / or semantic segmentation on image 5-1, 5-2, 5-3, ..., 5-n. In other words, a U-Net model, a CNN, DCNN model, an FCN model, a pyramid feature attention network model, and the like can be trained to generate map 115-1, 115-2, 115-3, ..., 115-m as a saliency map or a semantic segmentation map. The saliency map or a semantic segmentation map can be used as a filter indicating portions of the image 5-1 , 5-2, 5-3, ..., 5-n that can be difficult for the human eye to perceive or be invisible when displayed on the display of the wearable device.
[0079] The machine learning model 305 can be configured to generate the enhanced frame 10. The machine learning model 305 can be configured to generate the enhanced frame 10 based on the image 5-1, 5-2, 5-3, ..., 5-n and the map 115-1, 115-2, 115-3, ..., 115-m. In some implementations, camera 105-1 can be an RGB video camera, and the machine learning model 305 can be configured to merge or fuse a portion of an image captured by one or more of camera 105-2, 105-3. ..., 105-n. The portion of an image captured by one or more of camera 105-2, 105-3, ..., 105-n can be based on filtering image captured by one or more of camera 105-2, 105-3, ..., 105-n using map 115-1, 115-2, 115-3, .... 115-m. The portion of an image captured by one or more of camera 105-2. 105-3, ..., 105-n can be difficult for the human eye to perceive or be invisible when rendered as an RGB frame or image. Therefore, the machinelearning model 305 can be configured to merge or fuse (e.g., overlay) a portion of, for example, a thermal image or an infrared image with the RGB image. The portion of, for example, the thermal image or the infrared image can represent an object that can be difficult for the human eye to perceive or be invisible when rendered as an RGB frame or image making the object visible as an infrared or thermal image overlaying the RGB frame.
[0080] In some implementations, the machine learning model 305 can be a neural network (e g., a U-Net based neural network). For example, the input layer of the machine learning model 305 can take in the RGB image, infrared camera image, depth image, and thermal camera image, as well the 115-1, 115-2, 115-3, ..., 115-m. The machine learning model 305 can have multiple hidden layers with increasing number of filters and pooling layers to reduce the spatial dimension of the input images. The output layer of the machine learning model 305 can generate the enhanced image 10 that highlights objects that are invisible from human eyes.
[0081] Example 1. FIG. 5 is a block diagram of a method of generating an enhanced frame of a video according to an example implementation. As show n in FIG. 5, in step S505 receiving a first image and a second image (e.g., of a plurality of images) representing a frame of a video, the first image including a first image portion that is determined based on a visual criterion or criteria (or a human visual criterion or criteria). Alternatively (or in addition), the first image can include a portion that is determined to be indiscernible by the human eye based on a visual criterion or criteria. The first image and the second image (e.g., of a plurality' of images) can be captured using sensors of a wearable device and the images can be automatically received upon an indication that an enhanced frame is to be generated. The images can be images captured subsequently. The portion that is determined to be indiscernible by the human eye in the first image can be a first image portion. For example, the first image portion can be indiscernible by the human eye when viewing an RGB image. In step S510 generating a map based on the second image (e.g., the second image of the plurality of images). The first and second image can be separate images. In step S515 merging the first image and the second image using the map as an enhanced frame representing the frame of the video, the enhanced frame including the first image and a second image portion from the second image, the second image portion corresponding to the first image portion. For example, the first image portion can be indiscernible by the human eye when viewing an RGB image. However, the second image portion can bediscernible when viewing, for example, an infrared image, a thermal image, and / or the like. In step S520 rendering the enhanced frame.
[0082] Example 2. The method of Example 1, wherein the first image and the second image (e.g., of a plurality of images) can include two or more of an RGB image, an infrared camera image, a depth camera image, and a thermal camera image.
[0083] Example 3. The method of Example 1, wherein the map can be generated using one of a saliency analysis or a semantic segmentation.
[0084] Example 4. The method of Example 1, wherein the map can be generated using a machine learning model trained based on a labelled dataset including saliency maps that indicate which objects in a scene are relevant to human perception.
[0085] Example 5. The method of Example 1, wherein the map can be generated using a machine learning model including a first neural network configured to receive the plurality of images as stacked image tiles and to generate a set of feature maps for the plurality of images and a second neural network configured to refine boundaries of an object based on low-level features of the set of feature maps. Stacked image tiles can be a plurality of tiles that together form an image. Each layer of the stack can include, for example, incremental increases of granularity'. For example, a bottom layer of the stack can include a coarse granularity' and a top layer of the stack can include a fine granularity. Each layer of the stack can represent a depth. For example, a bottom layer of the stack can include an object(s) and / or feature(s) that is further away from the camera and atop layer of the stack can include an object(s) and / or feature(s) that is closer to the camera. Each layer of the stack can represent a color. For example, a bottom layer of the stack can include a first color and a top layer of the stack can include am nth color. The layers can be in what is often called a z-order.
[0086] Example 6. The method of Example 1, wherein the map can be generated using a machine learning model including a first neural network configured to receive the plurality of images as stacked image tiles and to generate a set of feature maps for the plurality of images, a second neural network configured to identify context information associated with high-level features of the set of feature maps, and a third neural network configured to select a scale and receptive field for a region of interest in the set of feature maps based on the context information.
[0087] Example 7. The method of Example 6, wherein the region of interest can include the portion that is indiscernible by the human eye.
[0088] Example 8. The method of Example 1, wherein the plurality of imagescan be received from sensors of a wearable device configured to display see-through video. See-through video can be is video that is projected onto a display that allows a user to see what is shown on the display and what is in the background. For example, see-through video can be projected onto a display of a wearable device where the user can see both the video and the real-world.
[0089] Example 9. The method of Example 1 , wherein the enhanced frame can include a first image and a portion of a second image, the portion of the second image can include the portion that is indiscernible by the human eye, and the portion of the second image can overlay the portion that is indiscernible by the human eye in the first image.
[0090] Example 10. The method of Example 9, wherein the first image can be an RGB image.
[0091] Example 11. The method of Example 9, wherein the first image can be an RGB image and the second image is an infrared image.
[0092] Example 12. The method of Example 9, wherein the first image can be an RGB image and the second image is a thermal image.
[0093] Example 13. The method of Example 1, wherein the enhanced image can be rendered on a display of a wearable device.
[0094] Example 14. FIG. 6 is a block diagram of a method of a machine learning model according to an example implementation. As shown in FIG. 6, in step S605 receiving training data, the training data includes multi-modal images, saliency maps indicating objects of interest and a ground truth image. In step S610 resizing and normalizing the multi-modal images. In step S615 generating a map(s) based on the multi-modal images using a first machine learning model. In step S620 generating an enhanced frame based on the multi-modal image and the map using a second machine learning model. In step S625 training the first machine learning model based on the maps, the saliency maps, and a first loss function. In step S630 training the second machine learning model based on the enhanced frame and a second loss function.
[0095] Example 15. The method of Example 14, wherein the training data can include RGB images, infrared camera images, depth images, and thermal camera images for the same scene.
[0096] Example 16. The method of Example 14, wherein generating the enhanced frame can include merging (or fusing) a first image of the multi-modal image and a portion of a second image of the multi-modal image.
[0097] Example 17. The method of Example 16, wherein the portion of the second image can be generated based on the map(s).
[0098] Example 18. The method of Example 14, further including evaluating the performance of the trained first machine learning model and the trained second machine learning model using second test data and comparing the results of an enhanced frame generated by the second machine learning model to the original multimodal images to determine the effectiveness of the fusion.
[0099] Example 19. The method of Example 14, wherein the first machine learning model includes a first neural network configured to receive the plurality of images as stacked image tiles and to generate a set of feature maps for the plurality of images and a second neural network configured to refine boundaries of an object based on low-level features of the set of feature maps.
[0100] Example 20. The method of Example 14, wherein the first machine learning model includes a first neural network configured to receive the plurality of images as stacked image tiles and to generate a set of feature maps for the plurality of images, a second neural network configured to identify context information associated with high-level features of the set of feature maps, and a third neural network configured to select a scale and receptive field for a region of interest in the set of feature maps based on the context information.
[0101] Example 21. The method of Example 14, wherein the enhanced frame can include a first image and a portion of a second image, the portion of the second image can include the portion that is indiscernible by the human eye, and the portion of the second image can overlay the portion that is indiscernible by the human eye in the first image.
[0102] Example 22. A method can include any combination of one or more of Example 1 to Example 21.
[0103] Example 23. A non- transitory computer-readable storage medium comprising instructions stored thereon that, when executed by at least one processor, are configured to cause a computing system to perform the method of any of Examples 1-22.
[0104] Example 24. An apparatus comprising means for performing the method of any of Examples 1-22.
[0105] Example 25. An apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and thecomputer program code configured to, with the at least one processor, cause the apparatus at least to perform the method of any of Examples 1-22.
[0106] FIG. 7 illustrates a pictorial and block diagram of a video see-through system according to an example implementation. The see-through system can be a system of a wearable device. The see-through system can be incorporated in the wearable device and / or a companion device of the wearable device. The wearable device can be a virtual reality (VR) device, an augmented reality (AR) device, an AR / VR device, a mixed reality device, a head mounted display, and the like. As shown in FIG. 7, the video see-through system can include the machine learning model 205 and the machine learning model 305. The machine learning model 205 can be configured to generate map(s) 725. 730, 735 based on image(s) 705. 710, 715. 720. The machine learning model 305 can be configured to generate an enhanced frame 740 based on image(s) 705, 710, 715, 720 and map(s) 725, 730, 735.
[0107] The image(s) 705, 710, 715, 720 can be multi-modal images. For example, the image(s) 705, 710, 715. 720 can be generated using a video sensor(s) including RGB video, infrared camera, time-of-flight depth camera, thermal camera, event camera, and / or the like. In some implementations, the image(s) 705, 710, 715, 720 can be associated with one frame of a video to be displayed on a display of the wearable device. In other words, image(s) 705, 710, 715, 720 can be frame and / or image data used to generate a frame.
[0108] Therefore, the machine learning model 205 can be configured to generate map(s) 725, 730, 735 based on RGB video data, infrared image data, time-of- flight depth image data, thermal image data, event image data, and / or the like. Therefore, the machine learning model 205 can be configured to generate a map based on multi-modal images. Further, the map(s) 725, 730, 735 can be athermal map, a depth map, a map based on one sensor (or camera), a map based on two or more sensors (or cameras), a map based on all cameras, and / or the like.
[0109] In some implementations, the map(s) 725, 730, 735 can be a saliency map or. for example, a saliency thermal map, a sahency depth map, and / or the like. In some implementations, the map(s) 725, 730, 735 can be a semantic segmentation map or, for example, a semantic segmentation thermal map, a semantic segmentation depth map, and / or the like. In some implementations, the map(s) 725, 730, 735 can be a distance map or, for example, a distance thermal map. a distance depth map, and / or the like.
[0110] The machine learning model 305 can be configured to generate an enhanced frame 740. The machine learning model 305 can be configured to generate enhanced frame 740 based on the image(s) 705, 710, 715, 720. The machine learning model 305 can be configured to generate an enhanced frame 740 based on the map(s) 725, 730, 735. The machine learning model 305 can be configured to generate an enhanced frame 740 based on the image(s) 705, 710. 715, 720 and the map(s) 725, 730, 735. Therefore, the machine learning model 305 can be configured to generate an enhanced frame 740 based on multi-modal images (e.g., image(s) 705, 710, 715, 720) and the map(s) 725, 730, 735.
[0111] The machine learning model 205 can be configured to generate map(s) 725, 730. 735 using, for example, saliency analysis and / or semantic segmentation. In some implementations, the saliency analysis and / or semantic segmentation can be based on a neural network or machine learning model. For example, the saliency analysis can be based on a U-Net model, a deep convolutional neural network (CNN) or DCNN model, a fully convolutional network (FCN) model, a pyramid feature attention network model, and the like. For example, a U-Net model, a CNN. DCNN model, an FCN model, a pyramid feature attention network model, and the like can be trained to perform saliency analysis and / or semantic segmentation on image(s) 705, 710, 715, 720. In other words, a U-Net model, a CNN, DCNN model, a FCN model, a pyramid feature attention network model, and the like can be trained to generate map(s) 725, 730, 735 as a saliency map or a semantic segmentation map. The saliency map or a semantic segmentation map can be used as a filter indicating portions of the image(s) 705, 710, 715, 720 that can be difficult for the human eye to perceive or be invisible when displayed on the display of the wearable device.
[0112] The machine learning model 305 can be configured to generate an enhanced frame 740. The machine learning model 305 can be configured to generate an enhanced frame 740 based on the image(s) 705, 710, 715, 720 and the map(s) 725, 730, 735. In some implementations, image 705 can be an RGB image, and the machine learning model 305 can be configured to merge or fuse a portion of image 705 with a portion of one or more of image(s) 710, 715, 720. The portion of one or more of image(s) 710, 715, 720 can be based on filtering one or more of image(s) 710, 715, 720 using map(s) 725, 730. 735. The portion of image(s) 710, 715, 720 can be difficult for the human eye to perceive or be invisible when rendered as image 705 as an RGB frame or image. Therefore, the machine learning model 305 can be configured to merge orfuse (e.g., overlay) a portion of, for example, a thermal image or an infrared image with the RGB image. The portion of, for example, the thermal image or the infrared image can represent an object that can be difficult for the human eye to perceive or be invisible when rendered as an RGB frame or image making the object visible as an infrared or thermal image overlaying the RGB frame.
[0113] In some implementations, the machine learning model 305 can be a UNet-based neural network where a set of training data that contains source videos from various cameras and the desired selective fused video as the output can be used to train the U-Net based neural network to predict a fused video frame as an enhanced frame including portions of a frame that should be perceived by a user that may not have without the enhancement.
[0114] FIG. 8 is a front view of an example head mounted wearable device according to an example implementation. In some implementations, the head mounted wearable device can be a virtual reality (VR) device, an augmented reality (AR) device, an AR / VR device, a mixed reality (MR) device, smartglasses, and / or the like including a video see-through system. The example described in FIG. 8 is of a smartglasses 800 implementation.
[0115] The example smartglasses 800 can be configured to sense or capture multi-modal images and generate a feature map or saliency map using the multi-modal images. The multi-modal images can include an image (e.g., RGB image) including a portion that is indiscernible by the human eye. The multi-modal images and the feature map or saliency map can be used in an image merging or fusing operation that includes a first image (e.g., RGB image) with an overlay of a portion of a second image (e.g., infrared image) that includes the portion that is indiscernible by the human eye in an image mode that can be perceived by the human eye. The example smartglasses 800 can be configured to display the fused image or frame as an enhanced frame of a video on a display of the smartglasses 800.
[0116] The example smartglasses 800 includes a frame 810. The frame 810 includes a front frame portion 820. and a pair of temple arm portions 824 rotatably coupled to the front frame portion 820 by respective hinge portions 822. The front frame portion 820 includes rim portions 814 surrounding respective optical portions in the form of lenses 816, with a bridge portion 818 connecting the rim portions 814. The temple arm portions 824 are coupled, for example, pivotably or rotatably coupled, to the front frame portion 820 at peripheral portions of the respective rim portions 814. Insome examples, the lenses 816 are corrective / prescription lenses. In some examples, the lenses 816 are an optical material including glass and / or plastic portions that do not necessarily incorporate corrective / prescription parameters.
[0117] In some examples, the smartglasses 800 can include a display device (not shown) that can output visual content, so that the visual content is visible to the user. In some implementations, the display device can be provided in one or both of the two arm portions 824 to provide for binocular output of content. In some examples, the display device can be a transparent near eye display (e.g., of a video see-through system). In some examples, the display device may be configured to project light from a display source onto a portion of teleprompter glass functioning as a beam splitter seated at an angle (e.g., 30-45 degrees). The beam splitter may allow for reflection and transmission values that allow the light from the display source to be partially reflected while the remaining light is transmitted through. Such an optic design may allow a user to see both physical items in the world, for example, through the lenses 816, next to content (for example, digital images, user interface elements, virtual content, and the like) output by the display device. In some implementations, waveguide optics may be used to depict content on the display device. The digital images can be rendered at an offset from the physical items in the w orld due to errors in head pose data. Therefore, example implementations can use a harmonic exponential filter in a head pose signal processing pipeline to correct for the errors in the head pose data.
[0118] In some implementations, the display device can output visual content, so that the visual content is visible to the user. In some implementations, the visual content can include (and / or be) an image, a frame of a video, and / or a user interface (UI) (e.g., including an image or frame of a video). The UI can include at least one UI element. A UI element can include screens, windows, buttons, menus, toggles, icons, and / or other visual elements that a user uses to interact with the smartglasses 800. Typically, a pointer device (e.g., a mouse) can be used to direct cursors and make selections on a UI operating on a computing device. However, a pointer device may not be practical for use on a wearable device. Therefore, a gesture can be used to interact with the UI operating on the smartglasses 800. The gesture can be generated based on an eye-gaze characteristic(s) associated with gaze tracking of the user of the smartglasses 800 eyes and a head movement of the user of the smartglasses 800.
[0119] In some implementations, the UI can be head-locked, in that the UI remains in the same location on the display when the user moves the user’s head and / orthe smartglasses 800 move. For example, overlaid AR content can move along with the user’s head movements while staying in the user’s field of view (FOV). As an example, in a head-locked visual UI, the position and orientation of the overlaid AR content in the AR display depends only on the position and orientation of the viewer’s head (as if the AR object or content were rigidly attached by a rigid rod to the viewer’s head) independent of any features of the real-world scene. The overlaid AR content may move across the real-world scene in the AR display instantly in response to viewer’s head movements for all types (e.g., all amplitudes, frequencies, and directions, 6 degree of freedom (6DoF) movements) of the viewer’s head motions. For example, a circular or sinusoidal frequency head motion will cause the overlaid AR content to move across the real-world scene at the same circular or sinusoidal frequency as the head motion. Further, for example, a tilt or roll head motion will cause the overlaid AR content to tilt away from a gravity vertical at the same angle as the head motion tilt or roll away from the gravity vertical.
[0120] In some examples, the sensing system 812 may include various sensing devices and the control system (not shown) may include various control system devices (e.g., included in one or both of the tw o arm portions 824) including, for example, one or more processors operably coupled to the components of the control system. In some examples, the control system may include a communication module providing for communication and exchange of information between the smartglasses 800 and other external devices. The sensing system 812 can include sensors 802, 804, 806, 808 (e.g., cameras). For example, the sensing system 812 can include video sensors, including RGB video, infrared cameras, time-of-flight depth cameras, thermal cameras, event cameras, and / or the like used to generate an enhanced frame and / or image. In some implementations, the control system can include one or more processors configured to generate the enhanced frame (described in more detail above).
[0121] In some examples, the smartglasses 800 includes one or more of an audio output device (such as, for example, one or more speakers), a sensing system 812, a control system (or at least one processor), and an outward facing image sensor, or world-facing camera (e g., sensor 806). In an example implementation, the outward facing image sensor, or world-facing camera (e.g., sensor 806) can be used to generate image data used to generate head pose data. The world-facing camera (e.g., sensor 806) can include an inertial measurement unit (IMU). Alternatively (or in addition to), the IMU can be a separate module included in the smartglasses 800. For example, the IMUcan be included in one of the two arm portions 824. The IMU can include a set of gyros configured to measure rotational velocity and an accelerometer configured to measure an acceleration of the world-facing camera (e.g., sensor 806) as the camera moves with the head and / or body or the user.
[0122] Example implementations can include a non-transitory computer- readable storage medium comprising instructions stored thereon that, when executed by at least one processor, are configured to cause a computing system to perform any of the methods described above. Example implementations can include an apparatus including means for performing any of the methods described above. Example implementations can include an apparatus including at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform any of the methods described above.
[0123] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0124] These computer programs (also known as programs, software, softw are applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine- readable medium’" “computer-readable medium” refers to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0125] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (a LED (light-emitting diode), or OLED (organic LED), or LCD (liquid crystal display) monitor / screen) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0126] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (“LAN”), a wide area network (“WAN”), and the Internet.
[0127] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship.
[0128] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the specification.
[0129] In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
[0130] Further to the descriptions above, a user may be provided with controls allowing the user to make an election as to both if and when systems, programs, orfeatures described herein may enable collection of user information (e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user’s current location), and if the user is sent content or communications from a server. In addition, certain data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity may be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.
[0131] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes and equivalents will now occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the scope of the implementations. It should be understood that they have been presented by way of example only, not limitation, and various changes in form and details may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or subcombinations of the functions, components and / or features of the different implementations described.
[0132] While example implementations may include various modifications and alternative forms, implementations thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that there is no intent to limit example implementations to the particular forms disclosed, but on the contrary, example implementations are to cover all modifications, equivalents, and alternatives falling within the scope of the claims. Like numbers refer to like elements throughout the description of the figures.
[0133] Some of the above example implementations are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of the operations may be performed in parallel, concurrently or simultaneously. In addition, the order of operations may be re-arranged. The processes may be terminated when their operations are completed but may also haveadditional steps not included in the figure. The processes may correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0134] Methods discussed above, some of which are illustrated by the flow charts, may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine or computer readable medium such as a storage medium. A processor(s) may perform the necessary tasks.
[0135] Specific structural and functional details disclosed herein are merely representative for purposes of describing example implementations. Example implementations, however, be embodied in many alternate forms and should not be construed as limited to only the implementations set forth herein.
[0136] It will be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of example implementations. As used herein, the term and / or includes any and all combinations of one or more of the associated listed items.
[0137] It will be understood that when an element is referred to as being connected or coupled to another element, it can be directly connected or coupled to the other element or intervening elements may be present. In contrast, when an element is referred to as being directly connected or directly coupled to another element, there are no intervening elements present. Other words used to describe the relationship between elements should be interpreted in a like fashion (e.g, between versus directly between, adjacent versus directly adjacent, etc.).
[0138] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of example implementations. As used herein, the singular forms a. an and the are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms comprises, comprising, includes and / or including, when used herein, specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one ormore other features, integers, steps, operations, elements, components and / or groups thereof.
[0139] It should also be noted that in some alternative implementations, the functi ons / acts noted may occur out of the order noted in the figures. For example, two figures shown in succession may in fact be executed concurrently or may sometimes be executed in the reverse order, depending upon the functionality / acts involved.
[0140] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which example implementations belong. It will be further understood that terms, e.g., those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0141] Portions of the above example implementations and corresponding detailed description are presented in terms of software, or algorithms and symbolic representations of operation on data bits within a computer memory. These descriptions and representations are the ones by which those of ordinary skill in the art effectively convey the substance of their work to others of ordinary' skill in the art. An algorithm, as the term is used here, and as it is used generally, is conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of optical, electrical, or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0142] In the above illustrative implementations, reference to acts and symbolic representations of operations (e.g., in the form of flowcharts) that may be implemented as program modules or functional processes include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types and may be described and / or implemented using existing hardware at existing structural elements. Such existing hardware may include one or more Central Processing Units (CPUs), digital signal processors (DSPs), applicationspecific-integrated-circuits, field programmable gate arrays (FPGAs) computers or the like.
[0143] It should be bome in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, or as is apparent from the discussion, terms such as processing or computing or calculating or determining of displaying or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical, electronic quantities within the computer system’s registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.
[0144] Note also that the software implemented aspects of the example implementations are typically encoded on some form of non-transitory program storage medium or implemented over some type of transmission medium. The program storage medium may be magnetic (e.g., a floppy disk or a hard drive) or optical (e.g., a compact disk read only memory, or CD ROM), and may be read only or random access. Similarly, the transmission medium may be twisted wire pairs, coaxial cable, optical fiber, or some other suitable transmission medium known to the art. The example implementations are not limited by these aspects of any given implementation.
[0145] Lastly, it should also be noted that whilst the accompanying claims set out particular combinations of features described herein, the scope of the present disclosure is not limited to the particular combinations hereafter claimed, but instead extends to encompass any combination of features or implementations herein disclosed irrespective of whether or not that particular combination has been specifically enumerated in the accompanying claims at this time.
Claims
WHAT IS CLAIMED IS:
1. A method comprising: receiving a first image and a second image representing a frame of a video, the first image including a first image portion that is determined based on a visual criterion; generating a map based on the second image; merging the first image and the second image using the map as an enhanced frame representing the frame of the video, the enhanced frame including the first image and a second image portion from the second image, the second image portion corresponding to the first image portion; and rendering the enhanced frame.
2. The method of claim 1 , wherein the first image and the second image are included in a plurality of images that include two or more of an RGB image, an infrared camera image, a depth camera image, and a thermal camera image.
3. The method of claim 1 or claim 2, wherein the map is generated using one of a saliency analysis or a semantic segmentation.
4. The method of claim 1 or claim 2, wherein the map is generated using a machine learning model trained based on a labelled dataset including a saliency map that indicates which objects in a scene are relevant to human perception.
5. The method of claim 2, wherein the map is generated using a machine learning model including: a first neural network configured to receive the plurality of images as stacked image tiles and to generate a set of feature maps for the plurality’ of images; and a second neural network configured to refine boundaries of an object based on a low-level feature of the set of feature maps.
6. The method of claim 2. wherein the map is generated using a machine learning model including:a first neural network configured to receive the plurality of images as stacked image tiles and to generate a set of feature maps for the plurality of images: a second neural network configured to identify context information associated with a high-level feature of the set of feature maps; and a third neural network configured to select a scale and receptive field for a region of interest in the set of feature maps based on the context information.
7. The method of claim 6. wherein the region of interest includes the first image portion that is visually indiscernible by the human eye.
8. The method of any of claim 2 to claim 7, wherein the plurality of images are received from a plurality of sensors of a wearable device configured to display see- through video.
9. The method of any of claim 1 to claim 8, wherein the enhanced frame includes the first image and the second image portion, , and the second image portion overlays the first image portion in the first image.
10. The method of claim 9. wherein the first image is an RGB image.
11. The method of claim 9, wherein the first image is an RGB image and the second image is an infrared image.
12. The method of claim 9, wherein the first image is an RGB image and the second image is a thermal image.
13. The method of any of claim 1 to claim 12, wherein the enhanced frame is rendered on a display of a wearable device.
14. A non-transitory computer-readable storage medium comprising instructions stored thereon that, when executed by a processor, are configured to cause a computing system to perform the method of any of claims 1-13.
15. An apparatus comprising means for performing the method of any of claims 1- 13.
16. An apparatus comprising: at least one processor; and at least one memory including computer program code; the at least one memory' and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform the method of any of claims 1-13.
Citation Information
Patent Citations
CNN-based saliency detection system and method
CN112927209A