Method and system for an ultra-wideband sensor-based scene editing in an electronic device
Patent Information
- Application Number
- PCT/IB2026/052859
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure IB2026052859_01102026_PF_FP_ABST
Abstract
Description
DESCRIPTIONTITLE OF INVENTION : METHOD AND SYSTEM FORAN ULTRA- WIDEBAND SENSOR-BASED SCENE EDITING IN AN ELECTRONIC DEVICE TECHNICAL FIELD
[0001] The present disclosure relates to the field of image out-painting. More particularly, the present disclosure relates to a method and system for an ultra-wideband (UWB) sensor-based scene editing in an electronic device.BACKGROUND ART
[0002] Image out-painting is a process of extending the borders of an existing image to create a larger composition. The image out-painting typically involves generating new parts of the image beyond its original frame, making the scene look more expansive or complete. However, the existing image out-painting systems and methods trained with a multitude of images face the challenge of generating realistic images due to lack of awareness of the actual surroundings. Further, limited field of view (FOV [-90deg, +90deg]) of smartphone cameras restricts the amount of scene information captured, limiting the effectiveness of image out-painting techniques. Additionally, there are some artefacts introduced while generating the out-paintings such as disconnected image parts and unrealistic supplemental objects.
[0003] Figure 1 discloses image out-painting, in accordance with related art. As shown, an image A represents the actual scene, and an image B is a frame of the scene captured by a camera of the electronic device. Now, using the existing methods and systems, an image C is generated by placing a random image during the image out-painting. Thus, the existing image out-painting systems and methods often start from noise or rely on simplistic contextual models, resulting in inconsistent or unrealistic results. These methods struggle to accurately understand the scene context or preserve the visual coherence of the generated content. There is a growing demand for more accurate and context-aware image out-painting techniques that may produce visually appealing results with minimal manual intervention.
[0004] Therefore, there exists a need for an ultra-wideband (UWB) sensor-based scene editing in an electronic device that seamlessly integrates with the existing electronic device and enhances the overall quality of the image.SUMMARY OF INVENTION
[0005] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the present disclosure. This summary is neither intended to identify key or essential inventive concepts of the present disclosure nor is it intended for determining the scope of the present disclosure.
[0006] According to an embodiment of the present disclosure, a method of generating an ultra-wideband (UWB) sensor-based scene editing in an electronic device is disclosed. The method comprises providing first data of a scene captured from a camera of the electronic device into a diffusion model, the camera having a first field of view (FoV) of the scene. The method comprises providing second data of the scene captured using the UWB sensor of the electronic device into the diffusion model, the UWB sensor having a second FoV in addition to the first FoV. Further, the method comprises obtaining an edited image by complementing the first data using the second data in the diffusion model, the edited image reflecting, in a scene captured from the first FoV, spatial information of one or more objects located within the second FoV and outside the first FoV.
[0007] According to another embodiment of the present disclosure, a method of an ultra-wideband (UWB) sensor-based scene editing in an electronic device. The method comprises capturing a scene using a camera of the electronic device, the camera having a first field of view (FoV) of the scene. The method comprises determining a first context of the scene from the first field of view. The method comprises capturing the scene using the UWB sensor of the electronic device, the UWB sensor having a second FoV in addition to the first FoV. Further, the method comprises determining a second context of the scene from a plurality of angle of arrivals (AO As) of a plurality of UWB signals reflected from one or more objects present in the second FoV of the scene. The method comprises obtaining an edited image by complementing, via a diffusion model, the second context to the first context, the edited image reflecting, in the scene captured from the first FoV, the second context of the one or more objects that are located outside the first FoV.
[0008] According to another embodiment, a system for an ultra-wideband (UWB) sensor-based scene editing in an electronic device is disclosed. The system comprises one or more processors and a memory coupled with the one or moreprocessors. The one or more processors are configured to capture a scene using a camera of the electronic device to obtain the image, the camera having a first field of view (FoV) of the image. The one or more processors are configured to determine a first context of the scene from the first field of view. The one or more processors are configured to capture the scene using the UWB sensor of the electronic device, the UWB sensor having a second FoV in addition to the first FoV. Further, the one or more processors are configured to determine a second context of the scene from a plurality of angle of arrivals (AO As) of a plurality of UWB signals reflected from one or more objects present in the second FoV of the scene. Furthermore, the one or more processors are configured to obtain an edited image by complementing, via a diffusion model, the second context to the first context, the edited image reflecting, in the scene captured from the first FoV, the second context of the one or more objects that are located outside the first FoV.
[0009] To further clarify the advantages and features of the present disclosure, a more particular description of the present disclosure will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the present disclosure and are therefore not to be considered limiting of its scope. The present disclosure will be described and explained with additional specificity and detail with the accompanying drawings.BRIEF DESCRIPTION OF DRAWINGS
[0010] These and other features, aspects, and advantages of the present disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:
[0011] Figure 1 discloses image out-painting, in accordance with related art;
[0012] Figure 2 illustrates a pictorial diagram depicting an exemplary environment for an ultra-wideband (UWB) sensor-based scene editing in an electronic device, in accordance with an embodiment of the present disclosure;
[0013] Figure 3 illustrates a block diagram of an architecture of the system for the UWB sensor-based scene editing in an electronic device, in accordance with an embodiment of the present disclosure;
[0014] Figure 4 illustrates a block diagram to obtain a UWB wavelet embedding, in accordance with an embodiment of the present disclosure;
[0015] Figure 5 illustrates a block diagram showing UWB transformation performed on decluttered data, in accordance with an embodiment of the present disclosure;
[0016] Figures 6A-6B illustrate an exemplary distance estimation using UWB, in accordance with an embodiment of the present disclosure;
[0017] Figure 7 illustrates a block diagram for exemplary angle of arrival estimation, in accordance with an embodiment of the present disclosure.
[0018] Figure 8 illustrates an exemplary calculation of full range FoV using UWB, in accordance with an embodiment of the present disclosure.
[0019] Figure 9 illustrates an example of infusing the UWB Wavelet Embedding into the raw image, in accordance with an embodiment of the present disclosure;
[0020] Figures 10A-10B illustrate an example of UWB wavelength infusion, in accordance with an embodiment of the present disclosure;
[0021] Figures 11 A-l ID illustrate an example of one time calibration of AoA, in accordance with an embodiment of the present disclosure;
[0022] Figures 12A-12B illustrate an example of training a pair of original and out-painted image, in accordance with an embodiment of the present disclosure;
[0023] Figure 13 illustrates an architecture of a variational autoencoder, in accordance with an embodiment of the present disclosure;
[0024] Figure 14 illustrates a diffusion model, in accordance with an embodiment of the present disclosure;
[0025] Figures 15 A-l 5C illustrate an example of latent space sampling of an input, in accordance with an embodiment of the present disclosure;
[0026] Figures 16A-16B illustrate an exemplary out-painted image, in accordance with an embodiment of the present disclosure;
[0027] Figure 17 illustrates a flow chart showing a method of generating the UWB sensor-based scene editing in an electronic device, in accordance with an embodiment of the present disclosure;
[0028] Figure 18 illustrates a flow chart showing a method of the UWB sensorbased scene editing in the electronic device, in accordance with an embodiment of the present disclosure; and
[0029] Figures 19A-19B illustrate examples of implementation of image out-painting in different scenarios, in accordance with an embodiment of the present disclosure.
[0030] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale.
[0031] Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present disclosure so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.DESCRIPTION OF EMBODIMENTS
[0032] For the purpose of promoting an understanding of the principles of the present disclosure, reference will now be made to the various embodiments and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the present disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the present disclosure as illustrated therein being contemplated as would normally occur to one skilled in the art to which the present disclosure relates.
[0033] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the present disclosure and are not intended to be restrictive thereof.
[0034] Whether or not a certain feature or element was limited to being used only once, it may still be referred to as “one or more features” or “one or more elements” or “at least one feature” or “at least one element.” Furthermore, the use of the terms “one or more” or “at least one” feature or element do not preclude there being none of that feature or element, unless otherwise specified by limiting language including, but not limited to, “there needs to be one or more...” or “one or more elements is required.”
[0035] Reference is made herein to some “embodiments.” It should be understood that an embodiment is an example of a possible implementation of any features and / or elements of the present disclosure. Some embodiments have been describedfor the purpose of explaining one or more of the potential ways in which the specific features and / or elements of the proposed disclosure fulfil the requirements of uniqueness, utility, and non-obviousness.
[0036] Use of the phrases and / or terms including, but not limited to, “a first embodiment,” “a further embodiment,” “an alternate embodiment,” “one embodiment,” “an embodiment,” “multiple embodiments,” “some embodiments,” “other embodiments,” “further embodiment”, “furthermore embodiment”, “additional embodiment” or other variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, one or more particular features and / or elements described in connection with one or more embodiments may be found in one embodiment, or may be found in more than one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although one or more features and / or elements may be described herein in the context of only a single embodiment, in the context of more than one embodiment, or in the context of all embodiments, the features and / or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and / or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.
[0037] Any particular and all details set forth herein are used in the context of some embodiments and therefore should not necessarily be taken as limiting factors to the proposed disclosure.
[0038] The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components proceeded by “comprises... a” does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.
[0039] Hereinafter, it is understood that terms including “unit” or “module” at the end may refer to the unit for processing at least one function or operation and may be implemented in hardware, software, or a combination of hardware and software.
[0040] Unless otherwise defined, all terms, and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.
[0041] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted to not unnecessarily obscure the embodiments herein. Also, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments. The term “or” as used herein, refers to a non-exclusive or unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
[0042] As is traditional in the field, embodiments may be described and illustrated in terms of blocks that carry out a described function or functions. These blocks, which may be referred to herein as units or modules or the like, are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionally be driven by firmware and software. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.
[0043] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any alterations, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings. Although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.
[0044] For the sake of clarity, the first digit of a reference numeral of each component of the present disclosure is indicative of the Figure number, in which the corresponding component is shown. For example, reference numerals starting with digit “1” are shown at least in Figure 1. Similarly, reference numerals starting with digit “2” are shown at least in Figure 2.
[0045] Ultra-Wideband (UWB) sensors in smartphones are a cutting-edge technology that serve both communication and radar functions thus enabling precise location tracking, faster data transfer, and new forms of interaction between devices. In communication, UWB facilitates high-speed, short-range data transmission and offers enhanced indoor positioning capabilities. For instance, asset tracking and real-time location services. UWB also transfers large amounts of data at lower power consumption, supporting wireless communication. In addition to communication, UWB sensors provide radar-like capabilities for object detection, motion sensing, and even through-wall sensing. The UWB sensor data contains the information about spatial distribution of mass in the direction of the transmission.
[0046] Owing to its wider field of view (FoV), UWB sensor fetches additional information about the objects in the surroundings of the image target that is unavailable to the camera. To elaborate, UWB antenna in the UWB sensor includes the FoV of [-90° to +90°] in the azimuth direction and [-50° to +40°] in the elevation direction, with a lOdB attenuation limit. As a result, the UWB sensor captures a wider FoV compared to the main camera, thus providing the additional information. This extended coverage allows the UWB sensor to capture objects within and outside the camera's FoV. The UWB sensor brings spatial distribution and depth information, which allows for a more accurate generation of image extensions. The UWB sensor provides depth information, adding a third dimension to the scene,thus, enabling more realistic out-painting. Additionally, the integration of the UWB sensors with existing camera infrastructure is seamless and cost-effective, thereby, minimizing the need for extensive upgrades or equipment replacements, making it highly accessible for a wide range of users. Another key benefit is the UWB sensor’s low-light performance, as reflections are unaffected by lighting conditions. This ensures that the out-painting task remains accurate even in poorly lit environments, offering consistent results across both well-lit and dimly lit settings. In view of the above, this additional information when used to train the foundation models for image out-painting may lead to more realistic, logical out-painted images which have much lesser artefacts. Further, the UWB Radio Detection And Ranging (RADAR) reflections obtained from UWB sensors carry the same information independent of the lighting conditions. Thus, the models trained with UWB reflection acts as additional input for enhancing the less-illuminated areas of the image perform better.
[0047] An object of the present disclosure is to provide accurate and context-aware image out-painting techniques to generate visually appealing results with minimal manual intervention.
[0048] An object of the present disclosure is to provide a solution that seamlessly integrates with the existing electronic device and enhances the overall quality of the image.
[0049] Another object of the present disclosure is to infuse UWB sensor data with image sensor data for providing extended spatial context beyond the camera's field of view.
[0050] A further object of the present disclosure is to determine a positional embedding technique related to spatial understanding and scene reconstruction.
[0051] Yet another object of the present invention is to optimize spatial data captured using UWB sensor for propagation of information and textures into a diffusion model.
[0052] The method and system of the present disclosure involves leveraging UWB sensors for contextual initialization in image out-painting, providing a more accurate and visually appealing result. By incorporating spatial data from UWB sensors, the present disclosure enhances the contextual understanding of the scene, resulting in more coherent and realistic out-painted image.
[0053] Embodiments of the present disclosure will be described below in detail with reference to the accompanying drawings.
[0054] Figure 2 illustrates a pictorial diagram depicting an exemplary environment 200 for an ultra-wideband (UWB) sensor-based scene editing in an electronic device, in accordance with an embodiment of the present disclosure.
[0055] As shown in the figure, an electronic device 202 with an in-built camera 204 and a plurality of UWB sensors (206-1, 206-2, 206-3, 206-4) which communicates with a system 208 to generate more accurate and context-aware out-painted image 210.
[0056] The system 208 may include a software, a hardware, a combination of software or hardware, an in-built application on an electronic device 202 or an application to be installed and operated on the electronic device 202 in communication with a network interface. The system 208 may also be available via cloud-based server and available remotely from the electronic device 202.
[0057] In the embodiment when the system 208 is located outside the electronic device 202, a network interface (not shown) may be configured to provide network connectivity and enable communication between the system 208 and the electronic device 202. The network connectivity may be provided via a wireless connection or a wired connection. For example, the network connectivity may be provided via cellular technology, such as 3rd Generation (3G), 4thGeneration (4G), 5thGeneration (5G), pre-5G, 6thGeneration (6G), Bluetooth, Local Area Network (LAN), Wi-Fi, cable, or any other wired / wireless communication technology.
[0058] Figure 3 illustrates block diagram of an architecture of the system 208 for the UWB sensor-based scene editing in an electronic device, in accordance with an embodiment of the present disclosure. The system 208 may include a memory 304, one or more processors 302 (hereafter referred to as the processor 302), one or more modules 306 and a data unit 308. In an exemplary embodiment, the processor 302 may be operatively coupled to each of the memory 304, the modules 306, and the data unit 308. The processor 302 may execute a software program, such as code generated manually (i.e., programmed) to perform the desired operation. The processor 302 may implement various techniques such as, but not limited to, image processing, data extraction, Artificial Intelligence (Al), Machine Learning (ML), Deep Learning (DL) and so forth to achieve the desired objective.
[0059] In some embodiments, the memory 304 may be communicatively coupled to the at least one processor 302. The memory 304 may be configured to store data, instructions executable by the at least one processor 302. In one embodiment, the memory 304 may communicate via a bus within the system 208. The functions, acts or tasks illustrated in the figures or described may be performed by the programmed processor 302 for executing the instructions stored in the memory 304. The functions, acts or tasks are independent of the particular type of instructions set, storage media, processor or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code and the like, operating alone or in combination.
[0001] The modules 306, amongst other things, include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement data types. The modules 306 may also be implemented as, signal processor(s), state machine(s), logic circuitries, and / or any other device or component that manipulates signals based on operational instructions. Further, the modules 306 may be implemented in hardware, instructions executed by a processing unit, or by a combination thereof. The module(s) 306 enable the system 208 to perform the features / functions of the present disclosure, as discussed and explained in detail in conjunction with Figures 4-15 in the forthcoming paragraphs.
[0060] In an embodiment, the data unit 308 serves, amongst other things, as a repository for storing data processed, received, and generated by one or more of the modules 306.
[0061] In another embodiment of the present disclosure, the processor 302 via the modules 306 is configured to execute machine-readable instructions (software) which perform the working of the system 208 within the scope of the present invention as described in forthcoming paragraphs.
[0062] In an embodiment, the system 208 is configured to capture a scene using a camera of the electronic device to obtain the image, the camera having a first field of view (FoV) of the image. The system 208 is configured to determine a first context of the scene from the first field of view. Further, the system 208 is configured to capture the scene using the UWB sensor of the electronic device, the UWB sensor having a second FoV in addition to the first FoV. The system 208 is configured to determine a second context of the scene from a plurality of angle ofarrivals (AO As) of a plurality of UWB signals reflected from one or more objects present in the second FoV of the scene. Furthermore, the system 208 is configured to complement, via a diffusion model, the second context to the first context to edit the scene of the image captured from the first field of view.
[0063] In an embodiment, prior to determining the second context, the system 208 is configured to receive the plurality of UWB signals from the UWB sensor. The system 208 is configured to remove temperature variations of the plurality of UWB signals received in Channel Impulse Responses (CIRs). The system 208 is configured to de-clutter the plurality of UWB signals by removing effect of coupling in the CIRs. The coupling is direct reception of signals from UWB transmitter by UWB receiver. The system 208 is configured to process the decluttered plurality of UWB signals to extract magnitude data and phase data. Further, the system 208 is configured to remove one or more spurious peaks in the CIRs generated by noise from the de-cluttered plurality of UWB signals to generate denoised plurality of UWB signals to be used for determining the second context.
[0064] In an embodiment, prior to removing one or more spurious peaks, the system 208 is configured to process, using at least one of an interpolation method, a phase un-wrapping method and a magnitude de-noising method, the extracted magnitude data and the phase data to obtain transformed magnitude data and phase data.
[0065] In an embodiment, the system 208 is configured to extract the plurality of Ao As of the de-noised plurality of UWB signals in the CIRs for the one or more objects appearing in an image. The system 208 is configured to filter the AoAs of interest related to the one or more objects appearing in a particular cell of the image in the CIRs to obtain FoV limited de-noised UWB CIR signal. The system 208 is configured to calculate a UWB wavelet embedding based on the FoV limited denoised UWB CIR signal to be infused into the image. Further, the system 208 is configured to prepare one or more training datasets by infusing the UWB wavelet embedding into the image using a sinusoidal schedule to obtain a UWB wavelet embedding infused image.
[0066] In an embodiment, the system 208 is configured to transform the UWB wavelet embedding infused image into a latent space using a variational autoencoder. Further, the system 208 is configured to de-noise the UWB waveletembedding infused image by training U-Net model for wavelet embedding estimation. In addition, the system 208 is configured to transform the de-noised UWB wavelet embedding infused image from the latent space into an out-painted image using the variational auto-encoder.
[0067] In an embodiment, to complement the second context to the first context, the system 208 is configured to infuse the image obtained from the electronic device with the UWB wavelet embedding to generate a plurality of versions of training samples in latent space using the variational auto-encoder. The system 208 is configured to feed the plurality of versions of the training samples to the trained denoising U-net model to generate a final out-painted image in latent space. Further, the system 208 is configured to transform the final out-painted image from the latent space into the image using the decoder of the variational auto-encoder to obtain the edited scene of the image.
[0068] The present disclosure utilizes the UWB sensing data (obtained from UWB sensor) alongside electronic device captured images, thereby enabling a wider FoV and richer scene understanding and enhancing the accuracy and realism of out-painted images.
[0069] Further, starting the out-painting task from UWB data instead of noise provides a contextual foundation, resulting in more coherent and visually appealing expanded images. In addition, the integration of UWB sensing with generative models revolutionizes image out-painting by leveraging scene and material understanding, achieving improvement in image quality and realism compared to existing methods.
[0070] A detailed working and explanation of the system 208 will be explained through various components of Figures 4-15 in the forthcoming paragraphs.
[0071] In an embodiment, Figure 4 illustrates a block diagram 400 to obtain a UWB wavelet embedding, in accordance with an embodiment of the present disclosure.
[0072] At block 402, the processor 208 may receive UWB RADAR data from UWB sensor.
[0073] Subsequently, at block 404, the processor 208 may compensate the received UWB Radar data for temperature dependence received in Channel Impulse Responses (CIRs).
[0074] In an embodiment, temperature variations introduce a drift in the clutter present in the received UWB RADAR data. For example, a drift is modelled as 1-tap Infinite Impulse Response (UR) filter. In the UR filter, an instantaneous error (T) may be represented as h(tm~l, T) - href (T), filtered error e(r) may be represented as (1 - a) • e + a • em(r) and a drift compensated CIR he (tm, T) may be represented as h (tm-1, T) - e (T) . Further, tap coefficient in the filter may be platform dependent. Thus, a drift compensation system is used to remove it and stabilize the CIRs.
[0075] At block 406, the processor 208 may de-clutter the UWB signals by removing effect of coupling (via direct reception) in the CIRs, between transmitter and receiver to obtain de-cluttered data. The coupling is direct reception of signals from UWB transmitter by UWB receiver
[0076] To understand the effect of coupling and remove it, it is observed that the strength of the coupled path depends on the isolation between transmitter and receiver (on-chip, PCB and antenna). This isolation is also a function of the directivity of the transmitter and receiver antennas. Since the antennas used have a large field of view, and since the distance between the transmitter and receiver antennas are only about half wavelength, coupled path is much stronger than the reflections. In view of the above, since the clutter remains constant over time, it is estimated as the running mean of a CIR tap and subtracted from CIR under consideration.
[0077] At block 408, the processor 208 may transform the de-cluttered data into magnitude and phase, thereby making the UWB data more useful.
[0078] At block 410, the processor may remove spurious peaks in CIRs generated by noise. To remove the spurious peaks, methodology such as Cell Averaging Constant False Alarm Rate (CA-CFAR) is used to identify the most relevant peaks and get rid of potential false alarms. Thus, the output of Figure 4 is the de-noised UWB data at block 412, which is used as an input to the pre-training module.
[0079] Figure 5 illustrates a block diagram 408 showing UWB transformation performed on decluttered data, in accordance with an embodiment of the presentdisclosure. The processor 208 may process, using at least one of an interpolation method, a phase un-wrapping method and a magnitude de-noising method, the extracted magnitude data, and the phase data to obtain transformed magnitude data and phase data.
[0080] At block 502, the processor 208 may receive the de-cluttered data of the UWB signal for further processing.
[0081] At block 504, the processor 208 may process de-cluttered UWB signal to extract magnitude and phase of the de-cluttered data of the UWB signal.
[0082] At block 506, the processor 208 may perform interpolation method on the de-cluttered data. This optional block is used to improve the distance resolution of the data if distance resolution is insufficient. The interpolation distance resolution of the de-cluttered UWB data typically refers to the accuracy with which positions or distances are estimated between two or more UWB measurements.
[0083] At block 508, the processor 208 may perform phase unwrapping of the decluttered data of the UWB signal to avoid phase wrap around phenomenon at +-pi / 2 and avoid information loss. Phase unwrapping is a technique used to resolve ambiguities caused by the inherent periodic nature of phase measurements which may lead to loss of information, especially when the phase change exceeds these boundaries (+- pi / 2), and to incorrect interpretation of the signal phase.
[0084] At block 510, the processor 208 may remove spurious peaks by “running mean” to identify genuine peaks and thus, obtaining transformed magnitude and phase data at block 512. The term “running mean” indicates smoothening the magnitude of the decluttered UWB signal and filtering out spurious peaks that may be caused by noise or random fluctuations. By applying a running mean, the data is averaged over a sliding window, which helps identify genuine peaks by reducing the impact of isolated spikes that do not correspond to actual features of interest.
[0085] Further, magnitude de-noising may not be enough to remove all spurious peaks at block 510. As discussed, the CA-CFAR methodology is applied to identify the most relevant peaks and get rid of potential false alarms. The CA-CFAR works on an adaptive threshold mechanism based on the energy at neighbouring cells in the image.
[0086] In CA-CFAR, the noise level is estimated by averaging the signal levels of a number of surrounding cells (neighbours) around the cell (a "cell" typically refersto a specific time or frequency bin where the received signal is grouped or recorded). The average noise level is then used to set a threshold, and if the signal in the test cell exceeds this threshold, a target is detected. The threshold is determined by the noise power, scaled by a constant factor (called the threshold factor), which ensures the false alarm rate remains constant. Thus, for a peak to be identified as a “true peak”, its magnitude is required to be above the adaptive threshold.
[0087] Figures 6A-6B illustrate an exemplary distance estimation using UWB, in accordance with an embodiment of the present disclosure. In Figure 6B, a person is standing in the spherical ring around the electronic device with inner radius of 30cm and outer radius of 45cm. Figure 6A shows data collected at Antenna 1 for test scenario, where the 3rdrange bin shows high value.
[0088] In an embodiment, the UWB receivers receive the Channel Impulse Responses (CIRs) as shown in Figure 6A. At each time instant, a vector of values known as “range bins” is received. The size of the vector is configurable within the manufacturer margins. The first value in the vector indicates the reflection from any objects(s) present in a spherical space around the phone with radius 15 cm. The first value will be high, if there is an object present in this space, else low. The second value in the vector indicates the reflection from any objects(s) present in a spherical ring space around the phone with outer radius 30 cm and inner radius of 15cm. The second value will be high, if there is an object present in this space, else low. Thus, by identifying the location of the highest values in a given CIR, the distance of the object from the electronic device may be calculated.
[0089] In an embodiment, in the present disclosure once the decluttered and denoised UWB data is obtained, training corpus for the iterative UWB wavelet embedding prediction is to be prepared. Thus, to obtain the training corpus, the decluttered and de-noised UWB data and image data is received as input.
[0090] At step 1, the “Angle of arrival” (AoA) for every genuine CIR peak is estimated and CIRs with Ao As of interest for a given cell of pixels is filtered out. A detailed explanation of angle of arrival estimation will be explained in Figures 7-8 in the forthcoming paragraphs.
[0091] At step 2, wavelet transform of the FoV limited de-noised UWB CIR data to be infused into the image data is calculated. A detailed explanation of wavelet transform calculation will be explained in the forthcoming paragraphs.
[0092] At step 3, a training corpus is created by infusing the UWB Wavelet Embedding into the raw image using a sinusoidal schedule.
[0093] Further, a UWB antenna has aFoV of [-90deg, +90deg] in azimuth direction and [-50deg, +40deg] in elevation for a lOdB attenuation limit. Thus, to extract the UWB Radar CIRs carrying information about the objects present in a given cell of pixels of the image, their AoAs is required to be estimated and appropriate CIRs is to be selected. This is achieved by translating the phase difference between multiple receivers for the data obtained from same distances into an Angle of Arrival (AoA) estimate. A step-by-step explanation of the angle of arrival estimation is mentioned in Figure 7.
[0094] Figure 7 illustrates a block diagram 700 for exemplary angle of arrival estimation, in accordance with an embodiment of the present disclosure.
[0095] At block 702, CIRs from the receiver is received. Thus, the AoA estimation routine will be run for all bins.
[0096] At block 704, the bins where the estimated AoA matches with that of the cell of pixels under consideration are extracted.
[0097] AT blocks 706 and 708, the UWB data, carrying the information about the objects present in the cell of pixels under consideration exclusively, is identified.
[0098] At block 710, a phase difference of arrival (PdoA) to AoA look up table is generated as a part of a dedicated calibration process a priori.
[0099] Figure 8 illustrates an exemplary calculation of full range FoV using UWB, in accordance with an embodiment of the present disclosure. As shown, Rays 1 and 4 indicate the transmitted signal, Antenna 0 (ANT0) transmits the UWB RADAR signal waves as per its designed FoV, Ray 2 indicates the reflection received by Antenna 3 (ANT3) and the Rays 3 and 5 indicate the reflections received by Antenna 1 (ANTI) and the Ray 6 indicates that received by Antenna 2 (ANT2).
[0100] Further, due to the additional distance (d3 - d2) travelled by the Ray 3 as compared to Ray 2 reflected from the person, a phase difference is created between them. This is explained Figure 7, to estimate the AoA.
[0101] In addition, due to the additional distance (d5 - d6) travelled by the Ray 5 as compared to Ray 6 reflected from the glass with water, a phase difference is created between them. This is exploited to estimate the AoA.
[0102] In an embodiment, the UWB sensor power consumption is guided by the receiver antenna gain selected by the UWB chip firmware in an automatic manner. The UWB sensor has a mechanism to select an appropriate number for the receiver gain among the pre-configured ones. The UWB sensor is driven by the Signal to Noise Ratio (SNR) received at the receiver for a sample transmission. This prevents unnecessarily high power consumption by the system.
[0103] For example, if there is sky in the image, then it is possible that there is an object of interest such as aeroplane or bird or fireworks, which are outside the field of view of camera, however, may be captured by UWB. Therefore, the system receives data from its entire FoV. Based on the UWB wavelet embedding learned by the system during the iterative training phase, the system is to generate the most accurate and semantically correct out-painting for the given scenario during inference. Additionally, the UWB reflections from far objects are learned by the system during the training phase and they are rendered appropriately into the out-painted image during inference.
[0104] Now, in an embodiment, to calculate wavelet transform, it is understood that a wavelet transform decomposes the original signal and down-samples it at various rates to allow examining it at various magnification scales. The wavelet transform may be described mathematically as follows.
[0105] Where, a is the scaling factor, b is the shift parameter, f is the signal, f<p is the wavelet transform and Cjkare the wavelet coefficients. The wavelet coefficients are calculated for all receiver antennas available on the device after FoV limiting.
[0106] Figure 9 illustrate an example of infusing the UWB Wavelet Embedding into the raw image, in accordance with an embodiment of the present disclosure. As shown, image A is the image captured along with UWB information by the electronic device. Image B indicates processing of the image cell by cell. Further, image C indicates processing of image in cell (1,1) by cropping image and estimating UWB signal data corresponding to cell (1,1). Then, in image D the UWBsignal is converted to UWB wavelets for cell (1,1) which is the data set for cell (1,1). Similar steps are followed for the entire image.
[0107] Figures 10A-10B illustrate an example of UWB wavelength infusion, in accordance with an embodiment of the present disclosure.
[0108] Figure 10A shows using data set for each cell obtained in Figure 9 and again iterating the process cell by cell to infuse UWB wavelet with an image for unmasked area. As shown, ‘A’ and ‘B’ are image and UWB wavelet for cell (1,1) respectively.
[0109] Similarly, Figure 10B shows original image at t=0, at t=l, UWB wavelet is slightly infused in the image and at t=T, UWB wavelet is completely infused in the image.
[0110] Further, a given cell of pixels in an image is infused with its UWB wavelet data at certain time step t=ti, as explained below:
[0111] Let ft. be the fraction of the UWB wavelet coefficients to be added to the original camera image at time step tj.Then, ftiis calculated as,fti= fr+ 0.5
[0112] Where, f0is the initial fraction of UWB wavelet coefficients to be added (usually close to 0) fTis the final fraction of UWB wavelet coefficients to be added (usually close to 1).
[0113] Thus, using equation (1), the infused image at any time step ti, as shown below.Ati= A + fti* UWBwav
[0114] Where, A is cell of pixels under consideration of the original image,t. is the cell of pixels of the image infused with the UWB wavelet coefficients at time step tj and UWBwavis the UWB wavelet coefficients for the cell of pixels under consideration.
[0115] Further, For= T, ft. = fT. Thus, the ft. varies from 0 to 1 over a sinusoidal curve. Due to the nature of the plot, the fraction of UWB wavelets to be infused increases gradually which enhances the speed of convergence during training.
[0116] Further, the combined images of t=0, t=l, ..., t=T are generated by adding the UWB wavelet embedding to the expanded image, step by step, fractionally. Forexample, if T=4, combined Image at t=l may have [fT+ 0.5 * ( / o— / T) = 0.1464 fraction of UWB wavelet embeddings added to it.combined image at t=2 may have [fT+ 0.5 * ( / o— / T) = 0.5 fraction of UWB wavelet embeddings added to it. Thecombined image at t=3 may have [fT+ 0.50.8535 fraction of UWB wavelet embeddings added to it. The combined image at t=4 may have [fT+ 0.5 = 1.0 fraction of UWBwavelet embeddings added to it. The aforesaid values are derived with the assumption of ( / 0= 0 and fT= 1). Usually, T is selected to be very high (of the order of 100).
[0118] Thus, to summarise the training data preparation or training corpus, multiple images are captured of a multitude of objects or subjects in a variety of surroundings and the corresponding UWB Radar data is recorded simultaneously for data collection. Then, UWB Radar data is recorded for all the receiver antennas available on the device. Thereafter, the image is split into cells of size NxN pixels each.
[0119] It is important to note that UWB sensor is able to identify different liquids since permittivity values spans across large spectrum. Thus, distinguishing different liquids using permittivity is possible with high confidence using UWB sensor.
[0120] The next step is one time determination of angle of arrival in which each pixel of the image is stimulated by the light rays reflected from objects or received from a source of light present at a specific 2D (azimuth and elevation) angle, which is referred as “angle of arrival” (AoA). Due to the adjacency of the UWB Radar receivers and the Camera, this AoA gets translated to a different one for the UWB receiver antenna. In this step, this correspondence is established using simple geometry given the exact locations of the camera and the receiver antenna. Instead of a pixel, a cell of NxN pixels is used as a unit as explained earlier.
[0121] In an embodiment, to form a UWB Radar time series of an object or subject in the field of view, the object is tracked over time for its distance as well as its appearance / disappearance. In an example embodiment, the object is tracked usingKalman filter and data association for the data corresponding to same object may be done with Hungarian method.
[0122] Further, once the UWB CIR taps corresponding to the object tracked over time are available from all the receiver antennas, the corresponding AoA is measured as a function of the PDoA between two receiver data at a time.
[0123] Figures 11 A-l ID illustrate an example of one time calibration of AoA, in accordance with an embodiment of the present disclosure.
[0124] During calibration phase, a test object is placed at an angle such that its image is formed in the top-right corner of the camera image (Figure 11 A). The AoAs (azimuth and elevation) subtended by the object at the centre of the UWB receivers is computed using data received at multiple UWB receivers. At the same time, the location of the test object inside the captured image is noted. This pair of Object AoA (measured using UWB) and the location of the image of the object inside the camera image is noted.
[0125] Further, a test object is placed at an angle such that its image is formed in the bottom-right corner of the camera image (Figure 1 IB). The AoAs (azimuth and elevation) subtended by the object at the centre of the UWB receivers is computed using data received at multiple UWB receivers. At the same time, the location of the test object inside the captured image is noted. This pair of Object AoA (measured using UWB) and the location of the image of the object inside the camera image is noted.
[0126] Similar procedure is repeated for the cases of the test objects appearing in the top-left and bottom -left corners of the camera image, as shown in Figures 11C-11D. This information (4 pairs of corner AoAs and the location inside camera images) is stored in a database. During inference, based on the AoAs calculated for any given object, its location inside camera image is computed using the proportional relation.
[0127] To calculate the location, let us assume, if a and e are azimuth and elevation AoAs based on UWB, t, b, 1 and r are top, bottom, left and right extreme values of AoAs obtained from calibration and w and h are the image width and height in pixels.
[0128] Then, Horizontal location of object in camera view = ((a - 1) / (r - 1)) * w and Vertical location of object in camera view = ((e - b) / (t - b)) * h.
[0129] For example, in a scenario after calibration, UWB based Azimuth AoA varies between -15 and 15 degrees corresponding to the leftmost and the rightmost camera pixels and the elevation AoA varies between -45 and 45 degrees corresponding to the bottommost and the topmost camera pixels. For the object under consideration with given AoAs, its location in the camera image may be determined as follows:
[0130] Horizontal location = (10 - (-15) / (15 - (-15))) * 480 = 400 pixel; and Vertical location = (25 - (-45)) / (45 - (-45)) * 640 = 498 pixel.
[0131] For UWB object tracking and AoA measurement, in an embodiment, for a given scenario of two antennas in any one of the orientations (azimuth or elevation), the signal arrives at one of the antennas with a delay with respect to the other one which is reflected in the phases of the received signals. It is characterized in the form of PDoA. The UWB Radar data collected from a receiver antenna provides the range (distance) profile of the field of view.
[0132] For the range profile at a time instant, the presence of an object is indicated by one of the CIR taps. A measurement of AoA of an object between two horizontally adjacent antennas yields azimuth part of the AoA. Similarly, a measurement of AoA between two vertically adjacent antennas yields elevation part of the AoA. These two measurements, when combined, indicate the cell of pixels at which the object is viewed. Thus, a time series of the CIR taps corresponding to this pixel may be formed which is used in further processing.
[0133] Further, a wavelet transform is used to analyse the frequency content of a time series at different time resolutions. It outputs a 1 -dimensional description of the time series. One such description is obtained for data from every receiver antennas. Further, the wavelet coefficients comprise of a) coarse level coefficients and b) detail coefficients. The wavelet transform add information about the role of the cell with respect to its neighbouring cells.
[0134] Thereafter, ‘T’ images are generated by adding a specific fraction of the wavelet coefficients to the area to be out-painted. To enable a meaningful distribution of the transform coefficients, a sinusoidal profile is used to determine the fraction of the total coefficients to be added to the cell. A sinusoidal profile avoids adding unnecessarily less fraction of UWB at the beginning (small t) and very high fraction of UWB at the end (close to T).
[0135] In an embodiment, for de-noising U-Net trained for wavelet embedding estimation, at step 1, the UWB wavelet embedding infused image data is transformed into latent space using encoder part of variational auto-encoder.
[0136] At step 2, the de-noising U-Net model is trained for estimating the UWB Wavelet Embedding infused into the image.
[0137] At step 3, the out-painted image data is transformed from latent space into the original image domain using decoder part of variational auto-encoder.
[0138] Further, the U-Net model has a loss function, and it is minimized during training. For image generation or out-painting tasks, the loss function typically measures the difference between the predicted (generated) output and the actual (ground truth) image. The goal is to adjust the model's parameters (weights) so that this difference (or loss) becomes as small as possible. The present disclosure uses a combination of Mean Absolute Error loss (LI) or Mean Squared Error loss (L2) and perceptual loss to ensure both pixel accuracy and perceptual quality of the out-painted image. The UWB-infused embeddings guide the model to generate contextually accurate content outside the original frame.
[0139] In addition, the UWB data helps minimize the loss function by guiding the generative process to produce realistic and contextually accurate out-painted regions. The training dataset consists of pairs of original images and corresponding out-painted images. Additionally, the UWB wavelet embeddings, representing spatial data beyond the camera's view, are infused into the latent space during training. The U-Net encoder compresses both the image and UWB wavelet embedding data into the latent space. This latent representation contains crucial features from both the image and the UWB data, allowing the model to understand the spatial context beyond the image’s boundary.
[0140] During training, after the model generates an out-painted image, the loss is calculated by comparing the generated image to the ground truth image. Pixel Loss (LI or L2): Measures how close the generated pixels are to the real pixels. Perceptual Loss: Compares the generated image and the ground truth at a feature level, ensuring that the generated content looks realistic and is contextually aligned with the input image.
[0141] Further, the UWB wavelet embedding provides additional context in the latent space, making the model more aware of what’s outside the camera’s field ofview. As the U-Net decoder reconstructs the image, it uses the UWB-infused latent space to guide the generation of the out-painted regions. If the generated output fails to align with the spatial data provided by the UWB embedding, the loss function will be high. The model adjusts its weights to minimize this loss, learning to generate better contextually accurate out-painted images as training progresses.
[0142] The gradient descent algorithm updates the weights of the model, pushing the U-Net to use the UWB information more effectively at each step, resulting in smaller differences between the generated images and the ground truth.
[0143] Figures 12A-12B illustrate an example of training a pair of original and out-painted image, in accordance with an embodiment of the present disclosure. As shown, image A is the original image and image B is the out-painted image.
[0144] Figure 13 illustrates an architecture of a variational autoencoder, in accordance with an embodiment of the present disclosure. A variational Autoencoder is an Artificial neural network trained to generate a latent space representation of the input image at the output of the encoder part and to generate the input image back at the output of the decoder part. As the decoder regenerates the original image from the latent space representation, the encoder learns to extract the features related to the information about the prominent trends in the image and ignore the noise. Also, the latent space vector is small in size as compared to the original image thus reducing the number of computations required to process the image. Therefore, the latent space representation generated by the encoder is used iteratively in the diffusion process to train the de-noising U-net.
[0145] Further, for every sample available to train, ‘T’ images with varying amount of wavelet coefficients of the UWB time series added at each cell appropriately are available. For training, the “T=0” image is first converted to a latent space using a variational auto-encoder. This is done to reduce number of operations required for the learning process.
[0146] In the variational autoencoder, the encoder reduces the high-dimensional image and UWB data into a lower-dimensional latent space that holds the core information needed for reconstruction. The decoder takes this compressed latent representation and reconstructs the expanded image. The autoencoder is responsible for learning how to represent both the visible image and the UWB data in acompressed form and then reconstructing them together. Thus, ensuring that the out-painted areas generated by the UWB data are coherent with the original image.
[0147] For training, the latent space representation of the image is fed to a denoising U-net. It is a pixel level classifier network which employs convolutional structures to characterize the noise present in the image. To highlight the spatial relations between multiple regions of the image, a positional embedding which encodes the position of the wavelet coefficients in the image, is added for the cross attention layer to determine the context. At the output of the U-net, an estimation of the UWB wavelet embedding in the latent space vector is received. Using the decoder part of the VAE used earlier, it is converted back to the image domain.
[0148] In addition, the images generated by the de-noising U-net are in the latent space of the variational autoencoder. To transform them into human viewable format, they need to be transformed from the latent space representation. The decoder part of the variational auto-encoder is used to perform this operation.
[0149] Figure 14 illustrates a diffusion model, in accordance with an embodiment of the present disclosure. The diffusion model utilizes UWB data to enhance the diffusion mechanism in image out-painting. Thus, enhancing the model's ability to propagate information and generate realistic textures and details beyond the captured frame. The diffusion model integrates UWB (Ultra-Wideband) sensor data into the generative process, leveraging its extended field of view (FOV) and spatial awareness capabilities. Thus, allowing for more comprehensive scene reconstruction and contextually accurate out-painting. Further, the diffusion model implements positional embedding techniques that have not been traditionally used in image out-painting thereby enhancing the model's ability to spatially contextualize the scene, in turn improving the accuracy and coherence of generated images.
[0150] To obtain the Out-painted image for a given input camera image and UWB data, step 1 is infusing the camera image with the UWB wavelet embedding and transforming the ‘T’ images created in this manner into latent space using encoder part of variational auto-encoder.
[0151] At step 2, feeding the ‘T’ images to the trained de-noising U-net to generate the out-painted image in the latent space.
[0152] At step 3, transforming the out-painted image data from latent space into the original image domain using decoder part of variational auto-encoder.
[0153] Thus, during the inference phase, the input camera image and the UWB data are subjected to the pre-processing process as during the training phase. The UWB data from all receiver antennas is conditioned to remove the clutter, transform the raw data into magnitude and phase, optionally interpolate and remove unwanted spurious peaks. The clean UWB data is then FoV limited using the AoA information to extract objects present in all the designated cells of pixels. The FoV limited data is transformed into wavelet domain and UWB wavelet embedding is generated. The original camera image is then infused with the UWB wavelet embedding in a step-wise manner (Total ‘T’ steps) using a scheduler to create ‘T’ UWB wavelet embedding infused images. Finally, the ‘T’ images are transformed into the latent space using the encoder part of the variational auto-encoder already trained on a multitude of images. The transformed images are fed to the next system including de-noising U-net.
[0154] At inference step, the camera image and UWB signal is captured and a process of data preparation similar to that in the training part is repeated to generate UWB wavelet transform coefficient infused part to be out-painted. ‘T’ such images are generated using the sinusoidal profile used earlier.
[0155] Given the image at t=T, it generates an estimate of the wavelet coefficients at t=T and on subtracting from the image at t=T, gives an estimate of original Image at t=0. The Wavelet coefficient infused image estimate generated by the training cell when scaled by (T-l) / T and added to the original image estimate returns an estimate of the image at t=T-l.
[0156] The training cell is executed iteratively with the output of the previous execution fed to the new one. At the completion, under the constrains set by the UWB data, a refined estimate of the out-painted image is generated. The out-painted images are transformed from latent space to the original image domain using the decoder part of the variational auto-encoder.
[0157] Figures 15A-15C illustrate an example of latent space sampling of an input, in accordance with an embodiment of the present disclosure.
[0158] Figure 15A illustrates first iteration of the latent space sampling of the input. The input is a latent representation of image and UWB embedding. At firststep, an inference for the input is generated at t=T, i.e., Atrusing the trained “Denoising U-net”. After the first step of inference, the denoising U-net generates the estimate of the UWB embedding present in the image (JNFtT). The estimate of the UWB embedding present in the image may be subtracted from the input to get the first estimate of the out painted image in latent space= ^tT- INFtT). The input to the next step is formed by adding the UWB embedding scaled by ftT 1to estimate a fraction of the UWB embedding instead of the whole UWB embedding.
[0159] Figure 15B illustrates second iteration of the latent space sampling of the input. Initially, an inference for the input (AtT ifrom the previous inference) is generated using the trained “Denoising U-net”. Then, the denoising U-net generates the estimate of the UWB embedding present in the image (JNFtT i). This image may be subtracted from the input to get the second estimate of the out painted image in latent space (AtT-2Further, the input to the next step is formed by adding the UWB embedding scaled by ftT-2to estimate a smaller fraction of the UWB embedding instead of the current one.
[0160] Figure 15C illustrates last iteration of the latent space sampling of the input. Initially, the input (dtifrom the previous inference) using the trained “Denoising U-net”. After the last step of inference, the denoising U-net generates the estimate of the UWB embedding present in the image (JNFtl). This image may be subtracted from the input to get the final estimate of the out painted image in latent space (Ato= Ati- INFti). Thus, Atois the final out painted image estimate generated by the denoising U-net in the latent space. This final out painted image estimate is then fed to the decoder part of the variational auto-encoder. This is due to the reason that the images generated by the de-noising U-net are in the latent space of the variational autoencoder. To transform them into human viewable format, they are required to be transformed from the latent space representation. The decoder part of the variational auto-encoder is used to perform this operation.
[0161] Figures 16-16B illustrate an exemplary out-painted image, in accordance with an embodiment of the present disclosure. As shown in Figure 16A, image ‘A’ is the original scene and image ‘B’ is the image captured by a user using the electronic device. Further, image ‘C’ is the captured image along with the UWBsignals and image ‘D’ is gallery view of the captured image. In the image ‘D’, the user is displayed an expand image button and on clicking expand button user is provided with an option to select how much image does the user want to expand and accordingly the out-painting process begins.
[0162] As shown in Figure 16B, image ‘A’ shows expansion selected by the user is white dotted lines and image ‘B’ shows the overall expanded out-painted image.
[0163] Further, image ‘C’ shows left-side expansion selected by the user and image ‘D’ shows the left-side expanded out-painted image.
[0164] Lastly, image ‘E’ shows right-side expansion selected by the user and image ‘F’ shows the right-side expanded out-painted image.
[0165] In yet another embodiment, as illustrated in Figure 17, a flow chart showing the method 1700 of generating the UWB sensor-based scene editing in an electronic device, in accordance with an embodiment of the present disclosure.
[0166] At step 1702, the method 1700 comprises providing a first data of a scene captured from a camera of the electronic device into a diffusion model, the camera having a first field of view (FoV) of the scene.
[0167] At step 1704, the method 1700 comprises providing second data of the scene captured using the UWB sensor of the electronic device into the diffusion model, the UWB sensor having a second FoV in addition to the first FoV.
[0168] At step 1706, the method 1700 comprises complementing the first data using the second data in the diffusion model to edit the scene captured from the first field of view.
[0169] In an embodiment, prior to providing the second data, the method 1700 further comprises receiving a plurality of UWB signals from the UWB sensor. The method 1700 further comprises removing temperature variations of the plurality of UWB signals received in Channel Impulse Responses (CIRs). The method 1700 further comprises de-cluttering the plurality of UWB signals by removing effect of coupling in the CIRs, wherein the coupling is direct reception of signals from UWB transmitter by UWB receiver. Further, the method 1700 comprises processing the de-cluttered plurality of UWB signals to extract magnitude data and phase data. Furthermore, the method 1700 comprises removing one or more spurious peaks in the CIRs generated by noise from the de-cluttered plurality of UWB signals togenerate de-noised plurality of UWB signals to be used for determining the second data.
[0170] In an embodiment, prior to removing one or more spurious peaks, the method comprises processing, using at least one of an interpolation method, a phase un-wrapping method and a magnitude de-noising method, the extracted magnitude data and the phase data to obtain transformed magnitude data and phase data.
[0171] Figure 18 illustrates a flow chart showing a method of the UWB sensorbased scene editing in the electronic device, in accordance with an embodiment of the present disclosure.
[0172] At step 1802, the method 1800 comprises capturing a scene using a camera of the electronic device, the camera having a first field of view (FoV) of the scene. The method 1800 comprises determining a first context of the scene from the first field of view. The method 1800 comprises capturing the scene using the UWB sensor of the electronic device, the UWB sensor having a second FoV in addition to the first FoV. The method 1800 comprises determining second context of the scene from a plurality of angle of arrivals (AO As) of a plurality of UWB signals reflected from one or more objects present in the second FoV of the scene. Further, the method 1800 comprises complementing, via a diffusion model, the second context to the first context to edit the scene captured from the first field of view.
[0173] In an embodiment, prior to determining the second context, the method 1800 comprises receiving a plurality of UWB signals from the UWB sensor. The method 1800 comprises removing temperature variations of the plurality of UWB signals received in Channel Impulse Responses (CIRs). The method 1800 comprises de-cluttering the plurality of UWB signals by removing effect of coupling in the CIRs, wherein the coupling is direct reception of signals from UWB transmitter by UWB receiver. The method 1800 comprises processing the decluttered plurality of UWB signals to extract magnitude data and phase data. Further, the method 1800 comprises removing one or more spurious peaks in the CIRs generated by noise from the de-cluttered plurality of UWB signals to generate denoised plurality of UWB signals to be used for determining the second context.
[0174] In an embodiment, prior to removing one or more spurious peaks, the method 1800 comprises processing, using at least one of an interpolation method, a phase un-wrapping method and a magnitude de-noising method, the extractedmagnitude data and the phase data to obtain transformed magnitude data and phase data.
[0175] In an embodiment, the method 1800 comprises extracting the plurality of Ao As of the de-noised plurality of UWB signals in the CIRs for the one or more objects appearing in an image. The method 1800 comprises filtering the AoAs of interest related to the one or more objects appearing in a particular cell of the image in the CIRs to obtain FoV limited de-noised UWB CIR signal. Further, the method 1800 comprises calculating a UWB wavelet embedding based on the FoV limited de-noised UWB CIR signal to be infused into the image. Furthermore, the method 1800 comprises preparing one or more training datasets by infusing the UWB wavelet embedding into the image using a sinusoidal schedule to obtain a UWB wavelet embedding infused image.
[0176] In an embodiment, the method 1800 comprises transforming the UWB wavelet embedding infused image into a latent space using a variational autoencoder. The method 1800 comprises de-noising the UWB wavelet embedding infused image by training U-Net model for wavelet embedding estimation. Further, the method 1800 comprises transforming the de-noised UWB wavelet embedding infused image from the latent space into an out-painted image using the variational auto-encoder.
[0177] In an embodiment, complementing the second context to the first context, the method 1800 comprises infusing the image obtained from the electronic device with the UWB wavelet embedding to generate a plurality of versions of training samples in latent space using the variational auto-encoder. The method 1800 comprises feeding the plurality of versions of the training samples to the trained denoising U-net model to generate a final out-painted image in latent space. Further, the method 1800 comprises transforming the final out-painted image from the latent space into the image using the decoder of the variational auto-encoder to obtain the edited scene of the image.
[0178] In general, for all cases, the end-to-end steps of the present disclosure are followed to generate the UWB based out-painted image. After the steps, the present disclosure performs an additional step of estimating UWB SNR for all the cells of pixels. The SNR is a scientific indicator of the strength of the reflected data and hence its reliability. Thus, low SNR indicates unreliable data. For all the cells ofpixels, for which the UWB SNR falls below a certain pre-defined threshold, alternative techniques are used to complete the out-painted image. For example, for cells of pixels covering area farther than 20m, some of the techniques are explained below.
[0179] 1. AI-Based Generation: When UWB SNR is below a certain threshold, Al-driven techniques are applied to maintain the realism and continuity of the out painted image. These include:
[0180] A) Al Extrapolation & Inpainting: Al models such as GANs or diffusion models may extrapolate missing content by analyzing the visible parts of the image. These models use patterns, colors, and textures from the available scene to generate plausible content for areas beyond UWB’s reach. Image inpainting algorithms can also fill in missing regions using surrounding pixels, ensuring smooth transitions and realistic content in distant areas.
[0181] B) Blending with Background: For scenes with natural elements (such as sky, horizon, or distant landscapes), Al model may blend these features into the out-painted image based on scene recognition, ensuring continuity in distant sections where details are less critical.
[0182] C) Scene Understanding for Contextual Awareness: Al models equipped with scene understanding (e.g., semantic segmentation) may infer plausible distant content based on the visible parts of the image. For example, in landscapes, Al model may generate open fields, trees, or sky, ensuring consistency with the foreground context.
[0183] 2. User Input-Based Fine-Tuning: For greater customization and control, user input may be leveraged to refine the out-painting:
[0184] A) Manual Adjustments: Users may manually define or adjust content beyond 20 meters, allowing them to customize the out-painted areas according to their vision of the scene.
[0185] B) Interactive Fine-Tuning: By providing interactive tools, users may tweak the Al-generated content, ensuring that details such as specific objects or features in the distant regions match their expectations.
[0186] Further, the present disclosure provides the following technical advantages:
[0187] Larger Field of View: Due to the larger field of view of the UWB sensor, the present disclosure is able to capture information about the surroundings of thetarget object which is missing in the image recorded by the camera. This can be used as a guiding mechanism towards image out-painting.
[0188] Spatial distribution (3rdDimension): In the present disclosure, the UWB sensor is able to record the depth information from its Field of View. This leads to the generation of a more accurate extension of the original image.
[0189] Integration with Existing Infrastructure: The mechanism of the present disclosure of using camera image and UWB sensor encoding is seamlessly integrated with existing electronic device, thus minimizing the need for costly upgrades or replacements. This ensures easy and wide outreach to the smartphone users.
[0190] Low-light performance: In the present disclosure, as the UWB reflections are illumination-agnostic, the lighting conditions do not hamper the performance of the out-painting task. Thus, the images captured in ill-lit surroundings may be extended in a similar and very accurate manner to the well-lit cases.
[0191] Figures 19A-19B illustrate examples of implementation of image out-painting in different scenarios, in accordance with an embodiment of the present disclosure.
[0192] As illustrated in Figure 19A, the user is unable to walk back any further due to an obstacle such as a wall (as shown in image A of Figure 19A) or a busy street (as shown in image B of Figure 19A). In such a case the user is unable to take a photograph of a wider expanse, as desired. The present disclosure due to the wider FoV of the UWB sensor, is able to capture the details of the objects and background that is not visible in the camera image. The system and method of the present disclosure is able to expand the image in all directions (right, left, top, and bottom) using UWB sensor data. Thus, by combining the context of the UWB Radar data (obtained from the UWB sensor) with the image data context, a realistic image of the camera view augmented with the intended wider expanse may be generated. The present disclosure facilitates accurate prediction and generation of the content that exists outside the captured frame, filling in the missing parts of the scene.
[0193] Figure 19B illustrates an exemplary scenario of low-light condition. As shown in Figure 19B-A, the parts of the image captured using the electronic device may be underexposed or obscured by shadows, leading to a loss of detail and clarity. Traditionally, this requires manual editing or advanced hardware to capture better-lit images. Using the present disclosure, the user may crop the image to focus on the well-lit region, as shown in Figure 19B-B. This ensures that the correctly exposed portion of the image is retained, avoiding the poor-quality sections.
[0194] During the image capture, the UWB sensor collects additional spatial information beyond what the camera sees, including details of objects and materials in the poorly lit or completely dark areas. This UWB data is not affected by light levels, meaning it still contains crucial spatial context about the low-lit parts of the image, including shapes, objects, and distances. With the UWB data available, the present disclosure is able to regenerate the cropped, low-lit parts of the image. The present disclosure uses the spatial context from the UWB sensor to accurately fill in the missing parts while adjusting for lighting. This allows the system / method to generate those parts of the image with better lighting conditions and higher clarity, even though the original image had low-light issues. The result, as shown in Figure 19B-C, is a fully lit image where the low-lit or shadowed regions are replaced with contextually accurate, well -lit parts generated through the out-painting process. In summary, the process uses the well-lit part of the image and regenerates the low-lit sections using UWB data, producing a final image that appears naturally illuminated and coherent without requiring perfect lighting during capture.
[0195] Further, in another exemplary scenario, the present disclosure may generate an out painted image on a mid-range phone without ultra-wide camera. The mid-range smartphones often lack ultra-wide cameras, which limit users' ability to capture wide-angle shots, especially for landscape photography. Due to this, the users with only a primary camera may feel constrained when trying to capture more of a scene, due to lack of necessary hardware to do so. In such a scenario, the out painting feature of the present disclosure allows users to take a photo with their primary camera and then expand the image using UWB-based data integration. This feature extends the shot by filling in the missing areas, effectively simulating an ultra-wide image. Thus, using the present disclosure, the user may expand the view to the left, right, top, or bottom, bridging the hardware gap.
[0196] While specific language has been used to describe the present disclosure, any limitations arising on account thereto, are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein. Thedrawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment.
Claims
CLAIMS1. A method of generating an ultra-wideband (UWB) sensor-based scene editing in an electronic device, comprising:providing first data of a scene captured from a camera of the electronic device into a diffusion model, the camera having a first field of view (FoV) of the scene;providing second data of the scene captured using the UWB sensor of the electronic device into the diffusion model, the UWB sensor having a second FoV in addition to the first FoV;obtaining an edited image by complementing the first data using the second data in the diffusion model, the edited image reflecting, in a scene captured from the first FoV, spatial information of one or more objects located within the second FoV and outside the first FoV.
2. The method as claimed in claim 1, wherein prior to providing the second data, the method comprises:receiving a plurality of UWB signals from the UWB sensor; removing temperature variations of the plurality of UWB signals received in Channel Impulse Responses (CIRs);de-cluttering the plurality of UWB signals by removing effect of coupling in the CIRs, wherein the coupling is direct reception of signals from UWB transmitter by UWB receiver;processing the de-cluttered plurality of UWB signals to extract magnitude data and phase data; andremoving one or more spurious peaks in the CIRs generated by noise from the de-cluttered plurality of UWB signals to generate de-noised plurality of UWB signals to be used for determining the second data.
3. The method as claimed in claim 2, wherein prior to removing one or more spurious peaks, the method comprises:processing, using at least one of an interpolation method, a phase unwrapping method and a magnitude de-noising method, the extracted magnitude data and the phase data to obtain transformed magnitude data and phase data4. A method of an ultra-wideband (UWB) sensor-based scene editing in an electronic device, the method comprising:capturing a scene using a camera of the electronic device, the camera having a first field of view (FoV) of the scene;determining first context of the scene from the first field of view; capturing the scene using the UWB sensor of the electronic device, the UWB sensor having a second FoV in addition to the first FoV;determining second context of the scene from a plurality of angle of arrivals (AO As) of a plurality of UWB signals reflected from one or more objects present in the second FoV of the scene; andobtaining an edited image by complementing, via a diffusion model, the second context to the first context, the edited image reflecting, in the scene captured from the first FoV, the second context of the one or more objects that are located outside the first FoV.
5. The method as claimed in claim 4, wherein prior to determining the second context, the method comprises:receiving a plurality of UWB signals from the UWB sensor; removing temperature variations of the plurality of UWB signals received in Channel Impulse Responses (CIRs);de-cluttering the plurality of UWB signals by removing effect of coupling in the CIRs, wherein the coupling is direct reception of signals from UWB transmitter by UWB receiver;processing the de-cluttered plurality of UWB signals to extract magnitude data and phase data; andremoving one or more spurious peaks in the CIRs generated by noise from the de-cluttered plurality of UWB signals to generate de-noised plurality of UWB signals to be used for determining the second context.
6. The method as claimed in claim 5, wherein prior to removing one or more spurious peaks, the method comprises:processing, using at least one of an interpolation method, a phase unwrapping method and a magnitude de-noising method, the extracted magnitude data and the phase data to obtain transformed magnitude data and phase data.
7. The method as claimed in claim 5, further comprising: extracting the plurality of AoAs of the de-noised plurality of UWB signals in the CIRs for the one or more objects appearing in an image;filtering the AoAs of interest related to the one or more objects appearing in a particular cell of the image in the CIRs to obtain FoV limited de-noised UWB CIR signal;calculating a UWB wavelet embedding based on the FoV limited de-noised UWB CIR signal to be infused into the image; andpreparing one or more training datasets by infusing the UWB wavelet embedding into the image using a sinusoidal schedule to obtain a UWB wavelet embedding infused image.
8. The method as claimed in claim 7, further comprising: transforming the UWB wavelet embedding infused image into a latent space using a variational auto-encoder;de-noising the UWB wavelet embedding infused image by training U-Net model for wavelet embedding estimation; andtransforming the de-noised UWB wavelet embedding infused image from the latent space into an out-painted image using the variational auto-encoder.
9. The method as claimed in claim 7, wherein complementing the second context to the first context, further comprising:infusing the image obtained from the electronic device with the UWB wavelet embedding to generate a plurality of versions of training samples in latent space using the variational auto-encoder;feeding the plurality of versions of the training samples to the trained denoising U-net model to generate a final out-painted image in latent space; and transforming the final out-painted image from the latent space into the image using the decoder of the variational auto-encoder to obtain the edited scene of the image.
10. A system for an ultra-wideband (UWB) sensor-based scene editing in an electronic device, the device comprising:one or more processors;a memory coupled with the one or more processors, wherein the one or more processors are configured to:capture a scene using a camera of the electronic device to obtain the image, the camera having a first field of view (FoV) of the image;determine first context of the scene from the first field of view;capture the scene using the UWB sensor of the electronic device, the UWB sensor having a second FoV in addition to the first FoV;determine second context of the scene from a plurality of angle of arrivals (AO As) of a plurality of UWB signals reflected from one or more objects present in the second FoV of the scene; andobtain an edited image by complementing, via a diffusion model, the second context to the first context, the edited image reflecting, in the scene captured from the first FoV, the second context of the one or more objects that are located outside the first FoV.
11. The system as claimed in claim 10, wherein prior to determining the second context, the one or more processors are configured to:receive the plurality of UWB signals from the UWB sensor;remove temperature variations of the plurality of UWB signals received in Channel Impulse Responses (CIRs);de-clutter the plurality of UWB signals by removing effect of coupling in the CIRs, wherein the coupling is direct reception of signals from UWB transmitter by UWB receiver;process the de-cluttered plurality of UWB signals to extract magnitude data and phase data; andremove one or more spurious peaks in the CIRs generated by noise from the de-cluttered plurality of UWB signals to generate de-noised plurality of UWB signals to be used for determining the second context.
12. The system as claimed in claim 11, wherein prior to removing one or more spurious peaks, the one or more processors are configured to:process, using at least one of an interpolation method, a phase un-wrapping method and a magnitude de-noising method, the extracted magnitude data and the phase data to obtain transformed magnitude data and phase data.
13. The system as claimed in claim 11, the one or more processors are configured to:extract the plurality of AoAs of the de-noised plurality of UWB signals in the CIRs for the one or more objects appearing in an image;filter the AoAs of interest related to the one or more objects appearing in a particular cell of the image in the CIRs to obtain FoV limited deda-noised UWB CIR signal;calculate a UWB wavelet embedding based on the FoV limited de-noised UWB CIR signal to be infused into the image; andprepare one or more training datasets by infusing the UWB wavelet embedding into the image using a sinusoidal schedule to obtain a UWB wavelet embedding infused image.
14. The system as claimed in claim 13, the one or more processors are configured to:transform the UWB wavelet embedding infused image into a latent space using a variational auto-encoder;de-noise the UWB wavelet embedding infused image by training U-Net model for wavelet embedding estimation; andtransform the de-noised UWB wavelet embedding infused image from the latent space into an out-painted image using the variational auto-encoder.
15. The system as claimed in claim 13, wherein to complement the second context to the first context, the one or more processors are configured to:infuse the image obtained from the electronic device with the UWB wavelet embedding to generate a plurality of versions of training samples in latent space using the variational auto-encoder;feed the plurality of versions of the training samples to the trained denoising U-net model to generate a final out-painted image in latent space; and transform the final out-painted image from the latent space into the image using the decoder of the variational auto-encoder to obtain the edited scene of the image.