Method for reducing image variability under global and local illumination variations

By combining the mechanisms of human and insect visual systems, and employing color adaptation, LMC models, and color antagonistic coding, the instability of color caused by illumination changes and the influence of shadows in computer vision and robotics are resolved, achieving stable color discrimination and visual navigation accuracy.

CN121942007APending Publication Date: 2026-04-28OPTERAN TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
OPTERAN TECH LTD
Filing Date
2024-08-29
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve color constancy in computer vision and robotics without adjusting lighting or using dedicated light sensors, and the impact of shadows on visual navigation systems remains unresolved.

Method used

Using a bio-inspired algorithm that combines the mechanisms of human and insect visual systems, the algorithm estimates local and global illumination of images through color adaptation, LMC model and color antagonistic coding, and detects and removes shadows, thus achieving color constancy and shadow removal.

Benefits of technology

It improves the color discrimination and visual location recognition accuracy of image processing, enhances scene similarity assessment and resilience to changes in lighting conditions, and improves the positioning accuracy of robot navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121942007A_ABST
    Figure CN121942007A_ABST
Patent Text Reader

Abstract

A system, apparatus, and method for processing an image based on color constancy are provided herein. The method comprises the following steps: obtaining an image in a first color space; converting the image into data in a second color space; transforming the data using color adaptation; performing a first normalization on the transformed data, wherein the first normalization includes applying a dynamic spatial filtering technique to adjust the transformed data based on the light intensity; applying a set of filters to the normalized data, wherein the set of filters are convolved based on the normalized data associated with the image; performing a second normalization on the filtered data to obtain an illumination estimate of the image related to the filtered data; and outputting normalized data from the second normalization, where the normalized data maintains color constancy based on the illumination estimate, thereby removing illumination from the normalized data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a system, apparatus, and method for estimating the local and global illumination of an image and removing illumination through color correction to maintain the color constancy of the image and minimize the shadow effects of the image. Background Technology

[0002] Color constancy refers to the ability to perceive the color of an object regardless of changes in light source. This ability is reportedly inherent in species such as humans, fish, and bees. Color constancy evolved in these species to aid object recognition by consuming less memory, while maintaining equal or higher accuracy and more stable discrimination. Without color constancy, the color of an object under varying light variability becomes an unreliable factor, impairing the species' ability to accurately identify objects.

[0003] Numerous behavioral and neurobiological studies have explored the mechanisms behind color constancy. Two possible neural mechanisms have been proposed for color constancy: color adaptation and antagonistic processing. Color adaptation occurs at the level of a single photoreceptor, achieved by adjusting its activity according to relative light within a local area. Higher-level neural processing, namely dual-antagonistic cells in the early visual system (regions V1 and V4), may also be attributed to color constancy. On the other hand, antagonistic processing theory proposes that one member of a color pair inhibits the other.

[0004] In computer vision and robotics, particularly for robust color-based object recognition and tracking, a key requirement is the recording and memorization of reliable color cues unaffected by any external light variations. This is challenging when the light source is unknown and various lighting conditions exist simultaneously in the scene. Computing color constancy has led to different solutions for robust color coding, used to correct color casts in images to obtain typical images under white light.

[0005] In this regard, many biologically plausible and specific solutions have been proposed for calculating color constancy, ranging from simple algorithms based on low- or mid-level image feature statistics (such as gray-world models, white point algorithms, maximum red, green, and blue (RGB) calculations, gray-level shadows, and gray-level edges) to algorithms based on complex statistics and machine learning (including gamut mapping, Bayesian methods, and neural network-based methods). Although these solutions are computationally expensive, no solution has yet emerged that can achieve human-level color constancy regardless of lighting conditions and the type of light sensor. Currently, no solution can perfectly solve this problem in real-world natural scenes, especially one that simultaneously utilizes both color adaptation and antagonistic processing theories. Furthermore, no solution employs a process inspired by insect eyes—neural normalization through the response of neighboring photoreceptors—to estimate image illumination.

[0006] For these reasons, there remains an unmet need in the field of image processing: to computationally maintain human color constancy without adjusting lighting or using dedicated light sensors. This invention presents a novel bio-inspired solution to address this color constancy challenge, applying algorithms inspired by human and insect visual systems to provide stable color appearance and high color discrimination when processing images.

[0007] Furthermore, it is recognized that the success of robot navigation depends on the accuracy of visual location recognition and localization. However, shadows created by varying lighting conditions pose a significant challenge to these processes. Shadows originating from different light sources introduce differences in color, texture, and shape, causing inconsistencies in traditional visual features across different scenarios. These inconsistencies hinder accurate location matching and recognition, thereby impairing overall navigation performance. It should be understood that, in this context, improving recognition accuracy does not require accurate colors, but rather consistent colors.

[0008] To address this concern, in the context of color constancy, a method for detecting and removing shadows is needed to mitigate their adverse effects on vision-based navigation systems. Shadows cast by objects and structures under varying lighting angles and intensities profoundly alter the appearance of scenes and objects. This alteration introduces discrepancies into the visual data captured by robot sensors, making it difficult to accurately match features under diverse lighting conditions. Consequently, the reliability and accuracy of visual location recognition and localization are compromised, weakening the robot's navigation capabilities.

[0009] Therefore, it can be recognized that the methods for shadow detection and shadow removal described in this paper are of critical significance in the fields of object recognition and visual location recognition, especially for providing stable color appearance and high color discrimination when processing images. These methods significantly improve the ability of visual location recognition in robotics and computer vision by: 1. ensuring higher feature consistency; 2. enhancing scene similarity assessment; 3. strengthening resilience to changes in lighting conditions; and 4. ultimately achieving superior localization accuracy.

[0010] The embodiments and aspects described below are not limited to implementations that address any or all the disadvantages of the known methods described above. Summary of the Invention

[0011] This summary is provided to introduce, in a simplified form, some concepts further described in the following detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to define the scope of the claimed subject matter; variations and alternative features that contribute to the work of the invention and / or achieve substantially similar technical effects should be considered to fall within the scope of the invention disclosed herein.

[0012] As part of the Opteran vision framework, this disclosure provides a bio-inspired solution to the challenge of color constancy, for example, in the Opteran Development Kit (ODK) by adjusting the color distributions of two cameras to match overlapping application scenarios between them. In addition to other features described herein, the solution encompasses an algorithm / model for stable color appearance and high color discrimination, inspired by the visual systems of humans and major insects and tailored to the functional characteristics of the ODK camera system. Specifically, the model upon which the solution is based integrates three mechanisms: a color adaptation model, an LMC model (a model of thin-plate unipolar cells in insect visual lobes), and color antagonistic encoding, which are further described in the following sections.

[0013] Furthermore, bio-inspired solutions to the color constancy challenge include methods for estimating local illumination within partitioned regions of an input image. Also inspired by research on eye movement and selective attention in visual science, this method relies on three key components: selective attention mechanisms, the gray border hypothesis, and color normalization. These three components are integrated into a single algorithm / model to reduce local illumination.

[0014] Furthermore, this disclosure, in conjunction with illumination correction as described herein, provides another solution for image processing configured to detect and accurately remove shadows from input images caused by varying lighting conditions. These shadows can introduce inconsistencies in image features and significantly impact accurate perception. Therefore, shadow detection and removal play a crucial role in enhancing visual location recognition and localization methods for applications such as robot navigation. Specifically, the proposed solution integrates two mechanisms—shadow detection and shadow removal—as described herein, aiming to restore true colors and textures obscured by shadows inherent in the input image, providing a valuable complement to the Opteran vision framework.

[0015] In a first aspect, this disclosure provides a method or computer-implemented method for processing an image based on color constancy to remove illumination from the image, the method comprising: obtaining an image in a first color space; converting the image into data in a second color space; transforming the data using color adaptation; performing a first normalization on the transformed data, wherein the first normalization includes applying dynamic spatial filtering techniques to adjust the transformed data based on light intensity; applying a set of filters to the normalized data, wherein the set of filters is convolved based on the normalized data associated with the image; performing a second normalization on the filtered data to obtain an illumination estimate of the image associated with the filtered data; and outputting normalized data from the second normalization, wherein the normalized data maintains color constancy based on the illumination estimate, thereby removing illumination from the normalized data.

[0016] In a second aspect, this disclosure provides a method or computer-implemented method for reducing the influence of illumination on an image, the method comprising: receiving an input image; dividing the input image into multiple regions; analyzing the multiple regions based on color information and spatial location of pixels in each region; selecting a subset of regions affected by an illuminator from the multiple regions based on the analysis; identifying colored edges for at least the subset of regions; extracting color information from the colored edges; using the extracted color information to decompose the reflection component and illumination component of the input image; correcting the illumination of the input image based on the decomposed reflection component and illumination component; and outputting an illumination-corrected image.

[0017] In a third aspect, this disclosure provides a method or computer-implemented method for providing a shadowless image, the method comprising: receiving an input image in a first color space, wherein the input image includes at least one shadow region; generating a shadow region mask for the at least one shadow region; removing shadows from the shadow region of the input image based on the shadow region mask; and outputting a shadowless image.

[0018] In a fourth aspect, this disclosure provides an apparatus for processing an image to maintain the color constancy of the image, the apparatus comprising: at least one model configured to perform steps according to the first aspect, the second aspect and / or the third aspect and any aspect described herein.

[0019] In a fifth aspect, this disclosure provides a system for processing an image to establish color constancy of the image by removing illumination from the image, the system comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to perform the first aspect, the second aspect, and / or the third aspect, and any aspect described herein.

[0020] The methods described herein can be executed by software in a machine-readable form on a tangible storage medium, such as a computer program comprising computer program code means which, when run on a computer, is adapted to perform all the steps of any of the methods described herein, and wherein the computer program can be embodied on a computer-readable medium. Examples of tangible (or non-transitory) storage media include disks, thumb drives, memory cards, etc., and do not include propagation signals. The software can be adapted to execute on a parallel or serial processor, such that the method steps can be executed in any suitable order or simultaneously.

[0021] The methods described herein or computer-implemented methods can be described based on one or more models. Therefore, it should be understood that the term "method" can be used interchangeably with "model" where appropriate throughout the disclosure, although in some cases a model may not necessarily correspond to a single method, but may be incorporated as part of the method.

[0022] This application acknowledges that firmware and software can be valuable, separately tradable goods. It is intended to cover software that runs on or controls “dumb” or standard hardware to perform desired functions. It is also intended to cover software that “describes” or defines the configuration of hardware, such as HDL (Hardware Description Language) software used to design silicon chips or to configure general-purpose programmable chips to perform desired functions.

[0023] The optional features or options described herein can be suitably combined, as will be apparent to those skilled in the art, and can be combined with any aspect of the invention. Attached Figure Description

[0024] With reference to the following figures, embodiments of the present invention will be described by way of example, wherein:

[0025] Figure 1aThis is a flowchart of a model architecture for image processing based on various aspects of this disclosure;

[0026] Figure 1b This is a flowchart of image processing based on the model architecture according to various aspects of this disclosure;

[0027] Figure 2 This is a schematic diagram illustrating the application of the Spyder Checker (Macbeth ColourChecker color chart) model under different lighting conditions according to various aspects of this disclosure;

[0028] Figure 3 This is a schematic diagram comparing the pixel intensity distribution between an initial image and a corrected image according to various aspects of this disclosure;

[0029] Figure 4 This is a schematic diagram of the Fano factor (relative variance) of selected plots according to various aspects of this disclosure;

[0030] Figure 5 It is based on the environmental datasets of various aspects of this disclosure (and) Figure 9 A schematic diagram of the color distances between the color swatches in related set 1);

[0031] Figure 6 This is a schematic diagram of color matching between two ODK cameras that capture raw images and model-corrected images according to various aspects of this disclosure;

[0032] Figure 7 This is a schematic diagram of the Spyder Checker-Macbeth ColorChecker color chart based on various aspects of this disclosure;

[0033] Figure 8 This is a schematic diagram illustrating exemplary data acquisition settings based on various aspects of this disclosure;

[0034] Figure 9 This is a schematic diagram of example images from an environmental dataset (set 1) based on various aspects of this disclosure;

[0035] Figure 10 This is a schematic diagram of example images from another environmental dataset (Set 2) based on various aspects of this disclosure;

[0036] Figure 11 It is a schematic diagram comparing the pixel intensity distribution between an initial image and a corrected image, wherein the initial image is captured by a Pi camera according to various aspects of this disclosure;

[0037] Figure 12 This is a schematic diagram of the color distances between color palettes in the environmental dataset (Set 2) based on various aspects of this disclosure;

[0038] Figure 13 It is a block diagram of a computing device or apparatus suitable for implementing various aspects of this disclosure;

[0039] Figure 14 This is a schematic diagram illustrating the application of environmental models under different lighting conditions based on various aspects of this disclosure;

[0040] Figure 15a This is a flowchart of another model architecture for image processing based on various aspects of this disclosure;

[0041] Figure 15b It is a flowchart of image processing based on the other model architecture according to various aspects of this disclosure;

[0042] Figure 16 This is a schematic diagram of an exemplary output of image processing based on the other model architecture according to various aspects of this disclosure;

[0043] Figure 17 These are schematic diagrams of violin plots, which illustrate the color distance (ΔE) between images in each dataset for image processing using the other model architecture, according to various aspects of this disclosure.

[0044] Figure 18 This is a schematic diagram illustrating another model application for environments under different lighting conditions, based on various aspects of this disclosure;

[0045] Figure 19a This is a flowchart of another model architecture 1900 for image processing based on various aspects of this disclosure.

[0046] Figure 19b It is a flowchart of image processing based on the other model architecture according to various aspects of this disclosure;

[0047] Figure 20 These are schematic diagrams illustrating exemplary inputs and outputs of image processing based on the other model architecture according to various aspects of this disclosure; and

[0048] Figure 21 This is a schematic diagram illustrating exemplary inputs and outputs for image processing of different input images.

[0049] Common reference numerals are used in all the accompanying drawings to indicate similar features. Detailed Implementation

[0050] The embodiments of the invention are described below by way of example only. These examples illustrate suitable ways of practicing the invention as currently known to the applicant, although they are not the only ways of implementing the invention. The description sets forth the function of the example and the sequence of steps for constructing and operating the example. However, the same or equivalent functions and sequences can be implemented by different examples.

[0051] Color constancy is an important characteristic of color perception. It refers to the ability to perceive the color of an object without being affected by changes in the light source. This ability has been reported to exist in different species, such as humans, fish, and bees. Color constancy allows different species to perceive the color of an object relatively consistently under varying lighting conditions, or to identify objects regardless of lighting conditions. For example, at midday when the main light is daylight, and at sunset when the main light is red, objects appear green to us.

[0052] In fields such as computer vision and robotics, the mechanisms behind the concept of color constancy are therefore considered important. It is well known that by borrowing the concept of color constancy mechanisms, these mechanisms can be implemented to efficiently record and store reliable color cues that remain stable under various lighting conditions. Although several computational models have been proposed to realize color constancy, challenges remain in providing a more robust color constancy algorithm / model that is efficient and approximates human and animal-level vision.

[0053] The purpose of this invention is to address and overcome at least some of the challenges of existing computational models. This invention utilizes the neural mechanisms of color constancy in primates and insects to design a novel algorithm that maintains a stable color appearance under varying light sources. Evaluations of this invention using different datasets reveal that the proposed model reduces the intensity variations of color patches illuminated by different natural and artificial light sources (locally and globally).

[0054] Importantly, this invention enhances the camera's ability to distinguish colors by increasing the color distance between objects. This is useful for addressing the challenge of matching the overlap between two cameras in ODK by adjusting the camera's color distribution as shown in the figure. In short, this invention demonstrates some of the potential advantages of the Opteran vision framework by upgrading color encoding within the context of image processing.

[0055] One aspect of this invention is a biologically plausible algorithm for estimating illumination and a proposed stable color coding mechanism for visual object recognition in autonomous robots. The algorithm draws inspiration from accurate and fast color coding algorithms at the single neuron and neural network levels of the human and insect visual systems. In this paper, a multilayer neural network is proposed that combines three mechanisms from the proposed visual mechanisms for color constancy: retinal photoreceptor adaptation, lateral normalization between photoreceptors, and center-periphery spectral antagonism in early visual systems (i.e.,...). Figure 1a and Figure 1b The following is an overview of the proposed model, which serves as an exemplary embodiment or aspect of the invention. It should be understood that this aspect can be readily combined with other aspects of the model described herein.

[0056] It should be understood that the methods, algorithms, and / or models described herein may include one or more steps for partitioning an input image, original image, or training data / dataset into multiple partitions (this process is also referred to as image segmentation) before further processing of the input image, original image, or training data / dataset as described in this disclosure. Methods for image segmentation or partitioning images may include, but are not limited to, thresholding, region growing, edge-based segmentation, clustering, histogram-based bundling, k-means clustering, watershed, active contouring, ML-based segmentation using convolutional neural networks, graph-based segmentation, and superpixel-based segmentation. It is understood that any suitable image segmentation method, as well as some novel methods from the Opteran vision framework, can be used for global illumination, local illumination, shadow detection, and removal (obtaining shadow-free images) as described in the following sections.

[0057] Global Illumination

[0058] In this disclosure, a color space refers to an abstract space that has a specific organization of reproducible representations of colors. It should be understood that the color space can be arbitrary, that is, assigning physically realized colors to a set of physical color samples with corresponding specified color names, or being structured through mathematical rigor.

[0059] In this invention, the image is converted to a Long-Medium-Short (LMS) color space to simulate the L, M, and S cone cells in the human eye. Initially, gamma correction applied to the colors (Equation 1) is removed to generate a linear RGB image. Gamma correction, as used herein, refers to a non-linear operation used to adjust the brightness and contrast of an image. It involves applying a non-linear mapping to the pixel values ​​of an image to compensate for the inherent non-linear response of display devices such as monitors and televisions. Gamma correction can be used to control the brightness of an image. It helps ensure that the perceived brightness and contrast of an image remain consistent across different devices and viewing environments. Here, we assume that the original image includes gamma correction. The gamma correction of the original image can be removed, resulting in an image where each RGB value of the processed color is transformed into an equivalent linear RGB color space. Gamma correction transforms the color intensities from the physical world into a more uniform arrangement for human perception. Matrix transformations are used. We convert the linear RGB image into an LMS cone cell input.

[0060] In the next stage, we process LMC cells by applying a single transform T to the response of the photoreceptor (Equation 5). This transform normalizes the output of the photoreceptor relative to the collective activity of all photoreceptors in the same color channel. We ultimately propose an antagonistic coding model based on single-antagonistic and double-antagonistic cells in the human and insect visual systems (see Double-Antagonistic Model). By combining these stages of color processing, we can estimate the global illumination of the input image illuminated by an external light source. This allows us to correct for a wide range of colors in images under different lighting conditions due to different natural light sources and varying artificial light. For illustrative purposes, an inverse transform is applied... Equation 2 transforms the corrected image in the LMS color space to the sRGB space. In the following sections, we briefly describe the color constancy models and ideas behind each component: a) color adaptation, b) the LMC model, and c) the dual-antagonistic model of this invention. Further details will be discussed in the later "Detailed Implementation" section.

[0061] a) Color adaptation in the context of color constancy is defined as the ability of an animal to modulate retinal sensitivity in response to changes in light source. Color adaptation is a technique for explaining color constancy based on the ability of animals (including humans) to modulate the sensitivity of their photoreceptors to changes in light source. Color adaptation is closely related to the modulation of cone cell sensitivity in the retina in response to changes in light intensity.

[0062] An example of color adaptation technology applies an adapted version of Von Kries' color adaptation, based on the LMS cone cell sensitivity response function (LMS color space). The literature describing the general approach of Von Kries' color adaptation (Luo, M. Ronnier, “A review of chromatic adaptation transforms,” Review of Progress in Coloration and Related Topics, Vol. 30 (2000): 77–92) is incorporated herein by reference.

[0063] For the current adaptation, each cone cell increases or decreases its responsivity based on the spectral energy of the illuminator. Each cone cell increases or decreases its spectral sensitivity to adapt from the original illumination to the new illumination, thereby perceiving the object as a constant color. Based on LMS cone cell gain, von Klee's adaptation transforms one illuminator to another to maintain white consistency in both systems. Further extensions to von Klee's adaptation are also introduced. For example, several CAT transformations and the Bradford model are popular. Here, we implement some of these models by applying the matrix transformations shown in "Implementation of Chromatic adaptation transform".

[0064] b) The Thin-Layer Unipolar Cell Model (LMC Model): It has been hypothesized that the dendrites of insect thin-layer unipolar cells integrate visual information from adjacent photoreceptors. Therefore, they contribute to the spatial summation of visual information and normalize the photoreceptor responses using temporal encoding. This process also enhances the information capacity of neurons. Insects use this mechanism to improve their visual sensitivity at night. In this work, we demonstrate that this normalization process is equally beneficial in excluding illumination variations and improving color constancy by utilizing simple transformations of photoreceptor activity.

[0065] c) Single-antagonistic and dual-antagonistic cell models (S-DO antagonistic model): Antagonistic processing theory posits that, despite the increasing number of components reported in recent years, color perception is still controlled by the activity of three antagonistic color systems: red-green, blue-yellow, and white-black. This theory suggests that one member of a color pair inhibits the other. For example, we can see subtle yellow-green and subtle red-yellow hues, but we can never see subtle red-green or subtle yellow-blue hues. The following references are incorporated herein by reference: Shapley, Robert and Michael J. Hawken, “Color in the cortex: single-and-double-opponent cells,” Vision Research 51.7 (2011): 701-717; Hurvich, Leo M. and Dorothea Jameson, “An opponent-process theory of color vision,” Psychological Review 64 Vol. 6 No. 1 Part 1 (1957): 384; and Solomon, Samuel G. and Peter Lennie, “The machinery of color vision,” Nature Reviews Neuroscience 8.4 (2007): 276-286.

[0066] The output of photoreceptors propagates in a color-antagonistic manner via single antagonist cells in the retinal ganglion layer and the LGN, and dual antagonist cells in the V1 visual cortex. Single antagonist cells process color information through the center-periphery structure of their receptive field (RF). Different types of single antagonist cells exist, encoding red-green, blue-yellow, and black-white antagonistic color contrasts, respectively. For example, the RF structure of type II single antagonist cells is modeled as two Gaussian functions: a center for red excitation and a periphery for green inhibition. Furthermore, dual antagonist cells have been found in V1 in multiple studies. Detecting local color contrast is an important function of these cells. Dual antagonist cells simultaneously compute color antagonism and spatial antagonism. It has been suggested that dual antagonist cells may be the basis of illumination encoding. Interestingly, dual antagonist cells have also been observed in goldfish and bees. In this disclosure, we propose a simplified form of the S-DO antagonist model for color constancy, which combines the two other color constancy mechanisms mentioned above to improve the estimation of global illumination. The following references are cited and incorporated herein by reference: Conway, Bevil R. et al., “Advances in color science: from retina to behavior”, Journal of Neuroscience 30.45 (2010): 14955-14963; Shapley, Conway, Bevil R., David H. Hubei and Margaret S. Livingstone, “Color contrast in macaque V1”, Cerebral Cortex 12.9 (2002): 915-925.

[0067] One approach to evaluating color constancy models is to use angular error, which measures the angular distance between the estimated illumination and the actual illumination. Due to the difficulty in acquiring data corresponding to real-world illumination conditions, we decided to use a different method to quantitatively analyze the model's performance. In this way, pre-selected pixel blocks within segments on the ColourChecker, illuminated under varying lighting conditions, were analyzed before and after model application. The model was evaluated using three different datasets for the color constancy task: the Spyder Checkr 24 dataset, the Environmental dataset, and the Lab dataset. The Spyder Checkr 24, from datacolour, is a standard tool for color calibration. This dataset contains 34 color chart images illuminated under different natural and artificial light conditions. The Environmental dataset contains raw RGB paired images captured under varying lighting conditions (everyday natural light and artificial light from lamps) (see Data Acquisition). The Lab dataset includes RGB images of the lab space, also captured under varying lighting conditions and light sources, as shown in the attached figure.

[0068] To analyze color variability under different lighting conditions, a Fano factor is calculated for each channel of the selected patch. The Fano factor measures the relative variance of color intensity and shows the degree of variation relative to the population mean. The Fano factor is defined as follows: ,in and These are the variance and mean of the selected patches for each color channel i=R, G, or B, respectively.

[0069] The distance between two colors allows for a quantitative analysis of how far apart they are. The ΔE metric measures the degree of color change under different lighting conditions and validates the improvement in color discrimination achieved by our model / algorithm. It evaluates distance in the CIELab color space and represents the relative perceived magnitude of color difference. The larger the ΔE, the greater the distance between the colors.

[0070] Implementation details

[0071] With regard to this disclosure, it should be recognized that the following other literature (Stokes, Michael, “A standard default color space for the internet-srgb” http: / / www.w3.org / Graphics / Color / sRGB.html (1996); and Anderson, Matthew et al., “Proposal for a standard default color space for the internet-srgb”, Color and Imaging Conference, Vol. 1996, No. 1, Society for Imaging Science and Technology, 1996) are incorporated herein by reference.

[0072] Linear RGB: For each RGB value of the color to be processed in the input image, we transform these colors into the so-called linear RGB space by removing gamma correction using this formula:

[0073]

[0074] Where c srgb These are the (R, G, B) values ​​for each pixel in the image. Gamma correction can be reapplied later using the inverse formula:

[0075]

[0076] LMS Color Space: Because the spectral sensitivity of digital cameras differs from that of photoreceptors in human and insect vision, we transformed the input image from the linear RGB space to LMS cone cell input (L cone cells, M cone cells, and S cone cells) based on previous computational and biological experiments.

[0077]

[0078] in

[0079] Matrix M rgb→lms Map the pixels of each variation R(x,y), G(x,y), and B(x,y) of the input image Z(x,y) to l(x,y), m(x,y), and s(x,y) in the LMS space. The yellow component y(x,y) - m(x,y) + s(x,y) and the illumination component u(x,y) - l(x,y) + m(x,y) + s(x,y) can be simply calculated. Similarly, we can use the inverse matrix M... -1rgb→lms Transform the color information in the LMS color space back into the RGB color space.

[0080] Implementation of Color Adaptive Transformation: By applying a color adaptive transformation, we implement color adaptation for converting a source image to a target illumination. Color adaptation is performed using a reference white point. The transformed image contains an improved white range. Other colors are transformed based on the illumination transformation obtained through the white point transformation. Color adaptation for the image is implemented in two steps: step 1) estimating the global illuminator of the image; and step 2) converting the image to the target illuminator using the illuminator obtained in step 1.

[0081] Step 1) Gray-World Model: To estimate global illumination, we used the gray-world method. The gray-world assumption is a simple approach that assumes an image contains objects with different reflected colors, these reflected colors are uniform from minimum to maximum intensity, and therefore averaging over all pixels yields gray. An illumination estimate (a calculated or derived data representation of the illumination present in the image) is calculated by averaging all pixel values ​​over each channel. This illumination estimate will give an average color that approximates the color of the illuminated object. For an image with an equivalent color representation, the illumination estimate gives an average color of gray. The purpose of illumination estimation is to separate the effects of lighting from the intrinsic colors of objects in the image.

[0082]

[0083] Where R avg G avg and B avg It is the average value of each image channel (R, G, B).

[0084] Step 2) Transforming the image into the target image: First, we convert the sRGB values ​​of the image to the LMS cone space (see LMS color space). Then, by multiplying the LMS values ​​by one of the color adaptation transformations listed below, we transform the source image into the target irradiation. Below are some popular color adaptation matrix transformations, i.e., the standard von Klee method as shown below:

[0085]

[0086] The LMC model has been proposed, suggesting that adjacent photoreceptors in fruit flies are laterally connected via large monopolar cells (LMCs), which improves contrast encoding and facilitates the spatial summation of visual information. In fact, LMC cells tend to homogenize the distribution of photoreceptor responses to visual input by sending feedback to photoreceptors and altering their temporal responses. Therefore, it is recognized that the references (Laughlin, Simon, “A simple coding procedure enhances a neuron’s information capacity”, *Zeitschrift fur Naturforschung c*, Vol. 36, No. 9-10 (1981): 910-912; and Stockl, Anna Lisa, David Charles O’Carroll, and Eric James Warrant, “Hawkmoth lamina monopolar cells act as dynamic spatial filters to optimize vision at different light levels”, *Science Advances*, Vol. 6, No. 16 (2020): eaaz8645) are hereby incorporated in this paper by reference.

[0087] Therefore, following this normalization mechanism, we implemented LMC cells in our model / algorithm by modifying the photoreceptor response through linear transformation, such that all responses from low to high photoreceptor responses are within [0, R]. max They are evenly distributed within the range of ].

[0088] To implement a neural network version of LMC normalization, we can consider a simple form of temporal encoding of the photoreceptor by defining a time response matrix T, as follows: For each pixel (x,y) in an image I(x,y) of size m pixels × n pixels at one color channel R, G, or B: T((y - 1) m + x, I(x,y)) - 1, and for the remaining elements (i,j) of the matrix: T(i,j) - 0. In neuroscience, temporal coding is a phenomenon where the timing of a pulse signal indicates the value associated with that pulse signal. The temporal response matrix is ​​a sparse matrix that works similarly to temporal coding: each position along the first dimension of the matrix corresponds to a pixel, and the value of each pixel is indicated by its position along the second dimension of the matrix. The first dimension of the matrix has a length mn equal to the product of the pixel array dimensions m and n of the image. The second dimension of the matrix has a length N corresponding to the range of possible pixel values ​​I(x,y) for each color channel. In other words, each row of matrix T represents the activity of each photoreceptor, such that the activity peak shifts left or right based on the pixel intensity. For simplicity, we only take a value of 1 at the activity peak and a value of 0 at the rest of the temporal response range.

[0089] We assume that LMC neurons communicate via LMC(j) = U T(i,j) is used to calculate the integral of all activities at each time unit j, where U=1 is a 1×mn dimensional unit vector. Then, LMC accumulates the hierarchical signal R. 归一化 = LMC Q (where Q is an N×N dimensional upper triangular matrix) is sent as the normalized signal to the next layer. Overall, the normalized image is obtained using the following formula:

[0090]

[0091] Here, the superscript ' indicates matrix transpose. 归一化 It is mn×1 dimensional and can be reshaped into an image of size m×n as the final output of this stage.

[0092] This form of normalization functions as a dynamic spatial filter, which produces lateral suppression among receptors exposed to high light intensity and spatial summation among receptors exposed to low light intensity. Since the matrix T is sparse, we can leverage this sparsity to improve computational cost.

[0093] Dynamic spatial filtering refers to a form of normalization of data in a color space, where the data represents the receptor's representation of the visual environment. Dynamic spatial filtering effectively simulates and modifies the responses of receptors exposed to high and low light intensities, where these receptors are affected by lateral inhibition from adjacent receptors and spatial summation, respectively.

[0094] Dual-antagonist model: The first stage of color-sensitive cells consists of single-antagonist cells in the LGN. These single-antagonist cells encode color information within their center-periphery receptive field (RF) in red-green, blue-yellow, and black-white antagonistic ways. If we consider the L channel to be red, M to be green, and S to be blue, we can create an antagonistic color model using three components in the LMS space: 1) achromatic = I + m + s, 2) yellow-blue - I + m - s', and 3) red-green - lm. For simplicity, instead of using a Gaussian function, we construct the RF as a square, called RF1. i (x,y),i=1:6, in Figure 1a Presented in and corresponding to Figure 1b Therefore, the response of a single antagonistic cell to image l(x,y) is calculated as r. i so (x,y) - I(x,y) RF i (x,y), where This represents convolution. These represent red-green, blue-yellow, and achromatic monoantagonist cells, respectively. In these expressions, the symbol "+" indicates excitation and inhibition. Following the physiological connections in V1 of the early visual cortex, we can construct V1 biantagonist cells using the outputs of two monoantagonist cells with different scales K. The RF. Therefore, the response of the dual-antagonistic cells can be calculated as:

[0095]

[0096]

[0097]

[0098] Where k controls the relative contribution of the RF periphery. We assume that the next layer in the visual cortex (V4) contributes to a single global illumination vector E - (e^(i+1)) based on the global color state of each color component L, M, and S. 1、 e2 and e3) are encoded. Therefore, we transform the output of the dual antagonistic cells to the LMS space as follows:

[0099]

[0100] The vector estimate E is estimated by the pooling function f(.) as follows:

[0101]

[0102] Here, f(.) signifies a typical neural computation of maximizing or summing each color channel across the entire image. Finally, the input image in LMS space is corrected by dividing it by the illumination vector E. For image representation purposes, the inverse matrix is ​​then used. Equation 2 transforms the corrected image to the sRGB space.

[0103] Data Acquisition: Different image datasets were acquired using an all-in-one Raspberry Pi camera (Pi camera) and an ODK camera, as described in the following sections. The acquired datasets were used to train and evaluate the model's performance in estimating global illumination of natural light from different artificial light sources.

[0104] Color Dataset: Small color chart data was acquired on a Raspberry Pi Camera V2 with the same 8-megapixel Sony IMX219 (8 MP camera sensor) image sensor as the ODK. This single camera setup features a fixed focal length (not a fisheye) lens. The camera was placed in front of a DataColor SpyderCheckr24 (24 color patches and gray card) facing a window for camera calibration under natural light. Some images were supplemented with yellow or white LED light to enhance the lighting variation.

[0105] Pi Camera: This image acquisition was repeated using a Pi camera under consistent lighting conditions. In this experiment, 10 preset Auto White Balance (AWB) options (Off, Auto, Daylight, Cloudy, Shade, Tungsten, Fluorescent, Incandescent, Flash, and Horizon) and 13 preset Exposure settings (Off, Auto, Night, Night Preview, Backlight, Spotlight, Sports, Snow, Beach, Long Exposure, Fixed FPS, Image Stabilization, and Fireworks) were sequentially changed to ensure each combination was achieved. The process was automated using Python scripts to control the Pi camera module, change settings, capture, and save images.

[0106] Environmental Dataset: The experimental setup was configured using an ODK with two built-in IMX219 cameras, as shown in the figure. Both boxes were modified to have color charts on their visible sides for color consistency correction. The included colors were: black, white, gray, red, yellow, green, blue, and pink (visible from at least one image). The ODK was configured to save raw camera images (non-4pi, etc.) to rosbag every three minutes. To capture variations in lighting conditions, the experiment was conducted near a window, allowing natural light to illuminate the scene. Additionally, office lights were turned on / off at random intervals. To add structured light and shadows to the scene, three lights were placed in the setup. These were smart bulbs, controlled via an app to turn on and off every 3 / 7 / 11 minutes, respectively. Two lights were set to yellow light, and one to white light, to increase scene variation. The lights were positioned to generate shadows, highlights, and lighting patches, thus constructing a complex scene. Due to the fisheye lens's field of view, a blank white sheet of paper was placed within the shared field of view of both cameras for subsequent camera correction analysis. By limiting image capture to once every three minutes, capturing ODK images at three-minute intervals, a ROS topic can be created that can be subscribed to and logged. Once logged, the images are replayed in RQt for inspection and saved as PNG images. The original CSV values ​​are also preserved.

[0107] It can be recognized that the image contains a standard data color chart ( Figure 7 ) or several colored boards placed on the wall ( Figure 9 Datacolor's Spyder Checkr 24 is a color chart containing multiple colored squares in a grid. Figure 7 The data color chart shown in the figure contains a series of spectral reflectance patches that represent the most likely intensity range suitable for many uniform lighting conditions. Because the color test chart contains a uniform color range, we can use this chart to estimate the illumination of images used at least in part by this invention.

[0108] Another aspect of the invention is a method that functionally utilizes a gray-edge assumption to estimate the local illumination within these regions of a partitioned input image. This gray-edge assumption presupposes that the average edge difference within the scene window is achromatic. The resulting local illumination estimate can be used to correct the illumination of the input image using retinex-based correction, thereby effectively reducing local illumination and achieving accurate color correction for the input image.

[0109] This method addresses the challenges posed by localized illumination from multiple light sources. By combining a selective attention mechanism and the gray-edge assumption, our proposed solution demonstrates promising results in mitigating the effects of localized illumination and improving color constancy across various scenarios, as can be seen from the results illustrated in the figures. Importantly, this method can be combined with any of the other models / methods described herein. Through combination, these methods / models / approaches of the present invention can address a wide range of application needs and effectively handle complex lighting conditions encountered in real-world environments.

[0110] In summary, this aspect of the invention (and in combination with other aspects) provides valuable insights into the field of color constancy and offers a robust methodology for enhancing the color accuracy of images affected by localized illumination. This method is committed to further improvement and may be extended to address additional lighting challenges and broaden its applicability across various fields of computer vision and robotics.

[0111] Localized lighting

[0112] To further improve the color constancy of images, this disclosure provides another image processing method that combines any and all of the methods described herein. This method addresses at least some of the aforementioned challenges, such as those arising from multiple light sources, by removing local (color) illumination from the input image.

[0113] In short, this method is proposed to address local illumination of an input image, and it can be combined with any of the methods described herein to further improve color constancy. For example, the output of the local illumination method can be used as an image in a first color space to remove global illumination from the original image.

[0114] Inspired by research on eye movement and selective attention in visual science, this method relies on three key components: I) selective attention mechanisms, II) the gray-edge hypothesis, and III) color normalization. These three components are integrated into a single algorithm / model to reduce local illumination. This combined algorithm represents a significant innovation from previous research, providing a novel and efficient process that yields the improved results shown in the figure.

[0115] Through experiments and the results shown, this invention demonstrates its ability to rival state-of-the-art methods, with the added advantage of superior computational efficiency. Most impressively, our model / algorithm achieves significant progress in mitigating the effects of localized illumination, realizing higher accuracy in estimating color constancy compared to existing methods. In fact, this represents a significant leap forward in the field of color constancy and opens up exciting new possibilities for an area that has remained largely unexplored until now.

[0116] Implementation details

[0117] With regard to this disclosure, it will be appreciated that the following other references (Brainard, David H. and Brian A. Wandell, “Analysis of the retinex theory of colorvision”, JOSA A, Vol. 3, No. 10 (1986): 1651-1661; and Land, Edwin H. and John J. McCann, “Lightness and retinex theory”, JOSA A, Vol. 61, No. 1 (1971): 1-11) are incorporated herein by reference.

[0118] Our proposed algorithm / model combines selective attention and the gray-edge assumption, and is applied to each preprocessed image. The model analyzes edge information within subsets of regions of interest (patches). We assume that the average reflectance difference within a patch is achromatic. Therefore, within these regions, edge information is extracted according to Retinex theory to estimate color constancy (see [link to Retinex]). Figure 15a Integrate the following components to produce the resulting image.

[0119] 1) Selective Attention: In this step, we consider a simple form of selective attention—"Region-based analysis"—which involves dividing the image into multiple regions (i.e., patches) and analyzing each region individually. The color constancy algorithm can focus on regions that are likely influenced by a single illuminator. This method assumes that each region is illuminated by a single color. To detect these regions (or subsets thereof), we segment the input image using the k-means clustering algorithm, an unsupervised machine learning algorithm designed to divide a dataset into a predetermined number (k) of distinct, non-overlapping clusters. In the context of image segmentation, k-means is used here to group similar pixels together based on their color or intensity values, making them more likely to be under the same illumination. In this process, subsets of regions are selected by considering color information (as a numerical representation describing the intensity or amplitude of different color channels) and the spatial location of the pixels (or their position in the image or coordinate system). This achieves better segmentation results. The algorithm assigns each pixel to one of the k clusters based on the similarity of the color values ​​of the k clusters. The number of tiles is merely a free parameter of the model, which can be adjusted based on the size of the input image. However, instead of using clustering algorithms (i.e., k-means), this mechanism can be improved through selective attention, such as saliency maps.

[0120] The integration of saliency-based selective attention mechanisms can be used to improve and enhance model performance. Unlike the assumption of sequentially analyzing gray edges across all regions, using saliency maps guides our attention to the most informative and visually salient regions in the image (referred to here as salient regions). By utilizing saliency information, we can prioritize regions that are more likely to contain reliable color information and are less affected by local illumination variations caused by multiple light sources. Using saliency maps enables more efficient and targeted analysis, potentially improving the accuracy, speed, and efficiency of this invention.

[0121] 2) Edge Detection: To extract the color of edges (referred to as colored edges in this paper), we can employ the Canny edge detection algorithm, a technique widely used in image processing for edge detection. This algorithm involves several steps to identify color variations within an image. First, the image is smoothed to reduce noise. Then, the gradient magnitude and orientation are calculated. Non-maximum suppression is applied to refine the edges, and thresholding is used to select the most robust edges. However, for testing purposes, we can explore alternative methods using simple gradient operators to estimate image gradients. A straightforward option is to utilize the Sobel operator, as follows:

[0122]

[0123]

[0124] By convolving the two operators on the grayscale image Ig obtained from the input image I, we obtain the image gradient G on the orientation G and G on G. x and G y , as follows: G x = S x I g and G y – S y l g Then, the gradient magnitude and orientation are calculated as follows:

[0125]

[0126] Here, G x and G y Let represent the gradient values ​​in the x and y directions obtained from the Sobel operator, respectively. By obtaining these estimates of the first-order image derivatives, we can calculate the gradient magnitude to capture the intensity of color changes at edges, using the dot product between the image and the image gradient magnitudes, as shown below:

[0127]

[0128] For example, see Figure 16 The intermediate sub-image shows the edge information from the image. This information is crucial for subsequent steps in our model / algorithm and facilitates accurate estimation and correction of color variations caused by multiple light sources.

[0129] 3) Retinex-based Correction: To correct the color in the region of interest, a modified version of the Retinex-based method is employed, utilizing color information extracted from the edges within these regions. Assuming that the edges in the selected tile under colorless illumination are achromatic (gray edge assumption), and that any difference between the edge color in the data and the "true" achromatic color is caused by reflectance and / or illumination, we correct the tile's color by considering the reflectance and illumination of the identified edges, using a Retinex model limited to these edges. Here, our aim is to reduce the impact of local illumination variations caused by a single light source within the region.

[0130] Let r e (x,y),g e (x,y) and b e (x, y) represent the three color channels of the edge within the region of interest. The corrected colors... (x,y) (x,y) and (x, y) is obtained through the following equation:

[0131]

[0132]

[0133] Max b - (b e (x,y)), Max g - (g e (x,y)) and Max r - (r e (x, y) is calculated as the maximum RGB value of the edge within the region. These maximum values ​​are then used to correct the red color of the pixels within the region. (x,y) and blue color (x,y). Using a method similar to that observed in insect vision, green is treated as the "default" pseudo-achromatic color and remains unchanged, i.e.

[0134]

[0135] Understandably, the current algorithm changes the color of edgeless tiles to gray, which is not ideal if we intend to recover the true color of objects. However, this limitation does not pose a significant problem in reducing light variability. Further efforts are needed to improve this limitation by defining scalable tiles in the attention mechanism to achieve maximum information content in each tile.

[0136] Another aspect of the present invention is a method for removing shadows from an input image. The method comprises two algorithms: a shadow detection algorithm and a shadow removal algorithm. These two algorithms work sequentially, wherein the first algorithm detects shadows and creates a shadow mask, and the second algorithm removes the detected shadows based on the shadow mask. The final output is a shadow-free image.

[0137] It is recognized that shadows in an input image can lead to reduced visibility, decreased contrast, and altered color distribution, which can affect the interpretability and aesthetic quality of the image and cause problems when the image is used in computer vision and robotics applications. In various applications where shadow detection and removal can be used (i.e., medical imaging or remote sensing), shadow removal can be an important step before any accurate analysis or measurement can be performed.

[0138] Shadow detection and removal can help reduce the number of false positives identified during the aforementioned applications. Shadow pixels can actually disrupt the way human color constancy is preserved. Shadow-free images have proven to be robust and suitable for further applications as input to illumination according to any of the aspects described herein in order to maintain human color constancy. Shadow detection and removal algorithms can be suitably performed on the output of the image processing procedure described herein.

[0139] Image without shadow

[0140] To further improve image quality, this invention may include a shadow detection and removal process. The result of this process is a shadow-free image. One advantage of removing shadows or having a shadow-free image is enhanced visibility of the input image. Simply put, shadows obscure details in an image, making it difficult to distinguish objects or features. By removing shadows, important details in the image become clearer and more visible, thereby improving image quality.

[0141] Furthermore, removing shadows from the input image improves image contrast. For example, shadow pixels in an image often cause reduced contrast between different parts of the image. Shadow removal helps restore a more balanced contrast, making objects stand out relative to their background. Shadows also introduce color variations and shifts due to changes in lighting conditions. Shadow removal makes the color representation more consistent across the entire image, preserving the image's color constancy or its natural appearance. For example, shadow removal can produce a more natural and evenly lit image that closely resembles how a scene would look under uniform lighting conditions.

[0142] In summary, shadow-free images offer significant advantages, making them suitable for a wide range of applications. They can be combined with other algorithms described herein to establish correct color constancy, thereby enhancing the quality and usability of raw images from one or more cameras, eliminating the negative impact of shadows, and resulting in improved visibility, contrast, color consistency, and aesthetics. The following steps can be performed to obtain shadow-free images. It should be understood that the algorithm is not limited to these steps and may encompass other steps described herein.

[0143] In one aspect of the shadow detection algorithm: an RGB image is used as input. The RGB image is converted to the Lab color space, and the L channel is extracted. A histogram of the L channel is generated and smoothed. Local minima on the smoothed histogram are identified. A threshold is determined based on the distance to the minima. A shadow mask is created by selecting pixels with L values ​​less than the threshold. The shadow mask is output, at least according to... Figure 20 As shown.

[0144] In one aspect of the shadow removal algorithm: an RGB image and a shadow mask are used as inputs to the algorithm. The RGB image is converted to the HSV color space and segmented into shadow and highlight regions. The image is segmented using the k-means algorithm. Segments are labeled based on region intersections. Texture features are used to calculate the distances between segments. The labeled segments are matched using these distances. Histogram matching is performed on the channels of the labeled segments. The adjusted shadow segments are merged to obtain corrected segments. The corrected segments are converted back to the RGB color space. The process is iterated over all shadow segments, and then all segments are merged.

[0145] Output an image without shadows.

[0146] Implementation details

[0147] In this paper, we propose a shadow detection and removal algorithm. The main objective of this method is to accurately identify shadow-prone areas in an image. Shadows typically appear darker and have a different color compared to the rest of the scene. By comprehensively analyzing the intensity variations of illuminated and shadowed areas within the LAB color space, we derive a simple thresholding method to identify shadow-affected pixels. This threshold-based method enables efficient detection without requiring extensive training.

[0148] After identifying shadows, our model / algorithm considers texture and image structure, adjusting the color information of shadowed areas to closely match illuminated areas, thereby achieving shadow removal. To determine the nearest illuminated area, we introduce a metric based on texture patterns to quantify region similarity. To improve metric accuracy, both shadowed and illuminated areas are segmented into smaller units, assuming each unit presents a single texture pattern.

[0149] The similarity assessment between segments incorporates various visual features, such as texture density, frequency, entropy, and inter-segment centroid distance. We employ histogram matching to align the pixel colors in shadow segments with the nearest lit segments, assuming that the visual features of the shadow and corresponding lit segments are comparable. The purpose of histogram matching is to eliminate the influence of shadows on the initial underlying texture. Processing all shadow segments in this manner yields a shadow-free image.

[0150] Integrating shadow detection and removal methods into a visual location recognition and localization framework promises to enhance robot navigation. By supporting visual feature consistency and mitigating the effects of shadows, these methods can improve the accuracy and reliability of navigation systems. Therefore, even when facing challenging lighting conditions, robots can confidently identify familiar locations and accurately estimate their positions.

[0151] Our proposed algorithm / model comprises at least two algorithms: a) a shadow detection algorithm and b) a shadow removal algorithm. These two algorithms work sequentially, with the first algorithm (steps 1 to 9) detecting shadows from the input RGB image and outputting a shadow mask. The shadow removal algorithm (steps 1 to 11) removes the detected shadows based on this shadow mask. The final output of both algorithms is a shadow-free image of the initial image. Exemplary steps of these two algorithms are described below:

[0152] In another aspect of the shadow detection algorithm, I) the shadow detection algorithm performs the following steps:

[0153] 1. Input: RGB image I

[0154] 2. Extract the dimensions (h, w) of the image.

[0155] 3. Convert the I color space to Lab color space and extract channels L, a, and b.

[0156] 4. Generate a histogram for channel L.

[0157] 5. Smooth the histogram curve using a Gaussian window F.

[0158] 6. Identify local minima on a smoothed histogram.

[0159] 7. If the distance between the index of the first minimum value and the index of the second minimum value is less than p, then the index of the first minimum value is selected as the threshold T; otherwise, the index of the second minimum value is selected as the threshold T (the parameter p is selected based on the image size).

[0160] 8. Generate a mask S by selecting pixels in image I whose L channel values ​​are less than a threshold T. I .

[0161] 9. Output: Shadow mask S I

[0162] In another aspect of the shadow removal algorithm, II) the shadow removal algorithm performs the following steps:

[0163] 1. Input: RGB image I and shadow mask S I

[0164] 2. Convert I to HSV color space and extract channels h, s, and v.

[0165] 3. Divide the RGB value of I into two parts: the shaded area I. s Heguang District I L

[0166] 4. Segment image I using the k-means algorithm:

[0167] a. Flatten the image: Convert a 2D RGB image into a 1D array of RGB tuples.

[0168] b. Choose the number of clusters (k)

[0169] c. Initialize cluster centers: Select from the RGB values ​​present in image I using a uniform distribution.

[0170] d. Assign pixels to clusters i = 1:k

[0171] e. Calculate the distance between the value of each pixel and the cluster center.

[0172] f. Assign each pixel to the nearest cluster center.

[0173] g. Update cluster centers

[0174] h. Duplicate distribution and central update (dg)

[0175] I. Check convergence

[0176] j. Perform post-processing to correct the allocation based on the neighborhood.

[0177] k. Add a 1D assignment array to the flattened image (a)

[0178] l. Divide image I into images Where for any i, j satisfy

[0179] 5. Assign the label "Shadow" (S) or "Light" (L) to each segment:

[0180] for

[0181] for

[0182] 6. Based on image texture features, calculate the distance between segments—the distance between the i-th segment and the j-th segment. :

[0183] a. Edge distance Calculate the distance between the edges extracted from each segment.

[0184] b. Color distance Calculate the average color distance (in Lab color space) between color channels a and b (of segment i and segment j).

[0185] c. Entropy distance Calculate the entropy distance between the i-th segment and the j-th segment.

[0186] e. Neighborhood distance : Pixel distance between the centers of the i-th segment and the j-th segment

[0187] f. Define distance , making

[0188] 7. For each shaded segment Find the section placed in the shade. minimum distance The light-receiving section at the location

[0189] 8. Perform histogram matching:

[0190] a. Shaded area It is divided into its color channels h, s, and v, and the corresponding histograms are obtained for each. , and

[0191] b. Shaded sections It is divided into its color channels h, s, and v, and the corresponding histograms are obtained for each. , and

[0192] c. Obtain the cumulative distribution function of all histograms obtained in a and b. For the shaded segment And for the light-receiving segment as and

[0193] d. will and Match to get ,Will and Match to get and will and Match to get If

[0194] e. Merge shaded segments All recovery components , and In order to obtain

[0195] f. will Convert to RGB color space.

[0196] 9. Iteratively perform steps 7 and 8 for all shaded segments.

[0197] 10. All color-corrected segments and light-receiving segment Merge into images

[0198] 11. Output: Image without shadows

[0199] The specific steps outlined in the shadow removal and detection algorithms described above, as well as other algorithms described herein, can vary depending on the environment or context in which the algorithm is used. Therefore, it should be recognized that various aspects are described herein with respect to the invention, some of which may not be limited to the described sequential steps, and certain steps may vary based on a specific application.

[0200] This disclosure delves into the criticality of handling lighting and shadows in the field of visual location recognition and localization for robot navigation. Our ongoing efforts involve implementing adaptive algorithms designed to handle lighting and shadows. The following figures further provide exemplary embodiments of the invention corresponding to the various aspects described herein.

[0201] Figure 1a This is a flowchart illustrating an example process 100 for achieving color constancy according to the present invention. Process 100 illustrates an algorithm for achieving stable color appearance and high color discrimination, inspired by the visual systems of humans and major insects and tailored to the functional characteristics of ODK cameras. The figure includes a flowchart showing the details of process 100 in estimating multi-illumination scenarios that integrate three mechanisms / models: a) color adaptation, b) LMC normalization model, and c) dual antagonistic model.

[0202] Figure 1b The diagram corresponds to Figure 1a The flowchart illustrates an example process 150 of the algorithm proposed in this disclosure. This flowchart depicts an image processing method based on a model architecture according to various aspects of this disclosure, which removes illumination from an image based on color constancy. The image processing method may include steps 151 to 165.

[0203] In step 151, the method obtains an image in a first color space. This first color space can be the RGB color space. In step 153, the method converts the image into data in a second color space. This second color space is the LSM color space.

[0204] In step 155, the method uses color adaptation to transform the data, wherein using color adaptation to transform the data further includes: estimating the illuminator of the data using a grayscale world model; and using the illuminator to transform the data into a target illuminator. This can be achieved by applying one or more matrix transformations to the target illuminator based on one or more cameras used to acquire the image.

[0205] In step 157, the method performs a first normalization on the transformed data, wherein the first normalization includes applying dynamic spatial filtering techniques to adjust the transformed data based on light intensity.

[0206] In step 159, the method applies a set of filters to normalized data, wherein the set of filters is convolved based on the normalized data associated with the image; the set of filters represents two layers of the vision system.

[0207] The filter set may include at least one center-periphery structure, which represents an RF encoding color antagonism. The RF may be red-green antagonism, blue-yellow antagonism, and achromatic antagonism.

[0208] The filter set may also include at least two filters positioned in series, such that one of the at least two filters receives input from the other filter.

[0209] In step 161, the method performs a second normalization on the filtered data to obtain an illumination estimate of the image associated with the filtered data. In step 163, the method outputs normalized data from the second normalization, wherein the normalized data maintains color constancy based on the illumination estimate, thereby removing illumination from the normalized data.

[0210] Optionally, in step 165, a standard method can be used to convert RGB to the LMS color space, wherein the method converts normalized data from the second normalization into an image in the first color space.

[0211] according to Figure 1a and Figure 1b When performing the first normalization, the method can also identify a data representation of the receptor response associated with the transformed data. This data representation will include a uniform representation of the receptor response across the entire image.

[0212] This data representation can also be applied to the transformed data, wherein the method integrates the data representation over a time period / frame and uses dynamic spatial filtering to normalize the transformed data based on the integrated data representation, which represents performing time-coded adjustments for each sensor for light intensity.

[0213] The second normalization may also include the step of correcting the filtered data using a light vector generated by a pooling function, Equation 7, wherein the filtered data can be divided by the light vector.

[0214] The pooling function can include:

[0215]

[0216] Where f (x,y) (.) represents the data representation for which the maximum typical neural computation is performed on the filtered data, and The output of the double antagonistic filter in the LMS color space is shown.

[0217] The method can receive a raw image from one or more cameras, and obtains the input to step 151 by removing gamma correction from the raw image, wherein the method converts the raw image into the image in a first color space.

[0218] Figure 2 This is a schematic diagram illustrating the application of the Spyder Checker model under different lighting conditions according to various aspects of this disclosure. Various examples related to model application are included, and results are shown. These are examples of corrected images obtained by applying the model's components to a sample color chart. As shown, the color condition of the Spyder Checker in the corrected image on the right is almost constant. This indicates that the variability of the color patches at the Spyder Checker on the left is significantly reduced.

[0219] The left column 201 of this figure represents instances of the initial images of the color chart under varying lighting conditions, with images of various lighting conditions shown in the corresponding rows. The different components of this model and the mixing / combination model (last column 209) exclude global illumination from these images. The middle column is based on... Figure 1a and Figure 1b Intermediate results for color adaptation (column 203), LMC normalized model (column 305), and dual antagonistic model (column 407). Color adaptation 203, LMC 205, and antagonistic model 207 exhibit synergistic effects when combined, as shown in combined model 209.

[0220] Specifically, sample images (left column 201) were captured using a Raspberry Pi V2 camera under different lighting conditions. The last column 209 represents the corrected image after global illumination was excluded using this combined model. Different components of this model and the hybrid model excluded global illumination from these images. Because the color morphology of the Spyder Checker on the corrected image on the right side of the figure is almost constant, this indicates that the variability of the color patches at the Spyder Checker on the left side is significantly reduced (compare the color patches of the Spyder Checker in the left and right columns). Therefore, this combined model produces a color chart with almost constant color patches. The same results were obtained from datasets captured by the front and rear cameras of the ODK system (see environmental dataset). Figure 9 and Figure 10 See simulation environment. Figure 14 ).

[0221] To quantitatively analyze the variability of the Spyder Checker intensity, the average pixel values ​​of five color patches (red, green, blue, yellow, and white) were studied before and after applying the model. A violin plot visualizes the probability density of each selected color patch (R, G, B) before and after the model. Figure 3 Due to the extreme variability of light, the colors in the initial image exhibit a stretched distribution compared to the corrected image. It can be seen that after applying this model, the distribution of corresponding RGB pixels narrows, and the color values ​​of the color patches increase. However, using the Fano factor to measure the relative variance of color (e.g., ...) Figure 4 (As shown).

[0222] Figure 3 This is a schematic diagram comparing the pixel intensity distribution between an initial image and a corrected image according to various aspects of this disclosure. Figure 2 The results of the model application are quantified and plotted. The plot compares the pixel intensity distribution between the initial image and the corrected image. Each row shows the intensity distribution of the red (left 301), green (middle 303), and blue (right 305) color channels of the same selected color patch from the labeled initial image (blue on the left) and the labeled corrected image (orange on the right). As shown in the figure, the color distribution of the blue 311, green 313, red 315, yellow 317, and white 319 color patches is sorted from top to bottom.

[0223] Figure 4 This is a schematic diagram of the Fano factor (relative variance) of the selected color patch according to various aspects of this disclosure. The bar chart shows the Fano factor of the selected color patch, which represents the normalized variability of each color channel of the five selected patches 411 to 419 of the initial image (blue bar) and the model-corrected image (orange bar) (red: left 401, green: middle 403, and blue: right 405). This indicates that the model reduces the variability of most color channels of the selected color patch.

[0224] Specifically, in this figure, the Fano factors for the (R, G, B) values ​​of the five color patches 411 to 419 appear smaller in the corrected image because this indicates that by applying this model to the image, the variability of the same pixel is reduced for all color patches in the image. Second dataset (Set 2, see...) Figure 11 Similar results were obtained.

[0225] Figure 5 It is based on the environmental datasets of various aspects of this disclosure (and) Figure 9 A schematic diagram of the color distances between color swatches in the instance-related set 1). Figure 5In the first row 501, an instance of the initial image 501a captured by the ODK camera and its corrected image 501b after model processing is shown. The initial image and the corrected image in the second row show their respective distances.

[0226] The matrix in the second row, 503, shows the color distances between the eight color swatches represented in the first row (left: initial image 503a, right: corrected image 503b). The color distance between each pair of color swatches is shown by the elements of the matrix.

[0227] Lines 3 (505) and 4 (507) show the results according to... Figure 12 The average color distance (mean value of the color distance matrix) of the eight color patches in all images in the dataset. The box plot in the last row 509 (i.e., the summary of the third and fourth rows) reveals that the model that forms the basis of this invention increases the color distance between color patches.

[0228] As explained in the previous section, ΔE is used to measure the color distance between color patches in the initial and corrected images. As shown in the figure, these matrices display the distances between all pairs of color patches. In this figure, the average color distance between selected color patches (i.e., the average value of the color distance matrices) is plotted as a function of the image indices. The average color distance between color patches in the corrected image is nearly twice as large as the color distance in the initial image, a result that can be interpreted using a box plot as shown in the figure. Using, for example, in... Figure 12 The dataset illustrated in the figure yielded similar results. This demonstrates that the present invention improves the separation between colors and potentially enhances the color discriminability of the ODK system.

[0229] Figure 6 This is a schematic diagram illustrating an example of a model application for image matching between two ODK cameras according to various aspects of this disclosure. In this figure, the top row 601 shows images captured by the front and rear IMX219 cameras established in the ODK system. The bottom row 603 shows the image after correction by the present invention. Using a target image incorporating Spyder Checker, the color distribution of the two images is matched to the color distribution of the target image.

[0230] In addition to improving color discriminability, this invention also enables color distribution matching between the front and rear ODK cameras. By doing so (performing color distribution matching), this invention helps address the challenge of overlap when generating cylindrical images from two ODK cameras by reducing the color distance between overlapping areas of the front and rear images. This is shown in / from the second row, where the colors of both the front and rear images are matched to the colors of the template image. However, the model remains sensitive to the quality of the template image. This template image can be captured from a Spyder Checker or Macbeth ColorChecker color chart under (full) white light, such as... Figure 7 As shown.

[0231] Figure 7 This is a schematic diagram of a Spyder Checker or Macbeth ColorChecker color chart based on various aspects of this disclosure. The Spyder Checker is a color chart from Datacolor Corporation containing a series of spectral reflectance patches representing the (most) possible intensity range suitable for many uniform lighting conditions. Models are applied under different lighting conditions to... Figure 3 It is provided in the form of a Spyder Checker color chart.

[0232] Figure 8 It is used to obtain the same as Figure 1 to Figure 7 A schematic diagram of an exemplary data acquisition setup for the relevant model results is provided. The positions of the ODK camera 805, colored objects 801, 807a / b, windows (natural light) 803a / b, and lamps (artificial light) 809a / b / c are shown in this setup. The setup includes a radiator object 801 between two windows 803a / b, with natural light directed towards multiple box objects 807a / b. Multiple lamps 809a / b / c are provided as artificial light. The distances between these objects are shown accordingly. The ODK camera is located at the center of the setup.

[0233] Figure 9 and Figure 10 This is a schematic diagram of example images from environmental datasets (Set 1 and Set 2) based on various aspects of this disclosure. In each image, five samples of images 901a, 903a, 1001a, and 1003a were captured from a color palette using an ODK camera under different lighting conditions (top: front camera 901, 1001; bottom: rear camera 903, 1003) (see About Figure 2 (Design of environmental data in SpyderCheckr). The second row, 901b, 903b, 1001b, and 1003b, shows the corresponding images corrected by applying the proposed / combined model.

[0234] Figure 11 This is a diagram comparing the pixel intensity distribution between an initial image captured using a Pi camera and a corrected image according to various aspects of this disclosure. In this diagram, each column shows the intensity distribution of the red (left 1101), green (middle 1103), and blue (right 1105) color channels for the same selected color patch from the initial image (blue, captured by the Pi camera) and the corrected image (orange). The color distributions of the red 1111, green 1113, blue 1115, yellow 1117, and white 1119 color patches are ordered from top to bottom row.

[0235] Figure 12 Based on all aspects of this disclosure and regarding Figure 5 A schematic diagram of the color distances between color palettes in the environmental dataset (Set 2). Figure 12 In the diagram, the first row (1201) shows images captured by the ODK camera and instances of them corrected by the model. The matrix in the second row (1203) shows the color distances between the eight color patches represented in the first row (1201) (left: initial image, right: corrected image). The color distance between each pair of color patches is shown by the elements of the matrix. The third row (1205) and the fourth row (1207) show the average color distances (mean values ​​of the color distance matrices) of the eight color patches for all images in the dataset (set 2). The box plot in the last column (i.e., the summarized third and fourth columns) reveals the increase in color distances between the color patches by the model.

[0236] Figure 13 This is a block diagram illustrating an example computing device / system 1300, which can be used to implement one or more aspects, combinations thereof, modifications thereof, and / or as referenced. Figures 1a to 12 as well as Figures 14 to 21 The content and / or aspects described herein. The computing device / system 1300 includes one or more processor units 1302, input / output units 1304, communication units / interfaces 1306, and memory units 1308, wherein the one or more processor units 1302 are connected to the input / output units 1304, communication units / interfaces 1306, and memory units 1308. In some embodiments, the computing device / system 1300 may be a server, or one or more servers networked together. In some embodiments, the computing device / system 1300 may be a computer or a supercomputer / processing facility or one or more aspects, combinations thereof, modifications thereof, and / or as described herein, suitable for processing or performing the system, apparatus, method, and / or process. Figures 1a to 12 as well as Figures 14 to 21The content described herein and / or the hardware / software as described in the various aspects herein. The communication interface 1306 can connect the computing device / system 1300 to one or more services, devices, server systems, cloud-based platforms, or systems for implementing topic databases and / or knowledge graphs via a communication network to implement the invention as described herein. The memory unit 1308 can store one or more program instructions, code, or components, such as (by way of example only, but not limited to) and referenced... Figures 1a to 12 as well as Figures 14 to 21 The operating system and / or code / components, additional data, applications, application firmware / software, and / or additional program instructions, code, and / or components associated with implementing the function, and / or one or more functions associated with one or more methods and / or processes of the device, service, and / or server hosting the process / method / system; apparatus, mechanisms, and / or systems / platforms / architectures, combinations thereof, modifications thereof, and / or as described herein, for implementing the invention as described herein. Figures 1a to 12 as well as Figures 14 to 21 At least one of the contents described herein.

[0237] Figure 14 Based on all aspects of this disclosure and regarding Figure 2 This diagram illustrates the application of the model to the simulated environment under different lighting conditions. Here, sample images (initial image 1401) shown in the left column are selected from the simulated environment at different brightness levels [10, 30, 100, 500, 1000] displayed in six rows from top to bottom. The different components of models 1403, 1405, 1407, and the combined model (last column 1409) exclude the global illumination of these images. It can be observed that the color condition of the corrected image on the right is almost constant.

[0238] Figure 15a This is a flowchart of another model architecture for image processing concerning color constancy, particularly for preserving the local illumination of the input image. The diagram illustrates process 1500 of the proposed model. In process 1500, a patch 1503 is selected from the input image 1501 via a sequential mechanism. The colors of the edges within patch 1503 are calculated using Canney edge detection 1505. Colored edges 1507 are extracted. The colors of these edges are then used to estimate the local illumination 1509, and the colors of the selected patches are finally corrected by following Retinex theory. The output is the illumination-corrected image 1513. The colors of all patches are corrected through a sequential scanning process.

[0239] Figure 15b Based on Figure 15aThe diagram shows a flowchart of image processing for another model architecture. This diagram illustrates a method 1550 for reducing the influence of illumination on an image. The method includes at least the following steps.

[0240] In step 1551, the method receives an input image. This input image may be a raw image obtained from one or more cameras. It may also be an image in a color space as described herein.

[0241] In step 1553, the input image is divided into multiple regions. During this process, the input image is segmented using a k-means clustering algorithm based on the similarity of the color or intensity values ​​of the pixels from the input image, resulting in the creation of multiple regions.

[0242] In step 1555, the multiple regions are analyzed based on the color information and spatial location of the pixels in each region. The analysis is performed by identifying salient regions from the multiple regions using a clustering algorithm; a cluster map is generated based on the identified salient regions; and the salient regions are analyzed to select a subset of regions affected by the irradiation.

[0243] For example, a saliency map covering multiple regions can be applied. This saliency map is used to identify salient regions from these regions, constraining them into subsets of areas. These salient regions are further analyzed to select a subset of regions affected by the irradiation.

[0244] In step 1557, a subset of the regions is selected from the plurality of regions. This selection is influenced by the irradiation body based on the analysis in the previous step.

[0245] In step 1559, colored edges are identified for at least a subset of the region. This can be achieved using edge detection, particularly using the Canny edge detection algorithm configured to select multiple edges based on threshold intensity. For improved efficiency, the Sobel operator can also be used during edge detection.

[0246] In step 1561, color information is extracted from the colored edges;

[0247] In step 1563, the extracted color information is used to decompose the reflection and illumination components of the input image, which are two main factors contributing to the appearance of the object's color in the image. The reflection component, also known as surface reflectance or albedo, represents the inherent color and material properties of the object's surface, while the illumination component refers to the lighting conditions under which the object is observed, as described in the previous section. In this step, these components are decomposed.

[0248] In step 1565, the illumination of the input image is corrected based on the decomposed reflection and illumination components.

[0249] In step 1567, the method outputs an image after illumination correction.

[0250] Optionally or additionally, the method may include: identifying regions with colored edges from at least one subset of regions of the input image; and correcting the illumination of the input image based on the identified regions with colored edges.

[0251] Figure 16 This is a schematic diagram of an exemplary output of image processing based on the other model architecture. The top subfigure 1601 shows the initial image, while the middle subfigure 1603 shows the colored edges extracted during the process of this model. The last subfigure 1605 represents the final output after applying a Retinex-based correction process.

[0252] Figure 17 These are schematic diagrams of violin plots 1700, which illustrate the color distance (denoted by ΔE) between images in each dataset used for image processing. ΔE is a measure of the similarity or difference between images before and after the correction process. The larger the ΔE, the greater the distance between colors. Here, we evaluate the distance in the CIELab color space and represent the relative perceived magnitude of the color difference.

[0253] This figure shows five different violin plots (1701, 1703, 1705, 1707, and 1709) corresponding to datasets 1 through 5. Performance is evaluated using these five datasets shown in the violin plots. Two lab datasets are identical to those used in the previous global color constancy model. Three new lab datasets have been added, covering richer local color constancy. In the new datasets, different regions of the image are represented by data from... Figure 18 The example images show different artificial lighting.

[0254] Each violin plot depicts the color distance (ΔE) between images in each dataset. The left, middle, and right subplots show the ΔE measurements between the initial image (left), the image corrected using a global color constancy model (middle), and the image corrected using both local and global color constancy models (right), respectively. It is assumed that the images in each dataset were captured from the same location under different lighting conditions, including both global and local illumination.

[0255] In summary, the effectiveness of the model was evaluated by measuring the color distance between the images before and after model implementation using the ΔE metric. This analysis was conducted for both local and global color constancy models as described in this paper, and the results are plotted in the figures. Here, a lower ΔE value indicates better color constancy performance while reducing the difference between predicted and true pixel values. The violin plot on the left shows the probability distribution of ΔE values ​​between the initial images, illustrating the inherent variability of color appearance. The two violin plots below show the color distance between the corrected images. The middle plot corresponds to the application of the global color constancy model, while the left plot represents the results obtained after applying the local color constancy model.

[0256] As clearly shown in the figure, when processing images, the local color constancy model complements the global color constancy model to establish a stable color appearance and high color discrimination. The combination of these models significantly improves performance, resulting in even lower variability across all five datasets. These findings highlight the effectiveness of the local color constancy model in achieving more consistent color representations under different lighting conditions.

[0257] Figure 18 This is a schematic diagram illustrating the application of the model to environments under different lighting conditions. The left subplot 1801 shows example inputs and outputs of the color constancy model. The top subplot 1801a shows the initial image randomly selected from the dataset as input. The middle subplot 1801b and the bottom subplot 1801c depict the corrected images obtained after applying the global color constancy model and the local color constancy model, respectively.

[0258] Right side 1803 shows example images captured by ODK under different lighting conditions. These example images provide a visual representation of the various lighting conditions present in any exemplary dataset, which can be used to generate the output shown in left side 1801 (i.e., bottom subplot 1803c). These instances (top 1801a, middle 1801b, and bottom 1801c subplots) demonstrate the variability and complexity of lighting environments. The dataset thus consists of images captured under a range of local lighting conditions. This dataset was carefully selected to include environments with multiple light sources, ensuring our model / algorithm was tested in realistic and challenging lighting scenarios.

[0259] Figure 19aThis is a flowchart of another model architecture for image processing. As shown, this model architecture has two steps: a first step (top) for detecting shadows and a second step (bottom) for removing the detected shadows. The first step detects shadows by analyzing an initial RGB image 1901 converted in the LAB color space 1903, thereby generating a histogram. Then, a threshold-based method 1905, as described herein, is applied to filter the converted image in the LAB color space in conjunction with the histogram. Shadow pixels 1907 are selected to generate a shadow area mask 1909 as output.

[0260] In the second step, the image is first segmented into multiple regions using the k-means algorithm. The yellow curve in the figure indicates the boundary of each segment within the image. As shown, the segmented image 1911 is divided into light regions (here referred to as the light segment region) and shadow regions (here referred to as the shadow segment region). Considering the texture features of the input image as described herein, the target shadow segment 1913b from the first step is removed (based on the shadow region mask 1909) by adjusting the color to match the light region. Yellow pixels represent pixels assigned to the shadow region by the shadow detection algorithm. The similarity metric D identifies the nearest light region 1913a to the target shadow segment 1913b based on the texture features / pattern of the target shadow segment 1913b. Histogram matching 1913b aligns the color of the shadow segment with the color of the corresponding lit segment. Subsequently, after histogram matching between the shadow segment and its nearest lit segment, the correct shadow segment is generated, and the correct shadow segments are merged 1917 to form the output—a shadowless image 1919 of the initial RGB image.

[0261] In applications such as robot navigation, the use of shadowless images does indeed improve accuracy and reliability by enhancing visual consistency. For example, in the context of robot navigation, using shadowless images as input to local color constancy models and / or global color constancy models as described herein, or applying shadowless images to their output, allows the robot to confidently identify familiar locations and achieve accurate position estimation, even under challenging lighting conditions.

[0262] Figure 19b It is a flowchart of image processing based on the other model architecture according to various aspects of this disclosure;

[0263] Figure 20This is a schematic diagram of exemplary input and output for image processing based on the other model architecture. In this diagram, initial image 2001 shows an example input image, which is transformed 2003 into a histogram in the "lab" color space and then smoothed. Intermediate histogram 2005 shows the result of the smoothed image of the L channel (p < 10). Here, the smoothed histogram is generated from the "L" channel of the image converted to the "Lab" color space. The solid red lines represent the thresholds marking the shadow and highlight pixels.

[0264] According to the histogram, the first minimum value 2005a is between 0 and 5 (indicated by the red dashed line), while the threshold T2005b is 10. Following this, pixel selection 2007 occurs, generating a shadow mask 2009a. An image with the shadow mask S 2009a is shown and can be compared to the true shadow 2009b. Finally, a version of the image after shadow removal 2011 is also shown in the figure.

[0265] Figure 21 This is a schematic diagram illustrating exemplary inputs and outputs for image processing of different input images. The diagram shows an initial image (top left 2101) and a resulting image (bottom left 2103) obtained from the steps of applying shadow detection to the initial image and subsequently removing shadows from that image. The diagram also shows a real image (top right 2105) and an image with a yellow shadow mask (bottom right 2107).

[0266] Regarding the foregoing, the present invention is further described herein as one or more of the following aspects and options. These aspects and options are based on Figures 1 to... Figure 21 Any of these aspects and certain options may be appropriately disclosed. These aspects and certain options may be combined with any other aspects and features described herein, as will be apparent to those skilled in the art of robotics and machine vision.

[0267] In one aspect, there is a method for processing an image based on color constancy to remove illumination from the image, the method comprising: obtaining an image in a first color space; converting the image to data in a second color space; transforming the data using color adaptation; performing a first normalization on the transformed data, wherein the first normalization includes applying a dynamic spatial filtering technique to adjust the transformed data based on light intensity; applying a set of filters to the normalized data, wherein the set of filters is convolved based on the normalized data associated with the image; performing a second normalization on the filtered data to obtain an illumination estimate of the image associated with the filtered data; and outputting normalized data from the second normalization, wherein the normalized data maintains color constancy based on the illumination estimate, thereby removing illumination from the normalized data.

[0268] In another aspect, a method for reducing the influence of illumination on an image is provided, the method comprising: receiving an input image; dividing the input image into multiple regions; analyzing the multiple regions based on color information and spatial location of pixels in each region; selecting a subset of regions affected by an illuminator from the multiple regions based on the analysis; identifying colored edges for at least the subset of regions; extracting color information from the colored edges; using the extracted color information to decompose the reflection component and illumination component of the input image; correcting the illumination of the input image based on the decomposed reflection component and illumination component; and outputting an illumination-corrected image.

[0269] In another aspect, there is a method for providing a shadowless image, the method comprising: receiving an input image in a first color space, wherein the input image includes at least one shadow region; generating a shadow region mask for the at least one shadow region; removing shadows from the shadow region of the input image based on the shadow region mask; and outputting a shadowless image.

[0270] In another aspect, there is an apparatus for processing an image to maintain the color constancy of the image, the apparatus comprising: at least one model configured to perform the steps according to any of the aspects described herein.

[0271] In another aspect, there is a system for processing an image to establish the color constancy of the image by removing illumination from the image, the system comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to perform any of the aspects described herein.

[0272] On the other hand, there is a non-transitory computer medium on which computer program instructions are stored, and a computer-implemented method according to any of the aspects described herein.

[0273] As an option, the normalized data from the second normalization is converted into an image in a first color space. As another option, performing the first normalization further includes: identifying a data representation of the receptor response associated with the transformed data, wherein the data representation includes a uniform representation of the receptor response on the image; and applying the data representation to the transformed data. As another option, applying the data representation further includes: integrating the data representation over a time period; and normalizing the transformed data based on the integrated data representation using a dynamic spatial filtering technique, which represents a time-coded adjustment for each receptor for light intensity. As another option, the second normalization includes correcting the filtered data using a light vector generated using a pooling function, wherein the filtered data is divisible by the light vector. As another option, the pooling function includes:

[0274]

[0275] Where f (x,y) (.) represents the data representation for which the maximum typical neural computation is performed on the filtered data, and The output of a dual-antagonistic filter in a second color space is shown. Alternatively, the method includes: receiving a raw image from one or more cameras; removing gamma correction from the raw image; and converting the raw image to the image in a first color space. Alternatively, the first color space is a red-green-blue color space. Alternatively, the second color space is a long-medium-short color space. Alternatively, the use of color adaptation to transform the data further includes: estimating the illuminator of the data using a gray-world model; and converting the data to a target illuminator using the illuminator. Alternatively, the conversion of the data to a target illuminator further includes: applying one or more matrix transformations to the target illuminator according to one or more cameras used to acquire the image. Alternatively, the set of filters includes at least one center-periphery structure representing a receptive field (RF) encoding color antagonism. Alternatively, the RF includes red-green antagonism, blue-yellow antagonism, and achromatic antagonism. Alternatively, the set of filters includes at least two filters cascaded, such that one of the at least two filters receives input from the other filter. As an alternative, the set of filters represents two layers of the vision system. As another alternative, the analysis of the plurality of regions further includes: identifying salient regions from the plurality of regions using a clustering algorithm; generating a cluster map based on the identified salient regions; and analyzing the salient regions to select a subset of regions affected by the illuminator. As another alternative, the cluster map constrains the plurality of regions into a set number of partitions. As another alternative, dividing the input image into the plurality of regions further includes: segmenting the input image using a k-means clustering algorithm based on the similarity of color or intensity values ​​of pixels from the input image. As another alternative, identifying colored edges for at least the subset of regions further includes: performing edge detection using a Canny edge detection algorithm configured to select multiple edges based on a threshold intensity. As another alternative, identifying colored edges for at least the subset of regions further includes: performing edge detection using a Sobel operator. As another alternative, it further includes: identifying regions with colored edges from the at least one subset of the regions of the input image; and correcting the illumination of the input image based on the identified regions with colored edges.

[0276] As an alternative, removing shadows from shadow areas of the input image based on a shadow area mask further includes: converting the input image to data in a third color space; dividing the input image data into multiple regions; segmenting the multiple regions into shadow segments and light segments according to the shadow area mask; determining the distance between shadow segments and light segments based on texture features of the input image; pairing shadow segments with light segments based on the determined distance; performing histogram matching on the paired shadow segments and light segments to produce color-adjusted segments; iteratively performing the pairing and histogram matching until each shadow segment has been matched; merging the color-adjusted segments to form a shadowless image; and converting the shadowless image to an image in a first color space. As an alternative, pairing shadow segments with light segments based on the determined distance further includes: identifying one or more light segments that are closest to each shadow segment; and pairing each shadow segment with the one or more identified light segments. As an alternative, it further includes: segmenting the data into light segment portions and shadow segment portions; and labeling each portion used for segmentation. As another option, texture features include: edge distance, color distance, entropy distance, neighborhood distance, and combinations thereof. As another option, generating a shadow mask for the at least one shadow region further includes: identifying the at least one shadow region from the input image; and generating a shadow mask based on the identified at least one shadow region. As another option, it further includes: converting the input image to data in a fourth color space; generating a histogram based on color channels extracted from the data; smoothing the histogram using a Gaussian window; identifying local minima on the smoothed histogram; determining a threshold based on the identified local minima and a parameter associated with the input image, wherein the parameter is selected based on the size of the input image; and selecting pixels to be masked based on the color channels of the converted image being less than the threshold.

[0277] In the embodiments, aspects, and examples of the present invention described above, algorithms, models, processes, methods, systems, and / or apparatuses may be implemented and / or comprise one or more cloud platforms, servers, or computing systems or devices. Servers may include a single server or a network of servers, and cloud platforms may include multiple servers or server networks. In some instances, the functionality of servers and / or cloud platforms may be provided by a network of servers distributed across geographical regions (such as a globally distributed server network), and users can connect to an appropriate server network based on their location, etc.

[0278] For clarity, the above description refers to embodiments of the invention discussed in a single-user context. It should be understood that in practice, the system can be shared by multiple users, and possibly by a very large number of users simultaneously.

[0279] The above embodiments can be configured to be semi-automatic and / or fully automatic. In some instances, the user or operator querying the system / process / method can manually indicate some steps of the process / method to be performed.

[0280] Embodiments of the present invention—Systems, processes, methods, and / or tools for querying any data structures described herein, according to the present invention and / or as described herein, can be implemented as any form of computing and / or electronic device. Such a device may include one or more processors, which may be microprocessors, controllers, or any other suitable type of processor for processing computer-executable instructions to control the operation of the device in order to collect and record routing information. In some instances, for example, in the case of using a system-on-a-chip architecture, the processor may include one or more fixed-function blocks (also referred to as accelerators) that implement part of the process / method in hardware (rather than software or firmware). Platform software, including an operating system or any other suitable platform software, may be provided at the computing-based device to enable the execution of application software on that device.

[0281] The various functions described herein can be implemented in hardware, software, or any combination thereof. If implemented in software, these functions can be stored or transmitted as one or more instructions or code on or through a computer-readable medium or a non-transitory computer-readable medium. A computer-readable medium can include, for example, a computer-readable storage medium. A computer-readable storage medium can include volatile or non-volatile, removable or non-removable media implemented using any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. A computer-readable storage medium can be any available storage medium that can be accessed by a computer. By way of example and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, flash memory or other memory devices, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Optical discs and platters used herein include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy discs, and Blu-ray discs (BDs). Furthermore, the propagation of signals is not included within the scope of computer-readable storage media. Computer-readable media also includes communication media, which includes any medium that facilitates the transfer of a computer program from one place to another. For example, a connection or coupling can be a communication medium. For instance, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, it is included in the definition of a communication medium. Combinations of the above should also be included within the scope of computer-readable media.

[0282] Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, hardware logic components that may be used may include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc.

[0283] Although illustrated as a single system, it should be understood that computing devices can be distributed systems. Therefore, for example, several devices can communicate via a network connection and collaboratively perform tasks described as being performed by computing devices.

[0284] Although illustrated as a local device, it should be understood that the computing device may be located remotely and accessed via a network or other communication link (e.g., using a communication interface).

[0285] As used herein, the term "computer" refers to any device that has processing power that enables it to execute instructions. Those skilled in the art will recognize that such processing power is incorporated into many different devices, and therefore the term "computer" includes PCs, servers, IoT devices, mobile phones, personal digital assistants, and many other devices.

[0286] Those skilled in the art will recognize that storage devices used to store program instructions can be distributed across a network. For example, a remote computer can store instances of processes described as software. A local or terminal computer can access the remote computer and download part or all of the software to run the program. Alternatively, the local computer can download software fragments as needed, or execute some software instructions at a local terminal and some software instructions at a remote computer (or computer network). Those skilled in the art will also recognize that, by utilizing conventional techniques known to them, all or part of the software instructions can be executed by dedicated circuitry, such as DSPs, programmable logic arrays, etc.

[0287] It should be understood that the above benefits and advantages may relate to one embodiment or several embodiments. These embodiments are not limited to those that solve any or all of the stated problems or those that have any or all of the stated benefits and advantages. Variations should be considered to be included within the scope of this invention.

[0288] Any reference to the term "a" refers to one or more of these terms. The term "comprising" is used herein to mean including the identified method steps or elements, but such steps or elements are not included in an exclusive list, and the method or apparatus may contain additional steps or elements.

[0289] As used herein, the terms “component” and “system” are intended to cover a computer-readable data storage device configured with computer-executable instructions that, when executed by a processor, enable certain functions to be performed. These computer-executable instructions may include routines, functions, etc. It should also be understood that a component or system may reside on a single device or be distributed across several devices. Furthermore, as used herein, the terms “exemplary,” “example,” or “implementation” are intended to mean “serving as an illustration or instance of something.” Additionally, within the scope of the use of the term “includes” in the detailed description or claims, this term is intended to be inclusive in a manner similar to the term “comprising,” as “comprising” is interpreted when used as a transitional word in the claims.

[0290] The accompanying drawings illustrate exemplary methods. While these methods are shown and described as a series of actions performed in a specific sequence, it should be understood and recognized that these methods are not limited by the order of that sequence. For example, some actions may occur in a different order than that described herein. Furthermore, one action may occur simultaneously with another. Moreover, in some cases, not all actions are required to implement the methods described herein.

[0291] Furthermore, the actions described herein may include computer-executable instructions, which may be implemented by one or more processors and / or stored on one or more computer-readable media. These computer-executable instructions may include routines, subroutines, programs, threads of execution, etc. Additionally, the results of the actions of these methods may be stored in a computer-readable medium, displayed on a display device, etc.

[0292] The order of steps in the methods described herein is exemplary, but these steps can be performed in any suitable order, or simultaneously where appropriate. Furthermore, steps can be added to or substituted in any method, or individual steps can be deleted from any method, without departing from the scope of the subject matter described herein. Aspects of any of the above instances can be combined with aspects of any other instance described to form further instances without losing the desired effect.

[0293] It should be understood that the above description of the preferred embodiments is given only as an example, and various modifications can be made by those skilled in the art.

[0294] The foregoing includes examples of one or more embodiments. It is certainly not possible to describe every conceivable modification and alteration of the above-described apparatus or method for the purposes of describing the foregoing aspects; however, those skilled in the art will recognize that many further modifications and substitutions of the aspects are possible. Therefore, the described aspects are intended to cover all such changes, modifications, and variations falling within the scope of the appended claims.

[0295] This application also includes the subject matter of the following clauses:

[0296] 1. A computer-implemented method for processing an image based on color constancy to remove illumination from the image, the method comprising:

[0297] Obtain the image in the first color space;

[0298] Convert the image into data in a second color space;

[0299] The data is transformed using color adaptation.

[0300] Perform a first normalization on the transformed data, wherein the first normalization includes applying dynamic spatial filtering techniques to adjust the transformed data based on light intensity;

[0301] A set of filters is applied to normalized data, wherein the set of filters is convolved based on the normalized data associated with the image;

[0302] A second normalization is performed on the filtered data to obtain an illumination estimate of the image associated with the filtered data; and

[0303] Output normalized data from the second normalization, wherein the normalized data maintains color constancy based on the illumination estimate, thereby removing the illumination from the normalized data.

[0304] 2. The method described in Clause 1 further includes:

[0305] The normalized data from the second normalization is converted into an image in the first color space.

[0306] 3. The method according to Clause 1 or 2, wherein the performance of the first normalization further comprises:

[0307] Identify a data representation of the receptor response associated with the transformed data, wherein the data representation includes a uniform representation of the receptor response on the image; and

[0308] The data representation is applied to the transformed data.

[0309] 4. The method according to Clause 3, wherein the application data representation further includes:

[0310] Integrating the data representation over a time period; and normalizing the transformed data using the dynamic spatial filtering technique based on the integrated data representation, the dynamic spatial filtering technique representing a time-coded adjustment for each receptor for the light intensity.

[0311] 5. The method according to any one of the preceding clauses, wherein the second normalization includes correcting the filtered data using a lighting vector generated using a pooling function, wherein the filtered data is divisible by the lighting vector.

[0312] 6. The method according to Clause 5, wherein the pooling function comprises:

[0313]

[0314] Where f (x,y) (.) represents the data representation for which the maximum typical neural computation is performed on the filtered data, and The output of the double antagonistic filter in the second color space is shown.

[0315] 7. The method according to any one of the foregoing clauses further includes:

[0316] Receive raw images from one or more cameras;

[0317] Remove gamma correction from the original image; and

[0318] The original image is converted into the image in the first color space.

[0319] 8. The method according to any one of the preceding clauses, wherein the first color space is a red-green-blue color space.

[0320] 9. The method according to any one of the preceding clauses, wherein the second color space is a long-medium-short color space.

[0321] 10. The method according to any one of the preceding clauses, wherein the use of color adaptation to transform the data further comprises:

[0322] The irradiance of the data is estimated using a gray-world model; and

[0323] The data is converted into a target irradiator using the irradiator.

[0324] 11. The method according to Clause 10, wherein converting the data into a target irradiation body further comprises:

[0325] One or more matrix transformations are applied to the target irradiation body using one or more cameras used to acquire the image.

[0326] 12. The method according to any one of the preceding clauses, wherein the set of filters includes at least one center-periphery structure, the at least one center-periphery structure representing a receptive field (RF) encoding color antagonism.

[0327] 13. The method according to Clause 12, wherein the RF includes red-green antagonism, blue-yellow antagonism, and achromatic antagonism.

[0328] 14. The method according to any one of the preceding clauses, wherein the set of filters comprises at least two filters positioned in series, such that one of the at least two filters receives input from the other filter.

[0329] 15. The method according to any one of the preceding clauses, wherein the set of filters represents two layers of the vision system.

[0330] 16. The method according to any one of the preceding clauses further includes:

[0331] Obtain the input image;

[0332] The input image is divided into multiple regions;

[0333] The multiple regions are analyzed based on the color information and spatial location of the pixels in each region;

[0334] A subset of regions affected by the irradiation agent is selected from the plurality of regions based on analysis;

[0335] Identify colored edges for at least a subset of the region;

[0336] Extract the color information from the colored edge;

[0337] The extracted color information is used to decompose the reflection and illumination components of the input image;

[0338] The illumination of the input image is corrected based on the decomposed reflection and illumination components;

[0339] Output the image after illumination correction.

[0340] 17. The method according to any one of the preceding clauses, wherein the obtained input image is a raw image received from one or more cameras, the image in the first color space is obtained before the first normalization, the normalized data is derived from the output of the second normalization, or the image in the first color space is obtained after the second normalization.

[0341] 18. The method according to any one of the preceding clauses, wherein the analysis of the plurality of regions further comprises:

[0342] Clustering algorithms are used to identify salient regions from the plurality of regions;

[0343] Cluster maps are generated based on identified salient regions; and

[0344] The salient regions are analyzed to select the subset of regions affected by the irradiation agent.

[0345] 19. The method according to any one of the preceding clauses further includes:

[0346] Obtain the input image in the first color space;

[0347] Identify shadow areas from the input image;

[0348] The shadow mask is generated based on the identified shadow areas;

[0349] Shadows are removed from the shadow areas of the input image based on the shadow area mask; and

[0350] Output an image without shadows.

[0351] 20. The method according to Clause 19, wherein the input image is an original image or an image in the first color space.

[0352] 21. The method according to clause 19 or 20, wherein removing shadows from shadow areas of the input image based on the shadow area mask further comprises:

[0353] The input image is converted into data in a third color space;

[0354] The data of the input image is divided into multiple regions;

[0355] The multiple regions are divided into shadow segments and light segments based on the shadow area mask;

[0356] The distance between the shadow segment and the light segment is determined based on the texture features of the input image;

[0357] The shadow segment region is paired with the light segment region based on a determined distance;

[0358] Perform histogram matching on pairs of shadow and light segments to produce color-adjusted segments;

[0359] The pairing and histogram matching are performed iteratively until every shaded segment has been matched;

[0360] Merge the color-adjusted segments to form a shadow-free image; and

[0361] The shadowless image is converted into an image in the first color space.

[0362] 22. The method according to clauses 19 to 21, wherein identifying the shadow region from the input image further comprises:

[0363] The input image is converted into data in a fourth color space;

[0364] A histogram is generated based on the color channels extracted from the data;

[0365] Use a Gaussian window to smooth the histogram;

[0366] Identify local minima on smoothed histograms;

[0367] A threshold is determined based on the identified local minimum and parameters associated with the input image, wherein the parameters are selected based on the size of the input image.

[0368] 23. The method according to Clause 22, wherein generating the shadow mask based on the identified shadow area further comprises: selecting pixels to be masked based on the fact that the color channels of the converted image are less than the threshold.

[0369] 24. An apparatus for establishing color constancy of an image, the apparatus comprising:

[0370] One or more cameras are used to capture the image in a first color space;

[0371] The processing unit is used to convert the captured image into corresponding data in the second color space;

[0372] A first model, a second model, and a third model are configured to sequentially process the data to establish color constancy, wherein...

[0373] The first model is configured to transform the data using color adaptation, and the second model is configured to perform a first normalization on the transformed data, wherein the first normalization includes applying dynamic spatial filtering techniques to adjust the transformed data based on light intensity.

[0374] The third model is configured to apply a set of filters to normalized data, wherein the set of filters performs convolution based on the normalized data associated with the image; and perform a second normalization on the filtered data to obtain an illumination estimate of the image associated with the filtered data; and

[0375] An output module configured to output normalized data from the second normalization, wherein the normalized data maintains color constancy based on the illumination estimate.

[0376] 25. The apparatus according to clause 24, wherein the processing unit is configured to perform the method according to any one of clauses 2 to 23.

[0377] 26. A system for processing an image to establish color constancy of the image by removing illumination from the image, the system comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to perform any of the methods described in clauses 1 to 23.

[0378] 27. A computer-implemented method for reducing the effect of illumination on an image, the method comprising:

[0379] Receive input image;

[0380] The input image is divided into multiple regions;

[0381] The multiple regions are analyzed based on the color information and spatial location of the pixels in each region;

[0382] A subset of regions affected by the irradiation agent is selected from the plurality of regions based on analysis;

[0383] Identify colored edges for at least a subset of the region;

[0384] Extract the color information from the colored edge;

[0385] The extracted color information is used to decompose the reflection and illumination components of the input image;

[0386] The illumination of the input image is corrected based on the decomposed reflection and illumination components;

[0387] Output the image after illumination correction.

[0388] 28. The method according to Clause 27, wherein the input image is received according to any method of Clauses 1 to 15.

[0389] 29. The method according to clause 27 or 28, wherein the analysis of the plurality of regions further comprises:

[0390] Clustering algorithms are used to identify salient regions from the plurality of regions;

[0391] Cluster maps are generated based on identified salient regions; and

[0392] The salient regions are analyzed to select the subset of regions affected by the irradiator based on the clustering graph.

[0393] 30. The method according to Clause 29, wherein the clustering graph constrains the plurality of regions into a set of number of partitions.

[0394] 31. The method according to clauses 27 to 30, wherein dividing the input image into multiple regions further comprises: segmenting the input image using a k-means clustering algorithm based on the similarity of color or intensity values ​​of pixels from the input image.

[0395] 32. The method according to clauses 27 to 31, wherein identifying colored edges for at least a subset of the region further comprises: performing edge detection using a Canny edge detection algorithm configured to select a plurality of edges based on threshold intensity.

[0396] 33. The method according to clauses 27 to 32, wherein identifying colored edges for at least a subset of the region further comprises: performing edge detection using the Sobel operator.

[0397] 34. The method according to clauses 27 to 33 further comprises: identifying regions having the colored edges from at least one subset of regions of the input image; and correcting the illumination of the input image based on the identified regions having the colored edges.

[0398] 35. An apparatus for processing an image to maintain the color constancy of the image, the apparatus comprising: at least one model configured to perform the steps of any method pursuant to clauses 27 to 34.

[0399] 36. A system for processing an image to establish color constancy of the image by removing illumination from the image, the system comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to perform any of the methods of clauses 27 to 34.

[0400] 37. A computer-implemented method for providing a shadowless image, the method comprising:

[0401] Receive an input image in a first color space, wherein the input image includes at least one shadow area;

[0402] Generate a shadow mask for the at least one shadow region;

[0403] Shadows are removed from the shadow areas of the input image based on the shadow area mask; and

[0404] Output an image without shadows.

[0405] 38. The method according to Clause 37, wherein the input image is received according to any one of the methods according to Clauses 1 to 15 and / or Clauses 27 to 34.

[0406] 39. The method according to clause 37 or 38, wherein removing shadows from shadow areas of the input image based on the shadow area mask further comprises:

[0407] The input image is converted into data in a third color space;

[0408] The data of the input image is divided into multiple regions;

[0409] The plurality of regions are divided into shadow segments and light segments according to the shadow area mask;

[0410] The distance between the shadow segment and the light segment is determined based on the texture features of the input image;

[0411] The shadow segment region is paired with the light segment region based on a determined distance;

[0412] Perform histogram matching on pairs of shadow and light segments to produce color-adjusted segments;

[0413] The pairing and histogram matching are performed iteratively until every shaded segment has been matched;

[0414] Merge the color-adjusted segments to form a shadow-free image; and

[0415] The shadowless image is converted into an image in the first color space.

[0416] 40. The method according to Clause 39, wherein pairing the shadow segment region with the light segment region based on a determined distance further comprises:

[0417] Identify the one or more light segments that are closest to each shadow segment; and

[0418] Each shadow segment is paired with one or more identified light segment regions.

[0419] 41. The method according to clause 39 or 40 further includes: segmenting the data into light segment portions and shadow segment portions; and marking each segment for segmentation.

[0420] 42. The method according to clauses 39 to 41, wherein the texture features include: edge distance, color distance, entropy distance, neighborhood distance, and combinations thereof.

[0421] 43. The method according to clauses 39 to 42, wherein generating a shadow mask for the at least one shadow region further comprises:

[0422] Identify the at least one shadow region from the input image; and

[0423] A shadow mask is generated based on the identified at least one shadow region.

[0424] 44. The method described under Articles 39 to 43 further includes:

[0425] The input image is converted into data in a fourth color space;

[0426] A histogram is generated based on the color channels extracted from the data;

[0427] Use a Gaussian window to smooth the histogram;

[0428] Identify local minima on smoothed histograms;

[0429] A threshold is determined based on the identified local minima and parameters associated with the input image, wherein the parameters are selected based on the size of the input image; and

[0430] The pixels to be masked are selected based on the fact that the color channels of the converted image are less than the threshold.

[0431] 45. An apparatus for processing an image to maintain the color constancy of the image, the apparatus comprising: at least one model configured to perform the steps of any method pursuant to clauses 37 to 44.

[0432] 46. ​​A system for processing an image to establish color constancy of the image by removing illumination from the image, the system comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to perform any of the methods of clauses 37 to 44.

Claims

1. A computer-implemented method for processing an image based on color constancy to remove illumination from the image, the method comprising: Obtain the image in the first color space; Convert the image into data in a second color space; The data is transformed using color adaptation. A first normalization is performed on the transformed data, wherein the first normalization includes weighting the transformed data based on the cumulative distribution of light intensity values ​​in the image; A set of filters is applied to normalized data, wherein the set of filters is convolved based on the normalized data associated with the image; A second normalization is performed on the image data in the second color space by obtaining an illumination estimate of the image associated with the filtered data and dividing the image data in the second color space by the illumination estimate; as well as Output normalized data from the second normalization, wherein the normalized data maintains color constancy based on the illumination estimate, thereby removing the illumination from the normalized data.

2. The method according to claim 1, further comprising: The normalized data from the second normalization is converted into an image in the first color space.

3. The method according to claim 1 or 2, further comprising: The transformed data is remapped to a sparse matrix, where each position in the first dimension corresponds to a pixel in the image, and the second dimension corresponds to a range of possible pixel values. The first normalization is performed on the remapped data.

4. The method according to any one of the preceding claims, wherein the second normalization comprises correcting the filtered data using a lighting vector generated using a pooling function, wherein the filtered data is divisible by the lighting vector.

5. The method according to claim 4, wherein the pooling function comprises: Where f (x,y) (.) represents the data representation for which the maximum typical neural computation is performed on the filtered data, and The output of the double antagonistic filter in the second color space is shown.

6. The method according to any one of the preceding claims further comprises: Receive raw images from one or more cameras; Remove gamma correction from the original image; as well as The original image is converted into the image in the first color space.

7. The method according to any one of the preceding claims, wherein the first color space is a red-green-blue color space.

8. The method according to any one of the preceding claims, wherein the second color space is a long-medium-short color space.

9. The method according to any one of the preceding claims, wherein the use of color adaptation to transform the data further comprises: The irradiance of the data is estimated using a gray-world model; as well as The data is converted into a target irradiator using the irradiator.

10. The method of claim 9, wherein converting the data into a target irradiation body further comprises: One or more matrix transformations are applied to the target irradiation body using one or more cameras used to acquire the image.

11. The method according to any one of the preceding claims, wherein the set of filters includes at least one center-periphery structure, the at least one center-periphery structure representing a receptive field (RF) encoding color antagonism.

12. The method of claim 11, wherein the RF comprises red-green antagonism, blue-yellow antagonism, and achromatic antagonism.

13. The method according to any one of the preceding claims, wherein the set of filters comprises at least two filters cascaded, such that one of the at least two filters receives input from the other filter.

14. The method according to any one of the preceding claims, wherein the set of filters represents two layers of the vision system.

15. The method according to any one of the preceding claims further comprises: Obtain the input image; The input image is divided into multiple regions; The multiple regions are analyzed based on the color information and spatial location of the pixels in each region; A subset of regions affected by the irradiation agent is selected from the plurality of regions based on analysis; Identify colored edges for at least a subset of the region; Extract the color information from the colored edge; The extracted color information is used to decompose the reflection and illumination components of the input image; The illumination of the input image is corrected based on the decomposed reflection and illumination components; Output the image after illumination correction.

16. The method of claim 15, wherein the obtained input image is a raw image received from one or more cameras, the image in the first color space is obtained before the first normalization, the normalized data is from the output of the second normalization, or the image in the first color space is obtained after the second normalization.

17. The method of claim 15 or 16, wherein the analysis of the plurality of regions further comprises: Clustering algorithms are used to identify salient regions from the plurality of regions; Cluster maps are generated based on the identified salient regions; as well as The salient regions are analyzed to select the subset of regions affected by the irradiator based on the clustering graph.

18. The method of claim 17, wherein the clustering graph constrains the plurality of regions into a set number of partitions.

19. The method according to any one of claims 15 to 18, wherein dividing the input image into a plurality of regions further comprises: The input image is segmented using a k-means clustering algorithm based on the similarity of the color or intensity values ​​of pixels from the input image.

20. The method according to any one of claims 15 to 19, wherein identifying colored edges for at least a subset of the region further comprises: Edge detection is performed using the Canny edge detection algorithm, which is configured to select multiple edges based on threshold intensity.

21. The method according to any one of claims 15 to 20, wherein identifying colored edges for at least the subset of the region further comprises: Edge detection is performed using the Sobel operator.

22. The method according to any one of claims 15 to 21, further comprising: Identify regions with the colored edges from at least one subset of regions in the input image; And to correct the illumination of the input image based on the identified regions with the colored edges.

23. The method according to any one of claims 15 to 22, wherein the analysis of the plurality of regions further comprises: Clustering algorithms are used to identify salient regions from the plurality of regions; Cluster maps are generated based on the identified salient regions; as well as The salient regions are analyzed to select the subset of regions affected by the irradiation agent.

24. The method according to any of the preceding claims, further comprising: Obtain the input image in the first color space; Identify shadow areas from the input image; The shadow mask is generated based on the identified shadow areas; The shadow is removed from the shadow area of ​​the input image based on the shadow area mask. as well as Output an image without shadows.

25. The method of claim 24, wherein the input image is an original image or an image in the first color space.

26. The method of claim 24 or 25, wherein removing shadows from shadow areas of the input image based on the shadow area mask further comprises: The input image is converted into data in a third color space; The data of the input image is divided into multiple regions; The plurality of regions are divided into shadow segments and light segments according to the shadow area mask; The distance between the shadow segment and the light segment is determined based on the texture features of the input image; The shadow segment region is paired with the light segment region based on the minimum distance between the shadow segment and the light segment. Histogram matching is performed on the shadow segment based on the corresponding light segment to generate a color-adjusted segment; Merge the color-adjusted segments to form a shadow-free image; and The shadowless image is converted into an image in the first color space.

27. The method of claim 26, wherein pairing the shadow segment region with the light segment region based on a determined distance further comprises: Identify the one or more light segments that are closest to each shadow segment; as well as Each shadow segment is paired with one or more identified light segment regions.

28. The method according to claim 26 or 27, further comprising: The data is divided into light segment and shadow segment; And a label for each segment used for splitting.

29. The method according to any one of claims 26 to 28, wherein the texture feature comprises: Edge distance, color distance, entropy distance, neighborhood distance, and their combinations.

30. The method according to any one of claims 24 to 29, wherein generating a shadow mask for the at least one shadow region further comprises: Identify the at least one shadow region from the input image; as well as A shadow mask is generated based on the identified at least one shadow region.

31. The method according to any one of claims 24 to 30, wherein identifying the shadow region from the input image further comprises: The input image is converted into data in a fourth color space; A histogram is generated based on the color channels extracted from the data; Use a Gaussian window to smooth the histogram; Identify local minima on smoothed histograms; A threshold is determined based on the identified local minimum and parameters associated with the input image, wherein the parameters are selected based on the size of the input image.

32. The method of claim 31, wherein generating the shadow mask based on the identified shadow region further comprises: The pixels to be masked are selected based on the fact that the color channels of the converted image are less than the threshold.

33. An apparatus for establishing color constancy of an image, the apparatus comprising: One or more cameras are used to capture the image in a first color space; The processing unit is used to convert the captured image into corresponding data in the second color space; A first model, a second model, and a third model are configured to sequentially process the data to establish color constancy, wherein... The first model is configured to transform the data using color adaptation, and the second model is configured to perform a first normalization on the transformed data, wherein the first normalization includes weighting the transformed data based on the cumulative distribution of light intensity values ​​in the image, and The third model is configured to apply a set of filters to normalized data, wherein the set of filters is convolved based on the normalized data associated with the image; Furthermore, a second normalization is performed on the image data in the second color space by obtaining an illumination estimate of the image associated with the filtered data and dividing the image data in the second color space by the illumination estimate; as well as An output module configured to output normalized data from the second normalization, wherein the normalized data maintains color constancy based on the illumination estimate.

34. The apparatus of claim 33, wherein the processing unit is configured to perform the method of any one of claims 2 to 14.

35. An apparatus for processing an image to maintain the color constancy of the image, the apparatus comprising: At least one model, said at least one model being configured to perform the steps of any of the methods according to claims 15 to 32.

36. A system for processing an image to establish the color constancy of the image by removing illumination from the image, the system comprising: At least one processor; And a memory storing instructions that, when executed by the at least one processor, cause the system to perform any of the methods of claims 1 to 32.

37. A computer-implemented method for reducing the effect of illumination on an image, the method comprising: Receive input image; The input image is divided into multiple regions; The multiple regions are analyzed based on the color information and spatial location of the pixels in each region; A subset of regions affected by the irradiation agent is selected from the plurality of regions based on analysis; Identify colored edges for at least a subset of the region; Extract the color information from the colored edge; The extracted color information is used to decompose the reflection and illumination components of the input image; The illumination of the input image is corrected based on the decomposed reflection and illumination components; Output the image after illumination correction.

38. The method of claim 37, wherein the input image is received by any method of claims 1 to 14 and / or claims 24 to 32.

39. The method of claim 37 or 38, wherein the analysis of the plurality of regions further comprises: Clustering algorithms are used to identify salient regions from the plurality of regions; Cluster maps are generated based on the identified salient regions; as well as The salient regions are analyzed to select the subset of regions affected by the irradiator based on the clustering graph.

40. The method of claim 38, wherein the clustering graph constrains the plurality of regions into a set of number of partitions.

41. The method according to claims 37 to 40, wherein dividing the input image into multiple regions further comprises: The input image is segmented using a k-means clustering algorithm based on the similarity of the color or intensity values ​​of pixels from the input image.

42. The method of claims 37 to 41, wherein identifying colored edges for at least a subset of the region further comprises: Edge detection is performed using the Canny edge detection algorithm, which is configured to select multiple edges based on threshold intensity.

43. The method of claims 37 to 42, wherein identifying colored edges for at least a subset of the region further comprises: Edge detection is performed using the Sobel operator.

44. The method according to claims 37 to 43, further comprising: Identify regions with the colored edges from at least one subset of regions in the input image; And to correct the illumination of the input image based on the identified regions with the colored edges.

45. An apparatus for processing an image to maintain the color constancy of the image, the apparatus comprising: At least one model, said at least one model being configured to perform the steps of any of the methods according to claims 37 to 44.

46. ​​A system for processing an image to establish the color constancy of the image by removing illumination from the image, the system comprising: At least one processor; And a memory storing instructions that, when executed by the at least one processor, cause the system to perform any of the methods of claims 37 to 44.

47. A computer-implemented method for providing a shadowless image, the method comprising: Receive an input image in a first color space, wherein the input image includes at least one shadow area; Generate a shadow mask for the at least one shadow region; The shadow is removed from the shadow area of ​​the input image based on the shadow area mask. as well as Output an image without shadows.

48. The method of claim 47, wherein the input image is received according to any one of the methods of claims 1 to 23 and / or claims 37 to 44.

49. The method of claim 47 or 48, wherein removing shadows from shadow areas of the input image based on the shadow area mask further comprises: The input image is converted into data in a third color space; The data of the input image is divided into multiple regions; The plurality of regions are divided into shadow segments and light segments according to the shadow area mask; The distance between the shadow segment and the light segment is determined based on the texture features of the input image; The shadow segment region is paired with the light segment region based on the minimum distance between the shadow segment and the light segment. Histogram matching is performed on the shadow segment based on the corresponding light segment to generate a color-adjusted segment; The pairing and histogram matching are performed iteratively until every shaded segment has been matched; Merge the color-adjusted segments to form a shadow-free image; and The shadowless image is converted into an image in the first color space.

50. The method of claim 49, wherein pairing the shadow segment region with the light segment region based on a determined distance further comprises: Identify the one or more light segments that are closest to each shadow segment; as well as Each shadow segment is paired with one or more identified light segment regions.

51. The method according to claim 49 or 50, further comprising: The data is divided into light segment and shadow segment; And a label for each segment used for splitting.

52. The method according to claims 49 to 51, wherein the texture feature comprises: Edge distance, color distance, entropy distance, neighborhood distance, and their combinations.

53. The method according to claims 49 to 52, wherein generating a shadow mask for the at least one shadow region further comprises: Identify the at least one shadow region from the input image; as well as A shadow mask is generated based on the identified at least one shadow region.

54. The method according to claims 49 to 53, further comprising: The input image is converted into data in a fourth color space; A histogram is generated based on the color channels extracted from the data; Use a Gaussian window to smooth the histogram; Identify local minima on smoothed histograms; A threshold is determined based on the identified local minima and parameters associated with the input image, wherein the parameters are selected based on the size of the input image; as well as The pixels to be masked are selected based on the fact that the color channels of the converted image are less than the threshold.

55. An apparatus for processing an image to maintain the color constancy of the image, the apparatus comprising: At least one model, said at least one model being configured to perform the steps of any of the methods according to claims 47 to 54.

56. A system for processing an image to establish the color constancy of the image by removing illumination from the image, the system comprising: At least one processor; And a memory storing instructions that, when executed by the at least one processor, cause the system to perform any of the methods of claims 47 to 54.