Electronic rearview mirror image processing method and electronic rearview mirror
By acquiring visible light and infrared light images through electronic rearview mirrors, dividing the area and performing fusion processing, the problem of unclear rearview mirror visibility in low light or inclement weather is solved, achieving high-quality image display under these conditions and ensuring driver safety.
Patent Information
- Application Number
- CN202510903480.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-11-11
AI Technical Summary
In low light, inclement weather, or special circumstances, the visible light reflection from vehicle rearview mirrors cannot effectively provide a clear rear view, limiting the driver's field of vision and recognition ability, thus affecting driving safety.
Visible and infrared images are acquired using an electronic rearview mirror, local areas are divided, regional fusion weights are determined based on pixel parameters, and image fusion is performed to generate a high-quality fused image to enhance the driver's rearview vision.
Provides a bright and clear rearview image in low light or inclement weather conditions, ensuring that the driver can accurately identify the external environment and improve driving safety.
Smart Images

Figure CN120931501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an image processing method for an electronic rearview mirror and an electronic rearview mirror. Background Technology
[0002] Vehicle rearview mirrors provide rear visibility by reflecting visible light from the external environment. However, in low light, fog, rain, snow, and other adverse weather conditions, as well as when observing pedestrians wearing dark clothing or animals with dark skin, the driver's field of vision and recognition ability are greatly limited, failing to effectively meet the needs of vehicle driving safety. Summary of the Invention
[0003] In view of the above problems, embodiments of the present invention are proposed to provide an electronic rearview mirror image processing method and an electronic rearview mirror that overcome or at least partially solve the above problems.
[0004] To address the aforementioned problems, in a first aspect of this invention, an embodiment of the invention discloses an electronic rearview mirror image processing method, comprising:
[0005] Acquire visible light and infrared light images captured by the electronic rearview mirror;
[0006] The visible light image is divided to generate multiple first local regions, and the infrared light image is divided to generate multiple second local regions, wherein the first local regions and the second local regions are associated with the same target object;
[0007] Based on the pixel parameters of the associated first and second local regions, the region fusion weights are determined.
[0008] Based on the region fusion weights, the associated first local region and second local region are fused together with respect to the target object to generate a fused image.
[0009] Optionally, the step of dividing the visible light image to generate multiple first local regions includes:
[0010] Identify the target object in the visible light image;
[0011] A first local region is determined based on the target object in the visible light image; and / or
[0012] The step of dividing the infrared image to generate multiple second local regions includes:
[0013] Identify the target object in the infrared image;
[0014] A second local region is determined based on the target object in the infrared image.
[0015] Optionally, the step of identifying the target object in the visible light image includes:
[0016] The target object in the visible light image is identified using a preset target recognition model; and / or,
[0017] The step of identifying the target object in the infrared image includes:
[0018] The target object in the infrared image is identified using a preset target recognition model;
[0019] The preset target recognition model is generated in the following manner:
[0020] The convolutional neural network is trained based on preset training samples, and the cross-entropy loss value for this training is determined.
[0021] Backpropagation is performed based on the cross-entropy loss value of the current training iteration to update the convolutional neural network;
[0022] The step of training the convolutional neural network based on the updated convolutional neural network using preset training samples and determining the cross-entropy loss value for the current training iteration continues until the cross-entropy loss value for the current training iteration is less than a preset loss threshold.
[0023] Optionally, the step of determining the first local region based on the target object in the visible light image includes:
[0024] Get the window size;
[0025] Centered on the target object in the visible light image, the image region corresponding to the window size is determined as the first local region; and / or,
[0026] The step of determining the second local region based on the target object in the infrared image includes:
[0027] Get the window size;
[0028] Centered on the target object in the infrared image, the image region corresponding to the window size is determined as the second local region.
[0029] Optionally, the step of determining the region fusion weight based on the pixel parameters of the associated first and second local regions includes:
[0030] Determine the region types of the associated first and second local regions;
[0031] The region fusion weight is determined based on the region type and the pixel parameters of the associated first and second local regions.
[0032] Optionally, the region fusion weight includes a first local region weight value and a second local region weight value; the step of fusing the associated first local region and second local region for the target object based on the region fusion weight includes:
[0033] The first local region is determined by combining the first local region and its weight value.
[0034] The fused pixel value of the second region is determined by combining the weight value of the second local region with the weight value of the second local region.
[0035] Based on the target object, a fused image is generated by combining the fused pixel values of the first region and the fused pixel values of the second region.
[0036] Optionally, the step of determining the region fusion weight based on the region type, the pixel parameters of the associated first local region and second local region includes:
[0037] When the area type is a road or building area, the entropy value of the first local area and the entropy value of the second local area are determined to be the pixel parameter;
[0038] Determine the ratio of the entropy value of the first local region to the entropy value of the second local region;
[0039] The weight values of the first local region and the second local region are determined based on the ratio.
[0040] Optionally, the step of determining the region fusion weight based on the region type, the pixel parameters of the associated first local region and second local region includes:
[0041] If the area type is not a road or building area, determine the first basic weight value of the first local area and the second basic weight value of the second local area;
[0042] The weight adjustment value is determined based on pixel parameters;
[0043] The first local region weight value is determined by adding the weight adjustment value to the first basic weight value, and the second local region weight value is determined by deducting the weight adjustment value from the second basic weight value.
[0044] Optionally, if the region type is a sky or a vegetated region, the pixel parameters include pixel attributes; and / or;
[0045] When the region type is a vehicle region, the pixel parameters include size and distance; and / or;
[0046] When the area type is a pedestrian area, the pixel parameters include the temperature value corresponding to the pixel.
[0047] In a second aspect, an embodiment of the present invention discloses an electronic rearview mirror, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the steps of the electronic rearview mirror image processing method as described above.
[0048] The embodiments of the present invention have the following advantages:
[0049] This invention acquires visible light and infrared light images from an electronic rearview mirror; divides the visible light image into multiple first local regions and the infrared light image into multiple second local regions, where the first and second local regions are associated with the same target object; determines region fusion weights based on the pixel parameters of the associated first and second local regions; and fuses the associated first and second local regions with respect to the target object based on the region fusion weights to generate a fused image. By identifying and dividing the visible light and infrared light images into different regions, and determining corresponding region fusion weights for each region, the invention allows for dynamic determination of region fusion weights based on different situations and targets. This enables the fusion of the visible light and infrared light images based on the region fusion weights, resulting in a high-quality fused image. This allows the driver to determine the external environment based on the fused image, ensuring driving safety. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating the steps of an embodiment of an electronic rearview mirror image processing method according to the present invention;
[0051] Figure 2 This is a flowchart illustrating the steps of another embodiment of the electronic rearview mirror image processing method of the present invention;
[0052] Figure 3 This is a schematic block diagram of the processing architecture of an electronic rearview mirror provided in an embodiment of the present invention. Detailed Implementation
[0053] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0054] Reference Figure 1 The diagram illustrates a flowchart of an embodiment of an electronic rearview mirror image processing method according to the present invention. The electronic rearview mirror image processing method may specifically include the following steps:
[0055] Step 101: Acquire the visible light image and infrared light image captured by the electronic rearview mirror;
[0056] Electronic rearview mirrors can be equipped with both visible light and infrared cameras. These cameras detect and sample the vehicle's external environment. The infrared camera can be a high-sensitivity one, capable of sensing infrared radiation in specific wavelengths. Its optical system has a suitable focal length and field of view to ensure effective coverage of the area behind and around the vehicle. For example, the field of view can be set between 60° and 120°, and the focal length is determined based on the vehicle's dimensions and installation location, typically between 5 and 15 mm. The visible light camera, with its high resolution and wide dynamic range, can clearly capture visible light images of the area behind the vehicle under various lighting conditions. Its installation location is adapted to the infrared camera to ensure good spatial consistency between the images captured by both. The system can acquire both visible light and infrared images from the electronic rearview mirror. The visible light image is the one captured by the visible light camera. The infrared image is the one captured by the infrared camera.
[0057] Step 102: Divide the visible light image to generate multiple first local regions, and divide the infrared light image to generate multiple second local regions, wherein the first local regions and the second local regions are associated with the same target object;
[0058] After obtaining the visible light and infrared images, they can be divided into regions. Since both visible light and infrared images are detected from the same location, the acquired areas are the same, and the same division method can be used. Multiple first local regions can be generated for the visible light image. These first local regions are local image regions within the visible light image. Similarly, multiple second local regions can be generated for the infrared image. These second local regions are local image regions within the infrared image. Because the visible light and infrared images acquire the same area and are divided in the same way, the first and second local regions, after being divided for the same target object, are location-related. That is, for the same acquisition area and the same target object, there are corresponding first local regions in the visible light image and second local regions in the infrared image.
[0059] Step 103: Determine the region fusion weights based on the pixel parameters of the associated first and second local regions;
[0060] The region fusion weights corresponding to each pair of associated first and second local regions can be determined based on the pixel parameters in each pair of associated first and second local regions.
[0061] Step 104: Based on the region fusion weight, the associated first local region and second local region are fused for the target object to generate a fused image.
[0062] Based on the region fusion weights corresponding to each pair of associated first and second local regions, the first and second local regions of each pair are fused for the same target object to determine the corresponding local region image. Then, all the fused images are stitched together based on position to generate a fused image. The fused image provides the driver with a relatively clear rear view, ensuring the driver's driving safety.
[0063] This invention acquires visible light and infrared light images from an electronic rearview mirror; divides the visible light image into multiple first local regions and the infrared light image into multiple second local regions, where the first and second local regions are associated with the same target object; determines region fusion weights based on the pixel parameters of the associated first and second local regions; and fuses the associated first and second local regions with respect to the target object based on the region fusion weights to generate a fused image. By identifying and dividing the visible light and infrared light images into different regions, and determining corresponding region fusion weights for each region, the invention allows for dynamic determination of region fusion weights based on different situations and targets. This enables the fusion of the visible light and infrared light images based on the region fusion weights, resulting in a high-quality fused image. This allows the driver to determine the external environment based on the fused image, ensuring driving safety.
[0064] Reference Figure 2 The diagram illustrates a flowchart of another embodiment of the electronic rearview mirror image processing method of the present invention. The electronic rearview mirror image processing method may specifically include the following steps:
[0065] Step 201: Acquire the visible light image and infrared light image captured by the electronic rearview mirror;
[0066] It can acquire visible light and infrared images from electronic rearview mirrors. These images can be digital, allowing for direct processing and improving efficiency.
[0067] After obtaining the infrared image, it can be filtered to eliminate sampling noise and improve processing accuracy. Specifically, median filtering can be applied to the infrared image, which involves sorting the pixel values in the neighborhood of each pixel and replacing the current pixel value with the median value, thereby removing noise. Subsequent processing is then performed based on the filtered infrared image.
[0068] Visible light images can also be filtered. Median filtering involves sorting the neighboring pixel values of each pixel in the visible light image and replacing the current pixel value with the median value, thus removing noise. Furthermore, preprocessing operations such as color correction and grayscale stretching can be performed on visible light images. Grayscale stretching uses a linear transformation formula to enhance image contrast and make image details clearer. The processed visible light image is then used for further processing.
[0069] Step 202: Divide the visible light image to generate multiple first local regions, and divide the infrared light image to generate multiple second local regions, wherein the first local regions and the second local regions are associated with the same target object;
[0070] The visible light image is divided into multiple first local regions. The infrared image is divided into multiple second local regions using the same method. For the same target object within the same location range, the first and second local regions are associated. The visible light and infrared images can be divided using the same method.
[0071] In an optional embodiment of the present invention, the step of dividing the visible light image to generate multiple first local regions includes: identifying a target object in the visible light image; and determining a first local region based on the target object in the visible light image.
[0072] First, target objects in the visible light image can be identified. These target objects include, but are not limited to, people, vehicles, roads, backgrounds, and other objects used by the driver to observe the external environment. Based on the target objects in the visible light image, the corresponding first local region is determined.
[0073] The step of identifying target objects in the visible light image includes: using a preset target recognition model to identify the target objects in the visible light image. A preset target recognition model can be used to identify target objects in the visible light image, with the visible light image as input information. For example, the preset target recognition model can be a convolutional neural network (CNN) structure. After processing the visible light image through multiple convolutional layers, pooling layers, and fully connected layers, the category information of the target objects in the image is output. Using a neural network model allows for fast and accurate automated identification of target objects, improving identification efficiency.
[0074] Accordingly, in an optional embodiment of the present invention, the step of dividing the infrared light image to generate multiple second local regions includes: identifying a target object in the infrared light image; and determining a second local region based on the target object in the infrared light image.
[0075] It can identify target objects in infrared images, and these target objects can all be the same as those in visible light images. Based on the target objects in the infrared images, a corresponding second local region is determined.
[0076] The step of identifying target objects in the infrared image includes: using a preset target recognition model to identify the target objects in the infrared image. A preset target recognition model can be used to identify target objects in the infrared image, with the infrared image as input information. For example, the preset target recognition model can be a convolutional neural network structure. After processing the infrared image through multiple convolutional layers, pooling layers, and fully connected layers, the category information of the target objects in the image is output. Using a neural network model allows for fast and accurate automated identification of target objects, improving identification efficiency.
[0077] Furthermore, the preset target recognition model for target object recognition in infrared light images and target object recognition in visible light images is generated in the following manner: training a convolutional neural network based on preset training samples to determine the cross-entropy loss value of the current training; performing backpropagation based on the cross-entropy loss value of the current training to update the convolutional neural network; and performing the step of training the convolutional neural network based on preset training samples and determining the cross-entropy loss value of the current training based on the updated convolutional neural network until the cross-entropy loss value of the current training is less than a preset loss threshold.
[0078] When training a predefined target recognition model, a convolutional neural network (CNN) is used as the foundation, employing a large number of infrared and visible light images labeled with different object categories (such as pedestrians, vehicles, animals, and obstacles) as predefined training samples. A cross-entropy loss function is used, where represents the true label of the sample and represents the probability distribution predicted by the model, to determine the cross-entropy loss value for each training iteration. Based on this cross-entropy loss value, the model parameters are continuously adjusted through backpropagation. For example, increasing the number of convolutional layers or adjusting the kernel size can improve the model's ability to recognize objects at different scales; this updates the CNN. Based on the updated CNN, training is performed using the predefined training samples, continuing the process of determining the cross-entropy loss value for each training iteration until it falls below a predefined loss threshold, thus ensuring the CNN training meets the requirements and improving the model's recognition accuracy.
[0079] Furthermore, the Kalman filter algorithm can be used to track identified target objects. The target object's state vector contains information such as position coordinates and velocity components. The Kalman filter's prediction equation predicts the state based on the state transition matrix and time intervals. By continuously updating the state vector, stable tracking of the target object is achieved, and its trajectory and velocity information are obtained. During tracking, the Kalman filter's state vector and covariance matrix are continuously updated based on the prediction results and the target object's position information in the actual image to adapt to changes in the target object's motion. This effectively improves the vehicle's rear-view perception capability in various environments, identifies target objects that are difficult to detect with ordinary electronic rearview mirrors, provides drivers with more comprehensive and accurate safety assistance information, and enhances vehicle driving safety.
[0080] In an optional embodiment of the present invention, the step of determining a first local region based on a target object in the visible light image includes: obtaining a window size; and determining the image region corresponding to the window size as the first local region, with the target object in the visible light image as the center.
[0081] The window size for region division can be determined. Taking the target object in the visible light image as the center, and in conjunction with the window size, the corresponding image region is determined, and this image region is defined as the first local region.
[0082] In an optional embodiment of the present invention, the step of determining the second local region based on the target object in the infrared light image includes: obtaining the window size; and determining the image region corresponding to the window size as the second local region, with the target object in the infrared light image as the center.
[0083] Accordingly, the window size for region division can be determined. Taking the target object in the infrared image as the center and matching the window size, the corresponding image region can be determined, and this image region can be defined as the second local region.
[0084] Step 203: Determine the region fusion weights based on the pixel parameters of the associated first and second local regions;
[0085] After determining the associated first and second local regions, the region fusion weight of each local region is determined based on the pixel parameters within those regions. This region fusion weight can include the first local region weight value of the first local region and the second local region weight value of the second local region. The sum of the first and second local region weight values is 1. This can be expressed as: the second local region weight value is σ. IR (x, y), weight value of the first local region σ VIS (x, y), and σ IR(x, y) + σ VIS (x, y) = 1.
[0086] In an optional embodiment of the present invention, the step of determining the region fusion weight based on the pixel parameters of the associated first local region and second local region includes:
[0087] Sub-step S2031: Determine the region type of the associated first local region and second local region;
[0088] First, the region types of the associated first and second local regions can be determined. The region type is the type of target object corresponding to the local region. Region types include, but are not limited to, roads, buildings, sky, vegetation, vehicles, and pedestrians.
[0089] Sub-step S2032: Based on the region type and the pixel parameters of the associated first and second local regions, determine the region fusion weight.
[0090] Based on different region types, the pixel parameters corresponding to the associated first and second local regions are used to calculate and determine the region fusion weights. By dynamically adjusting the weights during the fusion of infrared and visible light images, images that better meet practical needs and visual effects can be generated, enabling them to perform better in different application scenarios and possess greater adaptability.
[0091] In one example of the present invention, the step of determining the region fusion weight based on the region type, the pixel parameters of the associated first local region and the second local region includes: when the region type is a road or building region, determining the entropy value of the first local region and the entropy value of the second local region as the pixel parameters; determining the ratio of the entropy value of the first local region and the entropy value of the second local region; and determining the weight value of the first local region and the weight value of the second local region based on the ratio.
[0092] When the region type is a road or building area, i.e., determining which image, the visible light image or the infrared image, is more effective in representing the external environment, then the infrared image is given a higher weight. The local region R in the image... A The probability distribution of pixel values within (x, y) is P. A (k) (where k represents the pixel value), then the entropy H of this local region. A (x, y) = -Σ k P A (k)log2P A (k). Calculate the probability distribution P. A When (k), first count the number of times each pixel value appears in the local region, n. k ,Then The entropy values of the first and second local regions can be calculated, thus obtaining the entropy value H of the second local region. IR (x, y) and the entropy value H of the first local region VIS (x, y). These entropy values are used as pixel parameters. When H IR (x,y)>H VIS When H = (x, y), it indicates that the infrared image contains more information during fusion, thus giving the second local region a larger weight; conversely, when H = (x, y), it indicates that the infrared image contains more information during fusion, thus giving the second local region a larger weight; IR (x,y)>H VIS When (x,y), the first local region is given a larger weight. When H IR (x,y)=H VIS When (x, y), the weights can be distributed equally, i.e., σ IR (x,y)=σ VIS (x, y). The ratio of the entropy values of the first local region and the second local region can be calculated, and then the total weight 1 can be mapped to the weight values of the first and second local regions based on this ratio. For the road region R road By comparing the entropy values, the weight values of the second local region and the first local region can be obtained respectively. and For building area R building By comparing the entropy values, the weight values of the second local region and the first local region can be obtained as follows: and
[0093] In an optional embodiment of the present invention, the step of determining the region fusion weight based on the region type, the pixel parameters of the associated first local region and the second local region includes: when the region type is not a road or building region, determining a first basic weight value of the first local region and a second basic weight value of the second local region; determining a weight adjustment value based on the pixel parameters; increasing the weight adjustment value on the first basic weight value to determine the weight value of the first local region, and decreasing the weight adjustment value on the second basic weight value to determine the weight value of the second local region.
[0094] For areas that are not roads or buildings, a base weight value can be assigned to both the first and second local regions. This base weight value can be predetermined based on user needs or experiments. For example, the first and second base weight values can each be 0.5. Then, a weight adjustment value is determined based on pixel parameters. This adjustment value is used to adjust both the first and second base weight values, ensuring that both remain at 1 throughout the adjustment process. The weight adjustment value is increased on the first base weight value to determine the weight value of the first local region, and decreased on the second base weight value to determine the weight value of the second local region. This allows for dynamic adjustment of the first and second local region weight values based on different application scenarios, resulting in an accurate fused image.
[0095] In one example of the present invention, when the area type is a sky or a vegetation area, the pixel parameters include pixel attributes; the weight adjustment value can be determined based on the pixel attributes. If color and brightness information are emphasized, the weight value of the first local area is increased and the weight value of the second local area is decreased; if thermal radiation information is emphasized, the weight value of the second local area is increased and the weight value of the first local area is decreased. The content and scene to be focused on can be determined according to the driver's selection. The weight adjustment value is set to a preset increment λ, which can be preset according to needs, and is not limited in this example. The value range of the increment λ is [-1, 1]. When λ = 1, the weight value of the second local area is increased and the weight value of the first local area is decreased; when λ = -1, the weight value of the first local area is increased and the weight value of the second local area is decreased. First local region weight value
[0096] That is, in the case of a sky region scene, the weight value of the second local region. First local region weight value Where, λ sky This refers to the weight adjustment value for the sky region scene.
[0097] In a vegetated area scenario, the weight value of the second local region First local region weight value Where, λ vegetation This represents the weight adjustment value for the vegetation area scenario.
[0098] In one example of the present invention, when the region type is a vehicle region, the pixel parameters include size and distance. The weights are adjusted for the vehicle region based on the camera distance and image size. First, the size S can be calculated using the camera parameters and image position. vehicle Distance D is calculated using camera parameters and image position.vehicle Then adjust the weight value. Where S threshold and D threshold These are the size threshold and the distance threshold. In the vehicle region scenario, the weight value of the second local region... First local region weight value
[0099] In one example of the present invention, when the area type is a pedestrian area, the pixel parameters include the temperature value corresponding to the pixel. An adjustment weight value is determined based on the temperature value, thereby accurately displaying the pedestrian's location and better protecting the safety of people and vehicles. The pedestrian temperature can be determined to be T. pedestrian The ambient temperature is T environment Weight adjustment value Where T threshold This is the temperature threshold. In pedestrian zone scenarios, the weight value of the second local region... First local region weight value
[0100] Step 204: Based on the region fusion weight, the associated first local region and second local region are fused for the target object to generate a fused image;
[0101] Based on the region fusion weights of each pair of associated first and second local regions, the first and second local regions are fused for the target object, and the regions where each target object is located are combined to obtain a fused image.
[0102] The step of fusing the associated first local region and second local region based on the region fusion weight to generate a fused image includes: determining a first region fusion pixel value by combining the first local region with the first local region weight value; determining a second region fusion pixel value by combining the second local region with the second local region weight value; and generating a fused image based on the target object by combining the first region fusion pixel value and the second region fusion pixel value.
[0103] For the same local region, the first local region can be combined with its weight value to obtain a first region fused pixel value, which represents the state of the first local region after fusion. Similarly, the second local region can be combined with its weight value to obtain a second region fused pixel value, which represents the state of the second local region after fusion. Then, based on the same target object, the first and second region fused pixel values are fused to obtain the image of that local region. Finally, all local regions are fused to generate a fused image.
[0104] Expressed as an expression: the pixel values of the merged image
[0105]
[0106] Step 205: Obtain image overlay information;
[0107] After fusing the images, image overlay information can be obtained.
[0108] Step 206: The image overlay information is overlaid onto the fused image to generate an output image.
[0109] Image overlay information is superimposed onto the fused image to generate an output image that assists the driver in confirming the external environment. For example, information such as the category, location (expressed as relative distance and angle from the vehicle), and trajectory of identified target objects can be overlaid onto the image. Intuitive graphical and textual representations are used, such as using different colored boxes to label different categories of target objects, displaying the object's name and distance value next to the box. When a target object is in a dangerous area (e.g., too close or its trajectory pointing towards the vehicle), a warning is given through flashing or color changing. This allows the driver to quickly and accurately understand the external environment through the output image.
[0110] This invention fuses visible light and infrared light images to provide drivers with bright and clear output images at night or in low light conditions, or in adverse weather conditions such as dense fog, heavy rain, or heavy snow, thus providing drivers with a relatively clear rear view.
[0111] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0112] Reference Figure 3 This invention also provides an electronic rearview mirror, including: a processor 301, a memory 302, and a computer program stored in the memory and capable of running on the processor. When executed by the processor, the computer program implements the steps of the electronic rearview mirror image processing method described above. The electronic rearview mirror image processing method includes:
[0113] Acquire visible light and infrared light images captured by the electronic rearview mirror;
[0114] The visible light image is divided to generate multiple first local regions, and the infrared light image is divided to generate multiple second local regions, wherein the first local regions and the second local regions are associated with the same target object;
[0115] Based on the pixel parameters of the associated first and second local regions, the region fusion weights are determined.
[0116] Based on the region fusion weights, the associated first and second local regions are fused for the target object to generate a fused image.
[0117] Optionally, the step of dividing the visible light image to generate multiple first local regions includes:
[0118] Identify the target object in the visible light image;
[0119] A first local region is determined based on the target object in the visible light image; and / or
[0120] The step of dividing the infrared image to generate multiple second local regions includes:
[0121] Identify the target object in the infrared image;
[0122] A second local region is determined based on the target object in the infrared image.
[0123] Optionally, the step of identifying the target object in the visible light image includes:
[0124] The target object in the visible light image is identified using a preset target recognition model; and / or,
[0125] The step of identifying the target object in the infrared image includes:
[0126] The target object in the infrared image is identified using a preset target recognition model;
[0127] The preset target recognition model is generated in the following manner:
[0128] The convolutional neural network is trained based on preset training samples, and the cross-entropy loss value for this training is determined.
[0129] Backpropagation is performed based on the cross-entropy loss value of the current training iteration to update the convolutional neural network;
[0130] The step of training the convolutional neural network based on the updated convolutional neural network using preset training samples and determining the cross-entropy loss value for the current training iteration continues until the cross-entropy loss value for the current training iteration is less than a preset loss threshold.
[0131] Optionally, the step of determining the first local region based on the target object in the visible light image includes:
[0132] Get the window size;
[0133] Centered on the target object in the visible light image, the image region corresponding to the window size is determined as the first local region; and / or,
[0134] The step of determining the second local region based on the target object in the infrared image includes:
[0135] Get the window size;
[0136] Centered on the target object in the infrared image, the image region corresponding to the window size is determined as the second local region.
[0137] Optionally, the step of determining the region fusion weight based on the pixel parameters of the associated first and second local regions includes:
[0138] Determine the region types of the associated first and second local regions;
[0139] The region fusion weight is determined based on the region type and the pixel parameters of the associated first and second local regions.
[0140] Optionally, the region fusion weight includes a first local region weight value and a second local region weight value; the step of fusing the associated first and second local regions with respect to the target object based on the region fusion weight to generate a fused image includes:
[0141] The first local region is determined by combining the first local region and its weight value.
[0142] The fused pixel value of the second region is determined by combining the weight value of the second local region with the weight value of the second local region.
[0143] Based on the target object, a fused image is generated by combining the fused pixel values of the first region and the fused pixel values of the second region.
[0144] Optionally, the step of determining the region fusion weight based on the region type, the pixel parameters of the associated first local region and second local region includes:
[0145] When the area type is a road or building area, the entropy value of the first local area and the entropy value of the second local area are determined to be the pixel parameter;
[0146] Determine the ratio of the entropy value of the first local region to the entropy value of the second local region;
[0147] The weight values of the first local region and the second local region are determined based on the ratio.
[0148] Optionally, the step of determining the region fusion weight based on the region type, the pixel parameters of the associated first local region and second local region includes:
[0149] If the area type is not a road or building area, determine the first basic weight value of the first local area and the second basic weight value of the second local area;
[0150] The weight adjustment value is determined based on pixel parameters;
[0151] The first local region weight value is determined by adding the weight adjustment value to the first basic weight value, and the second local region weight value is determined by deducting the weight adjustment value from the second basic weight value.
[0152] Optionally, if the region type is a sky or a vegetated region, the pixel parameters include pixel attributes; and / or;
[0153] When the region type is a vehicle region, the pixel parameters include size and distance; and / or;
[0154] When the area type is a pedestrian area, the pixel parameters include the temperature value corresponding to the pixel.
[0155] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0156] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0157] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0158] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0159] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0160] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0162] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0163] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0164] The above provides a detailed description of the electronic rearview mirror image processing method and the electronic rearview mirror provided by the present invention. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An image processing method for an electronic rearview mirror, characterized in that, include: Acquire visible light and infrared light images captured by the electronic rearview mirror; The visible light image is divided to generate multiple first local regions, and the infrared light image is divided to generate multiple second local regions, wherein the first local regions and the second local regions are associated with the same target object; Based on the pixel parameters of the associated first and second local regions, the region fusion weights are determined. Based on the region fusion weights, the associated first and second local regions are fused together with respect to the target object to generate a fused image.
2. The method according to claim 1, characterized in that, The step of dividing the visible light image to generate multiple first local regions includes: Identify the target object in the visible light image; A first local region is determined based on the target object in the visible light image; and / or The step of dividing the infrared image to generate multiple second local regions includes: Identify the target object in the infrared image; A second local region is determined based on the target object in the infrared image.
3. The method according to claim 2, characterized in that, The step of identifying the target object in the visible light image includes: The target object in the visible light image is identified using a preset target recognition model; and / or, The step of identifying the target object in the infrared image includes: The target object in the infrared image is identified using a preset target recognition model; The preset target recognition model is generated in the following manner: The convolutional neural network is trained based on preset training samples, and the cross-entropy loss value for this training is determined. Backpropagation is performed based on the cross-entropy loss value of the current training iteration to update the convolutional neural network; The step of training the convolutional neural network based on the updated convolutional neural network using preset training samples and determining the cross-entropy loss value for the current training iteration continues until the cross-entropy loss value for the current training iteration is less than a preset loss threshold.
4. The method according to claim 2, characterized in that, The step of determining the first local region based on the target object in the visible light image includes: Get the window size; Centered on the target object in the visible light image, the image region corresponding to the window size is determined as the first local region; and / or, The step of determining the second local region based on the target object in the infrared image includes: Get the window size; Centered on the target object in the infrared image, the image region corresponding to the window size is determined as the second local region.
5. The method according to claim 1, characterized in that, The step of determining the region fusion weight based on the pixel parameters of the associated first and second local regions includes: Determine the region types of the associated first and second local regions; The region fusion weight is determined based on the region type and the pixel parameters of the associated first and second local regions.
6. The method according to claim 5, characterized in that, The region fusion weight includes a first local region weight value and a second local region weight value; the step of fusing the associated first and second local regions for the target object based on the region fusion weight to generate a fused image includes: The first local region is determined by combining the first local region and its weight value. The fused pixel value of the second region is determined by combining the weight value of the second local region with the weight value of the second local region. Based on the target object, a fused image is generated by combining the fused pixel values of the first region and the fused pixel values of the second region.
7. The method according to claim 5, characterized in that, The step of determining the region fusion weight based on the region type and the pixel parameters of the associated first and second local regions includes: When the area type is a road or building area, the entropy value of the first local area and the entropy value of the second local area are determined to be the pixel parameter; Determine the ratio of the entropy value of the first local region to the entropy value of the second local region; The weight values of the first local region and the second local region are determined based on the ratio.
8. The method according to claim 5, characterized in that, The step of determining the region fusion weight based on the region type and the pixel parameters of the associated first and second local regions includes: If the area type is not a road or building area, determine the first basic weight value of the first local area and the second basic weight value of the second local area; The weight adjustment value is determined based on pixel parameters; The first local region weight value is determined by adding the weight adjustment value to the first basic weight value, and the second local region weight value is determined by deducting the weight adjustment value from the second basic weight value.
9. The method according to claim 8, characterized in that, When the region type is sky or vegetation region, the pixel parameters include pixel attributes; and / or; When the region type is a vehicle region, the pixel parameters include size and distance; and / or; When the area type is a pedestrian area, the pixel parameters include the temperature value corresponding to the pixel.
10. An electronic rearview mirror, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the electronic rearview mirror image processing method as described in any one of claims 1-9.