Method and system for improving night definition of streaming media rearview mirror
By receiving visible light and near-infrared dual-channel video streams, combining vehicle IMU data and ambient light intensity for spatiotemporal layered noise reduction and multispectral adaptive fusion, contrast and detail enhancement, and using MEMS tunable filter arrays and lightweight GAN models for image quality optimization, the problems of optical component temperature drift and response lag in improving the nighttime clarity of streaming media rearview mirrors have been solved, achieving efficient improvement in nighttime image clarity.
Patent Information
- Application Number
- CN202510948393.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, streaming media rearview mirrors suffer from several drawbacks in improving nighttime clarity, including large temperature drift and low light efficiency of optical components, neglect of higher-order physical quantities in motion compensation, failure to correlate environmental physical field characteristics in multispectral fusion, excessive computational load of deep models, reliance on trial-and-error learning for image quality optimization, and delayed response to sudden scenes.
By receiving visible light-near infrared dual-channel video streams, combining vehicle-mounted IMU data and ambient light intensity, spatiotemporal layered noise reduction is performed, multispectral adaptive fusion and contrast-detail enhancement are carried out, and image quality is optimized using MEMS tunable filter array and lightweight GAN model. A state space and reward model are established for real-time image quality optimization.
It reduces temperature drift of optical components, improves image edge sharpness, reduces false positive rate at night, reduces overexposure rate in optical flow scenes and artifacts of glass reflection in rainy weather, and achieves real-time image quality improvement.
Smart Images

Figure CN120912486A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a kind of night definition enhancement method and system for realizing streaming rearview mirror, belong to image processing technical field. BACKGROUND
[0002] Streaming rearview mirror night definition enhancement technology refers to the detail visibility, dynamic range and real-time of vehicle-mounted rearview video are enhanced in low-illumination environment by multi-spectral imaging, motion compensation, image fusion and other methods, and its core goal is to solve the problem of blurred vision caused by insufficient lighting, vehicle jolt and sudden change of glare in night driving scene, and to improve driving safety.
[0003] Currently, the light loss of traditional polarization light splitting device is high, and the temperature drift leads to high spectral crosstalk rate at-40℃, which seriously restricts the basic imaging quality. The motion estimation based on inter-frame difference does not model the influence of jerk, resulting in dynamic blur residual rate exceeding the normal range in high-speed turning scene. The fixed weight multi-spectral fusion has abnormal overexposure rate in tunnel exit and other sudden change scenes. It lacks physical correlation optimization of optical flow field and information entropy. The parameter quantity of traditional GAN model exceeds 2 million, the delay of vehicle-mounted deployment exceeds 50 ms, and the rain glass reflection artifact suppression rate is low. Reinforcement learning needs to converge through 10,000 trials and errors. The image quality recovery time in glare scene is more than 2.3 seconds, which cannot meet the real-time safety decision.
[0004] Therefore, the optical assembly of the prior art has large temperature drift and low light efficiency, resulting in unreliable raw data quality. Motion compensation ignores high-order physical quantities. Multi-spectral fusion is not related to environmental physical field characteristics. Deep model calculation load is too high. Image quality optimization relies on trial-and-error learning. The response to sudden scenes is lagging. SUMMARY
[0005] The present application provides a kind of night definition enhancement method and system for realizing streaming rearview mirror, which is mainly aimed at reducing the defects such as large temperature drift and low light efficiency of optical assembly, resulting in unreliable raw data quality, motion compensation ignoring high-order physical quantities, multi-spectral fusion not correlating environmental physical field characteristics, deep model calculation load being too high, image quality optimization relying on trial-and-error learning, and response to sudden scenes being lagging.
[0006] To achieve the above purpose, the present application provides a kind of night definition enhancement method for realizing streaming rearview mirror, which comprises: receiving the night video stream of visible light-near infrared dual channel of rearview mirror, collecting vehicle-mounted IMU data and ambient light intensity of vehicle; combining the vehicle-mounted IMU data and the ambient light intensity, performing time-space domain layered noise reduction on the night video stream to obtain a basic noise reduction image; according to the ambient light intensity, performing multi-spectral adaptive fusion on the basic noise reduction image to obtain a high dynamic range image; performing contrast-detail enhancement on the high dynamic range image to obtain a detail-enhanced image; establishing a state space, an action space and a reward model of the detail-enhanced image, performing real-time picture quality tuning on the detail-enhanced image through the state space, the action space and the reward model to obtain a definition-enhanced image.
[0007] Optionally, the received visible light-near-infrared dual-channel night video stream of the rearview mirror comprises: obtaining a dual-spectrum sensor and a MEMS tunable filter array in the rearview mirror, wherein the MEMS tunable filter array is composed of a plurality of Fabry-Perot cavities; when light on the rear of the vehicle enters the MEMS tunable filter array, a preset voltage mode is applied to the MEMS tunable filter array to dynamically adjust the cavity spacing of the Fabry-Perot cavity based on the preset voltage mode to obtain an adjusted cavity spacing by using the following formula:
[0008] wherein, represents the adjusted cavity spacing, represents the initial cavity spacing of the Fabry-Perot cavity, , represents a temperature compensation coefficient, represents the preset voltage mode; After obtaining the adjusted cavity spacing, the dual-spectrum sensor detects the night video stream corresponding to the night light beam reflected by the MEMS tunable filter array.
[0009] Optionally, the vehicle-mounted IMU data and the ambient light intensity of the vehicle are collected, comprising: deploying a vehicle-mounted IMU at the center of the chassis of the vehicle; configuring an ambient light sensor on the top of the instrument panel of the vehicle; collecting vehicle-mounted IMU data and ambient light intensity of the vehicle through the vehicle-mounted IMU and the ambient light sensor, respectively.
[0010] Optionally, the vehicle-mounted IMU data and the ambient light intensity are combined to perform spatial-temporal domain layered noise reduction on the night video stream to obtain a basic noise-reduced image, comprising: calculating the inter-frame motion vector between adjacent frame video streams in the night video stream through the acceleration and the jerk in the vehicle-mounted IMU data; constructing an affine transformation matrix between adjacent frame video streams in the night video stream according to the inter-frame motion vector; performing motion adaptive inter-frame alignment on the night video stream by using the affine transformation matrix to obtain an aligned video stream; performing wavelet transform on each frame of the aligned video stream by using a Haar wavelet basis to generate low-frequency illumination information and high-frequency texture information of each frame of the video stream; generating a filtering radius of the low-frequency illumination information by using the ambient light intensity; setting a color domain standard deviation, a spatial domain standard deviation and a filtering window size of the low-frequency illumination information by using the filtering radius; performing bilateral filtering processing on the low-frequency illumination information according to the color domain standard deviation, the spatial domain standard deviation and the filtering window size to obtain bilateral filtering information; performing non-local mean denoising on the high-frequency texture information to obtain mean denoising information; generating an information recombination weight of the mean denoising information by using the ambient light intensity; recombining the bilateral filtering information and the mean denoising information by using the information recombination weight to obtain a basic denoising image.
[0011] Optionally, the step of performing multi-spectral adaptive fusion on the basic denoising image according to the ambient light intensity to obtain a high dynamic range image, comprises: performing multi-spectral separation on the basic denoising image to obtain a visible light image and a near-infrared image; calculating a dense optical flow field from the near-infrared image to the visible light image; extracting spectral entropy of the visible light image; calculating a dynamic fusion weight of the visible light image and the near-infrared image according to the dense optical flow field and the spectral entropy by using the following formula:
[0012] wherein, the dynamic fusion weight is represented by w, the ambient light intensity is represented by I, the spectral entropy is represented by H, the optical flow divergence of the dense optical flow field is represented by D, the coefficient for controlling exponential decay is represented by a, the small constant for preventing the denominator from being zero is represented by e, the time difference between the current frame and the reference frame is represented by t; performing multi-spectral adaptive fusion on the visible light image and the near-infrared image based on the dynamic fusion weight to obtain a high dynamic range image.
[0013] Optionally, before the step of performing contrast-detail enhancement on the high dynamic range image to obtain a detail-enhanced image, further comprising: Prepare an image training sample; Input the image training sample into a U-Net generator of an untrained lightweight GAN to output a predicted enhanced image by the U-Net generator; Input the predicted enhanced image and a real labeled image into a PatchGAN discriminator of the untrained lightweight GAN to obtain a discrimination probability matrix output by the PatchGAN discriminator; Calculate a binary cross-entropy loss value corresponding to the discrimination probability matrix and a prediction loss value corresponding to the predicted enhanced image; Complete model parameter training of the untrained lightweight GAN by the binary cross-entropy loss value and the prediction loss value to obtain a lightweight GAN; The prediction loss value includes a perception loss value and a texture loss value.
[0014] Optionally, the contrast-detail enhancement on the high dynamic range image to obtain a detail enhanced image includes: Perform contrast block processing on the high dynamic range image by using an adaptive histogram equalization method to obtain an intermediate enhanced image; Input the intermediate enhanced image into a U-Net generator of a lightweight GAN to output a detail enhanced image by the U-Net generator.
[0015] Optionally, the establishment of the state space, the action space and the reward model of the detail enhanced image includes: Obtain an ambient light intensity; measure a vehicle speed of a vehicle; Analyze a noise point level and an image sharpness of the detail enhanced image; Construct the ambient light intensity, the vehicle speed, the noise point level and the image sharpness into a state space; Set a radius adjustment parameter, a weight adjustment parameter and a contrast adjustment parameter of the detail enhanced image; Construct the radius adjustment parameter, the weight adjustment parameter and the contrast adjustment parameter into an action space; Construct a reward model of the detail enhanced image by using the following formula:
[0016] wherein, represents a reward model, represents a structural similarity index, represents a human visual score output by a lightweight CNN, and represents an instantaneous reward value.
[0017] Optionally, the real-time quality tuning of the detail-enhanced image through the state space, the action space and the reward model to obtain the definition-enhanced image comprises: obtaining a causal graph model corresponding to the state space; determining an optimal action from the action space through counterfactual query according to the causal graph model; using the reward model as a distillation supervision signal to train a lightweight CNN through the optimal action and the distillation supervision signal to obtain a trained CNN; adjusting parameters of the detail-enhanced image through the trained CNN to obtain the definition-enhanced image.
[0018] To solve the above problems, the application further provides a night definition enhancement system for implementing a streaming media rearview mirror, which comprises: a data acquisition module configured to receive a night video stream of a rearview mirror in a visible light-near infrared dual channel, and acquire vehicle-mounted IMU data and ambient light intensity of a vehicle; a video noise reduction module configured to combine the vehicle-mounted IMU data and the ambient light intensity to perform spatial-temporal domain layered noise reduction on the night video stream to obtain a basic noise-reduced image; an image fusion module configured to perform multi-spectral adaptive fusion on the basic noise-reduced image according to the ambient light intensity to obtain a high dynamic range image; an image enhancement module configured to perform contrast-detail enhancement on the high dynamic range image to obtain a detail-enhanced image; a quality tuning module configured to establish a state space, an action space and a reward model of the detail-enhanced image, and perform real-time quality tuning of the detail-enhanced image through the state space, the action space and the reward model to obtain a definition-enhanced image.
[0019] Compared with the problems described in the background art, the embodiment of the present application realizes high channel purity through MEMS filter voltage control, reduces traditional polarization light splitting light loss, further, the embodiment of the present application avoids the interference of the car window film through the dashboard light intensity detection to avoid the interference of the car window film, the embodiment of the present application makes the noise reduction PSNR improved through the sudden degree compensation high-speed over-bending dynamic blur, the embodiment of the present application makes the tunnel exit glare scene overexposure rate reduced through the optical flow-entropy fusion, the embodiment of the present application makes the rain day glass reflection artifact reduced through the texture loss function, further, the embodiment of the present application enhances the details in a short time through the lightweight GAN, so that the image edge sharpness is improved, the embodiment of the present application reduces the night misjudgment rate through the fusion of the professional driver's human visual score, further, the embodiment of the present application reduces the delay and power consumption through the real-time picture quality optimization of the detail enhancement image through the causal distillation CNN, so that the delay and power consumption are reduced. Therefore, the night clarity improvement method and system for realizing streaming rearview mirror provided by the embodiment of the present application can reduce the defects of optical component temperature drift, low light efficiency, resulting in unreliable original data quality, motion compensation ignoring high-order physical quantities, multispectral fusion not associating environmental physical field characteristics, deep model calculation load being too high, picture quality optimization relying on trial and error learning, and sudden scene response lag. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 The flowchart of the night clarity improvement method for realizing streaming rearview mirror provided by an embodiment of the present application is shown. Figure 2 The flowchart of the real-time picture quality optimization of the night clarity improvement method for realizing streaming rearview mirror provided by an embodiment of the present application is shown. Figure 3 The module diagram of the night clarity improvement system for realizing streaming rearview mirror provided by an embodiment of the present application is shown.
[0021] The purpose of the present application, functional characteristics and advantages will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0022] It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0023] The embodiment of the present application provides a method for improving night clarity of a streaming rearview mirror. The execution subject of the method for improving night clarity of the streaming rearview mirror includes but is not limited to at least one of electronic devices such as a server, a terminal and the like which can be configured to execute the method provided by the embodiment of the present application. In other words, the method for improving night clarity of the streaming rearview mirror can be executed by software or hardware installed in a terminal device or a server device. The server includes but is not limited to a single server, a server cluster, a cloud server or a cloud server cluster and the like.
[0024] Embodiment 1 Referring to Figure 1 FIG. 1 shows a flowchart of the method for improving night clarity of the streaming rearview mirror provided by an embodiment of the present application. In the embodiment, the method for improving night clarity of the streaming rearview mirror includes the following steps. S1, receiving a visible light-near infrared dual-channel night video stream of a rearview mirror, collecting vehicle-mounted IMU data and ambient light intensity of a vehicle.
[0025] The embodiment of the present application realizes high-channel purity through voltage control of a MEMS filter, and reduces traditional polarization light splitting light loss.
[0026] In an embodiment of the present application, the receiving of the visible light-near infrared dual-channel night video stream of the rearview mirror includes: acquiring a dual-spectrum sensor and a MEMS tunable filter array in the rearview mirror, wherein the MEMS tunable filter array is composed of a plurality of Fabry-Perot cavities; when light on the rear of the vehicle enters the MEMS tunable filter array, a preset voltage mode is applied to the MEMS tunable filter array, so as to dynamically adjust the cavity spacing of the Fabry-Perot cavity based on the preset voltage mode by using the following formula to obtain an adjusted cavity spacing.
[0027] wherein, represents the adjusted cavity spacing, represents an initial cavity spacing of the Fabry-Perot cavity, , represents a temperature compensation coefficient, represents the preset voltage mode; After obtaining the adjusted cavity spacing, the dual-spectrum sensor detects a night video stream corresponding to a night light beam reflected by the MEMS tunable filter array.
[0028] The dual-spectrum sensor refers to a sensor capable of simultaneously capturing visible light and infrared spectrum information, and can convert the received light signal into an electrical signal, which usually consists of two main parts, a visible light sensor (such as an RGB light sensing unit) and an infrared sensor (such as an IR light sensing unit), the MEMS tunable filter array refers to a micro-electro-mechanical system driven optical filter device, which changes the optical properties by voltage, for example, a Fabry-Perot cavity array, the single-pixel size is 3μm*3μm, the night video stream refers to the video data collected at night or under low light conditions, which contains visible light and near-infrared spectrum information, such a video stream is captured by a dual-spectrum sensor, which can provide more comprehensive environmental information, and is commonly used to improve the clarity and visible distance of a vehicle-mounted imaging system at night or in bad weather.
[0029] For example, the preset voltage mode applied to the MEMS tunable filter array is, for example, a first voltage mode (0-2.5V) that reflects 450-650nm visible light to a visible light sensor, and a second voltage mode (2.5-5V) that reflects 800-940nm near-infrared light to an infrared sensor.
[0030] In the present application, the temperature compensation related data is as shown in Table 1:
[0031] Further, the present application embodiment can accurately capture bumps through the chassis IMU, and avoid interference of the window film through the instrument panel light intensity detection.
[0032] In an embodiment of the present application, the collection of vehicle-mounted IMU data and ambient light intensity of the vehicle comprises: deploying a vehicle-mounted IMU at the center of the chassis of the vehicle; configuring an ambient light sensor on the top of the instrument panel of the vehicle; collecting vehicle-mounted IMU data and ambient light intensity of the vehicle through the vehicle-mounted IMU and the ambient light sensor respectively.
[0033] The vehicle-mounted IMU (Inertial Measurement Unit) data refers to a series of information related to the motion state of the vehicle collected by the IMU sensor installed on the vehicle, which usually includes three-axis acceleration (i.e. the acceleration of the vehicle in front, back, left, right and up) and three-axis angular velocity (i.e. the rotation angular velocity of the vehicle around the front, back, left, right and up axes), etc. Through these data, the motion parameters of the vehicle can be calculated, such as acceleration change rate and speed change rate, etc. The ambient light intensity refers to the light brightness level of the environment around the vehicle, and the lux is used as a unit to represent the size of the light intensity. The ambient light sensor is a device for detecting the light intensity of the surrounding environment, which is usually based on the photoelectric effect, that is, when light shines on the surface of the sensor, a corresponding electrical signal will be generated, and the strength of the electrical signal is proportional to the light intensity.
[0034] S2, combine the vehicle IMU data and the ambient light intensity, and perform hierarchical noise reduction in the time-space domain on the night video stream to obtain a basic denoising image.
[0035] The embodiment of the application improves the noise reduction PSNR by compensating for high-speed turning dynamic blur.
[0036] In an embodiment of the application, the combination of the vehicle IMU data and the ambient light intensity, and the hierarchical noise reduction in the time-space domain on the night video stream to obtain a basic denoising image, comprises: calculating the inter-frame motion vector between adjacent frame video streams in the night video stream by the acceleration and the jerk in the vehicle IMU data; constructing an affine transformation matrix between adjacent frame video streams in the night video stream according to the inter-frame motion vector; performing motion adaptive inter-frame alignment on the night video stream using the affine transformation matrix to obtain an aligned video stream; performing wavelet transform on each frame of the aligned video stream using a Haar wavelet basis to generate low-frequency illumination information and high-frequency texture information of each frame of the video stream; generating a filter radius of the low-frequency illumination information using the ambient light intensity; setting the color domain standard deviation, the spatial domain standard deviation and the filter window size of the low-frequency illumination information using the filter radius; performing bilateral filtering processing on the low-frequency illumination information according to the color domain standard deviation, the spatial domain standard deviation and the filter window size to obtain bilateral filter information; performing non-local mean denoising on the high-frequency texture information to obtain mean denoising information; generating information reorganization weights of the mean denoising information using the ambient light intensity; reorganizing the bilateral filter information and the mean denoising information by the information reorganization weights to obtain a basic denoising image.
[0037] Optionally, the calculation of the inter-frame motion vector between adjacent frame video streams in the night video stream by the acceleration and the jerk in the vehicle IMU data is: the acceleration is defined as the rate of change of velocity with respect to time:
[0038] The jerk is calculated by acceleration difference:
[0039] The time difference between the current frame and the reference frame is , for example, (30fps video), in variable acceleration motion, the displacement is determined by the acceleration and the jerk:
[0040] wherein, is the initial velocity, is the initial acceleration, The inter-frame motion vector is calculated by the following formula :
[0041] wherein, represents the instantaneous velocity component of the previous frame, represents the instantaneous acceleration component of the previous frame, represents the jerk component of the current frame; Further, the affine transformation matrix between adjacent frames in the night video stream is constructed according to the inter-frame motion vector, which is:
[0042] wherein, , i.e. the angular change is determined by the integral of jerk over time; Further, the motion adaptive inter-frame alignment of the night video stream is performed by using the affine transformation matrix, and the aligned video stream is obtained, which is: the affine transformation matrix obtained by calculation is used to transform each frame in the night video stream, so that they are spatially aligned, thereby compensating for the inter-frame deviation caused by the movement of the vehicle or the camera, and finally obtaining the aligned video stream. Further, the low-frequency illumination information and high-frequency texture information of each frame of the aligned video stream are generated by using the Haar wavelet basis to perform wavelet transform on each frame of the aligned video stream, which is: two-dimensional Haar wavelet transform (standard DWT) is applied to each frame, and low-frequency illumination components (LL) and high-frequency texture components (LH / HL / HH) are separated. Further, the filter radius for generating the low-frequency illumination information by using the ambient light intensity is:
[0043] wherein, is the filter radius, is the ambient light intensity, the unit of which is pixel / lux; Further, the color domain standard deviation, the spatial domain standard deviation, and the filter window size of the low-frequency illumination information are set by using the filter radius, which is:
[0044] wherein, the filter window size is usually set to 2r+1, the unit of which is ; Further, the bilateral filtering information is obtained by performing bilateral filtering processing on the low-frequency illumination information according to the color domain standard deviation, the spatial domain standard deviation, and the filter window size, which is: the filter radius determines the range of the filter effect. When performing bilateral filtering, the spatial distance is determined by determines the weight of the pixel in the spatial position, and the color similarity is determined by determines the weight of the pixel in the color value, further, the high-frequency texture information is denoised by a non-local mean method to obtain mean denoising information, for example, a weighted average value of similar image blocks in a 7*7 window (standard NLM algorithm), further, the information reorganization weight generated by the ambient light intensity to the mean denoising information is:
[0045] wherein, lux (extremely dark environment), then (high-frequency texture), lux (normal night), then (focus on low-frequency light), , is a learnable parameter, the unit is lux; further, the base denoising image is obtained by reorganizing the bilateral filtering information and the mean denoising information through the information reorganization weight:
[0046] wherein, is a base denoising image.
[0047] S3, according to the ambient light intensity, the base denoising image is adaptively fused by multi-spectrum to obtain a high dynamic range image.
[0048] The embodiment of the application reduces the overexposure rate of the tunnel exit glare scene by optical flow-entropy fusion.
[0049] In an embodiment of the application, the multi-spectrum adaptive fusion of the base denoising image according to the ambient light intensity to obtain a high dynamic range image comprises: multi-spectrum separation of the base denoising image to obtain a visible light image and a near-infrared image; calculating a dense optical flow field of the near-infrared image to the visible light image; extracting the spectral entropy of the visible light image; according to the dense optical flow field and the spectral entropy, the dynamic fusion weight of the visible light image and the near-infrared image is calculated by the following formula:
[0050] wherein, indicates the dynamic fusion weight, indicates the ambient light intensity, indicates the spectral entropy, indicates the optical flow divergence of the dense optical flow field, indicates a coefficient for controlling exponential decay, A small constant representing the prevention of the denominator being zero, the time difference between the current frame and the reference frame is ; Based on the dynamic fusion weight, the visible light image and the near-infrared image are adaptively fused to obtain a high dynamic range image.
[0051] Illustratively, the multispectral separation of the base denoising image to obtain a visible light image and a near-infrared image can be achieved by using an optical filter or a digital filtering technology, allowing light of a specific waveband to pass through, thereby separating the visible light and near-infrared images. Further, the calculation of the dense optical flow field from the near-infrared image to the visible light image can be achieved by the Lucas-Kanade method, the Horn-Schunck method, etc. The extraction of the spectral entropy of the visible light image divides the image gray scale into multiple intervals, and the frequency of occurrence of each gray scale pixel is counted, thereby calculating the spectral entropy.
[0052] S4, contrast-detail enhancement is performed on the high dynamic range image to obtain a detail-enhanced image.
[0053] The embodiment of the application reduces the rain day glass reflection artifact through a texture loss function.
[0054] In an embodiment of the application, before the contrast-detail enhancement is performed on the high dynamic range image to obtain a detail-enhanced image, the method further comprises: preparing an image training sample; inputting the image training sample into a U-Net generator of an untrained lightweight GAN to output a predicted enhanced image through the U-Net generator; inputting the predicted enhanced image and a real labeled image into a PatchGAN discriminator of the untrained lightweight GAN to obtain a discrimination probability matrix output by the PatchGAN discriminator; calculating a binary cross-entropy loss value corresponding to the discrimination probability matrix and a prediction loss value corresponding to the predicted enhanced image; and training model parameters of the untrained lightweight GAN through the binary cross-entropy loss value and the prediction loss value to obtain a lightweight GAN, wherein the prediction loss value comprises a perception loss value and a texture loss value.
[0055] The lightweight GAN refers to an improved version of a generative adversarial network (GAN), which significantly reduces the model parameter quantity and computational complexity through algorithm optimization and architecture innovation, while maintaining or improving the quality of generated images, supports running on low-power devices (such as mobile terminals, embedded systems), and completes high-resolution image training within a few hours on a single GPU, and the generated image is close to the level of traditional GAN in detail and authenticity, and the lightweight GAN can be realized by deep separable convolution, attention mechanism and the like, for example, the deep separable convolution can decompose the standard convolution operation, and reduce more than 60%~80% parameter quantity, and the attention mechanism dynamically allocates computing resources to key feature areas to reduce data analysis redundancy, and the image training sample refers to a historical intermediate enhanced image, and the calculation formulas of the perceptual loss value and the texture loss value are as follows:
[0056] wherein, is a perceptual loss value, is a texture loss value, is an expected value, denotes a pre-trained feature extraction network, denotes a real image (high-resolution image), denotes a generator network, denotes an input image of the generator, and the Gram matrix is calculated and defined as:
[0057] wherein, denotes the height and width of a feature map, denotes the value of the feature map F at the channel c, position (h, w), and Similarly, F is an image output by the pre-trained feature extraction network; Further, the model parameter training of the untrained lightweight GAN through the binary cross-entropy loss value and the prediction loss value is as follows: the loss value (the weighted sum result between the perceptual loss value, the texture loss value and the binary cross-entropy loss value) is used to update the model parameters of the lightweight GAN through the back propagation algorithm, and the iteration is repeated until the model converges, and finally the optimized lightweight GAN is obtained through training.
[0058] Further, the embodiment of the present application enhances details in a short time through the lightweight GAN, so that the image edge sharpness is improved.
[0059] In an embodiment of the present application, the contrast-detail enhancement of the high dynamic range image to obtain a detail-enhanced image comprises: performing contrast block processing on the high dynamic range image by using an adaptive histogram equalization method to obtain an intermediate enhanced image; and inputting the intermediate enhanced image into a U-Net generator of a lightweight GAN to output a detail-enhanced image by the U-Net generator.
[0060] S5, a state space, an action space and a reward model of the detail-enhanced image are established, and the detail-enhanced image is real-time quality adjusted by the state space, the action space and the reward model to obtain a definition-enhanced image.
[0061] The embodiment of the present application reduces the misjudgment rate at night by fusing the human visual score of a professional driver.
[0062] In an embodiment of the present application, the establishment of the state space, the action space and the reward model of the detail-enhanced image comprises: acquiring an ambient light intensity; measuring a vehicle speed of a vehicle; analyzing a noise level and an image sharpness of the detail-enhanced image; constructing the ambient light intensity, the vehicle speed, the noise level and the image sharpness as a state space; setting a radius adjustment parameter, a weight adjustment parameter and a contrast adjustment parameter; constructing the radius adjustment parameter, the weight adjustment parameter and the contrast adjustment parameter as an action space; and constructing a reward model by using the following formula:
[0063] wherein, represents the reward model, represents a structural similarity index, represents a human visual score output by a lightweight CNN, and represents an instantaneous reward value.
[0064] The state space comprises the ambient light intensity, the vehicle speed, the noise level and the image sharpness, the radius adjustment parameter is used for controlling a local range of the filter radius, directly affects the balance between detail preservation and noise suppression, the weight adjustment parameter is used for controlling a local range of the , and the contrast adjustment parameter is used for controlling a difference intensity of bright and dark areas of an image, for example, the difference intensity can be controlled by the formula , is an original brightness value of a pixel point in the image, is a trainable parameter, and the action space is {Delta r, Delta beta, Delta alpha}, and it needs to be explained that the reward model and the action space are used for subsequent CNN training, and therefore the parameters of the reward model and the action space here are historical training parameters, and the state space is a current parameter.
[0065] Optionally, the analysis of the noise level and image sharpness of the detail-enhanced image is as follows: the noise level is evaluated by calculating the standard deviation of the image pixel values. The larger the standard deviation, the more noise there is in the image. The intensity and sharpness of the image edges are evaluated by calculating the gradient of the image (such as using the Sobel operator or the Canny operator). Areas with large gradient amplitudes usually indicate that the image is sharper.
[0066] Furthermore, in this embodiment of the invention, the detail-enhanced image is optimized in real time using a causal distillation CNN to reduce latency and power consumption.
[0067] In one embodiment of the present invention, the step of performing real-time image quality optimization on the detail-enhanced image through the state space, the action space, and the reward model to obtain a sharper image includes: obtaining a causal graph model corresponding to the state space; and determining the optimal action from the action space based on the causal graph model through a counterfactual query.
[0068] in, Indicates the optimal action. Indicates an action, This indicates events with a structural similarity index of 0.9. Indicates state, Indicates the state Next action At that time, the probability that SSIM is greater than 0.9 Represents a specific state instance, This represents a specific instance of an action. This indicates the counterfactual query process. Symbols indicating causal inference; The reward model is used as a distillation supervision signal to train a lightweight CNN using the optimal action and the distillation supervision signal, resulting in a trained CNN. The parameters of the detail-enhanced image are then adjusted in real time using the trained CNN to obtain an image with improved sharpness.
[0069] For example, the causal graph model corresponding to the state space is obtained as follows: the causal graph model represents the causal relationship between various factors and is generated from historical states (not the state space). For example, the causal relationship between ambient light intensity (cause) and noise level (effect), the causal relationship between vehicle speed (cause) and noise level (effect), and the causal relationship between noise level (cause) and image sharpness (effect). These causal relationships need to point to SSIM, such as the causal relationship between ambient light intensity (cause), noise level (effect), and SSIM, thereby providing a basis for probability calculation. The delivery logic is provided, such as using PC algorithm, Bayesian network, etc. to build a causal graph model of the history state, further, the optimal action is determined, for example: = {Dr=+0.2, Db=-0.1, Da=+0.3}, the reward model is taken as a distillation supervision signal, a lightweight CNN is trained by a historical training sample (a history optimal action) and the distillation supervision signal, and a trained CNN is obtained, for example: the reward model is taken as a "teacher signal", and the CNN is forced to learn the human visual quality evaluation standard (high SSIM and human score), the optimal action is an optimal parameter adjustment value obtained by a counterfactual query, and is taken as a training label to directly constrain the CNN output, the history state is taken as input data of the CNN, and a loss function required is: ||a- ||+||R'-R||, wherein a and R' are respectively an action and a reward value output by the CNN, after the CNN is optimized by a gradient optimization algorithm, etc., the numerical value of the current state space is input to the trained CNN, the trained CNN outputs an adjustment parameter, and each parameter after update is calculated according to the adjustment parameter by the following formula: rt=rt-1+Dr, bt=bt-1+Db, at=at-1+Da, the parameter after update is used to process a next frame of the current detail enhancement image (a newly collected original input frame), and a definition enhancement image is obtained.
[0070] The related data of real-time quality tuning are as shown in Table 2:
[0071] Referring to Figure 2 , a flowchart of real-time quality tuning for implementing the night definition enhancement method of the streaming media rearview mirror is provided in an embodiment of the present application. Figure 2 In the foregoing, the original frame t refers to the newly collected original input frame in "using the parameter after update to process a next frame of the current detail enhancement image (a newly collected original input frame)", the application of the new parameter refers to processing the image by rt=rt-1+Dr, bt=bt-1+Db, at=at-1+Da, the processing principle is similar to that of the steps S2 to S5, and details are not repeated here, and the waiting frame refers to a next original input frame of the original frame t which is not processed.
[0072] Compared with the problems described in the background art, the embodiment of the application realizes high channel purity through MEMS filter voltage control, reduces traditional polarization light splitting light loss, further, the embodiment of the application avoids vehicle window film interference through dashboard light intensity detection by precisely capturing bumps through a chassis IMU, the embodiment of the application makes dynamic blur of high-speed over-bending through jerk compensation, so that the noise reduction PSNR is improved, the embodiment of the application reduces the overexposure rate of the tunnel exit glare scene through optical flow-entropy fusion, the embodiment of the application reduces rain day glass reflection artifacts through a texture loss function, further, the embodiment of the application enhances details in a short time through a lightweight GAN, so that the image edge sharpness is improved, the embodiment of the application reduces the night misjudgment rate through fusion of professional driver's human visual score, further, the embodiment of the application reduces delay and power consumption through real-time picture quality optimization of the detail enhancement image through causal distillation CNN. Therefore, the night clarity improvement method and system for realizing streaming rearview mirror provided by the embodiment of the application can reduce the defects of large optical component temperature drift, low light efficiency, resulting in unreliable original data quality, motion compensation ignoring high-order physical quantities, multispectral fusion not associating environmental physical field characteristics, high depth model calculation load, picture quality optimization relying on trial-and-error learning, and sudden scene response lag.
[0073] Embodiment 2: As Figure 3 shown, it is a functional module diagram of a night clarity improvement system for realizing streaming rearview mirror.
[0074] The night clarity improvement system for realizing streaming rearview mirror 300 can be installed in an electronic device. According to the implemented functions, the night clarity improvement system for realizing streaming rearview mirror can include a data acquisition module 301, a video noise reduction module 302, an image fusion module 303, an image enhancement module 304, and a picture quality optimization module 305. The modules of the application can also be referred to as units, which refer to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0075] In the embodiment of the application, the functions of each module / unit are as follows: The data acquisition module 301 is used to receive a night video stream of a rearview mirror in a visible light-near infrared dual channel, and to acquire vehicle on-board IMU data and ambient light intensity; The video noise reduction module 302 is used to combine the on-board IMU data and the ambient light intensity to perform spatial domain layered noise reduction on the night video stream to obtain a basic noise reduction image; The image fusion module 303 is used to perform multispectral adaptive fusion on the basic noise reduction image according to the ambient light intensity to obtain a high dynamic range image; The image enhancement module 304 is configured to perform contrast-detail enhancement on the high dynamic range image to obtain a detail-enhanced image. The picture quality tuning module 305 is configured to establish a state space, an action space and a reward model of the detail-enhanced image, and perform real-time picture quality tuning on the detail-enhanced image through the state space, the action space and the reward model to obtain a definition-enhanced image.
[0076] In detail, the modules in the night definition enhancement system 300 for implementing the streaming rearview mirror are used in the same way as the technical means for implementing the night definition enhancement method for streaming rearview mirror in the above-mentioned Figure 1 The same technical effects can be achieved, and thus no further description is provided herein.
[0077] It is apparent for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application, and although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application.
Claims
1. A method for implementing night-time clarity enhancement for streaming rearview mirrors, characterized in that, The method comprises: receiving a visible light-near infrared dual-channel night video stream of a rearview mirror, collecting vehicle-mounted IMU data and ambient light intensity of a vehicle; combining the vehicle-mounted IMU data and the ambient light intensity, performing spatio-temporal domain hierarchical noise reduction on the night video stream to obtain a basic noise-reduced image; according to the ambient light intensity, performing multi-spectral adaptive fusion on the basic noise-reduced image to obtain a high dynamic range image; performing contrast-detail enhancement on the high dynamic range image to obtain a detail-enhanced image; establishing a state space, an action space and a reward model of the detail-enhanced image, and performing real-time image quality optimization on the detail-enhanced image through the state space, the action space and the reward model to obtain a definition-enhanced image.
2. The method for implementing night-time clarity enhancement for streaming mirrors of claim 1, wherein, The receiving of the visible light-near infrared dual-channel night video stream of the rearview mirror comprises: obtaining a dual-spectrum sensor and a MEMS tunable filter array in the rearview mirror, wherein the MEMS tunable filter array is composed of multiple Fabry-Perot cavities; when light on the rear of the vehicle enters the MEMS tunable filter array, a preset voltage mode is applied to the MEMS tunable filter array to dynamically adjust the cavity spacing of the Fabry-Perot cavity based on the preset voltage mode, and an adjusted cavity spacing is obtained; wherein, represents the adjusted cavity spacing, represents the cavity spacing of the initial Fabry-Perot cavity, , represents a temperature compensation coefficient, represents a preset voltage pattern; after obtaining the adjusted cavity spacing, the dual-spectrum sensor detects the night video stream corresponding to the night light beam reflected by the MEMS tunable filter array.
3. The method for implementing night-time clarity enhancement for streaming mirrors of claim 1, wherein, The collection of the vehicle-mounted IMU data and the ambient light intensity of the vehicle comprises: deploying a vehicle-mounted IMU at the center of the chassis of the vehicle; configuring an ambient light sensor on the top of the instrument panel of the vehicle; collecting the vehicle-mounted IMU data and the ambient light intensity of the vehicle through the vehicle-mounted IMU and the ambient light sensor respectively.
4. The method for implementing night-time clarity enhancement for streaming mirrors of claim 1, wherein, The combination of the vehicle-mounted IMU data and the ambient light intensity to perform spatio-temporal domain hierarchical noise reduction on the night video stream to obtain a basic noise-reduced image comprises: calculating the inter-frame motion vector between adjacent frame video streams in the night video stream through the acceleration and the jerk in the vehicle-mounted IMU data; constructing an affine transformation matrix between adjacent frame video streams in the night video stream according to the inter-frame motion vector; performing motion adaptive inter-frame alignment on the night video stream using the affine transformation matrix to obtain an aligned video stream; performing wavelet transform on each frame of the aligned video stream using a Haar wavelet basis to generate low-frequency illumination information and high-frequency texture information of each frame of the video stream; generating a filtering radius of the low-frequency illumination information using the ambient light intensity; setting the color domain standard deviation, the spatial domain standard deviation and the filtering window size of the low-frequency illumination information using the filtering radius; performing bilateral filtering processing on the low-frequency illumination information according to the color domain standard deviation, the spatial domain standard deviation and the filtering window size to obtain bilateral filtering information; performing non-local mean denoising on the high-frequency texture information to obtain mean denoising information; generating information reorganization weights of the mean denoising information using the ambient light intensity; Recombine the bilateral filtering information and the mean denoising information by the information recombination weight to obtain a basic denoising image.
5. The method for implementing night-time clarity enhancement for streaming mirrors of claim 1, wherein, The multi-spectral adaptive fusion of the basic denoising image according to the ambient light intensity to obtain a high dynamic range image comprises: Multi-spectral separation of the basic denoising image to obtain a visible light image and a near-infrared image; Calculating the dense optical flow field from the near-infrared image to the visible light image; Extracting the spectral entropy of the visible light image; According to the dense optical flow field and the spectral entropy, the dynamic fusion weight of the visible light image and the near-infrared image is calculated by the following formula: wherein, represents a dynamic fusion weight, represents an ambient light intensity, represents a spectral entropy, represents a light flow divergence of a dense optical flow field, represents a coefficient for controlling exponential decay, represents a small constant for preventing the denominator from being zero, and the time difference between the current frame and the reference frame is ; Based on the dynamic fusion weight, the multi-spectral adaptive fusion of the visible light image and the near-infrared image is carried out to obtain a high dynamic range image.
6. The method for implementing night-time clarity enhancement for streaming mirrors of claim 1, wherein, Before the contrast-detail enhancement of the high dynamic range image to obtain a detail-enhanced image, it further comprises: Prepare an image training sample; Input the image training sample into the U-Net generator of the untrained lightweight GAN to output a predicted enhanced image by the U-Net generator; Input the predicted enhanced image and the real labeled image into the PatchGAN discriminator of the untrained lightweight GAN to obtain the discrimination probability matrix output by the PatchGAN discriminator; Calculate the binary cross-entropy loss value corresponding to the discrimination probability matrix and the prediction loss value corresponding to the predicted enhanced image; Complete the model parameter training of the untrained lightweight GAN through the binary cross-entropy loss value and the prediction loss value to obtain a lightweight GAN. The prediction loss value includes a perception loss value and a texture loss value.
7. The method for implementing night-time clarity enhancement for streaming mirrors of claim 1, wherein, The contrast-detail enhancement of the high dynamic range image to obtain a detail-enhanced image comprises: Using adaptive histogram equalization method to perform contrast block processing on the high dynamic range image to obtain an intermediate enhanced image; Input the intermediate enhanced image into the U-Net generator of the lightweight GAN to output a detail-enhanced image by the U-Net generator.
8. The method for implementing night-time clarity enhancement for streaming mirrors of claim 1, wherein, The establishment of the state space, action space and reward model of the detail-enhanced image comprises: Obtain the ambient light intensity; Measure the vehicle speed of the vehicle; Analyze the noise level and image sharpness of the detail-enhanced image; The ambient light intensity, vehicle speed, noise level and image sharpness are constructed as a state space; Set the radius adjustment parameter, weight adjustment parameter and contrast adjustment parameter; The radius adjustment parameter, weight adjustment parameter and contrast adjustment parameter are constructed as an action space; The reward model is constructed by the following formula: wherein, represents a reward model, represents a structural similarity index, represents a human visual score of the lightweight CNN output, represents an immediate reward value.
9. The method for implementing night-time clarity enhancement for streaming mirrors of claim 1, wherein, The real-time picture quality optimization of the detail-enhanced image through the state space, action space and reward model to obtain a clarity-enhanced image comprises: Obtain the causal graph model corresponding to the state space; According to the causal graph model, determine the optimal action from the action space through counterfactual query; The reward model is used as a distillation supervision signal to train a lightweight CNN through the optimal action and the distillation supervision signal to obtain a trained CNN; The parameters of the detail enhancement image are adjusted in real time by the trained CNN to obtain a definition enhancement image.
10. A system for implementing night-time clarity enhancement of a streaming mirror, the system comprising: The system comprises: A data acquisition module is configured to receive a visible-near-infrared dual-channel night video stream of a rearview mirror, acquire vehicle-mounted IMU data and ambient light intensity of a vehicle; A video noise reduction module is configured to combine the vehicle-mounted IMU data and the ambient light intensity, perform spatial and temporal domain layered noise reduction on the night video stream, and obtain a basic denoised image; An image fusion module is configured to perform multispectral adaptive fusion on the basic denoised image according to the ambient light intensity, and obtain a high dynamic range image; An image enhancement module is configured to perform contrast-detail enhancement on the high dynamic range image, and obtain a detail enhancement image; A quality tuning module is configured to establish a state space, an action space and a reward model of the detail enhancement image, perform real-time quality tuning on the detail enhancement image through the state space, the action space and the reward model, and obtain a definition enhancement image.