Region segmentation noise reduction method and system based on fusion imaging cooperation, terminal and medium

Through the coordinated noise reduction method between the event camera and the visible light camera, the area is segmented based on the motion confidence map and different strategies are used to solve the problem of poor background noise suppression effect of event cameras, the imaging effect of dynamic and static mixed scenes is improved, and the adaptive coordination between high temporal resolution and high spatial resolution is achieved.

CN120387950AActive Publication Date: 2025-07-29SHENZHEN UNIV

Patent Information

Application Number
CN202510741421.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-29
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

In the prior art, the background noise suppression effect of event cameras is poor, resulting in the frequent occurrence of "static area event noise residue" and "dynamic area texture blur" in dynamic and static mixed scenarios. Especially in intelligent driving and industrial precision detection, it is difficult to meet the needs of high temporal resolution and high spatial resolution.

Method used

By acquiring the event stream and visible light images under the same light field, the event stream confidence map and the visual frame confidence map are constructed, the motion confidence map is generated and divided into static, dynamic and transition areas. Different noise reduction strategies are used to process each area, including high-resolution reserved static areas, motion compensation dynamic areas and medium-weight processing transition areas, so as to achieve bidirectional coordinated noise reduction between the event stream and the visible light image.

Benefits of technology

The background noise suppression effect of event cameras is improved, the imaging quality in dynamic and static mixed scenes is improved, and the problems of excessive smoothing of static areas and noise residues in dynamic areas in traditional noise reduction methods are solved, and the adaptive balance between time resolution and spatial resolution is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387950A_ABST
    Figure CN120387950A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a fusion imaging collaborative region segmentation noise reduction method and system, a terminal and a medium, and the fusion imaging collaborative region segmentation noise reduction method comprises the steps: obtaining an event stream and a visible light image in the same light field; obtaining an event flow confidence map according to the event flow, and obtaining a visual frame confidence map according to the visible light image; obtaining a motion confidence map according to the event flow confidence map and the visual frame confidence map, and segmenting the motion confidence map into a static region, a dynamic region and a transition region; and performing noise reduction processing on the static region, the dynamic region and the transition region according to different strategies to obtain a target visible light image and a target event stream. According to the method, different noise reduction strategies can be adopted for processing the divided dynamic area, static area and transition area, the background noise suppression effect of the event camera is improved, the noise reduction effect is improved, and the imaging effect in a dynamic and static mixed scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and particularly to a regional segmentation and noise reduction method, system, terminal, and computer-readable storage medium for fused imaging collaboration. Background Art

[0002] In fields with extremely high requirements for spatio-temporal resolution such as intelligent driving and industrial precision inspection, the visual perception system needs to simultaneously meet the dual requirements of microsecond-level time resolution and sub-pixel-level spatial resolution. However, event cameras and traditional visual cameras (including RGB and grayscale cameras) show significant complementarity and contradictions in performance: Although event cameras have microsecond-level time resolution and can capture the polarity of brightness changes during high-speed movement in real time, they are extremely sensitive to ambient light changes. Sudden changes in light or minor fluctuations in static backgrounds are likely to trigger a large number of false trigger events, resulting in the proportion of noise events exceeding 30%, seriously interfering with the extraction of target information; Although traditional visual cameras can output high-resolution brightness / color images and accurately restore the details of static scenes, they are limited by the frame rate, have serious motion blur in high-speed motion scenes, and the noise increases significantly under low light conditions.

[0003] Existing technologies mostly focus on the one-way assistance of event cameras to traditional visual cameras, such as compensating for the motion blur of traditional visual cameras through event streams. However, the technical research on the feedback of traditional visual cameras to event camera noise reduction is not yet mature. There is a lack of effective means to suppress the background noise of event cameras. The existing fixed threshold filtering method has a noise filtering rate of less than 50% and is prone to deleting valid events; at the same time, the lack of a two-way collaboration mechanism leads to the dual problems of "residual event noise in static regions" and "texture blur in dynamic regions" in mixed static and dynamic scenes. In scenarios such as intelligent driving (such as highlighting obstacle detection) and industrial precision inspection (such as sub-micron defect identification), the need to capture high-speed moving time series and restore the high resolution of static backgrounds is urgent. However, due to the lack of a systematic solution for cross-modal collaborative noise reduction in existing technologies, it is difficult to cope with the problems of noise and blur superposition under complex lighting and high-speed movement.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main purpose of the present application is to provide a regional segmentation and noise reduction method, system, terminal, and medium for fused imaging collaboration, aiming to solve the problem that when event cameras assist traditional visual cameras unidirectionally in existing technologies, the background noise suppression effect of event cameras is poor, resulting in the problems of "residual event noise in static regions" and "texture blur in dynamic regions" often occurring in mixed static and dynamic scenes.

[0006] In the first aspect of the embodiments of the present application, a method for region segmentation and noise reduction with fused imaging collaboration is provided. The method for region segmentation and noise reduction with fused imaging collaboration includes the following steps: obtaining an event stream and a visible light image under the same light field; obtaining an event stream confidence map based on the event stream, and obtaining a visual frame confidence map based on the visible light image; obtaining a motion confidence map based on the event stream confidence map and the visual frame confidence map, and segmenting the motion confidence map into a static region, a dynamic region, and a transition region; performing noise reduction processing on the static region, the dynamic region, and the transition region respectively according to different strategies to obtain a target visible light image and a target event stream.

[0007] Optionally, in an embodiment of the present application, the obtaining an event stream confidence map based on the event stream specifically includes: determining the number of events of a pixel point and the global maximum number of events within a time window according to the event stream; obtaining a ratio of the number of events to the global maximum number of events; fusing the ratio of the number of events with the direction consistency to obtain an event stream confidence map; the calculation formula of the event stream confidence map is: ; where is the event stream confidence, is the number of events of the pixel point within the time window , is the global maximum number of events within the time window , is the time polarity of adjacent pixels.

[0008] Optionally, in an embodiment of the present application, the obtaining a visual frame confidence map based on the visible light image specifically includes: determining the gradient magnitude of the image and the global maximum gradient magnitude according to the visible light image; obtaining a ratio of the gradient magnitude to the global maximum gradient magnitude; fusing the ratio of the gradient magnitude with the motion reliability to obtain a visual frame confidence map; the calculation formula of the visual frame confidence map is: ; where is the visual frame confidence, is the gradient magnitude of the pixel point , is the global maximum gradient magnitude, is the backward error of the optical flow estimation.

[0009] Optionally, in an embodiment of the present application, obtaining the motion confidence map according to the event stream confidence map and the visual frame confidence map specifically includes: obtaining the dynamic feature weight of the event stream confidence map and the static feature weight of the visual frame confidence map; weighting and fusing the event stream confidence map and the visual frame confidence map according to the dynamic feature weight and the static feature weight to obtain the motion confidence map.

[0010] Optionally, in an embodiment of the present application, the fusion formula of the motion confidence map is: ; ; Wherein, is the motion confidence, is the event stream confidence, is the visual frame confidence, is the dynamic feature weight, is the static feature weight.

[0011] Optionally, in an embodiment of the present application, segmenting the motion confidence map into a static region, a dynamic region, and a transition region specifically includes: obtaining a static determination threshold and a dynamic determination threshold; taking the region in the motion confidence map where the motion confidence is less than the static determination threshold as the static region, taking the region where the motion confidence is greater than the dynamic determination threshold as the dynamic region, and taking the region where the motion confidence is greater than or equal to the static determination threshold and less than or equal to the dynamic determination threshold as the transition region.

[0012] Optionally, in an embodiment of the present application, the static region includes a spatially corresponding static image layer and a static event stream layer, the dynamic region includes a spatially corresponding dynamic image layer and a dynamic event stream layer, and the transition region includes a spatially corresponding transition image layer and a transition event stream layer; the step of performing noise reduction processing on the static region, the dynamic region, and the transition region respectively according to different strategies to obtain a target visible light image and a target event stream specifically includes: for the static region, retaining the high resolution of the static image layer to obtain an updated static image layer, and performing low-weight denoising processing on the static event stream layer to obtain an updated static event stream layer; for the dynamic region, performing motion compensation optimization on the dynamic image layer to obtain an updated dynamic image layer, and performing high-weight denoising processing on the dynamic event stream layer to obtain an updated dynamic event stream; for the transition region, performing detail gradient descent on the transition image layer to obtain an updated transition image layer, and performing medium-weight denoising processing on the transition event stream layer to obtain an updated transition event stream; according to the updated static image layer, the updated static event stream layer, the updated dynamic image layer, the updated dynamic event stream, the updated transition image layer, and the updated transition event stream, obtaining a spatially corresponding target visible light image and a target event stream.

[0013] A second aspect of the embodiments of the present application further provides a region segmentation and noise reduction system for fusion imaging collaboration, where the region segmentation and noise reduction system for fusion imaging collaboration includes: A data acquisition module, configured to acquire an event stream and a visible light image under the same light field; A confidence calculation module, configured to obtain an event stream confidence map according to the event stream and obtain a visual frame confidence map according to the visible light image; A region segmentation module, configured to obtain a motion confidence map according to the event stream confidence map and the visual frame confidence map, and segment the motion confidence map into a static region, a dynamic region, and a transition region; A partition noise reduction module, configured to perform noise reduction processing on the static region, the dynamic region, and the transition region respectively according to different strategies to obtain a target visible light image and a target event stream.

[0014] A third aspect of the embodiments of the present application further provides a terminal, where the terminal includes: a memory, a processor, and a region segmentation and noise reduction program for fusion imaging collaboration stored on the memory and executable on the processor, and when the region segmentation and noise reduction program for fusion imaging collaboration is executed by the processor, the steps of the region segmentation and noise reduction method for fusion imaging collaboration as described above are implemented.

[0015] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores a region segmentation and noise reduction program for fused imaging collaboration. When the region segmentation and noise reduction program for fused imaging collaboration is executed by a processor, the steps of the region segmentation and noise reduction method for fused imaging collaboration as described above are implemented.

[0016] Beneficial effects: The present application provides a region segmentation and noise reduction method, system, terminal and medium for fused imaging collaboration. By using spatially corresponding event streams and visible light images, and based on motion confidence, the field of view is divided into dynamic regions, static regions and transition regions. Different noise reduction strategies are adopted for different regions. The event stream compensates for the motion blur of the visible light image, and the multi-frame fusion of visible light suppresses the event stream noise. The threshold weights are adjusted through two-way feedback, thus solving the problems of excessive smoothing in the static region and noise residue in the dynamic region in traditional noise reduction. Furthermore, the background noise suppression effect of the event camera is improved, the noise reduction effect is improved, and the imaging effect in a mixed static and dynamic scene is enhanced. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic structural diagram of the fused imaging device of the present application; Figure 2 It is a flowchart of a preferred embodiment of the region segmentation and noise reduction method for fused imaging collaboration of the present application; Figure 3 It is a schematic diagram of the specific implementation steps of the entire execution process in a preferred embodiment of the region segmentation and noise reduction method for fused imaging collaboration of the present application; Figure 4 It is a structural diagram of a preferred embodiment of the region segmentation and noise reduction system for fused imaging collaboration of the present application; Figure 5 It is a structural diagram of a preferred embodiment of the terminal of the present application.

[0019] Description of the Reference Numerals: 100, data acquisition module; 200, confidence calculation module; 300, region segmentation module; 400, partition noise reduction module. Detailed Embodiments

[0020] To make the purpose, technical solutions and effects of this application clearer and more definite, the following will describe the technical solutions in the embodiments of this application clearly and completely in conjunction with the accompanying drawings in the embodiments of this application. The described embodiments are only possible technical implementations of this application, not all possible implementations. Based on the embodiments in this application, those skilled in the art can fully combine the embodiments of this application to obtain other embodiments without creative labor, and these embodiments are also within the protection scope of this application.

[0021] First, introduce the nouns involved in the embodiments of this application: RGB camera: Red-Green-Blue camera, which captures color images through a three-channel color sensor; CMOS / CCD: Complementary Metal Oxide Semiconductor / Charge Coupled Device, an image sensor type (CMOS has low power consumption and CCD has high imaging quality); ADC: Analog-to-Digital Converter, which converts analog electrical signals into digital image data; Event Camera: Based on Dynamic Vision Sensor (DVS) technology, it asynchronously records pixel brightness changes.

[0022] The following introduces traditional vision cameras and event cameras in related technologies: Traditional vision cameras include but are not limited to RGB cameras and grayscale cameras. Traditional vision cameras focus the light in the scene onto an image sensor (such as CMOS or CCD) through an optical lens. The sensor converts the optical signal into an electrical signal, and then the analog-to-digital converter (ADC) converts the electrical signal into digital image data. Traditional vision cameras usually capture image sequences at a fixed frame rate (such as 30 frames per second or 60 frames per second) to generate a continuous video stream. Technical advantages: High resolution, capable of capturing high-resolution static images or videos, suitable for scenes that require fine details; Mature algorithm support: There are a large number of mature image processing and computer vision algorithms (such as object detection, recognition, tracking, etc.); Strong versatility: Suitable for a variety of application scenarios, including monitoring, industrial inspection, autonomous driving, etc.; Low cost: The technology is mature, the price is relatively low, and it is easy to deploy on a large scale.

[0023] However, the dynamic range of traditional vision cameras is limited. In environments with too strong or too weak light, traditional cameras are prone to overexposure or underexposure problems; The temporal resolution of traditional vision cameras is low. Due to the fixed frame rate limitation, there is a certain time delay in traditional cameras, and they cannot capture fast-changing scenes in real time. When an object exceeds the camera frame rate, imaging quality problems such as blurred images will occur.

[0024] Event cameras are based on an asynchronous pixel-level light intensity change detection mechanism, which only records the pixel points with brightness changes and their timestamps. Its time resolution can reach the microsecond level, but it sacrifices spatial resolution. Technical advantages: high dynamic range, capable of capturing scenes with extremely large brightness changes, avoiding overexposure or underexposure; low power consumption, because only the pixels with brightness changes are recorded, the power consumption is much lower than that of traditional vision cameras; high time resolution, rapid response to fast-changing scenes, suitable for capturing high-speed dynamic events.

[0025] However, event cameras are for non-visual detection and can only output pixel change information, unable to directly detect the specific information of the target object; the resolution of event cameras is limited. Due to high-time-resolution detection, the resolution of event cameras is quite limited, and because they are sensitive to the light vector, the noise suppression effect is not good and the signal quality is poor.

[0026] In addition, due to the high cost of high-speed cameras, it is difficult to popularize while meeting high spatio-temporal resolution.

[0027] In related technologies, usually an event camera is used to assist an RGB camera (the event camera empowers visible light), and rarely does an RGB camera assist an event camera (it is impossible to feed back the event camera through a visible light camera), that is, the advantages of the two types of sensors are not fused through dynamic area segmentation and two-way collaboration technology, and the technical bottlenecks of spatio-temporal resolution collaboration and noise suppression still exist.

[0028] When an event camera performs one-way assistance to a traditional vision camera, the background noise suppression effect of the event camera is poor, resulting in the problems of "residual event noise in the static area" and "blurred texture in the dynamic area" in the static-dynamic mixed scene. In this application, through the spatially corresponding event stream and visible light image, based on the motion confidence, the field of view is divided into a dynamic area, a static area, and a transition area, and different noise reduction strategies are adopted for different areas. The event stream compensates for the motion blur of the visible light image (dynamic area), and the multi-frame fusion of visible light suppresses the event stream noise (static area), and two-way feedback adjusts the threshold weight (transition area), so as to solve the problems of excessive smoothing in the static area and noise residue in the dynamic area in traditional noise reduction, thereby improving the background noise suppression effect of the event camera, improving the noise reduction effect, and further enhancing the imaging effect in the static-dynamic mixed scene.

[0029] The technical solution of this application will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0030] Such as Figure 1As shown in the figure, an embodiment of the present application provides a fusion imaging device, which is applied to a region segmentation and noise reduction method for fusion imaging collaboration. The fusion imaging device includes a visible light camera, a beam splitting optical component (beam splitting prism), an event camera, an external trigger box, and a computer. The visible light camera and the event camera are coaxially arranged, and the visible light camera and the event camera are respectively arranged opposite to the beam splitting prism. The computer is respectively connected to the visible light camera and the event camera, and the external trigger box is respectively connected to the visible light camera and the event camera. By using a beam splitting prism to achieve the optical path coupling of the visible light camera and the event camera, the beam splitting prism evenly distributes the incident light to the two cameras, ensuring the consistency of their observation fields of view, that is, ensuring that the event stream collected by the time camera and the visible light image collected by the visible light camera are under the same light field. Through hardware synchronization and light field sharing, using a beam splitting prism and a synchronization trigger module, the event camera and the visible light camera synchronously capture the spatio-temporal alignment signals of the same light field, eliminating the cross-modal data registration error at the hardware level, and providing a spatio-temporally consistent input basis for subsequent collaborative noise reduction.

[0031] Aiming at the problems of redundant triggering of event streams in the prior art, serious interference of ambient stray light on event streams, failure to effectively combine the high time resolution and high spatial resolution characteristics of event cameras and traditional vision cameras, and excessive smoothing of static regions and noise residue in dynamic regions by traditional noise reduction methods, the region segmentation and noise reduction method for fusion imaging collaboration of the present application, that is, a dynamic region segmentation and noise reduction strategy based on bidirectional collaboration, generates a pixel-level motion confidence map by fusing the spatio-temporal derivative of the event stream and the frame gradient feature of the traditional vision camera, and divides the scene into static, dynamic, and transitional regions. In the static region, multi-frame fusion of the traditional vision camera is used to suppress noise, and the event stream is fed back to filter out the ambient light mis-triggered events; in the dynamic region, a short-time exposure image is reconstructed through the event stream and motion compensation fusion is performed with the traditional vision camera frame to suppress motion blur; in the transitional region, the wavelet threshold is dynamically adjusted based on the motion confidence to achieve a smooth transition of the noise reduction intensity. By constructing a bidirectional collaborative information interaction closed loop, adaptive noise reduction in complex scenes is achieved, effectively improving the accuracy and robustness of the visual perception system.

[0032] Compared with the related technology, the advantages of the present application are as follows: Cross-modal bidirectional collaboration mechanism: By constructing the optical path of the beam splitting prism, the light field synchronization of the event camera and the traditional vision camera is achieved, and a bidirectional information interaction closed loop is constructed, breaking through the traditional unidirectional assistance mode. Using the microsecond-level time resolution of the event camera to reconstruct a short-time exposure image to compensate for the motion blur of the traditional camera; generating a high signal-to-noise ratio background model through multi-frame fusion of visible light to filter out the ambient light noise in the event stream. The fusion weight is adjusted in real time according to the scene complexity to achieve an adaptive balance between time resolution and spatial resolution.

[0033] Three-level Region Segmentation and Adaptive Noise Reduction: Based on event flow density and motion confidence (fusing spatio-temporal gradient and optical flow error), the field of view is divided into dynamic, static, and transitional regions. The dynamic region is mainly driven by the event flow for motion compensation or amplification to eliminate traditional camera blur or improve the sensitivity to minute vibrations; the static region samples a dual verification mechanism (event trigger frequency + background model matching) to retain texture details; the transitional region dynamically adjusts the wavelet noise reduction threshold in combination with motion confidence. After wavelet transform, a threshold function and threshold are set to denoise the information, and then an inverse transform is used to obtain a fused image with dynamic adjustment, achieving a balance between noise reduction and edge preservation. It solves the problems of excessive smoothing in the static region and noise residue in the dynamic region in traditional noise reduction.

[0034] The region segmentation and noise reduction method for collaborative fusion imaging according to a preferred embodiment of the present application, as Figure 2 shown, the region segmentation and noise reduction method for collaborative fusion imaging includes the following steps: In step S101, an event flow and a visible light image under the same light field are acquired.

[0035] It should be noted that in the present application, through the hardware synchronization and spatio-temporal calibration of the event camera and the visible light camera, dynamic region segmentation is performed based on the motion confidence map, a two-way collaborative noise reduction strategy is designed for different regions, and real-time optimization of noise reduction parameters is achieved through a feedback adjustment module.

[0036] Specifically, during the data acquisition and synchronization process, the same light field is accurately distributed to the event camera and the visible light camera through a beam splitter prism to ensure that the scenes captured by both are completely synchronized in space and time; the event camera records pixel-level brightness changes with a microsecond-level time resolution. When the brightness change of a pixel exceeds a preset threshold, an event is triggered, recording the pixel coordinates (identifying the pixel position where the event is triggered), the timestamp (recording the exact time when the event occurs), and the polarity (indicating the direction of the brightness change, increase or decrease). A visible light image frame synchronized with the event flow is acquired (such as an RGB or grayscale image), and if the input is an RGB image, it is grayscale-converted to convert the color image into a grayscale image.

[0037] In step S102, an event flow confidence map is obtained according to the event flow, and a visual frame confidence map is obtained according to the visible light image.

[0038] In a possible implementation manner, according to the event flow, the number of events of a pixel point within a time window and the global maximum number of events are determined; according to the number of events and the global maximum number of events, the proportion of the number of events is obtained; the proportion of the number of events is fused with the direction consistency to obtain an event flow confidence map.

[0039] Specifically, a time window is set , which is used to count event activities within a short period of time; within each time window, count the number of triggered events for each pixel point P ; within the same time window, count the number of events for all pixel points and find the global maximum ; through the pixel high-frequency motion intensity formula , normalize the pixel-level event count to the range [0, 1]. It can be understood that the normalized value represents the relative activity intensity of the pixel within the local area. The closer the value is to 1, the more events the pixel has triggered in a short time, which may be the edge of a dynamic object or a high-speed motion area. The closer the value is to 0, the weaker the pixel activity, which may be a static background or a low-speed area. Then, perform motion direction consistency analysis to analyze the consistency of the temporal polarity (brightness change direction) of adjacent pixels. If adjacent pixels have the same-direction brightness change in a short time (such as getting brighter or darker simultaneously), they are considered to belong to the same motion trajectory; to calculate the temporal polarity of adjacent pixels , with a value range of [0, 1]. The closer the value is to 1, the higher the motion direction consistency of adjacent pixels. The closer the value is to 0, the stronger the direction randomness. Perform event stream confidence fusion to fuse the high-frequency motion intensity and direction consistency to generate event stream confidence.

[0040] The calculation formula for the event stream confidence map is: ; where, is the event stream confidence, is the number of events of pixel point within the time window , is the global maximum number of events (i.e., the total number of events) within the time window , is the temporal polarity of adjacent pixels, representing the proportion of the same-direction ratio, that is, the motion direction consistency, with a value range of [0, 1].

[0041] It should be noted that the 5ms time window is a compromise solution that balances time resolution and computational efficiency. If the window is too short, it may lead to statistical noise (too few events). If the window is too long, it may reduce the time resolution (unable to capture rapid changes). In practical applications, the window length can be dynamically adjusted according to the scenario. is used to normalize the event counts of different pixels to ensure the comparability of confidence levels between different pixels. For example, in a scenario with high global activity (such as a fast-moving object), is larger, and the normalized high-frequency motion intensity can suppress the difference in the absolute number of events and highlight the relative activity intensity. The number of events ​directly reflects the frequency of pixel-level brightness changes. In dynamic regions (such as object edges), the brightness changes frequently, and the number of events is high. Through normalization, the confidence of the event stream quantifies the relative intensity of pixel-level dynamic activities and provides key inputs for subsequent region segmentation and feedback regulation.

[0042] In one possible implementation, according to the visible light image, determine the gradient magnitude and the global maximum gradient magnitude of the image; according to the gradient magnitude and the global maximum gradient magnitude, obtain the gradient magnitude ratio; fuse the gradient magnitude ratio with the motion reliability to obtain the visual frame confidence map.

[0043] Specifically, calculate the gradient magnitude through the Sobel (Sobel Operator) operator (an edge detection algorithm in image processing). ; The gradient magnitude reflects the severity of local brightness changes in the image. The gradient magnitude represents the intensity of brightness changes at a pixel point. The larger the value, the more likely the pixel is located at the edge or in a region with rich texture (such as the object boundary). The smaller the value, the more likely the pixel is located in a smooth region (such as the background).

[0044] is the gradient component in the horizontal direction: ; is the gradient component in the vertical direction: ; where is the input image.

[0045] Traverse all pixel points in the entire image to find the maximum value of the gradient magnitude (the maximum value of the gradient magnitudes of all pixel points in the image); perform normalization , eliminate the influence of brightness differences between different images, convert the gradient magnitude to relative intensity, make the confidence comparable under different scenarios, and ensure that the confidence only reflects local relative changes through normalization.

[0046] Then further combine the optical flow estimation backward error , (Lucas-Kanade algorithm), which takes values in the range [0,1] after normalization, is used to measure the accuracy of optical flow estimation. The smaller the error, the higher the confidence, and quantifies the reliability of motion estimation; the final confidence takes into account both structural information (gradient, pixel point gradient magnitude) and motion information (optical flow, motion reliability) and provides a reliable basis for dynamic region segmentation.

[0047] The calculation formula for the visual frame confidence map is: ; Among them, is the visual frame confidence, is the pixel point gradient magnitude, is the global maximum gradient magnitude, is the optical flow estimation reverse error.

[0048] In step S103, according to the event flow confidence map and the visual frame confidence map, a motion confidence map is obtained, and the motion confidence map is segmented into a static region, a dynamic region, and a transition region.

[0049] In a possible implementation manner, a dynamic feature weight of the event flow confidence map and a static feature weight of the visual frame confidence map are obtained; according to the dynamic feature weight and the static feature weight, the event flow confidence map and the visual frame confidence map are weighted and fused to obtain a motion confidence map.

[0050] This application generates a pixel motion confidence map by fusing event flow features and traditional visual frame gradient information, and realizes block self-adaptive segmentation.

[0051] The pixel point motion confidence is the weighted fusion of the event flow motion confidence and the traditional visual frame motion confidence The fusion formula of the motion confidence map is: ; ; Among them, is the motion confidence, is the event flow confidence, is the visual frame confidence, is the dynamic feature weight, is the static feature weight. In this embodiment, is the fusion feature weight, which is determined by offline training to maximize the scene segmentation accuracy and can be adjusted according to specific scenarios.

[0052] It should be noted that the event flow dominates the dynamic characteristics. The event camera has a microsecond-level time resolution and is sensitive to fast motion. Therefore, a higher weight ( > ) is given to capture the dynamic region. The visual frame dominates the static characteristics. The visible light camera has a high spatial resolution and strong ability to capture static scene details, and a lower weight ( ) is given to suppress event flow noise. The weights and can be dynamically adjusted according to specific scenarios. For example, in a high-speed motion scenario, (such as = 0.7), to enhance the contribution of the event stream confidence; in low-light scenarios, it can reduce (such as = 0.5), to reduce the impact of event stream noise.

[0053] In a possible implementation, a static determination threshold and a dynamic determination threshold are obtained; the area in the motion confidence map where the motion confidence is less than the static determination threshold is used as a static area, the area where the motion confidence is greater than the dynamic determination threshold is used as a dynamic area, and the area where the motion confidence is greater than or equal to the static determination threshold and less than or equal to the dynamic determination threshold is used as a transition area.

[0054] Specifically, as Figure 3 shown, the three-layer area division rule. If is set as the static area, the motion is extremely low or there is no motion, and the original details of the visible light image are retained; if is set as the dynamic area, it is a high-frequency motion area, and the event stream features are preferentially retained; if is set as the transition area, the motion state is blurred, and spatio-temporal information is fused; among them, the static determination threshold , the dynamic determination threshold are determined by the maximum inter-class variance method, and the adaptive threshold can be adjusted through specific scenarios.

[0055] In step S104, noise reduction processing is performed on the static area, the dynamic area, and the transition area respectively according to different strategies to obtain a target visible light image and a target event stream.

[0056] It should be noted that in this application, the same light field is allocated to the event camera and the visible light camera through hardware synchronization. The visible light image and the event stream are simultaneously divided into dynamic areas, static areas, and transition areas and are spatially corresponding one by one. From the perspective of the visible light image, for the dynamic area part, due to the integration of the dynamic object features of the high-time-resolution event flow, the resolution is affected and the image details decline. However, different from the blurred imaging of dynamic objects before processing, the image compensates for the motion information and retains a certain motion trajectory, making the image features of dynamic objects appear; the pixels in the static area have no obvious change from the original visible light image; the pixels in the transition area show a decreasing trend in series, connecting the high-spatial-resolution features of the static area and the high-time-resolution features of the dynamic area, presenting an intermediate state. From the perspective of the event stream, the event stream in the dynamic area is given a high weight, that is, the degree of denoising is the lowest and the event stream feature points are retained most completely; the transition area is the second, and the static area is given a low weight and the degree of denoising is the highest to filter out stray light and non-attention events.

[0057] This application performs cross-modal collaborative dynamic partitioning and noise reduction of video streams and event streams, fuses the spatio-temporal density of the event stream and the visible light gradient features, combines a static background model (generated by multi-frame non-local means), and performs dynamic region segmentation. On this basis, the event stream compensates for the motion blur of the visible light camera (dynamic region), the visible light multi-frame fusion suppresses the noise of the event stream (static region), and the two-way feedback adjusts the threshold weights (transition region).

[0058] In a possible implementation, the static region includes a static image layer and a static event stream layer corresponding in space, the dynamic region includes a dynamic image layer and a dynamic event stream layer corresponding in space, and the transition region includes a transition image layer and a transition event stream layer corresponding in space. For the static region, the static image layer is retained with high resolution to obtain an updated static image layer, and the static event stream layer is denoised with low weight to obtain an updated static event stream layer; for the dynamic region, the dynamic image layer is optimized with motion compensation to obtain an updated dynamic image layer, and the dynamic event stream layer is denoised with high weight to obtain an updated dynamic event stream; for the transition region, the transition image layer undergoes detail gradient descent to obtain an updated transition image layer, and the transition event stream layer is denoised with medium weight to obtain an updated transition event stream; according to the updated static image layer, the updated static event stream layer, the updated dynamic image layer, the updated dynamic event stream, the updated transition image layer, and the updated transition event stream, the corresponding target visible light image and target event stream in space are obtained.

[0059] In the two-way feedback adjustment, the event stream drives the feedback. Utilizing the high temporal resolution advantage of the event camera, when the temporal density of the event stream in a certain region rises sharply within a short time, it indicates that there may be fast-moving objects in this region, that is, the above-mentioned dynamic region. By combining frame and event stream fusion, techniques such as motion amplification or motion compensation are used to optimize and denoise the fast-moving visible light images with a sharp decline in imaging effect, including but not limited to frame interpolation, contrast enhancement, and edge sharpening to increase the details of fast-moving objects. The visual quality drives the feedback. Utilizing the high spatial resolution advantage of the visible light camera, a static background model is constructed, the pixel differences between adjacent frames are compared, and the static region and the slowly changing region are distinguished. For the event stream information in this region, it is mostly noise events. By synchronizing the same light field to the event camera, weight distribution is performed on the event stream in the static region of the event information, including but not limited to reducing mis-triggered events by adjusting the internal event trigger threshold, to achieve noise reduction of the event stream information.

[0060] Specifically, in the dynamic region noise reduction strategy: From the perspective of the event stream, in the dynamic region, the event stream is given a high weight, which means the lowest degree of denoising, in order to retain the feature points of the event stream to the greatest extent; By retaining the high-frequency motion information of the event stream, the motion trajectory and details of dynamic objects can be accurately captured. From the perspective of the visible light image, for the motion blur that may occur in the dynamic region in the visible light image, technologies such as motion amplification, frame interpolation, contrast enhancement, and edge sharpening are used for optimization; By compensating for the motion information, the features of dynamic objects in the visible light image are made more prominent. Although the resolution may decrease slightly, the motion trajectory is retained. It can be understood that the dynamic region is the most active part of the scene. Therefore, the core of the noise reduction strategy is to retain the motion information. By retaining the event stream features with a high weight and combining the motion compensation technology of the visible light image, the resolution can be sacrificed to a certain extent in exchange for clearer features of dynamic objects.

[0061] In the static region noise reduction strategy: From the perspective of the event stream, in the static region, the event stream is given a low weight, which means the highest degree of denoising, in order to filter out stray light and non-attentive events; Noise suppression is achieved by adjusting parameters such as the event trigger threshold to reduce the mis-triggered event stream information in the static region. From the perspective of the visible light image, the static region maintains its original high spatial resolution in the visible light image without obvious changes; Utilizing the high spatial resolution advantage of the visible light camera, a static background model is constructed, and the static region and slowly changing regions are distinguished by comparing the pixel differences between adjacent frames. It can be understood that the static region is the part of the scene with the least or no motion. Therefore, the core of the noise reduction strategy is to suppress noise. By denoising the event stream information with a low weight and combining the high resolution and background modeling technology of the visible light image, the static region and the dynamic region can be accurately distinguished while maintaining the high-quality image of the static region.

[0062] In the transition region noise reduction strategy, from the perspective of the event stream, in the transition region, the event stream is given a medium weight to balance the temporal resolution and noise; It retains certain event stream features to capture possible motion information and reduces noise interference through denoising. From the perspective of the visible light image, the details in the transition region show a decreasing trend in a geometric progression in the visible light image to connect the high spatial resolution features of the static region and the high temporal resolution features of the dynamic region; By fusing spatio-temporal information, the transition region presents an intermediate state visually, avoiding the abruptness of modal switching. It can be understood that the transition region is the part of the scene with a blurred motion state, including the transition state from static to dynamic or from dynamic to static. Therefore, the noise reduction strategy needs to balance the temporal resolution and noise suppression. By processing the event stream information with a medium weight and combining the detail gradient descent technology of the visible light image, it is possible to reduce noise interference while retaining certain motion information and achieve a smooth transition of spatio-temporal information.

[0063] In another embodiment of the present application, the same optical field is synchronously allocated to the event camera and the visible light camera through hardware, and image processing is performed through cross-modal collaborative dynamic partitioning and noise reduction of the video stream-event stream, so as to achieve two-way noise reduction of the event stream-video stream.

[0064] In another embodiment of the present application, an attention mechanism is added to adjust the weight of the event stream to adjust the event trigger threshold.

[0065] Next, a region segmentation and noise reduction system for fusion imaging collaboration according to an embodiment of the present application will be described with reference to the accompanying drawings.

[0066] Figure 4 It is a structural diagram of a region segmentation and noise reduction system for fusion imaging collaboration according to an embodiment of the present application.

[0067] As Figure 4 shown, the region segmentation and noise reduction system for fusion imaging collaboration includes: a data acquisition module 100, a confidence calculation module 200, a region segmentation module 300, and a partition noise reduction module 400.

[0068] Specifically, the data acquisition module 100 is configured to acquire an event stream and a visible light image under the same optical field; The confidence calculation module 200 is configured to obtain an event stream confidence map according to the event stream and obtain a visual frame confidence map according to the visible light image; The region segmentation module 300 is configured to obtain a motion confidence map according to the event stream confidence map and the visual frame confidence map, and segment the motion confidence map into a static region, a dynamic region, and a transition region; The partition noise reduction module 400 is configured to perform noise reduction processing on the static region, the dynamic region, and the transition region respectively according to different strategies to obtain a target visible light image and a target event stream.

[0069] Figure 5 It is a structural diagram of a terminal provided by an embodiment of the present application. The terminal may include: A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.

[0070] When the processor 502 executes the program, it implements the region segmentation and noise reduction method for fusion imaging collaboration provided in the above embodiment.

[0071] Furthermore, the terminal further includes: A communication interface 503 for communication between the memory 501 and the processor 502.

[0072] The memory 501 is used to store a computer program executable on the processor 502.

[0073] The memory 501 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.

[0074] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, an Extended Industry Standard Architecture (EIS) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 5 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0075] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.

[0076] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0077] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned region segmentation and noise reduction method for fusion imaging collaboration is implemented.

[0078] An embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the region segmentation and noise reduction method for fusion imaging collaboration provided in any of the embodiments corresponding to the present application Figure 2 as described in the embodiments.

[0079] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0080] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0081] Any process or method description shown in a flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application pertain.

[0082] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable storage media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0083] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0084] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0085] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0086] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

[0087] It should be understood that the application of the present application is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present application.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A regional segmentation and noise reduction method for fusion imaging collaboration, characterized in that The regional segmentation and noise reduction method for fusion imaging collaboration includes: Obtain an event stream and a visible light image under the same light field; Obtain an event stream confidence map based on the event stream, and obtain a visual frame confidence map based on the visible light image; Obtain a motion confidence map based on the event stream confidence map and the visual frame confidence map, and segment the motion confidence map into a static region, a dynamic region, and a transition region; Perform noise reduction processing on the static region, the dynamic region, and the transition region respectively according to different strategies to obtain a target visible light image and a target event stream.

2. The regional segmentation noise reduction method for fusion imaging collaboration according to claim 1, wherein The specific steps of obtaining the event stream confidence map according to the event stream include: Determine the number of events of pixel points within a time window and the global maximum number of events according to the event stream; Obtain the proportion of the number of events based on the number of events and the global maximum number of events; Fuse the proportion of the number of events with the direction consistency to obtain an event stream confidence map; The calculation formula of the event stream confidence map is: ; Among them, is the event stream confidence, is the number of events of pixel within the time window , is the global maximum number of events within the time window , is the adjacent pixel time polarity.

3. The regional segmentation and noise reduction method for fusion imaging collaboration according to claim 1, wherein The specific steps of obtaining the visual frame confidence map according to the visible light image include: Determine the gradient magnitude of the image and the global maximum gradient magnitude according to the visible light image; Obtain the proportion of the gradient magnitude based on the gradient magnitude and the global maximum gradient magnitude; Fuse the proportion of the gradient magnitude with the motion reliability to obtain a visual frame confidence map; The calculation formula of the visual frame confidence map is: ; Among them, is the visual frame confidence, is the pixel gradient magnitude, is the global maximum gradient magnitude, is the optical flow estimation backward error.

4. The regional segmentation and noise reduction method for fusion imaging collaboration according to claim 1, wherein The specific steps of obtaining the motion confidence map according to the event stream confidence map and the visual frame confidence map include: Obtain the dynamic feature weight of the event stream confidence map and the static feature weight of the visual frame confidence map; Perform weighted fusion on the event stream confidence map and the visual frame confidence map according to the dynamic feature weight and the static feature weight to obtain a motion confidence map.

5. The regional segmentation and noise reduction method for fusion imaging collaboration according to claim 4, wherein The fusion formula of the motion confidence map is: ; ; Among them, is the motion confidence, is the event stream confidence, is the visual frame confidence, is the dynamic feature weight, is the static feature weight.

6. The regional segmentation noise reduction method for fusion imaging collaboration according to claim 4, characterized in that The specific steps of segmenting the motion confidence map into a static region, a dynamic region, and a transition region include: Obtain a static determination threshold and a dynamic determination threshold; Regard the region where the motion confidence in the motion confidence map is less than the static determination threshold as the static region, regard the region where the motion confidence is greater than the dynamic determination threshold as the dynamic region, and regard the region where the motion confidence is greater than or equal to the static determination threshold and less than or equal to the dynamic determination threshold as the transition region.

7. The regional segmentation noise reduction method for fusion imaging collaboration according to claim 6, wherein The static region includes a static image layer and a static event stream layer corresponding in space, the dynamic region includes a dynamic image layer and a dynamic event stream layer corresponding in space, and the transition region includes a transition image layer and a transition event stream layer corresponding in space; The specific steps of performing noise reduction processing on the static region, the dynamic region, and the transition region respectively according to different strategies to obtain a target visible light image and a target event stream include: For the static region, retain the high resolution of the static image layer to obtain an updated static image layer, and perform low-weight denoising processing on the static event stream layer to obtain an updated static event stream layer; For the dynamic region, perform motion compensation optimization on the dynamic image layer to obtain an updated dynamic image layer, and perform high-weight denoising processing on the dynamic event stream layer to obtain an updated dynamic event stream; For the transition region, perform detail gradient descent on the transition image layer to obtain an updated transition image layer, and perform medium-weight denoising processing on the transition event stream layer to obtain an updated transition event stream; According to the updated static image layer, the updated static event stream layer, the updated dynamic image layer, the updated dynamic event stream, the updated transition image layer, and the updated transition event stream, obtain the corresponding target visible light image and target event stream in space.

8. A regional segmentation and noise reduction system for fusion imaging collaboration, characterized in that, The region segmentation and denoising system for fusion imaging collaboration includes: A data acquisition module for acquiring an event stream and a visible light image under the same light field; A confidence calculation module for obtaining an event stream confidence map according to the event stream and obtaining a visual frame confidence map according to the visible light image; A region segmentation module for obtaining a motion confidence map according to the event stream confidence map and the visual frame confidence map, and segmenting the motion confidence map into a static region, a dynamic region, and a transition region; A partition denoising module for performing denoising processing on the static region, the dynamic region, and the transition region respectively according to different strategies to obtain a target visible light image and a target event stream.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a region segmentation and denoising program for fusion imaging collaboration stored on the memory and executable on the processor. When the region segmentation and denoising program for fusion imaging collaboration is executed by the processor, the steps of the region segmentation and denoising method for fusion imaging collaboration according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a region segmentation and denoising program for fusion imaging collaboration. When the region segmentation and denoising program for fusion imaging collaboration is executed by a processor, the steps of the region segmentation and denoising method for fusion imaging collaboration according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Space-time matching method for event camera and traditional optical camera

    CN114463399A

  • Noise reduction processing apparatus, method, program, and camera apparatus

    JP2007235319A

Cited By

  • End-to-end real-time small target unmanned aerial vehicle detection method based on event camera

    CN120564091A