Image processing method and device, computer device and readable storage medium

By combining multimodal image processing methods of RGB cameras and event cameras, and using feature fusion based on the spatial pixel correspondence between event data and grayscale data as well as the semantic relationship of visible light image data, the problem of low imaging quality of CMOS or CCD cameras in complex dynamic scenes is solved, and high-quality image restoration and enhancement are achieved.

CN122335564APending Publication Date: 2026-07-03ARCSOFT CORP LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ARCSOFT CORP LTD
Filing Date
2026-02-13
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing CMOS or CCD cameras have low image quality in complex dynamic scenes, making it difficult to capture high-quality images under high-speed motion and low-light conditions. Furthermore, multi-frame fusion solutions suffer from computational latency and storage overhead when deployed in real time on mobile devices.

Method used

A multimodal image processing method using RGB and event cameras is adopted. The first encoding fusion is performed by the spatial pixel correspondence between event data and grayscale data, and the second encoding fusion is performed by combining the pixel semantic relationship of visible light image data. This achieves multiple downsampling and upsampling of features, and feature fusion is performed using a multi-head mutual attention mechanism.

Benefits of technology

It improves the overall observation redundancy and spatiotemporal sampling density of the imaging system, solves the problem of low imaging quality in complex dynamic scenes, and achieves high-quality and robust image restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122335564A_ABST
    Figure CN122335564A_ABST
Patent Text Reader

Abstract

This application relates to an image processing method, apparatus, computer device, and readable storage medium. After acquiring visible light image features, event features, and grayscale features, the method fuses the event features and grayscale features based on the spatial pixel correspondence between the event data and grayscale data to obtain secondary fusion features. Then, based on the pixel semantic correspondence between the grayscale data and visible light image data, the method fuses the secondary fusion features and visible light image features to obtain primary fusion features. Finally, a fused feature image is obtained based on the primary fusion features. This method can improve the overall observation redundancy and spatiotemporal sampling density of the imaging system, solve the problem of low imaging quality in complex dynamic scenes, and improve the resolution of the fused feature image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method, apparatus, computer device, and readable storage medium. Background Technology

[0002] With the rapid development of sensing technology, imaging systems are gradually evolving towards high-speed motion, low-light conditions, and high-quality image restoration. Simultaneously, camera devices are increasingly facing challenges such as rapid and intense motion, low light, and complex dynamic environments. Therefore, imaging systems are required not only to have high spatial resolution to ensure texture detail but also higher temporal resolution to capture fast-moving information and achieve real-time processing on low-power hardware platforms. However, existing cameras based on Complementary Metal-Oxide-Semiconductor (CMOS) or Charge-Coupled Device (CCD) are susceptible to factors such as camera displacement, exposure time, insufficient scene lighting, and high-speed target motion during imaging, resulting in lower image quality.

[0003] Some studies have attempted to improve image quality through short-exposure multi-frame alignment and fusion. However, these methods are theoretically limited by discrete frame rate sampling, with a temporal resolution typically on the order of milliseconds, making it difficult to capture continuous changes during high-speed motion. Furthermore, multi-frame alignment is highly dependent on optical flow or feature matching accuracy, and is prone to registration errors and ghosting artifacts under conditions of rapid motion, occlusion, or sudden changes in illumination. In addition, multi-frame caching and complex alignment calculations significantly increase system storage overhead and computational latency, hindering real-time deployment on mobile devices or embedded low-power platforms.

[0004] There is currently no effective solution to the problem of low imaging quality in complex dynamic scenes in traditional technologies. Summary of the Invention

[0005] Therefore, it is necessary to provide an image processing method, apparatus, computer device, and readable storage medium that can improve the imaging quality in complex dynamic scenes in order to address the above-mentioned technical problems.

[0006] In a first aspect, this application provides an image processing method, the method comprising: Determine visible light image features based on visible light image data acquired by an RGB camera; Event features are determined based on event data acquired by the event camera; grayscale features are determined based on grayscale data acquired by the event camera; wherein, the RGB camera and the event camera constitute a dual-camera imaging system; Based on the spatial pixel correspondence between the event data and the grayscale data, the event features and the grayscale features are first encoded and fused to obtain secondary fused features; Based on the pixel semantic correspondence between the grayscale data and the visible light image data, the secondary fusion feature and the visible light image feature are fused using a second encoding method to obtain the primary fusion feature; A fusion feature image is obtained based on the main fusion features.

[0007] In some embodiments, after obtaining the main fusion features, the method further includes: The primary fusion feature and the secondary fusion feature are subjected to multiple downsampling and fusion processes to obtain the downsampled output feature. The downsampled output features are then subjected to corresponding upsampling processing to obtain the fused feature image.

[0008] In some embodiments, the step of performing multiple downsampling and fusion processes on the primary fusion feature and the secondary fusion feature to obtain the downsampled output feature includes: The secondary fusion features are downsampled to obtain the first downsampled features; the primary fusion features are downsampled to obtain the second downsampled features. The event features, the first downsampled features, and the second downsampled features are fused to obtain the downsampled output features.

[0009] In some embodiments, performing corresponding upsampling processing on the downsampled output features to obtain the fused feature image includes: Upsampling is performed based on the downsampled output features and the main fusion features to obtain the fusion feature image.

[0010] In some embodiments, the first coding fusion and / or the second coding fusion are implemented based on a multi-head mutual attention mechanism.

[0011] In some embodiments, the image processing method is implemented using a trained image processing model, which is trained based on sample images; the sample images include at least RGB sample data, and the method for obtaining the RGB sample data includes: Obtain the first original image data corresponding to the event training data, wherein the first original image data is RGB image data and the clarity of the first original image data meets the preset clarity condition. A transformation matrix is ​​constructed based on the image parameters of the first original image data and the camera parameters of the dual-camera imaging system; The first original image data is mapped using the transformation matrix to obtain RGB sample data suitable for the dual-camera imaging system.

[0012] In some embodiments, constructing the transformation matrix based on the image parameters of the first original image data and the camera parameters of the dual-camera imaging system includes: The principal point and focal length of the event camera are determined based on the image parameters. A random perturbation is generated, and the camera intrinsic parameter matrix of the dual-camera imaging system is determined based on the random perturbation, the camera principal point, and the camera focal length. The rotation matrix between the two cameras in the dual-camera imaging system is randomly determined; The transformation matrix is ​​determined based on the camera intrinsic parameter matrix and the rotation matrix.

[0013] In some embodiments, the image processing method is implemented using a trained image processing model, which is trained based on sample images; the sample images include at least blurred RGB data, and the method for obtaining the blurred RGB data includes: Acquire multiple second raw image data, wherein the exposure duration and resolution of the second raw image data meet the preset image acquisition conditions; Multiple original second image data are fused together to obtain the blurred RGB data.

[0014] In some embodiments, fusing multiple second original image data to obtain the blurred RGB data includes: Multiple second original image data are interpolated, and the interpolated multiple second original image data are fused to obtain the blurred RGB data.

[0015] In some embodiments, before determining event features based on event data acquired by the event camera, the method further includes: The event data is filtered based on the exposure parameters of the RGB camera and the event camera to determine the valid event data.

[0016] In some embodiments, the step of filtering the event data based on the exposure parameters of the RGB camera and the event camera to determine valid event data includes: Obtain the starting exposure time and the visible light image exposure duration when the RGB camera acquires the visible light image data; Based on the initial exposure time and the exposure duration of the visible light image, calculate the row exposure start time and row exposure end time for each row of pixels in the visible light image data; Obtain the exposure time of the event pixels when the event camera acquires the event data; The start and end times of the line exposure are compared with the exposure times of the event pixels, and the valid event data is determined based on the comparison results.

[0017] In some embodiments, determining event features based on event data acquired by the event camera includes: The superposition range of event data is determined based on the visible light image exposure duration when the RGB camera acquires the visible light image data; the superposition reference time of event data is determined based on the midpoint of the visible light image exposure duration. Obtain the exposure time of the event pixels when the event camera acquires the event data; Based on the superposition reference time, the event data of the event pixel exposure time within the superposition range are superimposed to obtain the event voxel; Feature extraction is performed on the event voxels to obtain the event features.

[0018] Secondly, this application provides an image processing apparatus, the apparatus comprising a dual-camera imaging system and a processor; The dual-camera imaging system includes an RGB camera and an event camera. The RGB camera acquires visible light image data, and the event camera acquires event data and grayscale data. The processor determines visible light image features based on the visible light image data; determines event features based on the event data; determines grayscale features based on the grayscale data; performs a first encoding fusion on the event features and the grayscale features according to the pixel correspondence between the event data and the grayscale data in space to obtain secondary fusion features; performs a second encoding fusion on the secondary fusion features and the visible light image features according to the pixel semantic correspondence between the grayscale data and the visible light image data to obtain primary fusion features; and obtains a fusion feature image based on the primary fusion features.

[0019] Thirdly, this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.

[0020] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0021] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described above.

[0022] The image processing method described above, after acquiring visible light image features, event features, and grayscale features, fuses the event features and grayscale features based on the spatial pixel correspondence between the event data and grayscale data to obtain secondary fusion features; then, based on the semantic correspondence between the pixel pairs between grayscale data and visible light image data, it fuses the secondary fusion features and visible light image features to obtain primary fusion features; finally, it obtains a fused feature image based on the primary fusion features. This method can improve the overall observation redundancy and spatiotemporal sampling density of the imaging system, solve the problem of low imaging quality in complex dynamic scenes, and improve the resolution of the fused feature image. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is an application environment diagram of an image processing method in one embodiment; Figure 2 This is a flowchart illustrating an image processing method in one embodiment; Figure 3 This is a schematic diagram of a multi-head mutual attention fusion module in one embodiment; Figure 4 This is a flowchart illustrating a method for determining the transformation matrix in one embodiment; Figure 5 This is a flowchart illustrating a method for filtering valid event data in one embodiment; Figure 6 This is a flowchart illustrating an image deblurring method in a preferred embodiment; Figure 7 This is a schematic diagram of the structure of an image processing model in a preferred embodiment; Figure 8a This is a schematic diagram of visible light image data in a preferred embodiment; Figure 8b This is a schematic diagram of grayscale data in a preferred embodiment; Figure 8c This is a schematic diagram of event data visualization in a preferred embodiment; Figure 8d This is a schematic diagram of the deblurred result in a preferred embodiment; Figure 9 This is a structural block diagram of an image processing device in one embodiment; Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0026] With the rapid development of mobile terminals, automotive vision systems, drone platforms, and robot perception technologies, imaging systems are gradually evolving towards high-speed motion, low-light conditions, and high-quality image restoration. In these application scenarios, camera devices often face rapid and intense motion, low light, and complex dynamic environments. Therefore, imaging systems are required not only to have high spatial resolution to ensure texture details but also higher temporal resolution to capture fast-moving information and to achieve real-time processing on low-power hardware platforms. However, existing CMOS or CCD cameras based on frame-based exposure mechanisms use a time-integration sampling model in their imaging principle. This model accumulates and averages the scene radiance at consecutive moments within the exposure time window. Therefore, when the scene or camera shifts during exposure, the same pixel will mix spatial information from multiple moments, resulting in significant motion blur. Furthermore, in low-light scenes, CMOS sensors are prone to generating significant image noise. To obtain stable image quality, it is common practice to increase the exposure time. However, for high-dynamic motion scenes, extending the exposure time further exacerbates the image blur.

[0027] Furthermore, due to limitations in hardware readout bandwidth, most commercial CMOS sensors employ a rolling shutter mechanism with line-by-line scanning readout. Different rows of pixels correspond to different exposure times, resulting in the coupling of spatial and temporal coordinates in the image. When the camera or target moves rapidly, the scene states sampled in different rows are inconsistent, introducing geometric distortions such as tilting, stretching, compression, and rolling shutter effects into the image. This type of degradation exhibits significant spatial inconsistency, with its blur kernel changing with position, no longer satisfying the translation invariance assumption in convolutional neural network models. Therefore, it is difficult to effectively model and compensate for using traditional image restoration algorithms or convolutional neural networks. In addition, for high-dynamic motion scenes, conventional CMOS cameras have limited acquisition frame rates, typically only 30 to 60 FPS, making it difficult to capture details of high-speed motion and providing limited reference information for single-frame image restoration.

[0028] To compensate for the lack of information in a single frame, some studies have attempted to achieve deblurring and reconstruction through short-exposure multi-frame alignment and fusion. This involves using redundant information from temporally adjacent frames to perform superposition and averaging to approximate a sharper image. However, these methods are theoretically limited by discrete frame rate sampling, with a temporal resolution typically on the order of milliseconds, making it difficult to capture continuous changes during high-speed motion. Furthermore, during multi-frame alignment, the image frame's accuracy is highly dependent on optical flow or feature matching, making it prone to registration errors and ghosting artifacts under conditions of rapid motion, occlusion, or sudden changes in illumination. In addition, multi-frame caching and complex alignment calculations significantly increase system storage overhead and computational latency, hindering real-time deployment on mobile devices or embedded low-power platforms. Therefore, multi-frame fusion schemes still have significant limitations in practical engineering applications.

[0029] Event cameras, which have emerged in recent years, employ an asynchronous triggering mechanism completely different from traditional frame-based sampling. They output event signals only when pixel brightness changes, thus achieving microsecond-level temporal resolution and extremely low data redundancy. Because they do not rely on time-integrated exposure, event cameras theoretically do not suffer from motion blur and possess higher dynamic range and lower power consumption, demonstrating unique advantages in high-speed motion sensing. However, event sensors only record the polarity and timestamp information of brightness changes, lacking absolute intensity and complete texture description. They show almost no response in static or low-texture areas, making it difficult to directly generate high-quality, visually appealing intensity images.

[0030] In summary, existing technologies have not yet provided a unified solution capable of simultaneously achieving clear imaging, accurate motion sensing, and real-time low-power operation in complex dynamic environments. To overcome inherent defects such as temporal integral blurring in frame-based cameras, rolling shutter geometric distortion, and brightness loss in event cameras, this application proposes a multimodal image processing method combining RGB cameras and event cameras. This method aims to construct a multi-sensor complementary imaging architecture by introducing clear observations and high-precision temporal information, thereby improving the overall observation redundancy and spatiotemporal sampling density of the imaging system. This transforms the originally ill-conditioned deblurring inverse problem into a fully constrained fusion estimation problem, achieving high-quality and robust image restoration and enhancement.

[0031] The image processing method provided in this application embodiment can be applied to, for example... Figure 1The application environment shown is illustrated. The dual-camera imaging system 10 includes an RGB camera 102 and an event camera 104. The RGB camera 102 acquires visible light image data; the event camera 104 acquires event data and grayscale data. The processor 11 communicates with both the RGB camera 102 and the event camera 104. A memory 12 stores data and can be integrated with the processor 11 in the terminal containing the dual-camera imaging system 10, or integrated on a server, or placed in the cloud or other network servers. The processor 11 determines visible light image features based on the visible light image data acquired by the RGB camera 102; determines event features based on the event data acquired by the event camera 104; determines grayscale features based on the grayscale data acquired by the event camera 104; performs a first encoding fusion of the event features and grayscale features according to the spatial pixel correspondence between the event data and grayscale data to obtain secondary fusion features; performs a second encoding fusion of the secondary fusion features and visible light image features according to the pixel semantic correspondence between the grayscale data and visible light image data to obtain primary fusion features; and obtains a fusion feature image based on the primary fusion features to improve image quality. The image processing method in this application can be applied to blurry image processing, artifact processing, etc. in long exposure, high dynamic range, or low light scenes.

[0032] The dual-camera imaging system 10 is applied to a terminal device, which can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and other devices that require image processing of the acquired images. For example, IoT devices can be smart TVs, smartphones, smart in-vehicle devices, etc.

[0033] In one embodiment, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating an image processing method in one embodiment. This embodiment uses the method applied to a terminal as an example for illustration. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the image processing method includes the following steps: Step S201: Determine visible light image features based on visible light image data acquired by the RGB camera; determine event features based on event data acquired by the event camera; determine grayscale features based on grayscale data acquired by the event camera.

[0034] The image processing method in this embodiment is applied to a dual-camera imaging system consisting of an RGB camera and an event camera. The event camera can be a vision sensor triggered by pixel-level brightness changes. By detecting changes in light intensity and outputting event data, it can capture dynamic information under conditions of high-speed motion and low light. For example, the event camera can employ a Dynamic Vision Sensor (DVS), an Asynchronous Time-based Image Sensor (ATIS), or a Dynamic and Active Vision Sensor (DAVIS), etc.

[0035] The dual-camera imaging system can be applied to mobile terminals such as cameras, mobile phones, and tablets, smart wearable devices (such as smartwatches, smart glasses, or AR, XR, and VR devices), or other imaging devices. This imaging device can be installed on various types of vehicles, such as passenger cars, SUVs, commercial vehicles such as trucks, tractors, and large buses, as well as special vehicles such as fire trucks and ambulances. The imaging device can also be installed on robots, specifically including but not limited to industrial robots equipped with robotic arms, AGVs, robot dogs, water robots, drones, service robots that provide convenience or assistance, humanoid robots, etc.

[0036] In this embodiment, a dual-camera imaging system acquires visible light image data, event data, and grayscale data. The RGB camera itself can be a single-camera camera or a dual-camera camera; it can be a high-resolution RGB rolling shutter camera or a global shutter camera. The event camera can simultaneously acquire event data and grayscale data; specifically, the event camera can be a low-resolution global shutter DAVIS camera. The DAVIS camera is capable of simultaneously capturing high-frame-rate event data and low-frame-rate grayscale data.

[0037] After acquiring event data, grayscale data, and RGB visible light image data, feature extraction can be performed simultaneously on the event data, grayscale data, and RGB visible light image data to obtain event features, grayscale features, and visible light image features, respectively. Alternatively, before extracting event features, the validity of the event data can be analyzed. After determining the valid event data, feature extraction can be performed based on the valid event data to obtain event features.

[0038] Step S202: Based on the spatial pixel correspondence between event data and grayscale data, perform a first encoding fusion on event features and grayscale features to obtain secondary fused features.

[0039] Step S203: Based on the pixel semantic correspondence between grayscale data and visible light image data, perform a second encoding fusion on the secondary fusion features and visible light image features to obtain the main fusion features; In this embodiment, during the fusion of event data, grayscale data, and visible light data, the event data and grayscale data are first fused to obtain secondary fusion features. Then, these secondary fusion features are fused with visible light image features to obtain the final fusion result, i.e., the primary fusion features. This is because, due to the different acquisition principles of event data and visible light image data, they cannot achieve a strict spatial correspondence. Furthermore, event data and visible light image data do not belong to the same data modality. Directly aligning and fusing event features and visible light image features is extremely difficult and the results are hard to guarantee.

[0040] In this embodiment, considering that the event data and grayscale data come from the same event camera, specifically, the grayscale data comes from the visual part of the event camera and the event data comes from the event part of the event camera, the event data and grayscale data are strictly aligned in the spatial dimension after being collected, and thus have a spatial pixel correspondence. Based on this, fusing the event features and grayscale features can save alignment costs and improve the effect of feature alignment fusion.

[0041] Meanwhile, visible light image data and grayscale data both belong to the image domain and are data under the same modality. Their pixels are semantically common. Therefore, after obtaining the secondary fusion features, the pixel semantic correspondence between grayscale data and visible light image data can be used as the matching benchmark between visible light image data and event data to achieve alignment and fusion between event data and visible light image data.

[0042] That is, in this embodiment, the pixel correspondence between grayscale data and event data in space is first utilized to fuse event features into grayscale features; then, taking advantage of the similarity between grayscale features and visible light image features in the data domain, based on the pixel semantic correspondence between the two, the secondary fusion features are aligned and fused into visible light image data, thereby realizing the alignment and fusion of event features into visible light image features, so as to improve the fusion quality and ensure the final image deblurring effect.

[0043] In this embodiment, the first encoding fusion and the second encoding fusion can be the same fusion method or different fusion methods. Specifically, the first encoding fusion and the second encoding fusion can be implemented based on neural networks, based on traditional machine learning models, or through simple data operations, such as concatenation, weighted summation, principal component analysis, etc., or through multi-scale transformations, such as Laplace transform or wavelet transform.

[0044] Step S204: Obtain the fusion feature image based on the main fusion features.

[0045] After obtaining the main fusion features, a fusion feature image can be obtained by reconstructing or generating based on these features. Specifically, fusion feature images can be obtained using methods such as decoders, neural network iterations, and conditional generation.

[0046] After acquiring event data and visible light image data, the event features need to be pixel-level aligned and fused into the visible light image features. Due to modal differences between event data and visible light image data, and the original spatial mismatch, common feature stitching and addition fusion methods are ineffective. This embodiment employs a two-stage cross-modal implicit fusion scheme. Through steps S201 to S204, after obtaining visible light image data, event data, and grayscale data, the dual-camera imaging system sequentially fuses the event data, grayscale data, and visible light image data based on the spatial pixel correspondence between grayscale data and event data, and the pixel semantic correspondence between grayscale data and visible light image data. This improves the overall observation redundancy and spatiotemporal sampling density of the imaging system, solves the problem of low imaging quality in complex dynamic scenes, and improves the clarity of the fused feature image, i.e., the final image.

[0047] In some embodiments, if the image data (including visible light image data, grayscale data, and event data) has a high resolution, the feature dimension will also increase, thus drastically increasing the computational load during image processing. Therefore, to achieve better feature fusion results, it is necessary to downsample the main fusion features.

[0048] Specifically, after obtaining the primary fusion features, multiple downsampling and fusion processes are performed on the primary and secondary fusion features to obtain downsampled output features. Then, the downsampled output features are subjected to corresponding upsampling processes to obtain a fused feature image. In this embodiment, the primary and secondary fusion features are downsampled separately, and then feature fusion is performed based on the multiple downsampled features to obtain downsampled output features. During the upsampling and downsampling processes, high-dimensional features have stronger expressive power, especially semantic feature expressive power; low-dimensional features generally represent detailed information such as texture. Therefore, using a layer-by-layer resolution reduction method can take both into account, and combined with upsampling, a result of the same size as the visible light image data can be recovered.

[0049] In some embodiments, obtaining the downsampled output features includes: downsampling the secondary fusion features to obtain a first downsampled feature; and downsampling the primary fusion features to obtain a second downsampled feature. The downsampling process in this embodiment can be implemented through sampling, interpolation, or a neural network model. The downsampling processes for the primary and secondary fusion features can be the same or different. After obtaining the first and second downsampled features, the event features, the first downsampled feature, and the second downsampled feature are fused to obtain the downsampled output features. This embodiment does not limit the fusion method; fusion can be based on neural networks, machine learning models, data computation, multi-scale transformation, etc.

[0050] Specifically, the method for fusing event features, first downsampling features, and second downsampling features to obtain downsampling output features can be as follows: based on the pixel correspondence between event data and grayscale data in space, the event features and the first downsampling features are fused to obtain the first fused feature; based on the pixel semantic correspondence between grayscale data and visible light image data, the first fused feature and the second downsampling feature are fused to obtain the downsampling output feature.

[0051] In this embodiment, event features are fused into the visible light image features during the downsampling process. This can preserve the temporal dimension features of the event data, which helps to remove blur and artifacts in the image later.

[0052] In some embodiments, performing corresponding upsampling processing on the downsampled output features to obtain a fused feature image includes: upsampling based on the downsampled output features and the main fusion features to obtain the fused feature image. Typically, downsampling image features can obtain high-dimensional, semantically rich features, but it loses shallow features such as texture details. In this embodiment, upsampling is performed jointly based on the downsampled output features and the main fusion features. While restoring image resolution, this not only preserves deeper features with stronger semantic expressive power, but also obtains richer information such as texture details expressed by relatively shallow features based on the main fusion features.

[0053] In some embodiments, the first encoding fusion and / or the second encoding fusion are implemented based on a multi-head mutual attention mechanism. Since this embodiment needs to process two types of features simultaneously (e.g., simultaneously processing event features and grayscale features, or simultaneously processing grayscale features and visible light image features), a mutual attention mechanism is adopted. Compared with the ordinary mutual attention mechanism, the multi-head mutual attention mechanism can obtain better feature representation capabilities.

[0054] Figure 3 This is a schematic diagram of a multi-head mutual attention fusion module according to an embodiment of this application, as shown below. Figure 3 As shown, event features and grayscale features are input into the first multi-head mutual attention module for encoding and fusion to obtain secondary fused features. Visible light image features and secondary fused features are input into the second multi-head mutual attention module for encoding and fusion to obtain primary fused features.

[0055] In some embodiments, the primary fusion feature can also be obtained through multiple iterations. Specifically, in the first iteration, the event feature and grayscale feature are encoded and fused through a first multi-head mutual attention module to obtain an aligned and fused secondary feature. The aligned and fused secondary feature and visible light image feature are encoded and fused through a second multi-head mutual attention module to obtain an aligned and fused primary feature. In the second and subsequent iterations, the grayscale feature is replaced by the aligned and fused secondary feature obtained in the previous iteration, and the secondary feature and event feature are encoded and fused again to obtain a new secondary feature; the visible light image feature is replaced by the aligned and fused primary feature obtained in the previous iteration, and the primary feature and secondary feature are fused and encoded. During the iteration, the event feature remains unchanged. After the iteration is completed, the secondary feature obtained in the last iteration can be recorded as the secondary fusion feature for downsampling calculation, and the primary feature obtained in the last iteration can be recorded as the primary fusion feature for downsampling calculation.

[0056] It should be noted that the image processing methods in the embodiments of this application can be implemented using a neural network model. For example, based on the acquired sample images, an image processing model is trained to obtain a trained image processing model, which includes a feature extraction module, a feature fusion module, and an image output module. The feature extraction module extracts features from the acquired visible light image data to obtain visible light image features, extracts features from the acquired event data to obtain event features, and extracts features from the acquired grayscale data to obtain grayscale features. The feature fusion module performs a first encoding fusion of the event features and grayscale features based on the pixel correspondence between the event data and grayscale data in space to obtain secondary fused features. The fusion module performs a second encoding fusion of the secondary fused features and visible light image features based on the pixel semantic correspondence between the grayscale data and visible light image data to obtain primary fused features. The image output module obtains a fused feature image based on the primary fused features.

[0057] In some embodiments, the sample images used for training the image processing model include input sample pairs and ground truth images. The input sample pairs include event training data, grayscale training data, and blurred RGB data, while the ground truth images are the clear RGB data corresponding to the blurred RGB data. However, due to the limited number of sample images in a dual-camera imaging system comprising an RGB camera and an event camera, the training effect of the neural network model is poor. Therefore, this application provides a method for generating partial sample images of a dual-camera imaging system comprising an RGB camera and an event camera.

[0058] Both blurred RGB data and sharp RGB data can be obtained using the following methods, specifically: Step 1: Obtain the first original image data corresponding to the event training data, wherein the first original image data is RGB image data and the clarity of the first original image data meets the preset clarity conditions.

[0059] In this embodiment, the first original image data can be either blurry RGB original data or clear RGB original data, and the clarity condition can be set according to specific needs. For example, if you want to obtain blurry RGB data in a sample image, the first original image data corresponds to blurry RGB original data, and its clarity is less than or equal to a preset clarity threshold. If you want to obtain clear RGB data in a sample image, the first original image data corresponds to clear RGB original data, and the clarity of the first original image data is greater than the preset clarity threshold.

[0060] In some embodiments, blurred RGB raw data can be obtained by multi-frame fusion or Gaussian blurring of clear RGB raw data.

[0061] In some embodiments, the event training data and the first original image data are obtained through the same single-lens camera, which can simultaneously acquire event data and corresponding RGB image data, thereby obtaining the event training data and the first original image data in this embodiment.

[0062] After obtaining the first raw image data, the image parameters of the first raw image data can be determined, specifically the image size parameters, including the width and height of the image.

[0063] Step 2: Construct a transformation matrix based on the image parameters of the first original image data and the camera parameters of the dual-camera imaging system. The image parameters can be obtained from the camera parameters or set according to requirements. The camera parameters of the dual-camera imaging system are used to determine the translation and rotation relationships between the two cameras. These camera parameters can be preset.

[0064] Step 3: Map the first original image data using a transformation matrix to obtain RGB sample data suitable for the dual-camera imaging system. This RGB sample data corresponds to the first original image data and includes blurred RGB data and / or sharp RGB data. When the first original image data is blurred RGB original data, the RGB sample data is blurred RGB data; when the first original image data is sharp RGB original data, the RGB sample data is sharp RGB data.

[0065] This embodiment uses image data acquired by a single camera to obtain RGB image sample data from a dual-camera imaging system. This saves on the cost of acquiring sample images, increases the number of sample images, and improves the training effect of the image processing model.

[0066] After acquiring both blurred RGB data and clear RGB images, the clear RGB raw data is converted to grayscale to obtain grayscale training data. Based on this, the blurred RGB data, grayscale training data, event training data, and clear RGB data can be used together to train an image processing model.

[0067] Furthermore, Figure 4 This application provides a method for determining a transformation matrix according to an embodiment of the present application, such as... Figure 4 As shown, the method for constructing a transformation matrix based on the image parameters of the first original image data and the camera parameters of the dual-camera imaging system includes the following steps: Step S401: Determine the principal point and focal length of the event camera based on the image parameters; The image parameters include the image width and height. Assuming the image width is h and the height is w, then the camera principal point c... x Let w / 2, c y Set to h / 2, the camera focal length f is set according to the following formula: ; Step S402: Generate random perturbation, and determine the camera intrinsic parameter matrix of the dual-camera imaging system based on the random perturbation, the camera principal point, and the camera focal length; In this embodiment, to simulate the viewing angle change effect caused by the relative mounting positions of the RGB camera and the event camera in a dual-camera imaging system, a random perturbation is added to the camera principal point. This random perturbation needs to be performed within a preset range, then the camera intrinsic parameter matrix... for: ; Step S403: Randomly determine the rotation matrix between the two cameras in the dual-camera imaging system; wherein the parameters of the rotation matrix need to be set within a preset range; In this embodiment, to simulate the relative mounting angle between the two cameras in a dual-camera imaging system, this embodiment simulates the three axes of the RGB camera and the event camera. Random values ​​are selected within a certain range, where the three axes represent the parameters of the relative rotation matrix between the RGB camera and the event camera. The relative rotation matrix R between the RGB camera and the event camera can then be expressed as: in, This represents the matrix multiplication operation.

[0068] Step S404: Determine the transformation matrix based on the camera intrinsic parameter matrix and the rotation matrix.

[0069] Specifically, the transformation matrix can be obtained using the following formula: in, for The inverse matrix.

[0070] Through steps S401 to S404 above, this embodiment uses multi-view geometry to map existing image data onto a virtual dual-camera perspective by constructing a random homography transformation matrix based on image data acquired by a single camera, thereby achieving the purpose of simulating dual-camera data and improving the efficiency of acquiring sample images in the dual-camera imaging system.

[0071] In some embodiments, the blurred RGB data can also be obtained through the following method: First, acquire multiple second original image data, wherein the second original image data is RGB image data, and its exposure time and resolution meet preset image acquisition conditions. Specifically, the second original image data is a short-exposure image with high sharpness, its exposure time needs to be less than or equal to a preset exposure time threshold, and its sharpness needs to be greater than or equal to a preset resolution threshold. Then, fuse the multiple second original image data to obtain blurred RGB data. Specifically, the multiple second original image data can be weighted and fused into a single blurred image to simulate a long exposure effect. This embodiment obtains blurred RGB data by fusing multiple clear short-exposure second original image data, which can simultaneously obtain blurred RGB data and clear RGB data during sample acquisition, improving the quality of sample training pairs.

[0072] Furthermore, in the process of fusing multiple second original image data to obtain blurred RGB data, frame interpolation can be performed on the multiple second original image data, and the interpolated second original image data can be fused to obtain blurred RGB data. In this embodiment, frame interpolation makes the inter-frame changes of multiple second original image data smoother, and the fused blurred training image is more natural and closer to the real blur effect.

[0073] In some embodiments, because the frame rate of the event camera is higher than that of the RGB camera, the visible light image data acquired by the RGB camera cannot correspond to the event data acquired by the event camera. To eliminate invalid event data, before determining event features based on the event data acquired by the event camera, the event data needs to be filtered according to the exposure parameters of both the RGB and event cameras to determine valid event data. The exposure parameters of the RGB camera include the starting exposure time and the exposure duration of the visible light image when the RGB camera acquires the visible light image data. In this embodiment, by eliminating invalid event data, the fusion quality of subsequent event features, grayscale features, and visible light image features can be improved.

[0074] Furthermore, this embodiment provides a method for filtering valid event data. Figure 5 This is a flowchart of a method for determining valid event data according to an embodiment of this application, such as... Figure 5 As shown, the method includes the following steps: Step S501: Obtain the starting exposure time and the exposure duration of the visible light image when the RGB camera acquires visible light image data; When an RGB camera acquires each frame of visible light image data, it simultaneously records the corresponding start exposure time and the visible light image exposure duration. Step S502: Calculate the row exposure start time and row exposure end time for each row of pixels in the visible light image data based on the start exposure time and the visible light image exposure duration. In this embodiment, the RGB camera is a rolling shutter camera. The readout time of each row of pixels is different. Therefore, in order to filter valid event data, it is necessary to calculate the row exposure start time and row exposure end time of each row of pixels. Step S503: Obtain the exposure time of the event pixels when the event camera collects event data; When an event camera acquires event data, each pixel has an exposure time. Step S504: Compare the start time and end time of the line exposure with the exposure time of the event pixel, and determine the valid event data based on the comparison result.

[0075] This involves setting time boundaries based on the start and end times of each row in the visible light image data, selecting valid signals from the event data accordingly, and discarding invalid event data that exceeds the exposure time of the visible light image.

[0076] Specifically, each event pixel has four elements: spatial coordinates x and y, timestamp t, and event polarity p (-1 or +1). After obtaining the visible light image data, the spatial range of each row of the image and the exposure time range are known. Event data and visible light image data can be matched based on pixel coordinates and pixel timestamps to select event data within the same range.

[0077] Through the above steps S501 to S504, since the dual-camera imaging system with an RGB camera and an event camera is in the same time system, this embodiment aligns the high frame rate event data rows to the visible light image data of the RGB camera based on the starting exposure time and visible light image exposure duration of the RGB camera and the exposure time of each pixel of the event camera, thereby achieving the alignment of event data and RGB visible light image data in the time dimension.

[0078] In some embodiments, event data needs to be preprocessed before obtaining event features. Specifically, this includes the following steps: Step 1: Determine the overlay range of event data based on the exposure duration of the visible light image when the RGB camera acquires the visible light image data; determine the reference time for overlaying event data based on the midpoint of the visible light image exposure duration.

[0079] For example, if the visible light exposure time is [t1, t2], then the overlay range of the event data can be directly used as [t1, t2], or it can be extended based on this range; for example, the midpoint of the visible light image exposure time can be set as the overlay reference time; Step 2: Obtain the exposure time of the event pixels when the event camera collects event data; Step 3: Based on the overlay reference time, overlay the event data whose exposure times of the event pixels are within the overlay range to obtain the event voxel. For example, overlay the event data whose exposure times of the event pixels are in the range [t1, t2] to obtain the event voxel.

[0080] After obtaining the event voxels, the preprocessing process is completed, and feature extraction can then be performed on the event voxels to obtain event features.

[0081] Typically, during image processing, the recovered image is the image at the midpoint of the RGB camera's exposure time. In this embodiment, before extracting event features, the event data is preprocessed based on the visible light image exposure time. The event data is superimposed based on the midpoint of the visible light exposure time, which can better represent the changes in the captured image during the exposure time.

[0082] In some embodiments, after obtaining the event data, valid event data is first filtered, and then preprocessed as described above. Finally, feature extraction is performed on the event voxels to obtain event features.

[0083] Preferably, the technical solution of this application is described using the application of this application in a dual-camera imaging system to remove image blur and improve image clarity as an example. Figure 6 This is a flowchart of image deblurring according to a preferred embodiment of this application, such as... Figure 6 As shown, the method includes the following steps: Step S601: Acquire visible light image data, event data, and grayscale data through a dual-camera imaging system.

[0084] The dual-camera imaging system consists of a high-resolution RGB camera and a low-resolution event camera, with the RGB camera serving as the main camera and the event camera as the secondary camera. In this embodiment, the RGB camera is preferably an RGB rolling shutter camera, which acquires RGB data as visible light image data. The event camera is preferably a global shutter DAVIS camera. It should be noted that the DAVIS camera can simultaneously capture high-frame-rate event data and low-frame-rate scene grayscale data. In this embodiment, both the visible light image data and the grayscale data are in RAW domain format.

[0085] Preferably, the RGB camera in this embodiment captures images with long exposure times, and because the exposure time is too long, the images are blurry and have low clarity. Similarly, the grayscale data captured by the event camera is also a blurry image with long exposure times.

[0086] Step S602: Filter the event data to obtain valid event data.

[0087] Specifically, based on the start exposure time of the RGB camera and the exposure duration of the visible light image, as well as the exposure time of each pixel of the event camera, the high frame rate event data rows are aligned to the visible light image data. For example, based on the start exposure time and the exposure duration of the visible light image, the row exposure start time and row exposure end time of each pixel in the visible light image data are calculated. Then, using the row exposure start time and row exposure end time as time boundaries, valid event data is selected accordingly, and invalid event data exceeding the visible light image exposure duration is discarded, thus achieving alignment between the event data and the RGB visible light image data in the time dimension.

[0088] Step S603: Preprocess the valid event data to obtain event voxels.

[0089] Specifically, the superposition range of event data is determined based on the exposure time of the visible light image; the superposition reference time of event data is determined based on the midpoint of the exposure time of the visible light image; the exposure time of the event pixels when the event camera collects event data is obtained; and the event data whose exposure time of the event pixels is within the superposition range are superimposed based on the superposition reference time to obtain the event voxel.

[0090] After obtaining the event voxels, the subsequent processing of the event voxels, grayscale data, and visible light image data can be achieved through a trained image processing model. In this embodiment, the image processing model is a neural network based on deep learning. Figure 7 This is a structural diagram of an image processing model according to a preferred embodiment of this application, such as... Figure 7 As shown, this neural network uses a three-level downsampling feature extraction method to extract image features.

[0091] Step S604: Input the preprocessed event voxels into the event feature extraction module for encoding to extract event features; Step S605: The RGB data from the main camera is encoded by the corresponding image feature extraction module to obtain RGB features; the grayscale data from the secondary camera is encoded by the corresponding image feature extraction module to obtain grayscale features. This step can be performed simultaneously with steps S602 to S604. Step S606: Input the event features, RGB features, and grayscale features into the multi-feature fusion module to align the event data with RGB data and fuse features, thereby obtaining multi-dimensional fused features and achieving spatial alignment between the event data and RGB data.

[0092] Figure 8 is a schematic diagram of the image processing effect according to a preferred embodiment of this application, exemplarily. Figure 8a It is a blurred visible light image captured by an RGB camera. Figure 8b It is a blurry grayscale image captured by the event camera. Figure 8c These are event camera images that have undergone visualization processing of event data acquired by the event camera. Figure 8d The image shown is a clear image obtained according to the image deblurring method of this embodiment. It is evident that the method in this embodiment can effectively improve the clarity of blurred images.

[0093] The multi-feature fusion module in this embodiment is implemented based on a multi-head mutual attention mechanism. The specific implementation process is as follows: Step 1: Since event data and grayscale data are strictly pixel-aligned in space, this embodiment uses a multi-head mutual attention module to encode grayscale features and event features to achieve their fusion. The resulting grayscale-event fusion feature (i.e., secondary fusion feature) has the characteristics of both grayscale data and event data.

[0094] Step 2: Using the multi-head mutual attention module again, the grayscale features in the grayscale-event fusion features are used as the matching benchmark to align and fuse the grayscale-event fusion features and RGB features to obtain multi-dimensional fusion features (i.e., main fusion features), thereby realizing the alignment of event features to visible light image features.

[0095] Meanwhile, this embodiment adopts a three-level progressive downsampling resolution feature extraction method, and the above-mentioned multi-feature fusion process is implemented in three different resolution feature stages.

[0096] Step S607: Input the multidimensional fusion features into the motion blur removal module with three-level progressive upsampling resolution to obtain the deblurred RAW domain image data (i.e., the fusion feature image).

[0097] The following describes the downsampling fusion and upsampling processes. The downsampling fusion process includes steps 1 to 6, and the downsampling process includes steps 7 to 9: Step 1, Feature Extraction: Grayscale data is input into the first-layer secondary camera feature extraction module to obtain grayscale features, denoted as P11; RGB data is input into the first-layer main camera feature extraction module to obtain RGB features, denoted as P12; event data is input into the event feature extraction module to obtain event features, denoted as P3. Step 2, First Multi-Feature Fusion: In the first-layer multi-feature fusion module, grayscale feature P11 and event feature P3 are fused to obtain fused feature E1. E1 is then fused with RGB feature P12 to obtain fused feature R1. Step 3, Downsampling Feature Extraction: Input the fused feature E1 into the second-layer sub-camera feature extraction module to obtain the downsampled sub-camera feature P21; input the fused feature R1 into the second-layer main camera feature extraction module to obtain the downsampled main camera feature P22; Step 4, Second Feature Fusion: In the second-layer multi-feature fusion module, the secondary camera feature P21 and the event feature P3 are fused to obtain the fused feature E2. E2 is then fused with the primary camera feature P22 to obtain the fused feature R2. Step 5: Downsampling feature extraction: Input the fused feature E2 into the third-layer sub-camera feature extraction module to obtain the downsampled sub-camera feature P31; input the fused feature R2 into the third-layer main camera feature extraction module to obtain the downsampled main camera feature P32; Step 6, Third Feature Fusion: In the third-layer multi-feature fusion module, the secondary camera feature P31 and the event feature P3 are fused to obtain the fused feature E3. E3 is then fused with the primary camera feature P32 to obtain the fused feature R3. It should be noted that the aforementioned multi-feature fusion modules can all be multi-head mutual attention fusion modules.

[0098] Step 7: Upsample the fused feature R3 to obtain the upsampled feature U1; Step 8: Upsample based on U1 and fused feature R2 to obtain upsampled feature U2; Step 9: Upsample based on U2 and fused feature R1 to obtain upsampled feature U3; Finally, the deblurred fused image is recovered using the image processing model based on U3.

[0099] Through steps S601 to S607, after obtaining visible light image data, event data, and grayscale data, the dual-camera imaging system sequentially fuses the event data, grayscale data, and visible light image data according to the spatial pixel correspondence between grayscale data and event data, as well as the pixel semantic correspondence between grayscale data and visible light image data. This improves the overall observation redundancy and spatiotemporal sampling density of the imaging system, solves the problem of low imaging quality in complex dynamic scenes, and improves the resolution of the fused feature image.

[0100] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0101] Based on the same inventive concept, this application also provides an image processing apparatus for implementing the image processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more image processing apparatus embodiments provided below can be found in the limitations of the image processing method described above, and will not be repeated here.

[0102] In one exemplary embodiment, such as Figure 9 As shown, an image processing apparatus is provided, including: a feature extraction module 91, a feature fusion module 92, and an image output module 93. The feature extraction module 91 is used to determine visible light image features based on visible light image data acquired by the RGB camera; determine event features based on event data acquired by the event camera; and determine grayscale features based on grayscale data acquired by the event camera; wherein the RGB camera and the event camera constitute a dual-camera imaging system; The feature fusion module 92 is used to perform a first encoding fusion of the event features and the grayscale features based on the pixel correspondence between the event data and the grayscale data in space to obtain secondary fusion features; and to perform a second encoding fusion of the secondary fusion features and the visible light image features based on the pixel semantic correspondence between the grayscale data and the visible light image data to obtain primary fusion features. The image output module 93 is used to obtain a fused feature image based on the main fusion features.

[0103] Through the aforementioned image processing device, after obtaining visible light image data, event data, and grayscale data, the dual-camera imaging system sequentially fuses the event data, grayscale data, and visible light image data based on the spatial pixel correspondence between grayscale data and event data, as well as the pixel semantic correspondence between grayscale data and visible light image data. This improves the overall observation redundancy and spatiotemporal sampling density of the imaging system, solves the problem of low imaging quality in complex dynamic scenes, and improves the resolution of the fused feature image.

[0104] In some embodiments, after obtaining the main fusion features, the method further includes: The primary fusion feature and the secondary fusion feature are subjected to multiple downsampling and fusion processes to obtain the downsampled output feature. The downsampled output features are then subjected to corresponding upsampling processing to obtain the fused feature image.

[0105] In some embodiments, the step of performing multiple downsampling and fusion processes on the primary fusion feature and the secondary fusion feature to obtain the downsampled output feature includes: The secondary fusion features are downsampled to obtain the first downsampled features; the primary fusion features are downsampled to obtain the second downsampled features. The event features, the first downsampled features, and the second downsampled features are fused to obtain the downsampled output features.

[0106] In some embodiments, performing corresponding upsampling processing on the downsampled output features to obtain the fused feature image includes: Upsampling is performed based on the downsampled output features and the main fusion features to obtain the fusion feature image.

[0107] In some embodiments, the first coding fusion and / or the second coding fusion are implemented based on a multi-head mutual attention mechanism.

[0108] In some embodiments, the image processing method is implemented using a trained image processing model, which is trained based on sample images; the sample images include at least RGB sample data, and the method for obtaining the RGB sample data includes: Obtain the first original image data corresponding to the event training data, wherein the first original image data is RGB image data and the clarity of the first original image data meets the preset clarity condition. A transformation matrix is ​​constructed based on the image parameters of the first original image data and the camera parameters of the dual-camera imaging system; The first original image data is mapped using the transformation matrix to obtain RGB sample data suitable for the dual-camera imaging system.

[0109] In some embodiments, constructing the transformation matrix based on the image parameters of the first original image data and the camera parameters of the dual-camera imaging system includes: The principal point and focal length of the event camera are determined based on the image parameters. A random perturbation is generated, and the camera intrinsic parameter matrix of the dual-camera imaging system is determined based on the random perturbation, the camera principal point, and the camera focal length. The rotation matrix between the two cameras in the dual-camera imaging system is randomly determined; The transformation matrix is ​​determined based on the camera intrinsic parameter matrix and the rotation matrix.

[0110] In some embodiments, the image processing method is implemented using a trained image processing model, which is trained based on sample images; the sample images include at least blurred RGB data, and the method for obtaining the blurred RGB data includes: Acquire multiple second raw image data, wherein the exposure duration and resolution of the second raw image data meet the preset image acquisition conditions; Multiple original second image data are fused together to obtain the blurred RGB data.

[0111] In some embodiments, fusing multiple second original image data to obtain the blurred RGB data includes: Multiple second original image data are interpolated, and the interpolated multiple second original image data are fused to obtain the blurred RGB data.

[0112] In some embodiments, before determining event features based on event data acquired by the event camera, the method further includes: The event data is filtered based on the exposure parameters of the RGB camera and the event camera to determine the valid event data.

[0113] In some embodiments, the step of filtering the event data based on the exposure parameters of the RGB camera and the event camera to determine valid event data includes: Obtain the starting exposure time and the visible light image exposure duration when the RGB camera acquires the visible light image data; Based on the initial exposure time and the exposure duration of the visible light image, calculate the row exposure start time and row exposure end time for each row of pixels in the visible light image data; Obtain the exposure time of the event pixels when the event camera acquires the event data; The start and end times of the line exposure are compared with the exposure times of the event pixels, and the valid event data is determined based on the comparison results.

[0114] In some embodiments, determining event features based on event data acquired by the event camera includes: The superposition range of event data is determined based on the visible light image exposure duration when the RGB camera acquires the visible light image data; the superposition reference time of event data is determined based on the midpoint of the visible light image exposure duration. Obtain the exposure time of the event pixels when the event camera acquires the event data; Based on the superposition reference time, the event data of the event pixel exposure time within the superposition range are superimposed to obtain the event voxel; Feature extraction is performed on the event voxels to obtain the event features.

[0115] Each module in the aforementioned image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0116] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores image processing-related data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an image processing method.

[0117] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0118] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0119] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0120] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0121] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0122] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchain. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logic devices based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited to these.

[0123] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0124] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An image processing method, characterized by, The method includes: Determine visible light image features based on visible light image data acquired by an RGB camera; Event features are determined based on event data acquired by the event camera; grayscale features are determined based on grayscale data acquired by the event camera; wherein, the RGB camera and the event camera constitute a dual-camera imaging system; Based on the spatial pixel correspondence between the event data and the grayscale data, the event features and the grayscale features are first encoded and fused to obtain secondary fused features; Based on the pixel semantic correspondence between the grayscale data and the visible light image data, the secondary fusion feature and the visible light image feature are fused using a second encoding method to obtain the primary fusion feature; A fusion feature image is obtained based on the main fusion features.

2. The image processing method according to claim 1, characterized in that, After obtaining the main fusion features, the method further includes: The primary fusion feature and the secondary fusion feature are subjected to multiple downsampling and fusion processes to obtain the downsampled output feature. The downsampled output features are then subjected to corresponding upsampling processing to obtain the fused feature image.

3. The image processing method according to claim 2, characterized in that, The process of performing multiple downsampling and fusion processes on the primary and secondary fusion features to obtain downsampled output features includes: The secondary fusion features are downsampled to obtain the first downsampled features; the primary fusion features are downsampled to obtain the second downsampled features. The event features, the first downsampled features, and the second downsampled features are fused to obtain the downsampled output features.

4. The image processing method according to claim 2, characterized in that, The step of performing corresponding upsampling processing on the downsampled output features to obtain the fused feature image includes: Upsampling is performed based on the downsampled output features and the main fusion features to obtain the fusion feature image.

5. The image processing method according to claim 1, characterized in that, The first encoding fusion and / or the second encoding fusion are implemented based on a multi-head mutual attention mechanism.

6. The image processing method according to claim 1, characterized in that, The image processing method is implemented by a trained image processing model, which is trained based on sample images. The sample image includes at least RGB sample data, and the method for obtaining the RGB sample data includes: Obtain the first original image data corresponding to the event training data, wherein the first original image data is RGB image data and the clarity of the first original image data meets the preset clarity condition. A transformation matrix is ​​constructed based on the image parameters of the first original image data and the camera parameters of the dual-camera imaging system; The first original image data is mapped using the transformation matrix to obtain RGB sample data suitable for the dual-camera imaging system.

7. The image processing method according to claim 6, characterized in that, The step of constructing the transformation matrix based on the image parameters of the first original image data and the camera parameters of the dual-camera imaging system includes: The principal point and focal length of the event camera are determined based on the image parameters. A random perturbation is generated, and the camera intrinsic parameter matrix of the dual-camera imaging system is determined based on the random perturbation, the camera principal point, and the camera focal length. The rotation matrix between the two cameras in the dual-camera imaging system is randomly determined; The transformation matrix is ​​determined based on the camera intrinsic parameter matrix and the rotation matrix.

8. The image processing method according to claim 1, characterized in that, The image processing method is implemented through a trained image processing model, which is trained based on sample images; the sample images include at least blurred RGB data, and the method for obtaining the blurred RGB data includes: Acquire multiple second raw image data, wherein the exposure duration and resolution of the second raw image data meet the preset image acquisition conditions; Multiple original second image data are fused together to obtain the blurred RGB data.

9. The image processing method according to claim 8, characterized in that, The step of fusing multiple second original image data to obtain the blurred RGB data includes: Multiple second original image data are interpolated, and the interpolated multiple second original image data are fused to obtain the blurred RGB data.

10. The image processing method according to claim 1, characterized in that, Before determining event features based on event data acquired by the event camera, the method further includes: The event data is filtered based on the exposure parameters of the RGB camera and the event camera to determine the valid event data.

11. The image processing method according to claim 10, characterized in that, The step of filtering the event data based on the exposure parameters of the RGB camera and the event camera to determine valid event data includes: Obtain the starting exposure time and the visible light image exposure duration when the RGB camera acquires the visible light image data; Based on the initial exposure time and the exposure duration of the visible light image, calculate the row exposure start time and row exposure end time for each row of pixels in the visible light image data; Obtain the exposure time of the event pixels when the event camera acquires the event data; The start and end times of the line exposure are compared with the exposure times of the event pixels, and the valid event data is determined based on the comparison results.

12. The image processing method according to claim 1, characterized in that, The determination of event features based on event data acquired by the event camera includes: The superposition range of event data is determined based on the visible light image exposure duration when the RGB camera acquires the visible light image data; the superposition reference time of event data is determined based on the midpoint of the visible light image exposure duration. Obtain the exposure time of the event pixels when the event camera acquires the event data; Based on the superposition reference time, the event data of the event pixel exposure time within the superposition range are superimposed to obtain the event voxel; Feature extraction is performed on the event voxels to obtain the event features.

13. An image processing apparatus, characterized in that, The device includes a dual-camera imaging system and a processor; The dual-camera imaging system includes an RGB camera and an event camera. The RGB camera acquires visible light image data, and the event camera acquires event data and grayscale data. The processor determines visible light image features based on the visible light image data; and determines event features based on the event data. Determine grayscale features based on the grayscale data; Based on the spatial pixel correspondence between the event data and the grayscale data, the event features and the grayscale features are first encoded and fused to obtain secondary fused features; based on the pixel semantic correspondence between the grayscale data and the visible light image data, the secondary fused features and the visible light image features are second encoded and fused to obtain primary fused features. A fusion feature image is obtained based on the main fusion features.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.