Hybrid event and frame sensor processing method and system, computer device and medium

By acquiring raw data from hybrid event and frame sensors and performing noise calibration and event probability coupling operations, the problem of uniformity in cross-modal calibration in hybrid sensors is solved, achieving high-precision and reproducible sensor data calibration, which is suitable for high-speed and high dynamic range scenarios.

CN122120634APending Publication Date: 2026-05-29HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
Filing Date
2026-02-03
Publication Date
2026-05-29

Smart Images

  • Figure CN122120634A_ABST
    Figure CN122120634A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image sensor calibration, and particularly relates to a processing method and system of a hybrid event and frame sensor, computer equipment and a medium; the method comprises the following steps: acquiring original data of the hybrid event and frame sensor; performing a noise calibration operation on an intensity frame to generate an active pixel sensor noise model and a net intensity signal; performing an event probability coupling operation on an event stream to generate an event visual sensor event probability model; and performing unified collaborative processing on the active pixel sensor noise model and the event visual sensor event probability model to output uniformly calibrated sensor data. In this way, a unified and reproducible cross-modal calibration framework can be provided to solve the technical problems of existing calibration technologies, such as weak physical coupling, insufficient calibration granularity and black box implementation, and to provide a unified, reproducible and high-precision sensor data calibration output for high-speed and high-dynamic-range mobile perception applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image sensor calibration technology, and in particular to a method, system, computer device, and medium for processing hybrid event and frame sensors. Background Technology

[0002] With the rapid development of mobile computing and wearable devices, hybrid event-frame sensors, by integrating active pixel sensors (APS) and event-based vision sensors (EVS) onto the same chip or system, exhibit advantages such as high dynamic range, microsecond-level temporal resolution, and low power consumption, and are widely used in high-speed, high-dynamic-range scenarios such as smart glasses. However, hybrid sensors face significant cross-modal noise and mismatch issues in practical deployments: APS suffers from line bias, black level, dark current, and readout noise, while EVS's trigger threshold, inter-pixel mismatch, and background noise are affected by brightness and circuit status, making it difficult to guarantee the consistency of sensor output. Currently, the lack of a unified, transparent, and reproducible joint calibration and processing pipeline results in poor reusability and low efficiency in algorithm simulation and verification, severely restricting the large-scale application of hybrid sensors in industrial and academic fields.

[0003] In the field of image sensor calibration, existing technologies have attempted to improve data reliability through noise modeling and systematic error correction. For example, prior art document 1 (application publication number CN110807812A) discloses a systematic error calibration method for digital image sensors based on a priori noise models. This method establishes a noise composition model and uses statistical methods to estimate parameters such as fixed-mode noise and thermal noise to achieve systematic error correction for single RAW (raw data directly captured by the image sensor without any processing) images. Although this method has achieved certain results in traditional image sensor calibration, it mainly targets single-modal APS sensors and fails to fully consider the cross-modal coupling characteristics of APS and EVS in hybrid sensors. Specifically, existing technologies have the following limitations: First, the calibration process is mostly limited to a single mode (e.g., only processing APS or EVS), lacking a physical mapping from APS brightness to EVS event probability, resulting in inconsistent parameters in hybrid systems. Second, for the manufacturing non-uniformity of hybrid layouts such as Quad-Bayer, existing solutions often use global parameter estimation, ignoring position-by-position / color statistical calibration, resulting in residual position-dependent errors. Furthermore, traditional Image Signal Processor (ISP) pipelines often use serial black-box modules with closed-source implementations (e.g., MATLAB-based encapsulation), making it difficult to maintain pixel-by-pixel alignment with the noise model, which is detrimental to simulation and reverse optimization. Finally, event probability modeling often deviates from the physical basis of pixel voltage / threshold, relying solely on empirical fitting, resulting in weak generalization ability and interpretability. These problems combine to make it difficult for existing technologies to achieve high-precision, high-efficiency noise calibration in hybrid sensor scenarios, failing to meet the stringent real-time and reliability requirements of high-speed wearable devices.

[0004] Therefore, existing noise model-based calibration techniques, when dealing with mixed event-frame sensors, suffer from problems such as weakened physical coupling, insufficient calibration granularity, and black-box implementation due to the lack of a unified and reproducible cross-modal calibration framework. As a result, the accuracy and efficiency of noise correction cannot meet the requirements of high-speed and high dynamic range scenarios. Summary of the Invention

[0005] To address the aforementioned shortcomings or drawbacks, this invention provides a method, system, computer device, and medium for processing hybrid event and frame sensors. It offers a unified and reproducible cross-modal calibration framework to solve the technical problems of weakened physical coupling, insufficient calibration granularity, and black-box implementation in existing calibration techniques.

[0006] This invention provides a method for processing hybrid event and frame sensors, comprising: Acquire raw data from the hybrid event and frame sensors, including intensity frames from the active pixel sensor and event streams from the event vision sensor.

[0007] Perform noise calibration on the intensity frames to generate an active pixel sensor noise model and net intensity signal.

[0008] Based on the net intensity signal, an event probability coupling operation is performed on the event stream to generate an event probability model for the visual sensor.

[0009] A pre-defined linear-hold image signal processing pipeline performs unified collaborative processing on the noise model of the active pixel sensor and the event probability model of the event vision sensor, outputting uniformly calibrated sensor data.

[0010] According to a second aspect, the present invention provides a processing system for hybrid event and frame sensors, comprising: The raw data acquisition module is used to acquire raw data from the hybrid event and frame sensors, including intensity frames from the active pixel sensor and event streams from the event vision sensor.

[0011] The noise calibration processing module is used to perform noise calibration operations on the intensity frames to generate an active pixel sensor noise model and net intensity signal.

[0012] The probability model generation module is used to perform event probability coupling operations on the event stream based on the net intensity signal to generate an event probability model for the visual sensor.

[0013] The processing result output module is used to perform unified collaborative processing on the active pixel sensor noise model and the event probability model of the event vision sensor through a preset linear image signal processing pipeline, and output uniformly calibrated sensor data.

[0014] According to a third aspect, the present invention provides a computer device comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform any of the hybrid event and frame sensor processing methods in the embodiments of the present invention.

[0015] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform any of the mixed event and frame sensor processing methods in the embodiments of the present invention.

[0016] The present invention provides a method for processing hybrid event and frame sensors, which is implemented through four core steps: acquisition, calibration, coupling, and collaborative processing. First, the raw data of the brightness-exposure-related variance model hybrid event and frame sensors are acquired. This raw data includes intensity frames from an active pixel sensor and event streams from an event vision sensor. Next, noise calibration is performed on the intensity frames of the brightness-exposure-related variance model to generate an active pixel sensor noise model and a net intensity signal. Then, based on the net intensity signal of the brightness-exposure-related variance model, an event probability coupling operation is performed on the event stream of the brightness-exposure-related variance model to generate an event probability model for the event vision sensor. Finally, a pre-defined linear hold image signal processing pipeline is used to perform unified collaborative processing on the brightness-exposure-related variance model active pixel sensor noise model and the event probability model of the event vision sensor, outputting uniformly calibrated sensor data.

[0017] In the overall technical solution, this invention addresses the problem of weakened physical coupling in the variance model related to brightness and exposure, as described in the background technology. It establishes a mapping relationship between the net intensity signal and the event trigger probability through event probability coupling operations, solving the technical defect in the prior art where the event model is disconnected from physical quantities such as pixel voltage and threshold, resulting in weak model generalization ability and interpretability. Addressing the problem of insufficient calibration granularity in the variance model related to brightness and exposure, this invention decomposes and subtracts various fixed noises such as line bias, black level, and dark current in the intensity frame through noise calibration operations, achieving fine-grained modeling and correction of inherent sensor noise. Finally, addressing the problem of achieving black-box processing, this invention uses a linear image signal processing pipeline to perform unified collaborative processing of the two models, ensuring the linear transparency of the entire calibration and processing flow, and avoiding the drawbacks of unreproducible and difficult-to-align verification inherent in traditional closed-source image signal processing (ISP) pipelines. Therefore, the technical solution of the present invention provides a unified and reproducible cross-modal calibration framework, which solves the technical problems of weakened physical coupling, insufficient calibration granularity and black box implementation in existing calibration technologies, and provides a unified, reproducible and high-precision sensor data calibration output for high-speed, high dynamic range mobile sensing applications. Attached Figure Description

[0018] Figure 1 This is a flowchart of a method for processing hybrid events and frame sensors according to an embodiment of the present invention; Figure 2 This is a diagram illustrating the working principle and output example of an EVS simulator based on statistical modeling triggering according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of a hybrid event and frame sensor processing system according to an embodiment of the present invention; Figure 4This is a block diagram of a computer device for implementing embodiments of the present invention. Detailed Implementation

[0019] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0020] During the development of this invention, the inventors, through extensive experiments and data analysis, revealed the intrinsic relationship between the physical characteristics of sensors and the noise statistical model: traditional single-modal calibration methods not only ignore the coupling mechanism of cross-modal signals, but also suffer from position-dependent error accumulation due to differences in manufacturing processes. Based on this relationship, the inventors innovatively proposed this technical solution, which uses the linear intensity signal of an active pixel sensor (APS) as a reference, and through dual-path processing that couples noise calibration with event probability, combined with a linearly preserved image signal processing (ISP) pipeline, achieves unified calibration and collaborative processing of hybrid sensor data, embodying the core concept of "physical-statistical co-modeling".

[0021] Specifically, through comparative experiments, the invention team discovered three common problems with traditional statistical channel models: first, APS and EVS signals lack a physical mapping of brightness-event probability; second, they lack sufficient support for position-by-position calibration of hybrid layouts such as Quad-Bayer; and third, the processing pipeline suffers from nonlinear black-box operations. These technical deficiencies lead to difficulty in reproducing calibration results and poor cross-platform adaptability. However, the unified noise calibration and reproducible pipeline proposed in this invention can improve calibration accuracy and reliability, achieving physical consistency alignment of sensor data. Experimental data shows that this method... The residual of the fixed term on the GEN2 sensor is less than goodness of fit of variance model ( The accuracy rate reached over 0.97, especially in high dynamic range (HDR) scenarios, where the event probability prediction accuracy improved by more than 30%.

[0022] Therefore, this invention provides a hybrid event and frame sensor processing method based on the first aspect, which can be applied to a hybrid visual perception system (hereinafter referred to as the "system"). This system can run on various computing platforms through local deployment or cloud collaboration to complete high-speed, high dynamic range visual data acquisition and calibration processing. Specifically, this system can be deployed in various hardware environments, including but not limited to embedded mobile devices (such as smart glasses and AR or VR headsets), edge computing nodes (such as industrial vision gateways and security edge servers), cloud computing platforms (such as visual data service centers), and hybrid architectures with end-edge-cloud collaboration. By flexibly adapting to hardware foundations with different computing resources and real-time requirements, it achieves full-scene coverage from lightweight terminal processing to high-performance cloud optimization.

[0023] like Figure 1 As shown, the method may include: Step S110: Obtain raw data from the hybrid event and frame sensor.

[0024] The raw data includes intensity frames from the active pixel sensor and event streams from the event vision sensor. Raw data refers to the unprocessed set of signals directly acquired from the hybrid event and frame sensor; intensity frames from the active pixel sensor refer to two-dimensional image data captured at fixed time intervals, with each pixel value representing the light intensity integral result; event streams from the event vision sensor refer to asynchronously output event sequences that record brightness changes, with each event containing pixel location, timestamp, and polarity information.

[0025] Specifically, the system can control the hybrid sensor through a synchronous triggering mechanism to acquire APS intensity frames and EVS event streams in parallel, and use a timestamp alignment module to ensure data temporal consistency. During acquisition, the system sets exposure time parameters (e.g., from 1 ms to 80 ms) and event detection thresholds (e.g., from 0.1 lux to 10 lux) to adapt to different lighting conditions.

[0026] For example, the system deploys a set in a smart glasses scenario. Hybrid sensor (APS resolution) EVS resolution Under indoor lighting conditions of 500 lux, intensity frames were acquired with an exposure time of 10 milliseconds, and event streams were recorded simultaneously to generate synchronization data packets (the amount of data acquired in a single acquisition is approximately 1.5 megabytes, where MB refers to megabytes).

[0027] In another embodiment, such as Figure 2 The diagram shows the working principle and output example of an event vision sensor (EVS) simulator triggered by statistical modeling. Figure 2 The left side (i) illustrates the sampling principle of the Q function, which is based on the complementary cumulative distribution function of the standard normal distribution. (As defined in Equation 12), the theoretical trigger probability calculated by the aforementioned event probability model will be... This is mapped to a specific binary trigger decision. Figure 2 The example points are clearly marked. and its corresponding This intuitively reveals how the sampling mechanism transforms continuous probability values ​​(such as 15.9%) into a random, but statistically precise, event-triggered signal. Figure 2 The right side (k) presents a simulated event stream synthesized based on this principle, where alternating red and blue bars represent positive and negative polarity events, respectively. Their temporal and spatial distribution patterns are determined by the statistical model and sampling process. This embodiment fully presents the closed-loop process from statistical probability model to discrete event synthesis: first, based on the fine parameters obtained from calibration (as shown in Tables c and d), the value of each pixel in a given net intensity signal is calculated. Trigger probability below The Q-function sampling mechanism then determines whether an event should be generated at that moment. This process ensures that the event stream output by the simulator not only matches the real sensor data in terms of macroscopic statistical characteristics (such as the relationship between event rate and brightness), but also maintains physically interpretable randomness in terms of microscopic temporal dynamics, thus achieving high-fidelity, repeatable, and fully controllable EVS data simulation. This simulator provides a crucial and flexible tool for developing and testing event-driven algorithms, enabling researchers to efficiently verify the robustness of algorithms under various lighting, motion, and noise scenarios without the need for hardware prototypes or limitations imposed by real data acquisition environments.

[0028] Step S120: Perform noise calibration on the intensity frame to generate an active pixel sensor noise model and net intensity signal.

[0029] Among them, noise calibration operation refers to the process of decomposing and quantifying the inherent noise of the sensor through statistical methods; active pixel sensor noise model refers to the parameterized function describing the relationship between noise variance and brightness and exposure time; net intensity signal refers to linear light intensity data after deducting fixed noise components.

[0030] Specifically, the system can collect multiple sets of data through dark field sequences (lens occlusion, illuminance <1 lux) and uniform illumination sequences (illuminance 10~500 lux). It first calculates the initial estimates of black level noise, line offset noise and dark current noise, and then uses iterative optimization algorithms (such as least squares fitting) to jointly correct the noise parameters, finally obtaining the net intensity signal.

[0031] For example, the system uses a GEN2 sensor to acquire 100 frames of images in a dark field, calculates the mean black level noise to be 10 digital units (DN), and then fits the dark current noise figure to 0.05 DN / ms under uniform illumination through linear regression. After three iterations, the residual converges to below 0.002, and the net intensity signal is output for subsequent processing. Here, DN is a digital unit that can be obtained based on the sensor's ADC (Analog-to-Digital Converter) calibration.

[0032] Step S130: Based on the net intensity signal, perform an event probability coupling operation on the event stream to generate an event probability model for the visual sensor.

[0033] Among them, the event probability coupling operation refers to the calculation process of establishing the mapping relationship between the brightness signal and the event trigger probability; the event probability model of the event vision sensor refers to the parameterized function (such as the sigmoid function) that describes the change of the event trigger probability with brightness. Alternatively, the event probability model can also be a parameterized function (such as the sigmoid function and its higher-order extensions) that uses the net intensity signal as the independent variable to describe the change of the event trigger probability with it; its establishment process is the mathematical modeling of the coupling relationship between the statistical characteristics of the event flow and the brightness signal.

[0034] Specifically, the system can simultaneously record net intensity signals and event counts under static calibration experiments in dark, medium-brightness (10-100 lux) and high-brightness (500-2000 lux) conditions. It then uses maximum likelihood estimation to fit the event probability model parameters (such as the slope and offset of the sigmoid function) and optimizes the model accuracy using gradient descent. Furthermore, under the aforementioned static calibration experimental conditions, the system can also simultaneously acquire the net intensity signal of each pixel. And its corresponding event triggering data. Based on a large amount of synchronization data, with net strength signal Using the maximum likelihood estimation method as input variables, fit an event probability model. parameters (For example, the slope, offset, and coefficients of higher-order terms of the sigmoid function), and iteratively improve the model fitting accuracy through optimization algorithms such as gradient descent, ultimately obtaining a high coefficient of determination. A pixel-level event probability model.

[0035] For example, the system independently fits a model for each color filter array (CFA) location, collects data for 60 seconds under 100 lux illuminance, triggers events 5000 times, and obtains the S-shaped function parameters. Model determination coefficient ( The value reached 0.95.

[0036] Next, in some other embodiments, the system can also independently fit an event probability model for each CFA location, collecting 60 seconds of static scene data under 100 lux illuminance, with 5000 event triggers. An extended sigmoid function model is employed. Perform parameter fitting to capture the net intensity signal. The higher-order nonlinear relationship between the event trigger probability and the model parameters is obtained by fitting the model using maximum likelihood estimation and gradient descent optimization (learning rate set to 0.001, 200 iterations). Model determination coefficient ( The residual mean was close to zero, reaching 0.95. This validated the model's accuracy and consistency across brightness conditions. This embodiment demonstrates the advantages of fine-grained calibration per CFA location, providing a unified and reliable event probability mapping for hybrid sensors.

[0037] Step S140: Perform unified collaborative processing on the active pixel sensor noise model and the event probability model of the event vision sensor through a preset linear hold image signal processing pipeline, and output uniformly calibrated sensor data.

[0038] Among them, the linear image signal processing pipeline refers to a signal processing link composed of a series of linear operation modules (such as black level correction and de-mosaic) that avoids nonlinear distortion; the uniformly calibrated sensor data refers to standard format data that, after noise correction and event model alignment, can be directly used for downstream tasks (such as deblurring and low-light enhancement).

[0039] Specifically, the system can perform black level correction (subtracting fixed noise), pixel layout rearrangement (aggregating Quad-Bayer into standard Bayer), and linear demosaic (using...) sequentially in a pipeline. Convolution kernel), automatic white balance (based on gray pixel method) and color correction matrix ( (Linear transformation), maintaining the linearity of the data throughout the process, and aligning with the noise model pixel by pixel.

[0040] For example, the system processes APS intensity frames (resolution). When the pipeline outputs linear RGB data (12-bit depth), it also integrates event probability model parameters to generate a standard HDF5 format file (approximately 15 megabytes in size), which can be directly input into a simulation platform (such as Python or NumPy environment) for algorithm verification.

[0041] Therefore, according to the above implementation method, the system first acquires raw data from the hybrid event and frame sensors, including intensity frames from the active pixel sensor and event streams from the event vision sensor. Next, noise calibration is performed on the intensity frames to generate an active pixel sensor noise model and a net intensity signal. Then, based on the net intensity signal, an event probability coupling operation is performed on the event stream to generate an event probability model for the event vision sensor. Finally, a pre-defined linear hold image signal processing pipeline performs unified collaborative processing on the active pixel sensor noise model and the event vision sensor event probability model, outputting uniformly calibrated sensor data.

[0042] Specifically, in this implementation, to address the physical coupling weakening problem mentioned in the background technology, a mapping relationship between the net intensity signal and the event triggering probability is established through event probability coupling operations. This solves the technical defect in the prior art where the event model is disconnected from physical quantities such as pixel voltage and threshold, resulting in weak model generalization ability and interpretability. To address the insufficient calibration granularity problem in the background technology, various fixed noises such as line bias, black level, and dark current in the intensity frame are decomposed and subtracted through noise calibration operations, achieving fine-grained modeling and correction of the sensor's inherent noise. To address the black-box implementation problem mentioned in the background technology, the dual models are processed in a unified and collaborative manner through a linear image signal processing pipeline, ensuring the linear transparency of the entire calibration and processing flow and avoiding the drawbacks of non-reproducibility and difficulty in alignment verification brought about by traditional closed-source image signal processing (ISP) pipelines. Therefore, the technical solution of this implementation provides a unified and reproducible cross-modal calibration framework, which solves the technical problems of weakened physical coupling, insufficient calibration granularity and black box implementation in existing calibration technologies, and provides a unified, reproducible and high-precision sensor data calibration output for high-speed, high dynamic range mobile sensing applications.

[0043] In another embodiment, as shown in Table a below: Hybrid vision sensor (APS) RGB array, EVS is The table details the specific parameters of the active pixel sensor (APS) noise model obtained after the calibration method for the RGB array. and different pixel positions within the CFA (e.g.) , (etc.) The fitting coefficients of each color term in the noise variance model. For example, for the Gr (green) channel at position... The pixel, whose noise term is related to illumination intensity ( The coefficient is The noise term related to the square of the exposure time ( The coefficient is These parameters are on the order of magnitude (from arrive The systematic changes in the symbols and symbols precisely quantify the pixel-level non-uniformity introduced by manufacturing processes, color filtering, and photoelectric response nonlinearity. This detailed set of parameters was obtained through the aforementioned process of "collecting massive amounts of data at 8 exposure times and 5 illuminance levels, and fitting the data using weighted least squares." It directly proves that this scheme can achieve fine-grained, high-precision calibration for each CFA position and color channel, rather than providing a global coarse estimate. The noise model built based on this set of parameters, when applied to subsequent image signal processing pipelines, can accurately estimate and compensate for the specific attributes of each pixel, thereby improving the signal-to-noise ratio and consistency of the hybrid sensor output data from the source, providing a better data foundation for downstream high-level vision tasks.

[0044] In another embodiment, as shown in Table b below, is the second-generation hybrid vision sensor (Gen2, with an APS of [missing information]). RGB array, EVS is The active pixel sensor (APS) noise calibration parameters for a monochrome array are obtained. This embodiment further demonstrates the broad applicability of the above calibration method to different pixel array architectures and event sensor configurations. As shown in the table below, for an architecture coupled with a monochrome event sensor and a color APS, this method can also obtain the noise calibration parameters of different spatial locations within each color channel (e.g., Gr channel). and Detailed noise model coefficients for the location. Taking the Gr channel as an example, its noise term, which is linearly related to the net intensity signal ( The coefficient varies in different locations (e.g., in...). place as , and place as This precisely quantifies the pixel response non-uniformity caused by manufacturing deviations. Meanwhile, the noise term, which is linearly related to exposure time ( The coefficient of ) is generally in The magnitude (e.g., the average value of the R channel is) The average value of channel B is ), significantly higher than The sensor correspondence clearly reveals that the Gen2 sensor's noise characteristics are more dependent on exposure time, highlighting the sensitivity and necessity of this calibration scheme in characterizing the physical characteristics of different sensors. The complete parameter set obtained in this embodiment was calibrated through the unified process of "large-scale data acquisition under multiple exposure times and illuminances, and using weighted least squares method". This further verifies that this scheme does not depend on specific sensor designs and can generate a complete, position- and color-related noise digital twin model for hybrid vision systems, thereby ensuring that subsequent processing algorithms can be optimized based on accurate noise priors, improving the overall imaging robustness and dynamic range of the system.

[0045] In another embodiment, as shown in Table c below is Specific parameters for noise calibration of the hybrid vision sensor event vision sensor (EVS), whose active pixel sensor (APS) is... RGB array, Event Vision Sensor (EVS) RGB array. This table precisely displays the event trigger probability model. In the middle, corresponding to the positions of different color filter arrays (CFA) ( The six fitting parameters of ) to Taking the parameters of the Gr channel at position (0,0) as an example, its higher-order coefficients ( Although relatively small, it is comparable to the coefficients of lower-order terms ( Together, these constitute a complete fifth-order polynomial mapping, which fully verifies that the event probability model adopted in this scheme has the ability to characterize the complex nonlinear relationship between brightness and event triggering. Comparing the parameters of different color channels reveals that its core gain term... It varies between 3.16 and 3.59, while the bias term Maintaining relative stability (approximately 6.8), this systematic difference accurately reflects the modulation effect of the transmittance and photoelectric conversion efficiency of different color filters on the event trigger threshold. This set of high-precision parameters was obtained through the aforementioned process of "collecting event data over a long period under constant illumination and independently fitting the data," marking the successful establishment of a pixel-level statistical event generation model coupled with physical characteristics. Integrating this parameter set into a hybrid sensor simulator ensures that the simulated event stream is highly consistent with the statistical characteristics of the real hardware output, thus providing a reliable data foundation for the development and testing of event-based vision algorithms and effectively solving the long-standing challenges of insufficient real data and inadequate simulation fidelity in this field.

[0046] In another embodiment, Table d below shows the event vision sensor (EVS) noise calibration parameters of the second-generation hybrid vision sensor (Gen2), whose active pixel sensor (APS) is... RGB array, while the Event Vision Sensor (EVS) is W (monochrome) array. This embodiment demonstrates that even when an event sensor employs a broad-spectrum response monochrome pixel design, this scheme can still establish a high-precision event triggering probability model. As shown in Table d, for monochrome (White) event pixels, the six key parameters of its event probability model ( to The core gain term was fully calibrated. The bias term is 3.65. The value is 6.90. Compared to the parameters of a color EVS array (as shown in Table c), the parameter set of a monochrome EVS exhibits higher consistency because it does not need to distinguish the light transmission characteristics of different color filters. This directly reflects the accurate mapping of hardware design differences in the statistical model. It is worth noting that its higher-order nonlinear term coefficients (such as...) The parameters, being of the same order of magnitude as the color array model, demonstrate that complex nonlinear relationships are prevalent in event triggering mechanisms, independent of color filtering. This set of parameters was obtained through the same standardized process described above—"acquiring event streams under fixed illumination and fitting an sigmoid function"—further confirming the universality and robustness of this calibration framework for different event sensor architectures (both color and monochrome). The event generation model built based on this parameter set can accurately simulate the response of a monochrome event sensor to changes in scene brightness, providing a crucial, physically interpretable simulation foundation for developing and evaluating event-driven algorithms that do not rely on color information (such as high-speed monocular vision odometry and dynamic scene analysis).

[0047] In some embodiments, a noise calibration operation is performed on the intensity frame to generate an active pixel sensor noise model and a net intensity signal, including: Perform fixed-term noise decomposition to subtract line bias noise, black level noise, and dark current noise from the intensity frame to obtain the net intensity signal.

[0048] Among them, fixed-term noise decomposition refers to the process of separating and removing the inherent systematic noise of the sensor that is independent of the scene brightness; line bias noise refers to the fixed-mode noise caused by the difference between the lines of the sensor readout circuit; black level noise refers to the offset of the base signal output by the sensor under no-light conditions; dark current noise refers to the charge accumulation noise generated by thermal excitation that is proportional to the exposure time.

[0049] Specifically, the system can statistically calculate the mean and variance of black level noise by acquiring multiple sets of dark-field images (e.g., shooting 50 frames continuously with the lens cap on). Then, by calculating the mean of pixels in each row of the image and comparing it with the overall image mean, a distribution map of row offset noise is obtained. Dark current noise is obtained by acquiring dark-field images at different exposure times (e.g., 1 ms, 10 ms, 50 ms) and fitting the slope of the linear relationship between pixel values ​​and exposure time. Finally, a linear subtraction model is used: From the original intensity frame After removing these noises, the net intensity signal is obtained. ,in For row bias noise, This is black level noise. The dark current coefficient, This refers to the exposure time.

[0050] For example, the system acquires dark-field sequences at 25 degrees Celsius using a hybrid sensor of model "XYZ123". The calculated average black level noise is 204.5 DN (Digital Number), the standard deviation of the line offset noise is 1.8 DN, and the dark current coefficient D is 0.15 DN / ms. For a pixel with an exposure time of 20 milliseconds and an original value of 2500 DN, after deducting fixed-term noise, its net intensity signal is calculated as follows: .

[0051] Based on the net intensity signal and exposure time parameters, a variance model fitting is performed to construct a variance model related to brightness and exposure. Then, the parameters are fitted using the weighted least squares method to generate an active pixel sensor noise model.

[0052] Among them, the variance model related to brightness and exposure refers to the mathematical expression used to describe the variation of sensor noise variance with pixel brightness (net intensity signal) and exposure time, which is usually in the form of a low-order polynomial; the weighted least squares method is an optimization algorithm that assigns different weights to data points based on their reliability (usually determined by the number of samplings) in parameter fitting.

[0053] Specifically, the system can be designed with a variance model that includes a linear term for brightness, a squared term, an exposure time term, and their interaction terms. Its general form is as follows: ; Where Var is the noise variance. Net intensity signal, For the exposure time, to These are the six model parameters to be determined. In the experimental design, to reliably estimate all six parameters, the system set up a control variable group far exceeding the number of parameters. Specifically, at five uniformly distributed illuminance levels (10, 50, 100, 200, 500 lux), 100 images were acquired for eight different exposure times (5, 10, 20, 30, 40, 50, 60, 80 ms), resulting in a total of [number missing] images. This design ensures that there are sufficiently rich and independent observation conditions to constrain all model parameters within the two-dimensional control space consisting of brightness and exposure time. After data acquisition, the system calculates the variance statistics of each pixel block under each (illuminance, exposure time) combination. Subsequently, using the square root of the number of samplings (frames) for each data point as the weight, a weighted least squares objective function is constructed, and the optimal model parameters are solved using a numerically stable matrix factorization algorithm (such as QR factorization, which is an algorithm that decomposes a matrix into an orthogonal matrix and an upper triangular matrix).

[0054] For example, after fitting through the above process, the parameters of a certain sensor noise model are obtained as follows: The coefficient of determination for the model fit ( The variance reached 0.98, indicating that the model can accurately predict the noise variance under different working conditions.

[0055] Therefore, according to the above implementation method, the system can accurately separate scene-independent inherent noise from the original intensity frame and establish a statistical model that accurately quantifies the change of noise with imaging conditions (brightness, exposure time). This provides a reliable data foundation and physical constraints for generating high-quality net intensity signals and subsequent cross-modal collaborative processing.

[0056] In some embodiments, the step of performing fixed-term noise decomposition includes: The initial estimate of black level noise is calculated based on the dark field sequence. The row bias noise is calculated by the row mean difference. Defective pixels are identified and dark current noise is calculated by least squares fitting.

[0057] Among them, the dark field sequence refers to a set of multiple frames of image data acquired under the condition that the sensor lens is completely blocked (illuminance <1 lux), which is used to estimate the base noise in the absence of light; the row mean difference refers to calculating the average value of each row of pixels in the image and comparing it with the mean of the whole image to extract the fixed pattern noise caused by the inter-row difference; the least squares fitting is a mathematical optimization method that estimates the model parameters by minimizing the sum of squared residuals, which is used here to fit the linear relationship between pixel response and exposure time.

[0058] Specifically, the system can be executed through the following process: First, acquire at least 50 frames of images under dark conditions and calculate the time average value of each pixel as an initial estimate of black level noise; second, calculate the row mean sequence for each frame of image and subtract the global mean to obtain the row bias noise map; finally, acquire images with multiple exposure times (e.g., 1 ms, 10 ms, 50 ms) under uniform illumination, fit the response curve of each pixel using the least squares method (the slope reflects the dark current coefficient, and the intercept reflects the fixed bias), and identify residual abnormal points (e.g., exceeding 3 times the standard deviation) as defective pixels.

[0059] For example, the system acquires 100 frames of images (resolution) in a dark field at 25 degrees Celsius using a hybrid sensor model "Eiger-V3". The calculated mean black level noise was 205.3 DN (digital unit, DN refers to Digital Number); the line mean difference showed that the maximum line offset was 2.1 DN; by fitting the dark current coefficient with least squares, a pixel had a response value of 100 DN at an exposure time of 5 ms and 400 DN at 20 ms, and the fitted dark current coefficient was 15 DN / ms, and pixels with residuals greater than 10 DN were marked as defect points.

[0060] An iterative optimization approach is used to jointly correct row bias noise, black level noise, and dark current noise. This includes establishing residual convergence criteria and updating noise parameter estimates through multiple iterations until a preset convergence threshold is met.

[0061] Among them, iterative optimization refers to a mathematical method that gradually approaches the optimal solution by updating parameters in a loop; residual convergence criterion refers to the threshold standard used to terminate the iteration during the optimization process, which is usually based on the change in parameters or the rate of change of the objective function value.

[0062] Specifically, after initializing the noise parameters, the system can execute the following iterative loop: In each iteration, the net intensity signal is first calculated using the current parameter estimate, and then the residual between the net signal and the actual observation is calculated; if the relative rate of change of the L2 norm of the residual is less than the threshold (e.g., 0.001), then convergence is determined; otherwise, the noise parameters are updated by gradient descent (e.g., black level noise is adjusted by a step size of 0.1DN, and row bias noise is updated smoothly), with a maximum of 10 iterations to avoid infinite loops.

[0063] For example, the system is initialized with black level noise of 200DN, row bias noise of 0DN, and dark current coefficient of 0.1DN / ms. After 3 rounds of iteration, the residual change rate drops from 0.05 to 0.0005, which is lower than the threshold of 0.001, and the optimization is terminated. The final parameters are corrected to black level noise of 203.5DN, row bias noise of 1.8DN, and dark current coefficient of 0.12DN / ms.

[0064] After completing the joint correction, the linear response characteristics of the net intensity signal are verified, including checking the linearity of the signal output at different exposure times and confirming that the residual distribution conforms to the preset statistical characteristics.

[0065] Among them, linear response characteristic refers to the physical property that the sensor output signal is proportional to the incident light intensity; linearity is determined by the coefficient of determination of the fitted signal-exposure time relationship (…). Quantification; statistical characteristics of residual distribution include the near-zero mean of residuals and the upper limit of standard deviation.

[0066] Specifically, the system can acquire net intensity signals of uniformly lit scenes at multiple exposure times (e.g., 1 ms, 20 ms, 80 ms), and calculate the goodness of fit between the signal and the exposure time using linear regression. The requirement is a value greater than 0.98; simultaneously, analyze the residual histogram to confirm that the mean is within ±0.5DN, the standard deviation is less than 2DN, and it conforms to a normal distribution (through...). The test result was p-value > 0.05.

[0067] For example, when the system was tested at an illuminance of 100 lux, the net intensity signal was 50 DN at an exposure time of 1 ms, 980 DN at 20 ms, and 3920 DN at 80 ms. Linear regression yielded... The mean of the residuals is The standard deviation is 1.5DN, and The value is 0.06, which meets the linearity verification requirements.

[0068] Therefore, according to the above implementation method, the system can accurately separate and correct the inherent noise of the sensor, ensure the stability of parameter estimation through iterative optimization, and verify the physical consistency of the net intensity signal, providing highly reliable input data for subsequent event probability coupling and image processing.

[0069] In some embodiments, the step of performing variance model fitting includes: Based on the net intensity signal and exposure time parameters, a polynomial variance model is established. The polynomial variance model includes a first-order brightness term, a second-order brightness term, an exposure time term, and a brightness-exposure time cross term.

[0070] Among them, the polynomial variance model refers to a parameterized mathematical expression that describes the relationship between noise variance and brightness signal and exposure time through a polynomial function; the first-order brightness term represents the linear correlation between variance and brightness signal; the second-order brightness term represents the nonlinear correlation between variance and the square of brightness; the exposure time term represents the linear relationship between variance and exposure time; and the interaction term between brightness and exposure time represents the influence of the interaction between brightness and exposure time on variance.

[0071] Specifically, the system can construct a variance model expression by collecting multiple sets of data at different brightness levels (e.g., net intensity signal range of 0~10000DN) and exposure times (e.g., 1 millisecond to 80 milliseconds): ; Where Var is the noise variance (unit: ), Net intensity signal (unit: DN). Exposure time (in milliseconds). Model parameters. to The dimensions must be matched: The unit is , The unit is DN. Dimensionless The unit is , The unit is , The unit is The system uses a weighted least squares framework to fit the parameters, with the square root of the number of samplings as the weights, to ensure that the model accurately captures the illuminance-exposure coupling effect.

[0072] In the experimental design, to meet the estimation requirements of six parameters, the system set up a set of control variables far exceeding the number of parameters. Specifically, at five illuminance levels... Below, regarding the 8 sets of exposure times Data was collected separately, with 100 images acquired for each condition, resulting in a total of [data missing]. This design ensures sufficient independent observations in the two-dimensional space of brightness-exposure time, avoiding underdeterminism. For example, within an illuminance range of 50 to 500 lux, the system uses the aforementioned 8 sets of exposure time data to fit the variance model parameters of a certain sensor as follows: ; When the net intensity signal is 2000 DN and the exposure time is 20 milliseconds, the prediction variance is calculated as follows: ; Model fit determination coefficient ( The accuracy was verified to be 0.98.

[0073] The weight coefficients of the weighted least squares objective function are set based on the square root of the number of samplings, and the model parameters of the weighted least squares objective function are solved by matrix factorization algorithm.

[0074] Among them, the weighted least squares objective function refers to the optimization objective function that introduces weights to balance the reliability of data points; the weight coefficients are set according to the square root of the number of samplings to reduce the error impact of low statistical data points; matrix factorization algorithm refers to the numerical method of solving linear systems by decomposing the design matrix, such as QR decomposition.

[0075] Specifically, the system can set weights. ,in For the first The objective function is constructed by sampling a number of data points: ; For large datasets generated by high-resolution sensors (e.g., for 8 sets of exposure times, resolution of...), The system utilizes sensors with over 63 million effective data points to solve large-scale linear least squares problems using a QR decomposition algorithm. This algorithm effectively avoids ill-conditioned conditions in the design matrix, ensuring the numerical stability of parameter estimation. The solution process includes constructing a sparse design matrix, block-based QR decomposition to reduce memory usage, and back-substitution calculation of the parameter vector.

[0076] For example, the system processes a resolution of The mixed sensor data generated approximately 63 million valid data points across 8 exposure times. Through an optimized QR decomposition algorithm (such as a block-based implementation based on Householder transform), the parameter estimation error can be controlled within a certain range on a computing platform equipped with a standard CPU (such as an Intel Xeon E5-2680 v4). Within this timeframe, the overall computation time is approximately 25 seconds. This demonstrates the system's high efficiency and numerical robustness when processing massive amounts of sensor data.

[0077] The goodness-of-fit index is calculated based on the fitting results to evaluate the accuracy of the model. The residual distribution characteristics are analyzed, and the consistency of the model prediction is verified under different illumination conditions to complete the steps of fitting the variance model.

[0078] Among them, the goodness-of-fit index refers to a statistical measure that quantifies the degree of matching between the model's predictions and the observed data, such as the coefficient of determination (COP). The residual distribution characteristics refer to the statistical properties of the difference between observed and predicted values, including the mean, standard deviation, and distribution pattern; model prediction consistency refers to the ability of a model to maintain prediction accuracy under different external conditions (such as illumination).

[0079] Specifically, the system can calculate the coefficient of determination. Evaluate the overall goodness of fit and analyze the residuals. The mean (close to 0) and standard deviation (less than a threshold, e.g.) Furthermore, the prediction error was verified under low illumination (10 lux), medium illumination (100 lux), and high illumination (500 lux) to ensure that the maximum relative error was less than 5%.

[0080] For example, after fitting, we get The mean of the residuals is The standard deviation is At 10 lux illuminance, the prediction variance is... The measured value is The relative error is 4%; at 500 lux, the prediction variance is... The measured value is The relative error is 2.2%, which meets the consistency requirements.

[0081] Therefore, according to the above implementation method, the system can establish a high-precision noise variance model, ensure the robustness of parameter estimation through weighted optimization and matrix factorization, and guarantee the reliability of the model in different scenarios based on statistical verification, providing an accurate description of the noise characteristics of sensor data and supporting the stability of subsequent event coupling and image processing processes.

[0082] In some embodiments, based on the net intensity signal, an event probability coupling operation is performed on the event stream to generate an event probability model for the visual sensor, including: Based on the configured multi-brightness static scene configuration, net intensity signals and event streams are acquired under dark, medium-brightness, and high-brightness conditions.

[0083] Among them, the multi-brightness static scene configuration refers to setting up a static experimental environment by controlling the light intensity at multiple fixed levels to cover the working range of the sensor; the dark field condition refers to a lightless environment with an illuminance of less than 1 lux (lux is a unit of illuminance, referring to the luminous flux per square meter); the medium brightness condition refers to a typical indoor lighting environment with an illuminance between 10 lux and 100 lux; and the high brightness condition refers to a strong light or outdoor lighting environment with an illuminance between 500 lux and 2000 lux.

[0084] Specifically, the system can use a synchronous triggering mechanism to acquire intensity frames from the Active Pixel Sensor (APS) and event streams from the Event Vision Sensor (EVS) in parallel using the same optical lens, ensuring that the data is aligned in time and space. During acquisition, a fixed exposure time series (e.g., 1 millisecond, 10 milliseconds, 50 milliseconds, where millisecond is a unit of time, referring to one-thousandth of a second) and an event detection threshold are set, and data is acquired for a sufficient duration (e.g., 60 seconds) for each condition to statistically analyze event probabilities.

[0085] For example, the system is deployed in smart glasses applications. The hybrid sensor, under dark conditions (illuminance <1 lux), acquires a net intensity signal with an average value close to 0 DN (Digital Number), and an event stream background event rate of less than 0.1 events / second. Under medium brightness conditions (illuminance 50 lux), the net intensity signal ranges from 100 DN to 1000 DN, with an event trigger rate of approximately 5 events / second. Under high brightness conditions (illuminance 800 lux), the net intensity signal reaches as high as 5000 DN, and the event trigger rate increases to 20 events / second. The acquired data is stored as a timestamp-aligned sequence file, with a single experiment data volume of approximately 10 megabytes (MB refers to megabytes, 1 MB equals...). byte).

[0086] An event probability parameterization model is established based on the correspondence between the net intensity signal and the event trigger probability. The event probability parameterization model is configured to describe the nonlinear mapping relationship between the brightness input and the probability output.

[0087] Among them, the event probability parameterization model refers to the model that formally describes the change of event trigger probability with net intensity signal through mathematical functions. It usually uses sigmoid functions (such as the standard normal function) to capture nonlinear relationships.

[0088] Specifically, the system can define the model as follows: ,in This represents the probability of an event being triggered (dimensionless). Represents net strength signal (unit: DN). For model parameters, It is an S-shaped function, in the form of This model adapts to the complex mapping between brightness and probability through polynomial terms. For example, for a given pixel, the probability of observing an event is 0.3 when the net intensity signal is 500 DN, and 0.7 when it is 1000 DN. The parameters are obtained after fitting. Then when At that time, calculate .

[0089] The maximum likelihood estimation method is used to fit the parameters of the event probability parameterized model, and the optimal model parameters are obtained by solving the gradient descent optimization algorithm.

[0090] Among them, Maximum Likelihood Estimation (MLE) is a parameter estimation method that solves for the optimal parameters by maximizing the probability of the observed data occurring; Gradient Descent is an iterative optimization algorithm that minimizes the loss function by updating the parameters along the negative gradient direction of the objective function.

[0091] Specifically, the system can assume that event generation follows a Bernoulli distribution and construct a log-likelihood function: ,in This is a binary event observation metric (1 indicates triggering, 0 indicates not triggering). Using a gradient descent algorithm (such as the Adam optimizer), with a learning rate of 0.001 and 200 iterations (epoch refers to the number of iterations), the likelihood function is monitored. Convergence is determined when the change is less than 0.0001 for 10 consecutive iterations. For example, the system uses 1000 data points (each sampled 100 times), with initial parameters set to... After 150 iterations, the algorithm converged, yielding the optimal parameters. The final likelihood function value increased to (From the beginning) ).

[0092] Based on the optimal model parameters, and combined with the independent calibration results of different color filter array positions and color channels, an event probability model for the event visual sensor is generated.

[0093] Among them, the color filter array (CFA) position refers to the arrangement pattern of color filters on the sensor surface (such as the R, G, and B pixel positions in a Bayer array); the color channel refers to the spectral response channel corresponding to different filters (such as red, green, and blue); independent calibration refers to parameter estimation for each position and channel individually to address manufacturing inhomogeneities.

[0094] Specifically, the system can be configured for each CFA location (e.g., each sub-pixel in a 4×4 period) and color channel (e.g. The above parameter fitting process is executed independently to generate a set of parameter tables (e.g., one set for each position). This data is then stored as a lookup table or configuration file. When the model outputs, the event probability is calculated based on the parameters corresponding to the pixel position index.

[0095] For example, for a Quad-Bayer sensor, the system is calibrated for 16 CFA locations (each location representing a color combination): at the red channel location, the fitting parameters are... At the green channel location, the parameters are: The final generated event probability model contains 16 sets of parameters, and the model file size is approximately 1 kilobyte (KB refers to kilobyte, 1KB equals...). byte).

[0096] Therefore, according to the above implementation method, the system can generate a high-precision event probability model for an event vision sensor by acquiring multi-brightness scene data, establishing a parameterized model, fitting an optimization algorithm, and fine-grained calibration. This model accurately captures the physical statistical relationship between brightness and event probability and adapts to differences in sensor manufacturing, providing a reliable description of event triggering characteristics for hybrid sensors.

[0097] In some embodiments, the maximum likelihood estimation method is used to fit the parameters of the event probability parameterized model, and the optimal model parameters are obtained by solving the gradient descent optimization algorithm, including: A parameter estimation objective function is constructed based on the probability distribution relationship between event-triggered observation data and model predictions. The objective function is configured to quantify the likelihood between the observation data and the model predictions.

[0098] Among them, the probability distribution relationship refers to the statistical characteristic that the event-triggered behavior conforms to the Bernoulli distribution (a discrete probability distribution describing a binary random variable); the parameter estimation objective function refers to the optimization objective constructed based on the maximum likelihood principle, usually in the form of a log-likelihood function.

[0099] Specifically, the system can define the event-triggered observation as a binary sequence. (1 indicates triggered, 0 indicates not triggered), the model predicts the event probability. ,in Net intensity signal, These are the model parameters. The objective function is constructed as a negative log-likelihood function: ; The smaller the value of this function, the higher the likelihood.

[0100] For example, the system uses 1000 observation data points (N=1000), of which 250 events are triggered and 750 are not. When the model predicts the probability... At that time, for a single triggering event ( The likelihood contribution of ) is For events that have not been triggered ( The contribution of 0) is The initial objective function value is calculated as follows: .

[0101] Initialize the parameter vector of the event probability parameterized model, and configure the optimization algorithm and its hyperparameters for the current optimization task. The optimization algorithm refers to the gradient descent optimization algorithm, and the hyperparameters include the learning rate parameter and the convergence threshold parameter.

[0102] Here, the parameter vector refers to the set of parameters of the model to be optimized, such as... The learning rate parameter controls the step size of gradient descent; the convergence threshold parameter is used to determine the termination condition of optimization.

[0103] Specifically, the system can set the initial value of the parameter vector. The Adam optimizer (Adaptive Moment Estimation, a gradient descent algorithm with an adaptive learning rate) was selected as the optimization algorithm, with a learning rate α = 0.001 and a convergence threshold of [missing information]. Maximum number of iterations .

[0104] For example, initialization parameters Optimizer settings learning rate First-order moment decay factor Second-order moment attenuation factor Convergence threshold .

[0105] The gradient descent optimization algorithm is used to perform iterative optimization calculations. The gradient information is used to update the parameter estimates of the event probability parameterized model and monitor the convergence status of the optimization process. The gradient information is obtained by taking the partial derivative of the parameter estimation objective function.

[0106] Here, gradient information refers to the vector of partial derivatives of the objective function with respect to each parameter; the iterative optimization calculation process refers to the calculation process of updating parameters in a loop; and the convergence state is monitored by the objective function value or the change in parameters.

[0107] Specifically, in the t-th iteration, the system performs the following: First, it calculates the gradient. ; in Then update the parameters. ,in , First and second moment estimates for the Adam optimizer; finally, the change in the objective function is calculated. .

[0108] For example, the gradient is calculated in the first iteration. After the update, the parameters become The objective function value decreased from 570.0 to 568.3, with a change of... .

[0109] In response to the detection that the optimization process has reached the preset convergence condition, the iterative optimization calculation process ends and the optimal model parameters are obtained.

[0110] The convergence condition refers to the pre-set optimization stopping criteria, including the threshold for the change in the objective function and the maximum number of iterations.

[0111] Specifically, the system monitors the relative change of the objective function across multiple consecutive iterations: When this value is less than the convergence threshold The optimization is terminated immediately upon failure; a maximum number of iterations is set to prevent infinite loops. After optimization terminates, the current parameter estimates are output as the optimal model parameters. For example, when the 150th iteration is reached, the objective function value is... The difference between the previous iteration and the previous iteration relative change The convergence condition is met. The optimal parameters are then obtained. The total calculation time was 1.5 seconds.

[0112] Therefore, according to the above implementation method, the system can achieve efficient and stable solution of event probability model parameters by constructing a probabilistic statistical objective function, configuring and optimizing hyperparameters, performing gradient descent iterations and monitoring convergence conditions, thus providing a precise foundation for modeling event triggering characteristics of hybrid sensors.

[0113] In some embodiments, the linear-preserving image signal processing pipeline includes a black level correction module, a pixel layout conversion module, a linear demosaic module, an automatic white balance module, and a color correction module that sequentially perform processing operations; the pipeline performs unified collaborative processing on the active pixel sensor noise model and the event probability model of the event vision sensor through a preset linear-preserving image signal processing pipeline, outputting uniformly calibrated sensor data, including: Fixed noise component subtraction is performed by the black level correction module, which removes fixed pattern noise from the original sensor data based on the line bias noise, black level noise and dark current noise parameters in the noise model.

[0114] Fixed-mode noise refers to spatially fixed noise patterns introduced during sensor manufacturing that do not change with the scene, including inter-row differences, substrate offset, and thermally excited noise.

[0115] Specifically, the system can apply linear subtraction to each pixel of the raw sensor data by reading noise model parameters (such as row bias noise map, average black current, and dark current coefficient): ,in This is the original pixel value (unit: DN, digital unit). Line bias noise (unit: DN). D represents black level noise (unit: DN), and D is the dark current coefficient (unit: DN / ms). Exposure time (in milliseconds).

[0116] For example, the system processes data from a sensor with a resolution of 1280×720, the raw pixel values... The noise parameters are , D=0.12DN / ms, exposure time Milliseconds. The corrected pixel value is calculated as follows: The entire image processing takes less than 5 milliseconds (on a standard CPU).

[0117] The pixel layout conversion module performs array format conversion, identifies the mixed pixel array arrangement pattern and establishes a mapping relationship, and converts the mixed pixel array into a standard color filter array format.

[0118] Among them, hybrid pixel array refers to a non-standard arrangement of sensor filter layout (such as a Quad-Bayer structure, where each 2×2 pixel block contains sub-pixels of the same color); standard color filter array format refers to the common Bayer array (a pattern in which RGB filters are arranged alternately).

[0119] Specifically, the system can detect the periodic structure of the source array using pattern recognition algorithms (such as template matching), establish mapping rules from source pixels to target pixels (such as the weighted average aggregation of four sub-pixels of the same color into a standard Bayer pixel in Quad-Bayer), and apply a linear weighting formula. ,in For sub-pixel values, This is the weight (usually 1).

[0120] For example, converting a Quad-Bayer array (2560×1440 resolution) to a standard Bayer array (1280×720 resolution), each 2×2 sub-pixel of the same color (with values ​​respectively) Aggregated into a single pixel value The converted data volume was reduced by 75%, and the processing speed reached 30 frames per second.

[0121] Color interpolation is performed by a linear demosaic module, and full-color image data is restored using a linear filter kernel based on the distribution pattern of the color filter array.

[0122] The linear filter kernel refers to a set of convolution weight matrices with fixed coefficients, which is used to calculate the missing color value by weighted summation, avoiding nonlinear operations (such as edge detection).

[0123] Specifically, the system can predefine 5×5 convolution kernels (e.g., for different color positions in the Bayer array, such as R, G, and B pixels) Algorithm kernel), through convolution operations The missing color is interpolated, where K is the filter kernel coefficient. The input pixel values ​​are used. Mirror padding is used for boundary processing to avoid artifacts.

[0124] For example, for a missing R value at pixel G, the system uses the standard Malvar kernel as follows: ; Next, a local 5×5 pixel value (range 0~255DN) is input, convolved, and the interpolated R value is output. Processing a 1280×720 image takes approximately 10 milliseconds.

[0125] The automatic white balance module performs color balance processing, analyzes color statistical characteristics within a specific color space, and applies linear scaling transformation.

[0126] Here, a specific color space refers to the color representation domain used for white balance analysis (such as logarithmic chromaticity space). (A space that separates luminance and chrominance through logarithmic transformation); linear scaling transformation refers to adjusting the gain of each color channel through a diagonal matrix.

[0127] Specifically, the system can convert image data to a logarithmic color space, detect gray pixels (with chromaticity values ​​close to neutral), statistically analyze their distribution, and calculate the values ​​for each channel. Gain coefficient Then apply scaling: ; in, The input pixel vector.

[0128] For example, at an illuminance of 500 lux (lux is a unit of illuminance), statistically analyze 10,000 gray pixels and calculate the gain. Input pixel values The output after balancing is The white balance error is less than 2%.

[0129] The color correction module performs color space conversion, applying a linear transformation matrix to convert the image data to the target color space.

[0130] Here, the linear transformation matrix refers to a... or A real matrix is ​​used to map the sensor's original color space to a standard space (such as...). A standard color space for displays; the target color space refers to an industry-standard color representation system (such as...). ).

[0131] Specifically, the system can load a pre-calibrated color correction matrix (CCM) and perform matrix multiplication. Perform the transformation, where M is a 3×3 transformation matrix and b is a bias vector (3×1). The input is a color vector. The matrix parameters are obtained by calibration using a standard color chart (such as 24-color ColorChecker).

[0132] For example, the system uses: ; Next, input pixels. The output is calculated as follows: Color error after conversion (CIELAB units, an international standard colorimetric system for quantifying color differences, officially established by the International Commission on Illumination in 1976).

[0133] Therefore, according to the above implementation method, the system can achieve high-precision correction, format standardization, color reconstruction and spatial transformation of sensor data through multi-module collaboration of the linear pipeline, ensuring that the output data is strictly aligned with the noise model and event model, and providing a unified and reliable calibration data foundation for downstream vision tasks.

[0134] Figure 3 This is a structural block diagram of a hybrid event and frame sensor processing system according to an embodiment of the present invention.

[0135] like Figure 3 As shown, the processing system for the hybrid event and frame sensor includes: The raw data acquisition module 210 is used to acquire raw data from the hybrid event and frame sensors, including intensity frames from the active pixel sensor and event streams from the event vision sensor.

[0136] The noise calibration processing module 220 is used to perform noise calibration operations on the intensity frame to generate an active pixel sensor noise model and a net intensity signal.

[0137] The probability model generation module 230 is used to perform event probability coupling operations on the event stream based on the net intensity signal to generate an event probability model for the visual sensor.

[0138] The processing result output module 240 is used to perform unified collaborative processing on the active pixel sensor noise model and the event probability model of the event vision sensor through a preset linear image signal processing pipeline, and output uniformly calibrated sensor data.

[0139] The specific functions and examples of each module and submodule of the device in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0140] According to embodiments of the present invention, the above-described method of the present invention can be applied to a computer device and a readable storage medium.

[0141] Figure 4 A schematic block diagram of an example computer device 600 that can be used to implement embodiments of the present invention is shown. The computer device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The computer device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0142] like Figure 4 As shown, the computer device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the computer device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0143] Multiple components in computer device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows computer device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0144] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a hybrid event and frame sensor processing method. For example, in some embodiments, a hybrid event and frame sensor processing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the computer device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the hybrid event and frame sensor processing method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured by any other suitable means (e.g., by means of firmware) to perform a hybrid event and frame sensor processing method.

[0145] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0146] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0147] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0149] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0150] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0151] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0152] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for processing hybrid event and frame sensors, characterized in that, include: Acquire the raw data from the hybrid event and frame sensor, the raw data including intensity frames from the active pixel sensor and event streams from the event vision sensor; Perform noise calibration on the intensity frame to generate an active pixel sensor noise model and a net intensity signal; Based on the net intensity signal, an event probability coupling operation is performed on the event stream to generate an event probability model for the visual sensor. The active pixel sensor noise model and the event probability model of the event vision sensor are processed in a unified and coordinated manner through a preset linear image signal processing pipeline, and the sensor data with unified calibration is output.

2. The method according to claim 1, characterized in that, The step of performing noise calibration on the intensity frame to generate an active pixel sensor noise model and a net intensity signal includes: Perform fixed-term noise decomposition to subtract line bias noise, black level noise, and dark current noise from the intensity frame to obtain the net intensity signal; Based on the net intensity signal and exposure time parameters, a variance model fitting is performed to construct a variance model related to brightness and exposure. Then, the parameters are fitted using the weighted least squares method to generate the active pixel sensor noise model.

3. The method according to claim 2, characterized in that, The step of performing fixed-term noise decomposition includes: The initial estimate of black level noise is calculated based on the dark field sequence, the row bias noise is calculated by the row mean difference, and the defective pixels are identified and the dark current noise is calculated by least squares fitting. The row bias noise, black level noise and dark current noise are jointly corrected by using an iterative optimization method, including establishing residual convergence judgment conditions and updating the noise parameter estimates through multiple iterations until the preset convergence threshold is met. After completing the joint correction, the linear response characteristics of the net intensity signal are verified, including checking the linearity of the signal output under different exposure times and confirming that the residual distribution conforms to the preset statistical characteristics.

4. The method according to claim 2, characterized in that, The step of performing variance model fitting includes: Based on the net intensity signal and exposure time parameters, a polynomial variance model is established, which includes a first-order brightness term, a second-order brightness term, an exposure time term, and a brightness-exposure time cross term. The weight coefficients of the weighted least squares objective function are set based on the square root of the number of samplings, and the model parameters of the weighted least squares objective function are solved by matrix factorization algorithm. The goodness-of-fit index is calculated based on the fitting results to evaluate the accuracy of the model. The residual distribution characteristics are analyzed, and the consistency of the model prediction is verified under different illumination conditions to complete the steps of fitting the variance model.

5. The method according to claim 1, characterized in that, The step of performing an event probability coupling operation on the event stream based on the net intensity signal to generate an event visual sensor event probability model includes: Based on the configured multi-brightness static scene configuration, the net intensity signal and event stream are collected under dark field conditions, medium brightness conditions and high brightness conditions; An event probability parameterization model is established based on the correspondence between the net intensity signal and the event trigger probability. The event probability parameterization model is configured to describe the nonlinear mapping relationship between luminance input and probability output. The maximum likelihood estimation method is used to fit the parameters of the event probability parameterization model, and the optimal model parameters are obtained by solving the gradient descent optimization algorithm. Based on the optimal model parameters, the event probability model of the event visual sensor is generated by combining the independent calibration results of different color filter array positions and color channels.

6. The method according to claim 5, characterized in that, The step of fitting parameters to the event probability parameterized model using the maximum likelihood estimation method and obtaining the optimal model parameters through the gradient descent optimization algorithm includes: A parameter estimation objective function is constructed based on the probability distribution relationship between event-triggered observation data and model predictions. The objective function is configured to quantify the likelihood between the observation data and the model predictions. Initialize the parameter vector of the event probability parameterization model, and configure the optimization algorithm and hyperparameter settings of the optimization algorithm for the current optimization task. The optimization algorithm refers to the gradient descent optimization algorithm, and the hyperparameters include the learning rate parameter and the convergence judgment threshold parameter. The gradient descent optimization algorithm is used to perform an iterative optimization calculation process, update the parameter estimates of the event probability parameterized model with gradient information, and monitor the convergence state of the optimization process. The gradient information is obtained by taking the partial derivative of the parameter estimation objective function. In response to the detection that the optimization process has reached the preset convergence condition, the iterative optimization calculation process ends, and the optimal model parameters are obtained.

7. The method according to claim 1, characterized in that, The linear-preserving image signal processing pipeline includes a black level correction module, a pixel layout conversion module, a linear demosaic module, an automatic white balance module, and a color correction module that sequentially perform processing operations. The pipeline performs unified collaborative processing on the active pixel sensor noise model and the event probability model of the event visual sensor through a preset linear-preserving image signal processing pipeline, outputting uniformly calibrated sensor data, including: The black level correction module performs fixed noise component subtraction, removing fixed pattern noise from the original sensor data based on the line bias noise, black level noise, and dark current noise parameters in the noise model. The pixel layout conversion module performs array format conversion, identifies the mixed pixel array arrangement pattern and establishes a mapping relationship, and converts the mixed pixel array into a standard color filter array format. The linear demosaic module performs color interpolation processing, and a linear filter kernel is used to restore the full-color image data according to the distribution pattern of the color filter array. The automatic white balance module performs color balance processing, analyzes color statistical characteristics within a specific color space, and implements linear scaling transformation. The color correction module performs color space conversion, applying a linear transformation matrix to convert image data to the target color space.

8. A processing system for hybrid event and frame sensors, characterized in that, include: The raw data acquisition module is used to acquire the raw data of the hybrid event and frame sensor, the raw data including intensity frames from the active pixel sensor and event streams from the event vision sensor; The noise calibration processing module is used to perform noise calibration operations on the intensity frame to generate an active pixel sensor noise model and a net intensity signal. The probability model generation module is used to perform an event probability coupling operation on the event stream based on the net intensity signal to generate an event visual sensor event probability model; The processing result output module is used to perform unified collaborative processing on the active pixel sensor noise model and the event probability model of the event vision sensor through a preset linear image signal processing pipeline, and output uniformly calibrated sensor data.

9. A computer device, comprising: At least one processor; and a memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-7.