A real-time photometric stereo vision method and device based on an event camera

By designing the illumination patterns of the event camera and using a null vector fusion method, combined with SVD and deep learning, real-time reconstruction of photometric stereo vision using the event camera was achieved. This solved the problems of complexity, time consumption, and insufficient information prediction in traditional methods, and improved the quality of normal estimation and the speed of data acquisition.

CN118781255BActive Publication Date: 2026-01-06PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410769824.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2026-01-06
Estimated Expiration
2044-06-14

AI Technical Summary

Technical Problem

Traditional photometric stereo vision methods based on frame cameras are complex and time-consuming, making it difficult to meet real-time requirements. Hybrid camera systems have insufficient prediction of information in overexposed areas and cannot predict information in underexposed areas, resulting in limited dynamic range reconstruction by existing methods.

Method used

By leveraging the unique properties of event cameras, a lighting pattern is designed. Lighting signals are synchronously acquired through the event camera, null vectors are constructed, and surface normals are solved by fusing them. Surface normal estimation is optimized by combining SVD and deep learning methods, and real-time reconstruction is performed using an event camera photometric stereo vision neural network.

Benefits of technology

It enables real-time photometric stereo vision using a single event camera, reducing the cost of multiple cameras and synchronization, improving data sampling density and normal estimation quality, and accurately acquiring information about high-contrast and high-specular-reflection objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118781255B_ABST
    Figure CN118781255B_ABST
Patent Text Reader

Abstract

A real-time photometric stereo vision method and device based on an event camera, belonging to the field of computer vision, includes: lighting pattern design; simultaneously acquiring lighting signals and synchronization signals generated by lighting changes in the test object using an event camera; using the synchronization signals to calibrate the lighting direction when other events occur, achieving synchronous acquisition; constructing null vectors directly related to the surface normals of the test object from each pair of consecutively occurring events; and fusing the null vectors to solve for the surface normals. This invention reconstructs the surface normals of the test object from the basic model triggered by the event camera, reducing the cost of multiple cameras and synchronization; it utilizes the high event resolution of the event camera to acquire data under continuously changing light sources, increasing data sampling density and improving the quality of normal estimation; and it leverages the high dynamic range of the event camera to improve data acquisition speed and quality through native high dynamic range and compressed event representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision technology, specifically relating to a real-time photometric stereo vision method and device based on an event camera. Background Technology

[0002] With the development of computer technology and the increasing computing power, machine learning and deep learning technologies have made rapid progress, leading to the application of computer vision technologies in various scenarios. These include functions such as face detection, image editing, and nighttime photography in mobile phone cameras; pedestrian detection and road recognition in autonomous driving; face recognition for mobile payments and station identity verification; and simultaneous localization and mapping (SLAM) tasks for robots. With the advent of the era of big data and artificial intelligence, more and more application scenarios require the support of computer vision technology. The massive amounts of video and image data that need processing highlight the importance of underlying visual tasks. Therefore, the irreplaceable nature of underlying image processing technologies and their significance for higher semantic-level tasks have attracted widespread attention. High dynamic range imaging, as a fundamental task of computational photography, is extremely important for the development of other computer vision technologies.

[0003] Photometric stereo is a technique that estimates the surface normal direction by analyzing images of an object illuminated from various directions. Its unique advantage lies in its ability to reconstruct high-resolution and accurate surface details, particularly under conditions of densely sampled illumination and Lambertian reflections. Traditional frame-based photometric stereo methods are often complex and time-consuming, typically requiring the capture of a series of exposure-bound images to synthesize a high dynamic range image, thereby accurately capturing specular reflection areas on the object's surface. To achieve dense sampling of illumination directions, traditional frame-based photometric stereo methods necessitate the continuous movement of the light source using a robotic arm, making data acquisition time-consuming and labor-intensive, severely hindering applications with real-time requirements.

[0004] The paper "Surface Enhancement Using Real-Time Photometric Stereo and Reflectance Transformation" discloses a photometric stereo vision method. This method first uses multiple electronically controlled light sources, positioned around the test object, to calibrate the light source positions. It then uses a frame camera that supports external triggering to photograph the object. Circuitry synchronizes the rapid switching of light sources and camera exposure. Furthermore, it uses eight LED light sources to switch illumination, with the camera shooting at 60 frames per second. The specific implementation process is as follows: during image capture, electronic circuitry controls the eight LED light sources to illuminate sequentially; the direction and brightness of all LED light sources are calibrated to generate an illumination matrix; an observation matrix is ​​constructed based on the brightness and irradiance values ​​obtained from a single pixel image; and the least squares error is optimized to calculate the surface normal. The method disclosed in the aforementioned paper requires capturing multiple LDR images with different exposures, demanding high shooting skills. A stable camera is essential to prevent shaking, and there must be no moving objects in the scene; otherwise, the final composite result may be misaligned or produce ghosting, significantly reducing the quality of the reconstruction.

[0005] The paper "EventFusion Photometric Stereo Network" discloses a hybrid camera system consisting of an event camera and a frame camera. Under continuous moving light sources, it simultaneously captures events generated during the light source's movement and absolute brightness and irradiance information captured by a traditional frame camera when the light source reaches key points. The system uses events to interpolate the absolute brightness, and then uses the interpolated absolute brightness to achieve photometric stereo vision. The specific implementation process is as follows: First, a hybrid camera system consisting of an event camera and a traditional frame camera is built, and the images from both cameras are aligned in pixels and time. Under continuous moving light sources, events and images are simultaneously acquired. An event interpolation network is used to fuse the event and image information to obtain a high-density irradiance observation map. This high-density irradiance observation map is then input into an existing photometric stereo vision network to obtain predicted surface normals. While the method disclosed in the paper does not require multiple photos for synthesis, only a single LDR image is needed. However, this method can only predict information about overexposed areas, not underexposed areas. Furthermore, the prediction of overexposed areas largely depends on the network's training performance, relying on the network's learning experience, leading to predictions that do not match reality and resulting in a very limited reconstructed dynamic range. Summary of the Invention

[0006] The purpose of this invention is to provide a real-time photometric stereo vision method and apparatus based on an event camera. This invention utilizes the unique properties of event cameras to perform real-time photometric stereo vision.

[0007] The technical solution adopted by this invention to solve the technical problem is as follows:

[0008] This invention provides a real-time photometric stereo vision method based on an event camera, comprising the following steps:

[0009] Step 1: Lighting mode design;

[0010] Step 2: Simultaneously acquire the illumination signal and synchronization signal generated by the change in illumination of the test object using an event camera; use the synchronization signal to determine the illumination direction when other events occur, thus achieving synchronous acquisition;

[0011] Step 3: Starting from each pair of consecutive events, construct a null vector that is directly related to the surface normal of the object being tested;

[0012] Step 4: Solve for surface normals by fusion of null vectors.

[0013] Furthermore, the lighting modes are circle lighting mode, hypotrochoid lighting mode, and DiLiGenT lighting mode.

[0014] Furthermore, in step three, assuming that during the synchronous acquisition of events by the camera, the event triggering threshold is C, and pixel x triggers a total of K events; the polarity of these events is p. k The event occurs at time t. k , t k-1 Let k = {1, 2, ..., K} represent the times when adjacent events of the same pixel occur; x (t k () represents the absolute pixel brightness at the time the event is triggered; according to the event-camera triggering model, the following equation holds:

[0015] I x (t k ) = exp(p k C)I x (t k-1 )

[0016] According to the ideal Lambert model, assuming the normal direction of the surface of the object being tested is n x The diffuse reflectance is a x The direction of illumination at the time of the event is L(t) k The following equation holds true:

[0017] max[0, a x (n x ·L(t k ))]=exp(p k C)max[0, a x (n x ·L(t k-1 ))]

[0018] When the direction of illumination changes in the same way, the diffuse reflectance of the tested object's surface does not affect the event triggering; therefore, remove a. x The following equation holds true:

[0019] max[0, (n x ·L(t k ))]=exp(p k C)max[0,(n x ·L(t k-1 ))]

[0020] Based on the event camera triggering model, events only occur at moments where pixel brightness changes; therefore, it is inferred that at the moment the event occurs, the derivative of the pixel's absolute brightness with respect to time must be non-zero; the max operator on both sides of the formula is used to describe shadows in Lambertian reflections; when the max value is 0, the derivative of the pixel's absolute brightness with respect to time is also 0, which is related to the event occurrence time t. k Contradiction; therefore, at any moment t when the event occurs k Both sides of the equation should be greater than 0, therefore the event signal does not contain redundant information about pixels in the shadow region; removing the max operator from both sides of the equation, the following equation holds true:

[0021] n x ·L(t k ) = exp(p k C)(n x ·L(t k-1 ))

[0022] After linear transformation, it becomes:

[0023] n x ·(L(t k )-exp(p k C)L(t k-1 ))=0

[0024] According to the above formula, a pair of consecutive events can be converted into a vector perpendicular to the normal direction of the surface of the object being tested, which is the null vector, and its specific form is as follows:

[0025] z k =L(t) k+1 )-exp(p k+1 C)L(t k )

[0026] Among them, t k+1 and p k+1 These represent the time and polarity of adjacent events occurring at the same pixel.

[0027] Furthermore, in step four, after the tested object completes one or more cycles of sampling according to the designed lighting pattern, each image pixel will obtain a continuous event sequence containing several events; each pair of adjacent events in these continuous event sequences is converted into a null vector, and each null vector describes part of the information of the noisy surface normal; then, the surface normal is solved by fusing multiple null vectors.

[0028] Furthermore, in step four, the surface normal is solved using the SVD method: the K-1 null vectors generated from pixel x are concatenated into a null vector matrix Z of size 3×(K-1). k Theoretically, at least three events need to be triggered by a pixel to form a null vector matrix with a rank of at least 2 before further solutions can be obtained; the solution objective is to optimize the following mean square error function:

[0029]

[0030] in, For estimating the surface normal; for solving The value, for Perform eigenvalue decomposition to solve for the corresponding... The eigenvector of the matrix with the smallest eigenvalue.

[0031] Furthermore, in step four, a deep learning method is used to solve for the surface normals: by modifying the input of the existing photometric stereo vision neural network based on the frame camera, it is retrained and transferred to obtain a photometric stereo vision neural network suitable for the event camera; the surface normals are solved by fusing the null vectors of the photometric stereo vision neural network suitable for the event camera.

[0032] Furthermore, the method for solving surface normals using the photometric stereo vision neural network for event cameras by fusing null vectors is EventPS-FCN: the entire illumination mode time is divided into N equal-length intervals, and all null vectors in each interval and each pixel are superimposed to form a null vector image; for different pixels in each null vector image, the illumination direction changes in the same way, and the null vector image and the corresponding illumination direction are used as inputs, and two temporal convolutional layers are added to the input layer of the original PS-FCN network to extract temporal features from events in adjacent time intervals; the other parts of the photometric stereo vision neural network for event cameras follow the original PS-FCN network structure, extracting feature images through a feature extraction network, merging the feature images obtained from all interval encodings together through Max Pooling, and then using a decoding network to estimate the surface normals.

[0033] Furthermore, the method for solving surface normals using the photometric stereo vision neural network for event cameras by fusing null vectors is EventPS-CNN: the number of channels in the event observation map is increased from 1 to 3, and each pixel represents a null vector in the corresponding illumination direction. In this way, all null vectors of each pixel are gathered in the event observation map, which contains more information about each pixel. The other parts of the photometric stereo vision neural network for event cameras retain the original CNN-PS network structure, and the EventPS-CNN retains more detailed information about each null vector.

[0034] Furthermore, the photometric stereo vision neural network suitable for the event camera is trained using synthetic data. The training data synthesis process is as follows:

[0035] (a) The 3D geometry in the dataset includes 10 geometry from the Blobby dataset and 15 geometry from the Sculpture dataset;

[0036] (b) For each geometry, add random rotation and scaling transformations, and add random BRDF textures;

[0037] (c) Render the HDR image sequence using a ray tracing renderer based on the lighting pattern;

[0038] (d) Use an event camera simulator to convert the rendered image sequence into event signals.

[0039] This invention provides a real-time photometric stereo vision device based on an event camera, comprising: a computer, an event camera, a rotating light source, a battery module, a sensor, a DC motor, a synchronous belt, and a counter; the battery module is connected to the rotating light source, and the rotating light source is connected to the DC motor via the synchronous belt, driving the rotating light source to rotate; the sensor is mounted at the end of the DC motor; the event camera, sensor, and counter are respectively connected to the computer, with the sensor connected to the counter; the computer includes OpenEB software, a CPU, and a GPU; before starting the device, the rotating light source is moved to its initial position, and the object to be tested is placed directly in front of the rotating light source, and the counter is reset. The system starts a DC motor to drive a rotating light source. Simultaneously, a sensor at the motor's end measures its rotation speed, and a counter counts the motor's rotations. A synchronization signal is transmitted from the computer to the event camera, which receives the synchronization signal and begins acquiring event signals. The computer uses OpenEB software to read the event camera's output, acquiring both the event signal and the synchronization signal's timestamps. The synchronization signal timestamps are interpolated to determine the position of the rotating light source at the time of the event. Within the computer, the CPU categorizes events according to their pixel location and transmits the categorized events to the GPU for high-concurrency event processing, enabling real-time acquisition.

[0040] The beneficial effects of this invention are:

[0041] 1. This invention starts from a basic model triggered by an event camera, deriving null vector information directly related to surface normals for each event, and reconstructing the surface normals of the tested object using only a single event camera for photometric stereo vision. Compared to hybrid camera solutions, this invention reduces the cost of multiple cameras and synchronization.

[0042] 2. This invention utilizes the high event resolution of an event camera to acquire data under continuously changing light sources, thereby increasing data sampling density and improving the quality of normal estimation.

[0043] 3. This invention uses an event camera as input for photometric stereo vision tasks. By leveraging the high dynamic range characteristics of the event camera, and through native high dynamic range and compressed event representation, it improves the speed and quality of data acquisition. At the same time, it can accurately acquire information about high-contrast and high-specular-reflection objects, improving the quality of normal estimation. Attached Figure Description

[0044] Figure 1 The present invention provides a flowchart of a real-time photometric stereo vision method based on an event camera.

[0045] Figure 2 This invention provides an illumination pattern diagram for a real-time photometric stereo vision method based on an event camera.

[0046] Figure 3 The network structure diagram for solving the surface normals of EventPS-FCN is provided.

[0047] Figure 4 The network structure diagram for solving surface normals in EventPS-CNN.

[0048] Figure 5 The present invention provides a structural block diagram of a real-time photometric stereo vision device based on an event camera. Detailed Implementation

[0049] The present invention will be further described in detail below with reference to the accompanying drawings.

[0050] This invention provides a real-time photometric stereo vision method based on an event camera, including the construction of a synchronization system for continuously changing illumination and the event camera, the calculation of null vectors, and a surface normal solution algorithm; as well as the design of the event camera surface normal solution algorithm, including the use of SVD to calculate the event camera photometric stereo vision algorithm in real time, and the event camera photometric stereo vision transfer method of existing traditional camera photometric stereo vision networks; and the methods for training and testing the neural network, and the synthesis of training data.

[0051] See Figure 1 As shown, the present invention provides a real-time photometric stereo vision method based on an event camera, which specifically includes the following steps:

[0052] Step 1: Lighting mode design;

[0053] In the photometric stereo vision method based on frame cameras, traditional frame cameras use discrete lighting to illuminate the test object. Unlike traditional photometric stereo vision, event camera photometric stereo vision requires continuously changing lighting to illuminate the test object.

[0054] like Figure 2 As shown, this invention designs the following three continuously varying lighting modes to meet different needs: circle lighting mode, hypotrochoid lighting mode, and DiLiGenT lighting mode.

[0055] Among them, the "circle" lighting mode requires the simplest mechanical structure and can achieve a high rotation speed to meet the needs of high-speed photometric stereo vision; the "hypotrochoid" lighting mode further introduces the change of rotating lighting latitude to improve the normal recovery quality of flat objects; the "DiLiGenT" lighting mode uses the outermost lighting direction in the existing DiLiGenT dataset and rotates along the rectangular area, which is compatible with existing data and facilitates comparison.

[0056] Step 2: Synchronous acquisition of events by the camera;

[0057] This invention uses a sensor and electronic synchronization circuit to calibrate the illumination direction for the acquired events. Specifically, the event camera's photometric stereo vision uses continuously varying illumination to quickly capture sparse event streams, while traditional photometric stereo vision uses discrete illumination to slowly capture images of dense scenes.

[0058] First, the relative position between the test object and the rotating light source is fixed mechanically, and the geometric coordinates of the rotating light source are calibrated. Then, the rotating light source is returned to its initial position, and a sensor (such as a Hall sensor) is used to measure the rotation angle of the DC motor. When the rotating light source rotates at a constant speed to a specified angle, a synchronization signal is sent to the event camera. The event camera simultaneously acquires the illumination signal and the synchronization signal generated by the test object due to changes in illumination. The synchronization signal is used to calibrate the illumination direction when other events occur, thus achieving synchronous acquisition.

[0059] Step 3: Constructing the null vector;

[0060] This invention constructs a null vector directly related to the surface normal of the tested object, starting from each pair of consecutively occurring events. Assume that during synchronous event acquisition by the camera, the event triggering threshold is C, and pixel x triggers a total of K events; the polarity of these events is p. k The event occurs at time t. k , t k-1 Let k be the time when adjacent events of the same pixel occur, where k = {1, 2, ..., K}; I x (t k () represents the absolute pixel brightness at the time the event is triggered. According to the event-camera triggering model, the following equation holds:

[0061] I x (t k ) = exp(p k C)I x (t k-1 )

[0062] According to the ideal Lambert model, assuming the normal direction of the surface of the object being tested is n x The diffuse reflectance is a x The direction of illumination at the time of the event is L(t) k If ), then the following equation holds:

[0063] max[0, a x (n x ·L(t k ))]=exp(p k C)max[0, a x (n x ·L(t k-1 ))]

[0064] Due to diffuse reflectance a x If it appears on both sides of the equation, remove 'a'. x This does not affect the validity of the equation. This means that under the same change in lighting direction, the diffuse reflectance of the tested object's surface will not affect event triggering. (Remove a) x Then, the following equation holds:

[0065] max[0, (n x ·L(t k ))]=exp(p k C)max[0,(n x ·L(t k-1 ))]

[0066] According to the event camera triggering model, events only occur at moments when there is a change in pixel brightness. Therefore, it can be inferred that at the moment the event occurs, the derivative of the pixel's absolute brightness with respect to time must be non-zero. The max operator on both sides of the formula describes shadows in Lambertian reflections. When the max value is 0, the derivative of the pixel's absolute brightness with respect to time is also 0, which is related to the event occurrence time t. k Contradiction. Therefore, at any moment t when the event occurs... k Both sides of the equation should be greater than 0. This property also means that the event signal does not contain redundant information about pixels in the shadow region. Removing the max operator from both sides of the equation, the following equation holds true:

[0067] n x ·L(t k ) = exp(p k C)(n x ·L(t k-1 ))

[0068] It can be transformed linearly into:

[0069] n x ·(L(t k )-exp(p k C)L(t k-1 ))=0

[0070] According to the above formula, a pair of consecutive events is transformed into a vector perpendicular to the normal direction of the surface of the object being tested. This vector is called the null vector, and its specific form is as follows:

[0071] z k =L(t) k+1 )-exp(p k+1 C)L(t k )

[0072] Among them, t k+1 and pk+1 These represent the time and polarity of adjacent events occurring at the same pixel.

[0073] Step 4: Solve for surface normals by fusion of null vectors;

[0074] After the tested object completes one or more sampling cycles according to the designed lighting pattern, each image pixel will obtain a continuous event sequence containing several events. Each pair of adjacent events in these continuous event sequences is transformed into a null vector, and each null vector describes partial information about the noisy surface normal. Next, multiple null vectors need to be fused to solve for the surface normal. To this end, this invention proposes two types of solving algorithms: SVD and deep learning.

[0075] For the SVD solution method, the K-1 null vectors generated from pixel x are concatenated into a null vector matrix Z of size 3×(K-1). k Theoretically, at least three events need to be triggered by a pixel to form a null vector matrix with a rank of at least 2 before further solutions can be obtained. The objective is to optimize the following mean square error function:

[0076]

[0077] in, The estimated surface normal. To solve... The value, for Perform eigenvalue decomposition to solve for the corresponding... The eigenvector of the smallest eigenvalue of a matrix is ​​obtained through a method called EventPS-OP.

[0078] For deep learning solutions, existing photometric stereo vision neural networks based on frame cameras are modified and retrained to obtain photometric stereo vision neural networks suitable for event cameras. The following explanation uses PS-FCN and CNN-PS network structures as examples.

[0079] The original PS-FCN network structure performs convolution operations on images under each illumination individually and merges multiple feature images through a max-pooling layer. This invention constructs a null vector image as input while maintaining the pixel-to-pixel connectivity required by the convolutional neural network, thus transferring the PS-FCN network to photometric stereo vision for event cameras. This method is called EventPS-FCN. Figure 3 As shown, the entire illumination mode time is first divided into N equal-length intervals. All null vectors within each interval and each pixel are then superimposed to form a null vector image. For each pixel in the null vector image, the illumination direction changes in the same way. The null vector image and the corresponding illumination direction are used as input. Since event features are sparser than image features, and there is a continuous relationship between adjacent time intervals, two temporal convolutional layers are added to the input layer of the original PS-FCN network to extract temporal features from events in adjacent time intervals. The rest of the photometric stereo vision neural network for event cameras follows the original PS-FCN network structure. Feature images are extracted through a feature extraction network, and the feature images obtained from all interval encodings are merged together using Max Pooling. Then, a decoding network, i.e., a surface normal regression network, is used to estimate the surface normals.

[0080] like Figure 4 As shown, the original CNN-PS network structure transforms the input image pixel by pixel into a 32×32 resolution event observation map, and then uses convolutional layers on the event observation map to process each pixel individually. Following the same approach, the transformation from event sequences captured by the event camera to null vectors is also performed pixel by pixel. The original CNN-PS network is adapted for event camera photometric stereo vision by modifying the event observation map. The number of channels in the event observation map is increased from 1 (grayscale image) to 3 (x, y, z coordinate components of the null vector). Each pixel represents a null vector in the corresponding illumination direction. In this way, all null vectors for each pixel are aggregated in the event observation map. The rest of the neural network for event camera photometric stereo vision retains the original CNN-PS neural network structure; this method is called EventPS-CNN. Compared to the time intervals in EventPS-FCN, the event observation map contains more information about each pixel. Therefore, EventPS-CNN retains more detailed information about each null vector.

[0081] Among them, for deep learning solution methods, in addition to the CNN-PS and PS-FCN structures already implemented in this invention, more advanced traditional camera photometric stereo vision networks can also be migrated to event camera photometric stereo vision versions.

[0082] Furthermore, this invention uses synthetic data to train the photometric stereo vision neural network suitable for event cameras. The training data synthesis process is as follows:

[0083] (a) The 3D geometry in the dataset includes 10 geometry from the Blobby dataset and 15 geometry selected from the Sculpture dataset;

[0084] (b) For each geometry, add random rotation and scaling transformations, and add a random BRDF texture.

[0085] This step is similar to the previous photometric stereo vision method based on deep learning;

[0086] (c) Based on the lighting pattern, use a ray tracing renderer to render an HDR image sequence with a sequence length of approximately 600 frames;

[0087] (d) Use an event camera simulator to convert the rendered image sequence into event signals.

[0088] See Figure 5 As shown, the present invention provides a real-time photometric stereo vision device based on an event camera, which specifically includes the following components:

[0089] The system comprises a computer, an event camera, a rotating light source, a battery module, a sensor, a DC motor, a synchronous belt, and a counter. The battery module is integrated with the rotating light source and powers it. The rotating light source is connected to the DC motor via a synchronous belt, which drives the light source to rotate. The sensor is mounted at the end of the DC motor. The event camera, sensor, and counter are all connected to the computer, with the sensor also connected to the counter. The computer contains OpenEB software, a CPU, and a GPU.

[0090] Before starting the device, the rotating light source is moved to its initial position, and the object under test is placed directly in front of the rotating light source. The counter is reset, and the DC motor is started to drive the rotating light source to rotate. Simultaneously, the number of rotations of the DC motor is measured by a sensor at its end, and the counter counts the rotations. A synchronization signal is transmitted to the event camera via the computer. The event camera receives the synchronization signal and begins acquiring event signals. The computer uses OpenEB software to read the output of the event camera, simultaneously acquiring the timestamps of both the event signal and the synchronization signal. Interpolation of the synchronization signal timestamps yields the position of the rotating light source when the event occurred. Within the computer, events are categorized on the CPU according to their pixel location. The categorized events are then transferred to the GPU for high-concurrency event processing, enabling real-time acquisition. Specifically, SVD calculations can be performed on the GPU to implement the EventPS-OP surface normal solution algorithm; a null vector image can be constructed on the GPU, and the EventPS-FCN surface normal solution algorithm can be implemented using PyTorch; and an event observation map can be constructed on the GPU, and the EventPS-CNN surface normal solution algorithm can be implemented using TensorFlow.

[0091] The event camera may specifically be a Prophesee EVK4 HD IMX636, but is not limited to this.

[0092] The rotating light source can be in the form of an LED light or a laser with a galvanometer, but is not limited to these. When using a laser with a galvanometer, the lighting pattern can be improved to more complex graphics.

[0093] The sensor may be a Hall sensor, but is not limited to it.

[0094] Specifically, the DC motor can be a 24V DC brushed motor (with a speed of up to 1800 RPM), but is not limited to this.

[0095] The counter can be an Arduino Uno Rev3, but is not limited to it.

[0096] Specifically, the GPU can be an NVIDIA GeForce GTX 1080, and the GPU code can be implemented using OpenCL.

[0097] This invention successfully implements a real-time photometric stereo vision algorithm based on an event camera, which can run on high-speed turntables and GPU devices, and outputs the surface normal information of the tested object in real time at a rate of more than 30 frames per second.

[0098] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A real-time photometric stereo method based on an event camera, characterized in that, The method comprises the following steps: Step one, light mode design; Step two, simultaneously collecting the light signal and synchronization signal generated by the tested object due to light change by using the event camera; using the synchronization signal to calibrate the light direction when other events occur, and realizing synchronous acquisition; Step three, starting from each pair of continuously occurring events, constructing a zero vector directly related to the normal of the surface of the tested object; Assume that the event camera synchronously captures the process, the event trigger threshold is C, and the pixel x triggers K events in total; the polarity of these events is p k , the event occurrence time is t k , t k-1 is the time when the adjacent events of the same pixel occur, k={1,2,…,K}; I x (t k ) is the absolute brightness of the pixel when the event is triggered; according to the event camera trigger model, the following equation holds: I x (t k )=exp(p k C)I x (t k-1 ) According to the ideal Lambertian model, assuming the surface normal direction of the tested object is n x , the diffuse reflectance is a x , and the light direction at the time of the event is L(t k ), the following equation holds: max[0, a(n · L(t))] = exp(pC) max[0, a(n · L(t))] + 1 x (n x ·L(t k ))] = exp(p k C)max[0, a x (n x ·L(t k-1 ))] In case the direction of illumination changes equally, the diffuse reflectivity of the surface of the object under test does not influence the event triggering, so a x The following equation holds: max[0,(n x ·L(t k ))] = exp(p k C)max[0,(n x ·L(t k-1 ))] According to the event camera trigger model, an event only occurs when there is a change in the pixel luminance; therefore, it is inferred that the derivative of the pixel absolute luminance with respect to time must be non-zero at the event occurrence time; the max operator on both sides of the equation is used to describe the shadow in the Lambertian reflection; when the max value takes 0, the derivative of the pixel absolute luminance with respect to time is also 0, which is the event occurrence time t k Contradiction; therefore, at any event occurrence time t k , both sides of the equation should be greater than 0, so the event signal does not contain redundant information of the pixels in the shadow area; remove the max operator on both sides of the equation, then the following equation holds: n x • L(t k ) = exp(p k C)(n x • L(t k-1 )) After linear transformation, it is: n x • (L(t k )-exp(p k C)L(t k-1 )) = 0 According to the above formula, a pair of continuous events is converted into a vector perpendicular to the normal direction of the surface of the tested object, i.e. a zero vector, and the specific form is as follows: z k = L(t k+1 )- exp(p k+1 C)L(t k ) where t k+1 and p k+1 are the time and polarity of the adjacent event, respectively. Step four, zero vector fusion to solve the surface normal.

2. The real-time photometric stereo method based on event camera according to claim 1, wherein, The light mode is a circle light mode, a hypotrochoid light mode and a DiLiGenT light mode.

3. The real-time photometric stereo method based on event camera according to claim 1, wherein, In step four, after the tested object completes one cycle or multiple cycles of sampling according to the designed light mode, each image pixel will obtain a continuous event sequence containing several events; each pair of adjacent events in the continuous event sequence is converted into a zero vector, and each zero vector describes part of the information of the noisy surface normal; then, the surface normal is solved by fusing multiple zero vectors.

4. The real-time photometric stereo method based on event camera according to claim 1, wherein, In step four, the SVD solving method is used to solve the surface normal: the K-1 null vectors converted by the pixel x are connected to form a null vector matrix Z with a size of 3 x (K-1) k In theory, at least 3 events triggered by the pixel are needed to form a null vector matrix with a rank of at least 2 for further solving; the solving target is to optimize the following mean square error function: wherein, is the estimated surface normal; is the solution to the numerical value of is eigenvalue decomposed, the eigenvector corresponding to the matrix of the smallest eigenvalue is solved.

5. The real-time photometric stereo method based on event camera according to claim 1, characterized in that, In step four, a deep learning solving method is used to solve the surface normal: by modifying the input of an existing photometric stereo vision neural network based on a frame camera, retraining, and migrating to obtain a photometric stereo vision neural network suitable for an event camera; The surface normal is solved by fusing the zero vectors by using the photometric stereo vision neural network suitable for the event camera.

6. The event camera based real-time photometric stereo method according to claim 5, wherein, The method for solving the surface normal by fusing the zero vectors by using the photometric stereo vision neural network suitable for the event camera is EventPS-FCN: dividing the entire light mode time into N equal intervals, superimposing all zero vectors in each interval and each pixel to form a zero vector image; for different pixels of each zero vector image, the light direction changes the same, and the zero vector image and the corresponding light direction are used as input, two convolution layers with time dimension are added to the input layer of the original PS-FCN network to extract time sequence features from adjacent time intervals; the other parts of the photometric stereo vision neural network suitable for the event camera follow the structure of the original PS-FCN network, the feature images are extracted through the feature extraction network, and all interval encoded feature images are merged together through Max Pooling, and then the decoding network is used to estimate the surface normal.

7. The real-time photometric stereo method based on event camera according to claim 5, characterized in that, The method for fusing nulling vectors to solve surface normals by using the photometric stereo neural network suitable for event cameras is EventPS-CNN: the number of channels of the event Observation Map is increased from 1 to 3, and each pixel represents a nulling vector in the corresponding light direction, in this way, all nulling vectors of each pixel are gathered in the event Observation Map, which contains more information of each pixel, and other parts of the photometric stereo neural network suitable for event cameras follow the original CNN-PS network structure, and the EventPS-CNN retains more detailed information about each nulling vector.

8. The event camera based real-time photometric stereo method of claim 5, wherein, The photometric stereo neural network suitable for event cameras is trained using synthetic data, and the training data synthesis process is as follows: (a) The three-dimensional geometry in the data set includes 10 geometries in the Blobby data set and 15 geometries in the Sculpture data set; (b) For each geometry, add a random rotation and scaling transform, and add a random BRDF texture; (c) According to the light mode, use the ray tracing renderer to render the HDR image sequence; (d) Use the event camera simulator to convert the rendered image sequence into event signals.

9. An event camera based real-time photometric stereo device for implementing the event camera based real-time photometric stereo method of claim 1, characterized in that, It includes: Computer, event camera, rotating light source, battery module, sensor, DC motor, synchronous belt and counter; The battery module is connected with the rotating light source, the rotating light source is connected with the DC motor through the synchronous belt, the rotating light source is driven to rotate through the DC motor and the synchronous belt, the sensor is installed at the end of the DC motor, the event camera, the sensor and the counter are connected with the computer respectively, the sensor is connected with the counter, and the computer is provided with OpenEB software, CPU and GPU; Before starting the device, move the rotating light source to the initial position, place the object to be tested in front of the rotating light source, reset the counter, start the DC motor to drive the rotating light source to rotate, measure the number of rotations of the DC motor through the sensor at the end of the DC motor while the DC motor is rotating, and count the number of rotations of the DC motor through the counter, transmit the synchronization signal to the event camera through the computer, at this time the event camera receives the synchronization signal and starts to collect the event signal; The computer uses OpenEB software to read the output of the event camera, simultaneously collects the time stamps of the event signal and the synchronization signal, interpolates the time stamp of the synchronization signal to obtain the position of the rotating light source when the event occurs; In the computer, the events are classified according to the pixel positions on the CPU, and the classified events are transmitted to the GPU, so that the events are high-concurrent processed, and real-time collection is realized.

Citation Information

Patent Citations

  • Data set making method and verification system based on photometric three-dimensional surface reconstruction

    CN116543247A

  • Metal surface welding slag detection system and method, electronic device and storage medium

    CN118169128A