Three-dimensional imaging method based on 3D and AI visual sensing visible light movement
Through the three-dimensional imaging method based on 3D and AI visual sensing, diffuse and specular reflection events are separated and fused, the problem of high computational complexity of traditional three-dimensional imaging methods is solved, and higher modeling accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510340999.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Due to the high computational complexity of traditional three-dimensional imaging methods, they limit the real-time nature of the system and put high requirements on hardware performance, making it difficult to accurately model the scene during motion.
Using a three-dimensional imaging method based on 3D and AI visual sensing visible light movements, the original event data is obtained through AI visual sensing, spatiotemporal feature extraction and anti-pole geometric constraint deconstruction, diffuse and specular reflection events are separated, and a complete three-dimensional model is obtained through multimodal fusion reconstruction.
It improves the modeling accuracy and stability of the scene during movement, reduces the error in imaging depth, and significantly reduces the distance accuracy error of the synthetic three-dimensional image, and is suitable for applications that require precise dimensions and position information.
Smart Images

Figure CN120219663A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a three-dimensional imaging method based on a 3D and AI vision sensing visible light module. Background Art
[0002] Traditional three-dimensional imaging mainly relies on hardware such as laser scanners and structured light projection devices in combination with complex software algorithms to obtain the surface information of an object and reconstruct its three-dimensional shape. However, these traditional methods often come with a high computational complexity, which not only limits the real-time performance of the system but also places higher requirements on the hardware performance. Summary of the Invention
[0003] This application provides a three-dimensional imaging method based on a 3D and AI vision sensing visible light module, which is used to improve the modeling accuracy and stability of a scene during movement.
[0004] In a first aspect, an embodiment of this application provides a three-dimensional imaging method based on a 3D and AI vision sensing visible light module, and the method includes: Obtaining original event data through an AI vision sensing visible light module; Performing spatio-temporal feature extraction on the original event data to obtain an event stream data matrix; Performing epipolar geometry constraint deconstruction according to the event stream data matrix to obtain a diffuse reflection event subset and a specular reflection event subset; Performing triangulation reconstruction according to the diffuse reflection event subset to obtain a diffuse reflection point cloud model; Using the diffuse reflection point cloud model as a virtual screen to perform polarized light measurement reconstruction on the specular reflection event subset to obtain a specular reflection point cloud model; Performing multi-modal fusion reconstruction according to the diffuse reflection point cloud model and the specular reflection point cloud model to obtain a fused three-dimensional model of the complete scene.
[0005] In a second aspect, an embodiment of this application provides a three-dimensional imaging device based on a 3D and AI vision sensing visible light module, and the device includes: A data acquisition module, configured to obtain original event data through an AI vision sensing visible light module; A feature extraction module, configured to perform spatio-temporal feature extraction on the original event data to obtain an event stream data matrix; A data decomposition module, configured to perform epipolar geometry constraint deconstruction according to the event stream data matrix to obtain a diffuse reflection event subset and a specular reflection event subset; A first reconstruction module, configured to perform triangulation reconstruction according to the diffuse reflection event subset to obtain a diffuse reflection point cloud model; A second reconstruction module, configured to use the diffuse reflection point cloud model as a virtual screen to perform polarized light measurement reconstruction on the specular reflection event subset, so as to obtain a specular reflection point cloud model; A model fusion module, configured to perform multimodal fusion reconstruction according to the diffuse reflection point cloud model and the specular reflection point cloud model, so as to obtain a fused three-dimensional model of the complete scene.
[0006] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes a memory and a processor; The memory is used to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the three-dimensional imaging method based on a 3D and AI vision sensing visible light module as described in any one of the embodiments of the present application.
[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the three-dimensional imaging method based on a 3D and AI vision sensing visible light module as described in any one of the embodiments of the present application.
[0008] An embodiment of the present application provides a three-dimensional imaging method based on a 3D and AI vision sensing visible light module. The method includes: obtaining original event data through an AI vision sensing visible light module; performing spatio-temporal feature extraction on the original event data to obtain an event stream data matrix; performing epipolar geometry constraint deconstruction according to the event stream data matrix to obtain a diffuse reflection event subset and a specular reflection event subset; performing triangulation reconstruction according to the diffuse reflection event subset to obtain a diffuse reflection point cloud model; using the diffuse reflection point cloud model as a virtual screen to perform polarized light measurement reconstruction on the specular reflection event subset to obtain a specular reflection point cloud model; performing multimodal fusion reconstruction according to the diffuse reflection point cloud model and the specular reflection point cloud model to obtain a fused three-dimensional model of the complete scene. Through the above method, the diffuse reflection event and the specular reflection event are separated from the acquired data. Through the combination reference of the two, the ability to capture details can be effectively improved, so as to obtain a more realistic and detailed three-dimensional model. For three-dimensional imaging of an object mixed with diffuse reflection characteristics and specular reflection characteristics, the error of the imaging depth is reduced, and the distance accuracy of the synthesized three-dimensional image is greatly reduced, which is crucial for applications that require accurate size and position information. Description of the Drawings
[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0010] Figure 1 It is a schematic flow chart of a three-dimensional imaging method based on a 3D and AI vision sensing visible light module provided by an embodiment of the present application; Figure 2 It is a schematic block diagram of a three-dimensional imaging device based on a 3D and AI vision sensing visible light module provided by an embodiment of the present application. Detailed implementation manners
[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0012] The flow chart shown in the drawings is only an example for illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order may change according to the actual situation.
[0013] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0014] It should be further understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0015] The AI vision sensing visible light module in the embodiments of the present application is an event-based sensor. Different from traditional frame rate image sensors, event-based image sensors do not capture the entire image frame at fixed time intervals, but only record the parts of the scene that change, that is, "events". This working mode enables event-based image sensors to have significant advantages in dealing with high-speed movements and changes.
[0016] To more clearly introduce the technical solution of this application, the technical solution of this application will also be introduced through specific embodiments below. It should be noted that the specific embodiments are used to expand and explain the technical solution of this application, rather than limiting this application.
[0017] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a three-dimensional imaging method provided by an embodiment of this application for a 3D and AI vision sensing visible light module. As Figure 1 shown, the specific steps of the three-dimensional imaging method based on the 3D and AI vision sensing visible light module include: S101 - S106.
[0018] S101. Obtain the original event data through the AI vision sensing visible light module.
[0019] Exemplarily, when obtaining the original event data, double scanning needs to be performed to obtain diffuse reflection events and specular reflection events. The event acquisition device has a binocular AI vision sensing visible light module, which observes the scene from different perspectives. During the scanning process, first perform diffuse reflection scanning. By adjusting the light source angle and intensity, the light is irradiated on the surface of the object to be measured at different angles, and the event data generated by diffuse reflection is captured. Subsequently, perform specular reflection scanning. By changing the light source position and polarizer angle, the specular reflection events generated on the object surface are obtained. Each event contains information such as pixel coordinates (x, y), timestamp t, and polarity p.
[0020] Through the double scanning strategy, the diffuse reflection characteristics and specular reflection characteristics of the object surface can be obtained respectively, providing complete original data support for subsequent 3D reconstruction. This method can better process objects with complex surface characteristics and improve the reconstruction accuracy.
[0021] In some embodiments, obtaining the original event data through the AI vision sensing visible light module specifically includes: S1011 - S1016.
[0022] S1011. Monitor the log intensity change of each pixel captured by the AI vision sensing visible light module to obtain pixel intensity change data.
[0023] Exemplarily, a reference brightness value L0 is set for each pixel point. The reference brightness value L0 is obtained by averaging the pixel points of 100 consecutive frames in the static state of the scene. The brightness value L of the pixel point at time t t is logarithmically operated with the reference brightness value L0 to obtain the logarithmic intensity ratio R, where R = log(L t / L0). The central difference method is used to perform the time derivative operation on the logarithmic intensity ratio R: dR / dt = (R t+Δt -R t-Δt / (2Δt), where Δt is the sampling time interval, and a 5×5 Gaussian filter is used to suppress noise in the time derivative sequence to obtain pixel intensity change data.
[0024] S1012. Perform threshold comparison judgment on the pixel intensity change data, and generate an event trigger signal including a timestamp, polarity, and coordinate position according to the judgment result.
[0025] Exemplarily, when performing threshold comparison judgment on the pixel intensity change data, based on the pixel intensity of each pixel point, the specific formula for adaptively calculating the threshold is: T p =μ p +kp×σ p , T n =μ n -k n ×σ n . Where μ p is the mean value of positive polarity events, σ p is the standard deviation, k p is the weighting coefficient, μ n is the mean value of negative polarity events, σ n is the standard deviation, k n is the weighting coefficient. When the pixel intensity change data is greater than T p , a positive polarity event with a value of +1 is generated. When the pixel intensity change data is less than T n , a negative polarity event with a value of -1 is generated. The event timestamp is accurately recorded at the microsecond level, and the 32-bit timestamp, 1-bit polarity flag, and 2 16-bit pixel coordinates of the event are packed to generate a 65-bit event trigger signal.
[0026] S1013. Accumulatively count the intensity changes of pixel positions according to the event trigger signal to obtain an event count matrix.
[0027] Exemplarily, a counting matrix with dimensions of H×W is established, where H and W are the height and width of the image respectively. Dual counters are set at each position of the counting matrix to record positive and negative polarity events respectively, and each event trigger signal is counted and accumulated at the corresponding pixel position according to the polarity. A time window is set, and the events within the window are weighted by exponential decay: w(t)=exp(-t / τ), where t is the event age and τ is the decay time constant. Pixel positions with count values greater than the preset threshold are marked to obtain an event count matrix including positive and negative event counts, time weights, and valid position marks.
[0028] S1014. Perform n-bin voxel grid partitioning on the event count matrix in the time dimension to obtain spatio-temporal discretized data blocks.
[0029] Exemplarily, the total observation time T is divided into n equal time periods, each with a length of ΔT = T / n. The following processing is performed on the event data within each time period ΔT: The space is divided into a regular voxel grid of 16×16×16, with the side length of each voxel being l = L / 16, where L is the scene size. The event density ρ = Ne / V is calculated within each voxel, where Ne is the number of events and V is the voxel volume. According to the event density, the voxel blocks are divided into three categories: When ρ > T1 (T1 = 100 events / mm³), it is divided into high-density blocks and subdivided using a 4×4×4 sub-grid; when T2 ≤ ρ ≤ T1 (T2 = 10 events / mm³), it is divided into medium-density blocks and subdivided using a 2×2×2 sub-grid; when ρ < T2, it is divided into low-density blocks and the original voxel size is maintained. The center coordinates, size, number of events, and polarity distribution of each density block are recorded. After classification according to the above classification method, a spatio-temporal discretized data block is obtained.
[0030] S1015. Perform a polarity cumulative projection on each 3D slice in the spatio-temporal discretized data block to obtain a sequence of two-dimensional projection frames.
[0031] S1016. Perform pixel value normalization processing on the sequence of two-dimensional projection frames to obtain the original event data stream.
[0032] Exemplarily, a projection matrix P with dimensions H'×W' is constructed, where H' is the height resolution (number of pixels) of the projection plane and W' is the width resolution (number of pixels) of the projection plane. The three-dimensional voxel grid within each time period Δt is projected onto the xy plane along the z-axis direction, and the events during the projection are depth-weighted: w(z) = exp(-αz), where w(z) is the weighted value of the event, α is the attenuation coefficient, and z is the voxel depth. The specific calculation formulas for respectively accumulating the weighted values of positive and negative polarity events are: ; where P(x, y) is the cumulative projection value at position (x, y) on the projection plane, Np(x, y, z): the cumulative number of positive polarity events at coordinate (x, y, z), Nn(x, y, z): the cumulative number of negative polarity events at coordinate (x, y, z).
[0033] After obtaining the cumulative projection value P(x, y), a Gaussian kernel function with a standard deviation of 1.5 is used for 3×3 neighborhood smoothing, projecting the event data in 3D space onto a 2D plane while retaining depth information, and performing non-maximum suppression to obtain a sequence of two-dimensional projection frames containing depth information.
[0034] S102. Extract spatio-temporal features from the original event data to obtain an event stream data matrix.
[0035] Exemplarily, perform time synchronization and spatial alignment preprocessing on the original event data. Use a spatio-temporal convolutional neural network to extract the temporal features of the event stream, including: temporal features such as the time interval and duration of event occurrence. At the same time, extract spatial features, such as the spatial distribution and local structure of the events. Organize the extracted features into an event stream data matrix in the form of a multi-dimensional tensor, where each pixel in the event stream data matrix contains a spatio-temporal feature vector corresponding to the position.
[0036] Through spatio-temporal feature extraction, discrete event data can be converted into a structured feature representation, facilitating subsequent geometric constraint analysis and 3D reconstruction. This step can effectively reduce data noise and improve the expressive ability of features.
[0037] In some embodiments, perform spatio-temporal feature extraction on the original event data to obtain an event stream data matrix, including: S1021 - S1028.
[0038] S1021. Perform temporal segmentation processing on the original event data, and divide the event data into multiple temporal bin sequences according to the time interval.
[0039] Exemplarily, perform temporal segmentation processing on the original event data, including: divide the total observation time T into N temporal bins according to a fixed time interval Δt = 10ms, and for each event within each temporal bin, calculate the time weight according to the occurrence time t i : w(t i ) = exp(-|t i - t c | / τ), where t c is the center time of the bin, τ is the decay constant, project the weighted events onto the H×W spatial grid to generate event voxels, each voxel records the number of events ne and the cumulative weight we, and normalize the event density of each voxel: ρ = ne × we / V, where V is the voxel volume, to generate a voxel sequence {V1, V2,..., VN} of N temporal bins.
[0040] S1022. Perform a downsampling operation on the temporal bin sequence to obtain a dimensionality-reduced state sequence.
[0041] Exemplarily, construct a 4-layer pyramid network structure, with the resolution ratio of each layer being 1:2:4:8. Perform 3×3 convolution operations in each layer of the network, with a stride of 2. Use LeakyReLU as the activation function with a slope of 0.1. Perform max pooling operations on the feature channel dimension, and stack the feature maps of each layer to obtain a multi-scale feature tensor F ∈ R (C×H'×W') , where C is the number of feature channels, and perform PCA dimensionality reduction on the feature tensor F to obtain a dimensionality-reduced state sequence S = {s1, s2,..., sN}.
[0042] S1023. Construct a global spatial dependence extractor based on the dimensionality-reduced state sequence, calculate the attention for the dimensionality-reduced state at each time step, and obtain the global spatial dependence features.
[0043] Exemplarily, the global spatial dependence extractor is a two-stream attention network architecture, including a main stream and an auxiliary stream. The main stream uses 3 residual blocks to extract features from the dimensionality-reduced state sequence. Each residual block contains two layers of 3×3 convolutions and a skip connection to obtain the main stream feature map. The auxiliary stream uses 1×1 convolutions to compress the channels of the dimensionality-reduced state sequence to obtain the auxiliary stream feature map.
[0044] Align the main stream feature map and the auxiliary stream feature map in position to obtain the aligned feature map F.
[0045] Calculate the position attention map according to the aligned feature map F. Specifically, calculate the position attention score A(p, q) according to the aligned feature map F, and construct the position attention map from the position attention score A(p, q). The specific calculation formula of A(p, q) is: A(p, q)=exp(θ(F p ) T ·φ(F q )) / Σexp(θ(F p ) T ·φ(F r )); where A(p, q) is the position attention score between position p and position q, F p is the feature vector at position p on the aligned feature map F, F q is the feature vector at position q on the aligned feature map F, θ(·) is the query transformation function implemented by 1×1 convolution, φ(·): the key-value transformation function implemented by 1×1 convolution, and exp(·) is the natural exponential function.
[0046] Perform channel attention calculation on the aligned feature map F: M = σ(W2·ReLU(W1·AvgPool(F))), where W1 and W2 are the weights of the fully connected layers, and σ is the sigmoid function.
[0047] Multiply the output features of the position attention and the channel attention to obtain the global dependence features.
[0048] S1024. Perform motion perception analysis on the global spatial dependence features, extract the difference information between adjacent states through subtraction operations, and obtain the motion feature map.
[0049] Exemplarily, construct a temporal difference network: calculate the temporal difference for the global dependence features at adjacent time steps: D t =G t -G t-1. Adaptive sampling is performed on the differential features using a 3×3 deformable convolution. The offsets δpi and weights wi of 9 sampling points are set, and the sampled position features are calculated: y(p)=Σw i ·x(p + δi). Channel group convolution is performed on the sampled features, with the number of groups being 8. The Swish activation function is used: f(x)=x·sigmoid(βx), where β is a learnable parameter, to obtain the motion feature map M∈R C ×H×W .
[0050] S1025. Spatial attention calculation is performed based on the motion feature map to generate a discriminative attention weight map.
[0051] Exemplarily, a multi - head attention module is constructed: The input feature is divided into h attention heads, and the feature dimension of each head is d = C / h. Queries, key - value pairs are calculated for each attention head: Q = W q M, K = W k M, V = W v M, where W q , W k , W v are learnable linear transformation matrices. The attention scores are calculated: score(Q, K)=softmax(QK T / ). The attention output is obtained: Ahead = score(Q, K)V. The multi - head outputs are concatenated and fused through a 1×1 convolution: A = Conv1×1([A1; A2;...; Ah]). LayerNormalization is used for feature normalization to generate the discriminative attention weight map W.
[0052] S1026. The state features are modulated based on the discriminative attention weight map to obtain enhanced state features.
[0053] Exemplarily, the attention weight map W is upsampled to the original feature resolution, and the spatial adaptive normalization coefficients are calculated: γ = f(AvgPool(W)), β = g(MaxPool(W)), where f(·) and g(·) are two - layer MLP networks. The input features are normalized: x'=(x - μ) / σ, and the affine transformation is applied: y = γx'+β. A gating mechanism is used to fuse the original features and the enhanced features: z = σ(Wz·[x, y]), r = σ(Wr·[x, y]), h = tanh(Wh·[r⊙x, y]), o=(1 - z)⊙x+z⊙h, to obtain the enhanced state feature E.
[0054] S1027. Cross - domain attention fusion is performed on the enhanced state features to generate a multi - scale feature representation.
[0055] The enhanced feature E is divided into two domains: spatial and channel. A 3-layer feature pyramid is constructed within each domain, and deformable convolution is used for feature extraction to calculate the intra-domain attention: A i =softmax(F i T ·W q ·W k ·F i ), calculate the cross-domain attention: A c =softmax(F s T ·W q ·W k ·F c ), fuse the intra-domain and cross-domain attention features: F' i =γ i ·A i ·F i +β i , F' c =γ c ·A c ·F c +β c , adopt a feature selection gating unit: g = sigmoid(W g ·[F' i , F' c ), o = g ⊙ F' i +(1 - g) ⊙ F' c ), to generate a multi-scale feature representation R.
[0056] S1028. Perform adaptive weighted balancing processing based on the multi-scale feature representation to obtain an event stream data matrix.
[0057] Exemplarily, calculate the feature statistics: μ c =AvgPool(R c ), σ c =StdPool(R c ), normalize the features of each scale c: R c' =(R c -μ c ) / (σ c +ε), learn the inter-scale balance weights: wc = softmax(MLP([μ c , σ c )), weighted fuse the multi-scale features: F = Σw c ·R c' , perform feature recalibration: s = sigmoid(W s ·F), F' = s ⊙ F, to obtain the event stream data matrix F'.
[0058] S103. Perform epipolar geometry constraint deconstruction based on the event stream data matrix to obtain a diffuse reflection event subset and a specular reflection event subset.
[0059] It should be noted that the fundamental matrix is a 3×3 matrix that describes the geometric relationship between two uncalibrated images in a binocular vision system. It contains the rotation and translation information between the two cameras, as well as the internal parameter information of the cameras. The essential matrix is a special case of the fundamental matrix, which describes the geometric relationship between calibrated (normalized) images. It only contains the information of the rotation matrix R and the translation vector t between the cameras, and does not contain the camera internal parameters.
[0060] Exemplarily, when performing epipolar geometry constraint deconstruction based on the event stream data matrix, it is necessary to establish an epipolar geometry relationship model between binocular event cameras. Use the fundamental matrix and the essential matrix to describe the geometric constraint relationship between two views. Then, according to the epipolar line constraint conditions, match and classify the event stream data. Specifically, for each event point, judge whether the event belongs to diffuse reflection or specular reflection according to its corresponding epipolar line position in another image and the spatio-temporal feature similarity of the event.
[0061] Through epipolar geometry constraint deconstruction, events can be reliably divided into two categories: diffuse reflection and specular reflection, providing a basis for subsequent separate reconstructions. This classification method based on geometric constraints has a strong physical basis and can improve the accuracy of classification.
[0062] In some embodiments, performing epipolar geometry constraint deconstruction based on the event stream data matrix to obtain a diffuse reflection event subset and a specular reflection event subset includes: S1031 - S1036.
[0063] S1031. Perform polarity correlation analysis on the event stream data matrix, calculate the cross-correlation coefficient of event polarities within a preset spatial neighborhood, and construct a polarity association graph according to the preset relationship threshold and the cross-correlation coefficient.
[0064] Exemplarily, calculate the polarity difference degree between the central event and other events in the neighborhood: D(i, j)=|P(i)-P(j)|, where P(i) and P(j) respectively represent the polarity values of events i and j. Construct a polarity similarity matrix S, and its elements are calculated as: S(i, j)=exp(-D(i, j) / σ²), where σ is the Gaussian kernel parameter; perform eigenvalue decomposition on the polarity similarity matrix S, and take the eigenvector corresponding to the largest eigenvalue as the local polarity main direction; construct a spatial polarity vector field based on the local polarity main direction; perform non-local mean filtering on the spatial polarity vector field to obtain a smoothed polarity field representation; calculate the cross-correlation coefficient between adjacent events: R(i, j)=<V i , V j > / ||V i ||·||Vj ||, where V i and V j are the polarity vectors corresponding to events i and j; perform binary processing on the cross-correlation coefficient according to a preset relationship threshold τ: C(i, j) = 1 if if R(i, j) > τ else 0; construct a polarity association graph G=(V, E) based on the binary result, where the vertex set V corresponds to events and the edge set E corresponds to event pairs with a correlation coefficient greater than the threshold.
[0065] S1032. Perform fundamental matrix estimation on the polarity association graph, and perform iterative calculations through the RANSAC algorithm and preset camera motion parameters to obtain the epipolar geometric transformation matrix.
[0066] Exemplarily, set the maximum number of iterations N and the inlier ratio threshold ε, and set the number of sampled points m and the inlier distance threshold d to perform the RANSAC iteration process: 1) Randomly sample m pairs of matching point pairs from the polarity association graph.
[0067] 2) Calculate the fundamental matrix F using the eight-point method: construct the coefficient matrix A and solve the equation Af = 0 through singular value decomposition.
[0068] 3) Perform rank constraint on F: achieve it by setting the minimum singular value to zero after SVD decomposition.
[0069] 4) Calculate the epipolar constraint error of all point pairs with respect to F.
[0070] 5) Count the number of inliers that satisfy the error threshold.
[0071] 6) If the inlier rate exceeds the threshold ε, update the optimal solution. Perform local optimization on the optimal fundamental matrix F: 7) Only use inliers to construct the augmented coefficient matrix.
[0072] 8) Use the Levenberg-Marquardt algorithm for nonlinear optimization.
[0073] 9) The optimization objective is to minimize the algebraic distance of the epipolar constraint error. Verify the geometric rationality of the fundamental matrix: 10) Calculate the essential matrix E = K2ᵀFK1.
[0074] 11) Decompose the essential matrix to obtain the camera motion parameters [R|t].
[0075] 12) Verify the reprojection error through triangulation. Output the optimized epipolar geometric transformation matrix F.
[0076] S1033. Calculate the epipolar distance of the event point pair according to the epipolar geometry transformation matrix, and perform event consistency verification based on a preset distance threshold to obtain a geometric consistency score.
[0077] Exemplarily, extract the first slice image and the second slice image from the event stream data matrix, which represent the states of the scene at two different times. Calculate the epipolar line l2 = F·p1 of p1 in the second slice image; calculate the geometric distance d1 = |p2ᵀ·l2| / sqrt(l 21 ² + l 22 ²). Calculate the epipolar line l1 = Fᵀ·p2 of p2 in the first slice image, and calculate the geometric distance d2 = |p1ᵀ·l1| / sqrt(l 11 ² + l 12 ²); Take the symmetric distance d = (d1 + d2) / 2 as the epipolar distance; Calculate the geometric consistency score based on the preset distance threshold γ; Perform soft threshold processing using the Sigmoid function: score = 1 / (1 + exp((d - γ) / α)). Where α is a smoothing factor used to adjust the steepness of the score curve; Perform spatial regularization on the consistency score: Construct a K-nearest neighbor graph to represent the spatial relationship of events; Apply the graph Laplacian operator for local smoothing to obtain a geometric consistency score matrix.
[0078] S1034. Perform bimodal clustering analysis on the geometric consistency score, and divide the events into a diffuse reflection candidate set and a specular reflection candidate set through a Gaussian mixture model.
[0079] Exemplarily, set the number of components K = 2, corresponding to diffuse reflection and specular reflection respectively, and randomly initialize the mean vectors μ1, μ2 and covariance matrices Σ1, Σ2. Set the mixing weights π1, π2; Iteratively execute the EM algorithm: E-step: Calculate the posterior probability of each event belonging to each component; M-step: Update the model parameters μ, Σ and π; Calculate the log-likelihood function value and check the convergence; Perform post-processing based on the clustering results: Apply spatial continuity constraints to merge spatially adjacent regions of the same class; Remove outliers and small-area connected regions; Re-estimate the probability of the boundary regions; Output the diffuse reflection candidate set and the specular reflection candidate set.
[0080] S1035. According to the spatial distribution of the diffuse reflection candidate set and the specular reflection candidate set, construct a local surface normal vector field and perform a normal vector consistency test to obtain a surface reflection characteristic classification result.
[0081] Exemplarily, extract the 3D point set within the local neighborhood, calculate the covariance matrix and perform eigenvalue decomposition. Take the eigenvector corresponding to the minimum eigenvalue as the normal vector estimate, and construct a hierarchical normal vector field. Use an octree structure to organize the space, estimate the normal vectors at different scales, propagate and fuse the normal vector estimates from bottom to top; perform a normal vector consistency check. Calculate the angle between adjacent normal vectors, detect the mutation regions of the normal vectors, mark the discontinuous boundaries, and output the classification result of the surface reflection characteristics.
[0082] S1036. Classify and organize the events based on the classification result of the surface reflection characteristics, and generate a diffuse reflection event subset and a specular reflection event subset.
[0083] Exemplarily, combine the normal vector field and the geometric consistency score, calculate the local curvature and the normal vector entropy, and extract the surface microstructure features; apply the classification rules: define the discrimination criteria for diffuse reflection and specular reflection; handle the attribution of the fuzzy regions; ensure the spatial continuity; generate the final event subsets: encode the diffuse reflection events and the specular reflection events respectively; record the spatio-temporal coordinates and the reflection attributes of the events; establish the topological relationship between the events.
[0084] S104. Perform triangulation reconstruction based on the diffuse reflection event subset to obtain a diffuse reflection point cloud model.
[0085] Exemplarily, based on the calibrated camera parameters and the binocular geometric relationship, calculate the disparity information of the feature points. Then, according to the triangulation principle, calculate the coordinates of the three-dimensional space points based on the disparity values and the camera parameters. Perform triangulation on all the matching point pairs to obtain the discrete point cloud representing the diffuse reflection surface of the object. Finally, perform filtering and smoothing processing on the point cloud to obtain a complete diffuse reflection point cloud model.
[0086] Through triangulation reconstruction, the geometric shape of the diffuse reflection surface of the object can be accurately restored, providing a geometric reference for the subsequent specular reflection reconstruction. This method has simple calculation and high reconstruction accuracy, and is suitable for processing diffuse reflection surfaces.
[0087] In some embodiments, performing triangulation reconstruction based on the diffuse reflection event subset to obtain a diffuse reflection point cloud model includes: sorting the diffuse reflection event subset according to the time stamps to obtain an event sequence arranged in chronological order; calculating the spatial distance and time difference of each event point according to the event sequence, calculating the triangulation relationship between every two adjacent event points based on the spatial distance and time difference to obtain three-dimensional coordinate information; performing point cloud reconstruction processing on the three-dimensional coordinate point set, and performing sparsification processing on the three-dimensional coordinate point set through the voxel grid method to obtain a diffuse reflection point cloud model.
[0088] S105. Use the diffuse reflection point cloud model as a virtual screen to perform polarization measurement reconstruction on the specular reflection event subset to obtain a specular reflection point cloud model.
[0089] Exemplarily, an optical path model of specular reflection is established, considering the relationship between incident light, reflected light, and surface normal vectors. Then, according to the intensity variation law of polarized light and combining the reflection intensity data collected at multiple polarization angles, the normal vectors of each point on the surface are calculated. Finally, based on the normal vector field and the position constraints provided by the diffuse reflection point cloud model, the fine structure of the specular reflection surface is reconstructed through an optimization algorithm.
[0090] Through polarization measurement reconstruction, the specular reflection characteristics and microscopic structure of the object surface can be accurately restored, complementing the deficiencies of diffuse reflection reconstruction. This method is particularly suitable for processing surfaces with high gloss.
[0091] In some embodiments, using the diffuse reflection point cloud model as a virtual screen, polarization measurement reconstruction is performed on a subset of specular reflection events to obtain a specular reflection point cloud model, including: performing Gaussian filtering on the diffuse reflection point cloud model to generate a smooth diffuse reflection surface model; performing projection mapping on the subset of specular reflection events according to the smooth diffuse reflection surface model to obtain an event projection map; calculating the polarization angles of each event point in the event projection map to obtain the polarization angle distribution of each event point; performing polarization information decoding according to the polarization angle distribution to obtain a specular reflection feature matrix; estimating the three-dimensional coordinates of the specular reflection feature matrix to obtain a three-dimensional coordinate set of specular reflection event points; and performing point cloud generation processing according to the three-dimensional coordinate set of specular reflection event points to obtain a specular reflection point cloud model.
[0092] Exemplarily, applying Gaussian filtering to the existing diffuse reflection point cloud model aims to generate a smooth surface model and reduce the influence of noise. This smooth surface will serve as the reference plane for subsequent projection. Projecting a subset of specular reflection events (which may be the reflection data captured by the sensor) onto the smooth diffuse reflection surface generates an event projection map that records the two-dimensional distribution of reflection points. Calculating the polarization angles of each projected event point forms a distribution map of polarization angles, and this polarization information contains important clues about the surface geometry. Decoding the polarization angle distribution generates a matrix containing specular reflection characteristics, which may include information such as reflection intensity and angle. Estimating the three-dimensional coordinates based on the specular reflection feature matrix obtains the spatial position information of the reflection points, and finally, a point cloud model of specular reflection is generated through these three-dimensional coordinates.
[0093] Through the above method, the diffuse reflection and specular reflection information are combined, and polarization measurement is used to provide additional geometric information, achieving a 2D to 3D reconstruction through multi-step processing.
[0094] S106. Perform multimodal fusion reconstruction according to the diffuse reflection point cloud model and the specular reflection point cloud model to obtain a fused three-dimensional model of the complete scene.
[0095] In the multi-modal fusion and reconstruction stage, first, the diffuse point cloud model and the specular reflection point cloud model are registered to align the two models in the same coordinate system. Then, based on the complementary characteristics of the two models, a fusion weight function is designed to adaptively adjust the fusion ratio of the two models according to the surface characteristics in different regions. For the overlapping regions, an optimization algorithm is used to solve the optimal fusion parameters, so that the fused model not only maintains geometric continuity but also can accurately express the reflection characteristics of the surface.
[0096] In some embodiments, multi-modal fusion reconstruction is performed according to the diffuse point cloud model and the specular reflection point cloud model to obtain a fused three-dimensional model of the complete scene, including: extracting diffuse reflection features and specular reflection features respectively according to the diffuse point cloud model and the specular reflection point cloud model to obtain a diffuse reflection feature matrix and a specular reflection feature matrix; performing mid-term fusion processing on the diffuse reflection feature matrix and the specular reflection feature matrix to obtain a fused feature matrix; inputting the fused feature matrix into a regressor for regression processing to obtain a preliminary three-dimensional structure; performing global optimization processing according to the preliminary three-dimensional structure, and using a preset graph cut algorithm to correct the consistency of features from different perspectives to obtain an optimized three-dimensional structure; performing detail enhancement processing on the optimized three-dimensional structure to generate a fused three-dimensional model of the complete scene.
[0097] Exemplarily, when extracting the diffuse reflection features, a deep learning network of PointNet++ is used to extract local features such as the normal vector and curvature of the point cloud, calculate the point cloud density and distribution features, and generate a diffuse reflection feature matrix. When extracting the specular reflection features, the high-reflectivity regions are detected, the specular reflection intensity and direction information are extracted, the geometric properties of the reflecting surface are analyzed, the reflection light path features are constructed, and a specular reflection feature matrix is obtained. During the intermediate fusion process, feature alignment is required. Specifically, spatial registration and temporal synchronization are performed, the attention mechanism is used to highlight important features, and the corresponding relationship between the diffuse reflection feature matrix and the specular reflection feature matrix is established. According to the above corresponding relationship, an adaptive weight allocation mechanism is applied for weighted summation, and at the same time, a complementary information enhancement module is introduced to construct a multi-level fusion network. In the regression process, an encoder-decoder architecture is used for network structure design, skip connections are added to retain detailed information, and residual blocks are used to improve the feature extraction ability. Combining the reconstruction error and the regularization term, the loss function design for the regression process is carried out, and by adding geometric constraint loss and introducing symmetry and continuity constraints, a preliminary three-dimensional structure is obtained. Based on the preliminary three-dimensional structure, the similarity metric between nodes is defined, an energy function is designed to optimize the similarity metric, and a feature map network structure is constructed. Multi-view feature matching, pose optimization, and loop detection are performed on the feature map network structure to obtain an optimized three-dimensional structure. To improve the accuracy, a subdivision surface algorithm is applied to perform edge sharpening on the optimized three-dimensional structure and add texture details. Specifically, the optimized three-dimensional structure is meshed and then hole filling is performed to optimize the topological structure.
[0098] Through the above mechanism, the characteristics of diffuse reflection and specular reflection are fully considered, and high-quality 3D reconstruction results are obtained through multiple stages of processing and optimization.
[0099] Please refer to Figure 2 , Figure 2 FIG. is a schematic block diagram of a three-dimensional imaging device based on a 3D and AI vision sensing visible light module. The three-dimensional imaging device 200 based on the 3D and AI vision sensing visible light module is used to execute the aforementioned three-dimensional imaging method based on the 3D and AI vision sensing visible light module. Among them, the three-dimensional imaging device 200 based on the 3D and AI vision sensing visible light module can be configured in a server.
[0100] Among them, the server can be an independent server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0101] As Figure 2 shown, a three-dimensional imaging device 200 based on a 3D and AI vision sensing visible light module includes: a data acquisition module 201, a feature extraction module 202, a data decomposition module 203, a first reconstruction module 204, a second reconstruction module 205, and a model fusion module 206.
[0102] The data acquisition module 201 is configured to acquire original event data through the AI vision sensing visible light module.
[0103] The feature extraction module 202 is configured to perform spatio-temporal feature extraction on the original event data to obtain an event stream data matrix.
[0104] The data decomposition module 203 is configured to perform epipolar geometry constraint deconstruction according to the event stream data matrix to obtain a diffuse reflection event subset and a specular reflection event subset.
[0105] The first reconstruction module 204 is configured to perform triangulation reconstruction according to the diffuse reflection event subset to obtain a diffuse reflection point cloud model.
[0106] The second reconstruction module 205 is configured to use the diffuse reflection point cloud model as a virtual screen to perform polarization measurement reconstruction on the specular reflection event subset to obtain a specular reflection point cloud model.
[0107] The model fusion module 206 is configured to perform multi-modal fusion reconstruction according to the diffuse reflection point cloud model and the specular reflection point cloud model to obtain a fused three-dimensional model of the complete scene.
[0108] An embodiment of the present application provides an electronic device, which includes a memory and a processor; the memory is used to store a computer program; the processor is configured to execute the computer program and implement a three-dimensional imaging method based on a 3D and AI vision sensing visible light module as described in any one of the embodiments of the present application when executing the computer program.
[0109] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement a three-dimensional imaging method based on a 3D and AI vision sensing visible light module as described in any one of the embodiments of the present application.
[0110] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or replacements, and these modifications or replacements should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A three-dimensional imaging method based on 3D and AI visual sensing visible light movement, characterized in that: The method comprises: Obtain raw event data through AI visual sensing visible light movement; Extracting spatiotemporal features from the original event data to obtain an event stream data matrix; Perform epipolar geometry constraint deconstruction according to the event stream data matrix to obtain a diffuse reflection event subset and a specular reflection event subset; Perform triangulation reconstruction according to the diffuse reflection event subset to obtain a diffuse reflection point cloud model; Using the diffuse reflection point cloud model as a virtual screen, performing polarization measurement reconstruction on the specular reflection event subset to obtain a specular reflection point cloud model; Multimodal fusion reconstruction is performed based on the diffuse reflection point cloud model and the specular reflection point cloud model to obtain a fused three-dimensional model of the complete scene.
2. The three-dimensional imaging method based on 3D and AI visual sensing visible light core as claimed in claim 1, characterized in that: The raw event data is obtained by using the AI visual sensing visible light movement, including: Perform log intensity change monitoring on each pixel photographed by the AI vision sensing visible light movement to obtain pixel intensity change data; Performing threshold comparison judgment on the pixel intensity change data, and generating an event trigger signal including a timestamp, polarity and coordinate position according to the judgment result; Accumulating and counting intensity changes of pixel positions according to the event trigger signal to obtain an event count matrix; Dividing the event count matrix into n-bin voxel grids according to the time dimension to obtain time-space discretization data blocks; Performing polar accumulation projection on each 3D slice in the spatiotemporal discretization data block to obtain a two-dimensional projection frame sequence; The pixel value normalization process is performed on the two-dimensional projection frame sequence to obtain the original event data stream.
3. The three-dimensional imaging method based on 3D and AI visual sensing visible light core as claimed in claim 2, characterized in that: The extracting of spatiotemporal features from the original event data to obtain an event stream data matrix includes: Performing time series segmentation processing on the original event data, dividing the event data into multiple time series bin sequences according to time intervals; Performing a downsampling operation on the time series bin sequence to obtain a reduced-dimensional state sequence; Constructing a global spatial dependency extractor according to the reduced dimensionality state sequence, performing attention calculation on the reduced dimensionality state at each time step, and obtaining a global spatial dependency feature; Performing motion perception analysis on the global space-dependent features, extracting difference information between adjacent states through subtraction operations, and obtaining a motion feature map; Perform spatial attention calculation according to the motion feature map to generate a discriminative attention weight map; Modulating the state feature based on the discriminative attention weight map to obtain an enhanced state feature; Performing cross-domain attention fusion on the enhanced state features to generate multi-scale feature representation; Adaptive weighted balancing processing is performed according to the multi-scale feature representation to obtain the event stream data matrix.
4. The three-dimensional imaging method based on 3D and AI visual sensing visible light core as claimed in claim 1, characterized in that: The performing epipolar geometry constraint deconstruction according to the event stream data matrix to obtain a diffuse reflection event subset and a specular reflection event subset includes: Performing polarity correlation analysis on the event stream data matrix, calculating the mutual correlation coefficient of event polarities in a preset spatial neighborhood, and constructing a polarity correlation graph according to a preset relationship threshold and the mutual correlation coefficient; Performing basic matrix estimation on the polar association graph, performing iterative calculations using a RANSAC algorithm and preset camera motion parameters to obtain an epipolar geometric transformation matrix; Calculating the epipolar distance of event point pairs according to the epipolar geometric transformation matrix, and performing event consistency verification based on a preset distance threshold to obtain a geometric consistency score; Performing a bimodal cluster analysis on the geometric consistency scores, and dividing the events into a diffuse reflection candidate set and a specular reflection candidate set by using a Gaussian mixture model; According to the spatial distribution of the diffuse reflection candidate set and the specular reflection candidate set, a local surface normal vector field is constructed, and a normal vector consistency check is performed to obtain a surface reflection characteristic classification result; The events are classified and sorted based on the surface reflection characteristic classification result to generate the diffuse reflection event subset and the specular reflection event subset.
5. The three-dimensional imaging method based on 3D and AI visual sensing visible light core as claimed in claim 1, characterized in that: The method uses the diffuse reflection point cloud model as a virtual screen to perform polarization measurement reconstruction on the specular reflection event subset to obtain the specular reflection point cloud model, including: Performing Gaussian filtering on the diffuse reflection point cloud model to generate a smooth diffuse reflection surface model; Performing projection mapping on the specular reflection event subset according to the smooth diffuse reflection surface model to obtain an event projection map; Calculating the polarization angle of the event projection image to obtain the polarization angle distribution of each event point; Perform polarization information decoding according to the polarization angle distribution to obtain a mirror reflection feature matrix; Performing three-dimensional coordinate estimation on the specular reflection feature matrix to obtain a three-dimensional coordinate set of specular reflection event points; Point cloud generation processing is performed according to the three-dimensional coordinate set of the specular reflection event point to obtain the specular reflection point cloud model.
6. The three-dimensional imaging method based on 3D and AI visual sensing visible light core as claimed in claim 1, characterized in that: The performing multimodal fusion reconstruction according to the diffuse reflection point cloud model and the specular reflection point cloud model to obtain a fused three-dimensional model of the complete scene includes: According to the diffuse reflection point cloud model and the specular reflection point cloud model, respectively extracting diffuse reflection features and specular reflection features to obtain a diffuse reflection feature matrix and a specular reflection feature matrix; Performing mid-term fusion processing on the diffuse reflection feature matrix and the specular reflection feature matrix to obtain a fused feature matrix; Inputting the fused feature matrix into a regressor for regression processing to obtain a preliminary three-dimensional structure; According to the preliminary three-dimensional structure, a global optimization process is performed, and a preset graph cut algorithm is used to perform consistency correction on features under different viewing angles to obtain an optimized three-dimensional structure; The optimized three-dimensional structure is subjected to detail enhancement processing to generate a fused three-dimensional model of the complete scene.
7. The three-dimensional imaging method based on 3D and AI visual sensing visible light core as claimed in claim 1, characterized in that: The step of performing triangulation reconstruction according to the diffuse reflection event subset to obtain a diffuse reflection point cloud model comprises: Sorting the diffuse reflection event subset according to timestamps to obtain a chronologically ordered sequence of events; Calculating the spatial distance and time difference of each event point according to the event sequence, and calculating the triangulation relationship between every two adjacent event points according to the spatial distance and the time difference to obtain three-dimensional coordinate information; Point cloud reconstruction is performed according to the three-dimensional coordinate point set, and the three-dimensional coordinate point set is sparsely processed by a voxel grid method to obtain the diffuse reflection point cloud model.