Four-eye structured light stereoscopic vision imaging method
By employing a four-view structured light stereo vision method and utilizing multi-view consistency verification and spatial filtering optimization algorithms, the occlusion and error compensation problems of traditional binocular and triocular systems in complex surface measurement are solved, achieving high-precision and robust 3D reconstruction results.
Patent Information
- Application Number
- CN202510898146.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-11-14
AI Technical Summary
Existing binocular and triocular structured light stereo vision systems are easily affected by occlusion, uneven lighting, and reflection interference when processing complex surfaces, resulting in missing or failed parallax information, and insufficient viewpoint configuration leads to limited error compensation capabilities.
A four-lens structured light stereo vision method is adopted, in which four cameras with fixed viewing angle differences are arranged in three-dimensional space. By synchronously acquiring structured light stripe image data, combined with multi-view consistency verification and spatial filtering optimization algorithms, high-precision three-dimensional point cloud data is generated and dense parallax field is constructed to achieve sparse occlusion compensation.
It significantly improves the integrity and robustness of 3D reconstruction, reduces information loss in occluded areas, enhances the robustness of edge and detail structure matching, and improves the accuracy and stability of 3D point clouds, making it suitable for high-precision industrial inspection and intelligent visual analysis.
Smart Images

Figure CN120953480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional imaging technology, and in particular to a four-eye structured light stereo vision imaging method. Background Technology
[0002] Structured light stereo vision is a typical active 3D imaging method, widely used in industrial measurement, surface inspection, and robot navigation. Its basic principle is: a specific pattern is formed on the surface of the target object using a structured light projection device; images of the deformed pattern are acquired from different angles using one or more imaging devices; and the 3D coordinates of the object's surface are obtained using geometric triangulation relationships, achieving non-contact, high-precision shape reconstruction.
[0003] Current mainstream solutions mostly employ binocular structured light systems, using two cameras to construct the geometry for stereo parallax measurement. However, in actual measurement, binocular systems are easily affected by factors such as local occlusion of the object surface, uneven illumination, or surface reflection interference, resulting in missing parallax information or matching failures. This is especially true when dealing with targets with complex steps, uneven structures, or reflective surfaces, often leading to problems such as local breaks in the 3D point cloud or contour distortion. To improve the measurement capability for complex surfaces, some studies have introduced tri-vision structures, adding an extra viewpoint to enhance the redundancy of image acquisition and the robustness of stereo matching. Although tri-vision systems alleviate occlusion and reflection problems to some extent, matching ambiguities may still exist in edge contour regions or non-co-view regions. Furthermore, the limited matching combinations formed by three viewpoints result in limited error compensation capabilities and susceptibility to factors such as local distortion or coded light interference.
[0004] Current research and practice show that stereo vision systems are limited in improving the integrity and stability of 3D reconstruction due to limitations in viewpoint coverage and the geometric configuration of redundant matching structures. How to introduce viewpoint configurations with higher structural stability and stronger occlusion robustness while maintaining system compactness, and construct structured light stereo vision methods with higher error tolerance and boundary compensation capabilities, has become a crucial technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] This application provides a four-eye structured light stereo vision imaging method, and the technical solution provided is as follows:
[0006] In a first aspect, this application provides a four-eye structured light stereo vision imaging method, the method comprising:
[0007] Four cameras with a fixed viewing angle difference are arranged in three-dimensional space, and multi-view image data with structured light stripes are acquired synchronously under the condition of uniform structured light projection.
[0008] Image preprocessing and stripe coding recognition are performed on the acquired four-view structured light image data to extract the position of the center line of the structured light stripes and its corresponding spatial projection information under each view.
[0009] Based on the spatial projection information of multi-view structured light stripes, high-precision three-dimensional point cloud data of target samples are generated using parallax calculation and three-dimensional reconstruction algorithms, and spatial registration and filtering optimization of point clouds are completed.
[0010] Based on the obtained 3D point cloud data, combined with the multi-view structured light stripe reflection intensity information, the point cloud and light intensity data are fused to generate a composite dataset containing spatial and spectral information.
[0011] A dense disparity field is constructed based on a composite dataset, and stereo consistency verification is performed to generate a high-fidelity 3D reconstruction model with complete topological relationships and sparse occlusion compensation capabilities.
[0012] In one specific implementation, the arrangement of four cameras with a fixed viewing angle difference in three-dimensional space includes:
[0013] The four cameras are configured in a non-coplanar manner and are symmetrically distributed around the target area. Their optical axes are oriented to form an observation cone with a certain angle between them around the target.
[0014] Each camera needs to undergo individual intrinsic parameter calibration, and extrinsic parameter calibration consistent with the global coordinate system is completed by combining the measured target.
[0015] In one specific implementation scheme, the simultaneous acquisition of multi-view image data with structured light stripes under uniform structured light projection conditions includes:
[0016] The image acquisition timing of the four cameras is controlled by a hardware synchronization mechanism, and a structured light coded pattern is projected onto the sample surface through a preset structured light projection device.
[0017] After image acquisition is completed, grayscale equalization and gamma correction are performed on the original image in sequence, and preliminary distortion correction is performed on the image according to the previously completed calibration parameters.
[0018] In one specific implementation scheme, the step of performing image preprocessing and stripe coding recognition on the acquired four-view structured light image data, and extracting the position of the center line of the structured light stripes and its corresponding spatial projection information under each view includes:
[0019] Perform normalization processing on each viewpoint image separately;
[0020] Based on the structured light coding mode adopted, the image is stripe decoding: if it is phase-shift stripe coding, the phase difference of multiple frames of different phase images needs to be calculated and fitted to obtain a continuous phase map; if it is gray-scale modulation coding or binary coding structured light, gray-scale threshold segmentation or pattern template matching method is used to identify each stripe region.
[0021] The center extraction algorithm is used to extract the pixel sequence of the center line of the continuous stripes and record their precise positions in their respective image coordinate systems.
[0022] The extracted stripe center lines are projected back to their respective spatial lines of view relative to the world coordinate system by calling the extrinsic parameters of each of the four perspectives.
[0023] In a specific feasible implementation, the process of generating high-precision 3D point cloud data of the target sample based on the spatial projection information of multi-view structured light stripes, using parallax calculation and 3D reconstruction algorithms, and completing the spatial registration and filtering optimization of the point cloud includes:
[0024] By constructing a matching model between image pairs and combining camera calibration parameters, the center lines of structured light stripes corresponding to the positions in multiple viewpoints are reconstructed by triangulation.
[0025] For pixel pairs between different viewpoints, the coordinates of their 3D intersection point are estimated using the principle of minimum distance between two rays.
[0026] A rigid registration algorithm based on feature alignment is used to unify all point clouds into a global coordinate system.
[0027] In one specific implementation scheme, the step of fusing point cloud and light intensity data based on the obtained 3D point cloud data and combining multi-view structured light stripe reflection intensity information to generate a composite dataset containing spatial and spectral information includes:
[0028] Based on the generated 3D point cloud coordinates, combined with the intrinsic and extrinsic parameter matrices and imaging models of cameras from various viewpoints, the 3D spatial points are back-projected onto the structured light image coordinate system of the corresponding viewpoint to obtain their position index in the image.
[0029] By extracting the grayscale value or reflection intensity value of the corresponding pixel in the original structured light stripe image through the position index, and mapping it back to the point cloud space, a one-to-one correspondence between three-dimensional geometric points and two-dimensional light intensity is achieved.
[0030] In one specific implementation scheme, the step of constructing a dense disparity field based on a composite dataset and performing stereo consistency verification to generate a high-fidelity 3D reconstruction model with complete topological relationships and sparse occlusion compensation capabilities includes:
[0031] Based on multiple subsets of point clouds generated from four perspectives, dense stereo matching technology is used to calculate the disparity field, estimate the disparity value pixel by pixel, and map it back to three-dimensional space.
[0032] By using multi-view consistency verification, the dense point set of the preliminary reconstruction is screened for global consistency, with a focus on handling overlapping, redundant, or drifting points caused by baseline differences or occlusion reconstruction between different cameras.
[0033] Secondly, this application provides a four-eye structured light stereo vision imaging system, which adopts the following technical solution:
[0034] A four-eye structured light stereo vision imaging system, comprising:
[0035] The image acquisition module is used to arrange four cameras with a fixed viewing angle difference in three-dimensional space and synchronously acquire multi-view image data with structured light stripes under uniform structured light projection conditions.
[0036] The spatial projection module is used to perform image preprocessing and stripe coding recognition on the acquired four-view structured light image data, and to extract the position of the center line of the structured light stripes and its corresponding spatial projection information under each view.
[0037] The point cloud generation module is used to generate high-precision three-dimensional point cloud data of target samples based on the spatial projection information of multi-view structured light stripes, using parallax calculation and three-dimensional reconstruction algorithms, and to complete the spatial registration and filtering optimization of the point cloud.
[0038] The data fusion module is used to fuse point cloud and light intensity data based on the obtained 3D point cloud data and combine multi-view structured light stripe reflection intensity information to generate a composite dataset containing spatial and spectral information.
[0039] The 3D imaging module is used to construct a dense parallax field based on a composite dataset and perform stereo consistency verification to generate a high-fidelity 3D reconstruction model with complete topological relationships and sparse occlusion compensation capabilities.
[0040] Thirdly, this application provides an electronic device, the device including a processor and a memory; the memory stores a program, the program being loaded and executed by the processor to implement a four-eye structured light stereoscopic vision imaging method as described in the first aspect.
[0041] Fourthly, this application provides a computer-readable storage medium storing a program that, when executed by a processor, is used to implement a four-eye structured light stereo vision imaging method as described in the first aspect.
[0042] In summary, the beneficial effects of this application include at least the following:
[0043] (1) By employing a quad-camera setup and structured light synchronous projection technology, multi-view, multi-baseline image acquisition and 3D reconstruction were achieved, effectively expanding the field of view coverage and significantly reducing the blind spots and information loss problems caused by insufficient viewing angles in traditional binocular or trinocular systems. Based on the fusion of rich viewing information, the integrity and continuity of 3D point cloud data were significantly improved, especially in complex shapes and edge areas, enabling more accurate spatial matching and detail capture, thus ensuring high-precision stereoscopic vision imaging effects and meeting the stringent requirements for minute defects and detailed structures in industrial inspection.
[0044] (2) By introducing multi-view consistency verification and spatial filtering optimization algorithms, noise suppression and error correction are performed on the initially reconstructed 3D point cloud, greatly improving the stability and robustness of the reconstructed data. This technology effectively reduces reconstruction errors caused by changes in illumination, surface reflection, and structured light stripe breaks, ensuring that the output 3D model has high precision and high reliability, adapting to various complex industrial application environments, realizing high-quality stereo vision inspection and analysis, and promoting the practical application of structured light imaging technology in precision manufacturing and intelligent inspection fields.
[0045] This invention utilizes four cameras with fixed viewing angle differences deployed in three-dimensional space to simultaneously acquire structured light fringe image data under uniform structured light projection conditions. The cameras are installed in a geometrically symmetrical manner, ensuring sufficient and complementary spatial parallax between each pair of viewing angles. After image acquisition, the system performs fringe decoding, pixel-level matching, parallax calculation, and multi-view 3D reconstruction operations, supplemented by multi-path redundant registration and consistency verification strategies to construct a complete and high-precision 3D point cloud model. In the post-processing stage, an occlusion area recovery algorithm and a spatial filtering mechanism are introduced to further improve the continuity and realism of dense point clouds in edge, abrupt, and reflective regions. Based on the above technical solution, this application effectively solves the problems of insufficient viewing angle coverage, missing information in occluded areas, and difficulty in edge structure matching in traditional binocular and trinocular structured light stereo vision imaging. Specifically, the quad-camera configuration provides richer visual information, significantly reducing the parallax matching blind spot caused by single-view occlusion, and effectively improving the integrity and continuity of stereo reconstruction. The multi-baseline 3D reconstruction strategy enhances the robustness of edge and detail structure matching, avoiding depth breaks and artifacts. Combined with post-processing techniques of multi-view consistency verification and spatial filtering, noise and matching errors are suppressed, improving the accuracy and stability of the 3D point cloud. The overall solution, while ensuring a compact system structure and efficient data acquisition, achieves high-precision, robust 3D stereo vision imaging in complex scenes, making it suitable for high-precision industrial inspection and intelligent visual analysis.
[0046] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating the four-eye structured light stereo vision imaging method in the embodiments of this application.
[0048] Figure 2 This is a structural block diagram of the four-eye structured light stereo vision imaging system in the embodiments of this application.
[0049] Figure 3 This is a block diagram of an electronic device for tetracular structured light stereoscopic vision imaging in an embodiment of this application. Detailed Implementation
[0050] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.
[0051] Optionally, this application uses the four-eye structured light stereo vision imaging method provided in various embodiments in an electronic device as an example for illustration. The electronic device is a terminal or server. The terminal can be a mobile phone, computer, tablet computer, etc. This embodiment does not limit the type of electronic device.
[0052] Reference Figure 1 This is a flowchart illustrating a four-eye structured light stereo vision imaging method according to an embodiment of this application. The method includes at least the following steps:
[0053] Step S101: Arrange four cameras with a fixed viewing angle difference in three-dimensional space, and synchronously acquire multi-view image data with structured light stripes under the condition of uniform structured light projection.
[0054] In step S101, to achieve multi-view imaging under stereo vision conditions, the physical installation and arrangement of four cameras are first completed in the three-dimensional workspace. The four cameras are non-coplanar and symmetrically distributed around the target area. Their optical axes form observation cones at certain angles around the target, typically set between 15° and 30° to balance parallax information and the coverage of overlapping image areas. The arrangement process must ensure that the relative positions of the cameras are fixed and rigidly secured using precision brackets to prevent minor displacements during subsequent imaging. Furthermore, each camera requires individual intrinsic parameter calibration to obtain information such as focal length, principal point position, and distortion parameters. This is combined with a measured target to complete extrinsic parameter calibration (rotation matrix and translation vector) consistent with the global coordinate system, achieving unified coordinate mapping between multiple viewpoints. This arrangement strategy ensures the multi-view system has sufficient parallax foundation while avoiding overlap of visually occluded areas, improving the density and integrity of subsequent spatial reconstruction.
[0055] After completing the geometric configuration, a unified hardware synchronization mechanism (such as an external trigger signal controller) controls the image acquisition sequence of the four cameras to ensure that each viewpoint acquires the target image within the exact same exposure time. To improve the resolution and robustness of 3D information extraction, a structured light coded pattern is projected onto the sample surface using a pre-set structured light projection device (such as a DLP projector or laser scanning module) during image acquisition. The structured light pattern can be grayscale sinusoidal stripes, phase-shift coding, or spatial frequency coding, to adapt to samples with different surface textures, color reflectivities, and morphological characteristics. During actual projection, to reduce ambient light interference and improve pattern clarity, this step employs a short exposure time and a high-sensitivity sensor, adjusting the ambient lighting or introducing a filtering mechanism to control the background light ratio when necessary. After image acquisition, grayscale equalization and gamma correction are performed sequentially on the original images to suppress noise and enhance the edge clarity of the structured light stripes. Simultaneously, preliminary distortion correction is performed on the images based on previously completed calibration parameters to unify the observation coordinate system of each camera.
[0056] This step ensures spatiotemporal consistency among multi-view images through a highly consistent and low-latency synchronous acquisition mechanism. Simultaneously, the structured light projection pattern and imaging equipment work in perfect harmony, laying a stable foundation for subsequent high-precision fringe decoding and spatial point cloud reconstruction. Compared to traditional binocular systems, the four-eye structure significantly expands the observation angle range and increases the density of spatial overlap areas. It also exhibits higher stability and redundancy when handling complex geometric structures and occluded topography, forming the fundamental data source for the entire 3D vision method.
[0057] Step S102: Perform image preprocessing and stripe coding recognition on the acquired four-view structured light image data, and extract the position of the center line of the structured light stripes and its corresponding spatial projection information under each view.
[0058] In step S102, after image acquisition, the structured light images acquired by the four cameras are preprocessed and structural information extracted. First, to ensure the stability of subsequent encoding and recognition, a standardized processing procedure is performed on each viewpoint image, including image grayscale normalization, illumination equalization correction, noise suppression, and local contrast enhancement. In the noise suppression stage, Gaussian filtering or adaptive median filtering can be selected based on the image texture complexity to maintain edge clarity, while Laplacian or Sobel operators are introduced to enhance stripe edges. Subsequently, according to the structured light encoding mode adopted, stripe decoding is performed on the image: if it is phase-shifted stripe encoding, the phase difference of multiple frames of different phase images needs to be calculated and fitted to obtain a continuous phase map; if it is grayscale modulation encoding or binary encoded structured light, grayscale threshold segmentation or pattern template matching methods are used to identify each stripe region. To enhance the accuracy of stripe centerline positioning, a sub-pixel-level center extraction algorithm is introduced based on the above decoding results, such as a method based on grayscale centroid or fitting curve approximation, to extract a continuous stripe centerline pixel sequence and record its precise position in its respective image coordinate system.
[0059] After extracting the centerline, the image coordinates need to be spatially mapped using camera calibration information to establish the projection relationship from pixels to three-dimensional space. To do this, the extrinsic parameters (rotation matrix and translation vector) of each of the four viewpoints are invoked to project the extracted fringe centerlines back to their respective spatial lines of sight relative to the world coordinate system. Because the structured light fringes deform on the sample surface, the spatial projection lines corresponding to the fringe centerlines at different viewpoints will intersect or overlap in space on the object's surface. To improve matching accuracy, a parallax constraint model needs to be introduced to check the geometric consistency of the fringe centerlines between different viewpoints. For example, fringe position matching and reprojection error calculation are performed between binocular / trinocular pairs, and combined with left-right consistency or depth confidence analysis, abnormally matched fringes or incomplete information caused by occlusion are eliminated.
[0060] Through this step, the structured light images from four perspectives not only completed stripe pattern recognition but also extracted highly consistent centerline spatial projection paths, providing a crucial foundation for subsequent spatial localization of corresponding points in three-dimensional space. Thanks to redundant observation from four perspectives and multi-directional stripe information fusion, this step effectively alleviates the reconstruction blind zone problem caused by occlusion or pattern overlap in traditional binocular systems, further enhancing the integrity and anti-interference capability of point cloud construction.
[0061] Step S103: Based on the spatial projection information of multi-view structured light stripes, high-precision three-dimensional point cloud data of the target sample is generated using parallax calculation and three-dimensional reconstruction algorithms, and spatial registration and filtering optimization of the point cloud are completed.
[0062] In step S103, based on the spatial projection information of the extracted structured light stripe centerlines from the four cameras, multi-view matching and disparity solving are first performed. Specifically, by constructing a matching model between image pairs and combining camera calibration parameters, the corresponding structured light stripe centerlines in multiple viewpoints are reconstructed through triangulation. To improve matching accuracy, a pixel-level + sub-pixel-level hybrid method is used for stripe point correspondence matching, and projection constraints (such as consistency of the same structured light stripe number) are introduced to limit the matching search range. In the actual reconstruction process, for pixel point pairs between different viewpoints, the three-dimensional intersection coordinates are estimated using the two-ray minimum distance principle, while mismatched points that do not meet the consistency constraints are eliminated. Since the four-view configuration provides multi-path cross-validation conditions, a multi-view joint matching optimization mechanism can be introduced to stably obtain effective three-dimensional points even in the presence of occlusion or stripe blurring areas.
[0063] After the initial generation of the 3D point cloud, spatial registration and quality optimization are required. First, considering the local offsets and overlapping areas in the 3D sub-point clouds constructed from different perspectives, a rigid registration algorithm based on feature alignment (such as SVD to solve the minimum mean square error model) is adopted to unify all point clouds into a global coordinate system. To enhance the accuracy of describing complex boundaries and detailed regions, a density-weighted strategy or surface normal constraints can be introduced during the registration stage to ensure a continuous transition of the registered point cloud boundary contours. After registration, to eliminate reconstruction noise caused by occlusion, reflection, and low-texture areas, a multi-scale statistical filtering method is introduced to remove or reconstruct outliers and regions with abrupt changes in local density. For sample structures with high edge accuracy requirements, normal consistency filtering or surface fitting smoothing algorithms can be further combined to optimize the curvature changes of the point cloud, forming a topologically coherent, high-quality point cloud representation.
[0064] Through the 3D reconstruction and optimization process in this step, we can finally obtain 3D point cloud data with continuous structure, clear boundaries, and accurate geometric information, laying a highly reliable spatial foundation for subsequent sample contour analysis, defect identification, or multimodal fusion processing. Compared with traditional binocular or single-line structured light systems, this scheme, by utilizing a four-eye joint reconstruction and parallax fusion mechanism, maintains high fidelity and robustness even when facing complex morphologies, curved surface occlusion, or multi-layered structures, significantly improving the integrity of 3D imaging and measurement stability.
[0065] Step S104: Based on the obtained 3D point cloud data, combined with the multi-view structured light stripe reflection intensity information, perform point cloud and light intensity data fusion processing to generate a composite dataset containing spatial and spectral information.
[0066] In step S104, after completing the 3D point cloud reconstruction, to further improve the visual representation capability and defect recognition accuracy of the target sample, this step fuses the spatial point cloud data with the reflection intensity information in the structured light image point by point. Specifically, firstly, based on the generated 3D point cloud coordinates, combined with the intrinsic and extrinsic parameter matrices and imaging models of each viewpoint camera, the 3D spatial points are back-projected onto the structured light image coordinate system of the corresponding viewpoint to obtain their position index in the image. Through this position index, the grayscale value or reflection intensity value of the corresponding pixel in the original structured light stripe image is extracted, and the intensity information can be mapped and assigned back to the point cloud space, realizing a one-to-one correspondence between 3D geometric points and 2D light intensity.
[0067] To improve the accuracy and robustness of data fusion, a weighted fusion strategy is introduced when there is overlapping sampling of reflection intensity from multiple viewpoints. Based on the incident angle of each camera, the influence of focal length distortion, and the probability of local texture occlusion, the optimal weight for each point is calculated, ultimately synthesizing a single reflection intensity value. Simultaneously, considering the coding properties of structured light fringes, this intensity information not only includes traditional image grayscale but may also contain implicit information such as modulation waveform phase and fringe frequency response. This can be expanded into a set of spectral / modulation feature vectors and appended to each point cloud point, constructing a composite point cloud data structure with multi-dimensional information expression capabilities.
[0068] The fused composite dataset not only contains X, Y, and Z spatial coordinates but also incorporates multi-channel visual features such as grayscale intensity, phase information, and structured light response, essentially constructing a set of extended fields with optical properties on the point cloud dimension. This composite point cloud data can be used by subsequent AI models for multimodal feature learning, improving the ability to identify subtle defects or complex texture differences. Furthermore, it provides richer contextual clues for defect localization, maintaining stable recognition accuracy and spatial awareness even in complex structures, blurred textures, and highly reflective scenes. This fusion strategy is particularly suitable for the accurate identification and measurement analysis of transparent / semi-transparent materials, layered samples, and mixed-material regions in practical industrial inspection.
[0069] Step S105: Construct a dense disparity field based on the composite dataset and perform stereo consistency verification to generate a high-fidelity 3D reconstruction model with complete topological relationships and sparse occlusion compensation capabilities.
[0070] In step S105, after acquiring composite point cloud data with fused light intensity features, in order to further improve the spatial continuity and depth accuracy of the three-dimensional imaging results, the reconstructed point cloud is subjected to densification and structural consistency compensation processing.
[0071] Specifically, firstly, based on multiple subsets of point clouds generated from four perspectives, a disparity field is calculated using dense stereo matching techniques (such as a multi-baseline stereo matching algorithm based on cost volume). The disparity value is estimated pixel by pixel and mapped back to 3D space. During dense matching, to address the depth blurring problem caused by low-texture areas, occluded areas, or scenes with repetitive textures, a cost aggregation and disparity confidence evaluation mechanism is introduced. Interpolation or difference correction is performed on low-reliability areas to improve the overall matching stability.
[0072] Next, a global consistency screening is performed on the initially reconstructed dense point set using multi-view consistency verification methods, focusing on overlapping, redundant, or drifting points caused by baseline differences or occlusion reconstruction between different cameras. This process includes methods such as three-view or four-view cross-validation, inter-point normal consistency calculation, and local curvature change analysis to ensure that the reconstructed points have topological continuity, reasonable depth gradient, and uniform normal distribution in spatial logic. For occluded areas or blind spots not illuminated by structured light fringes, this step combines the depth gradient and edge contour information of adjacent areas to perform missing compensation, constructing a topological completion model of the fitted surface, thereby forming a complete 3D model with sparse compensation capabilities.
[0073] The final high-fidelity 3D model significantly outperforms conventional binocular / tricular structured light reconstruction systems in terms of spatial density, geometric consistency, and edge integrity. By combining structured light geometric coding, four-view dense parallax, light intensity feature-assisted judgment, and consistency rules in joint modeling, the reconstruction quality of complex boundary regions (such as surface transitions and abrupt step changes) is enhanced, and reconstruction voids caused by occlusion interference and insufficient texture are effectively suppressed. This model can be directly used for downstream tasks such as high-precision 3D measurement, surface defect modeling, and structural analysis, providing a solid spatial data foundation for automated quality inspection and digital modeling in application scenarios.
[0074] In summary, by deploying four cameras with fixed viewing angle differences in three-dimensional space and simultaneously acquiring structured light fringe image data under unified structured light projection conditions, the cameras are installed in a geometrically symmetrical manner, ensuring sufficient and complementary spatial parallax between each pair of viewing angles. After image acquisition, the system performs fringe decoding, pixel-level matching, parallax calculation, and multi-view 3D reconstruction operations, supplemented by multi-path redundant registration and consistency verification strategies to construct a complete and high-precision 3D point cloud model. In the post-processing stage, an occlusion area recovery algorithm and a spatial filtering mechanism are introduced to further improve the continuity and realism of dense point clouds in edge, abrupt, and reflective regions. Based on the above technical solution, this application effectively solves the problems of insufficient viewing angle coverage, missing information in occluded regions, and difficulty in edge structure matching in traditional binocular and trinocular structured light stereo vision imaging.
[0075] Specifically, the quad-camera configuration provides richer visual information, significantly reducing the parallax matching blind spot caused by single-view occlusion, and effectively improving the integrity and continuity of stereo reconstruction. The multi-baseline 3D reconstruction strategy enhances the robustness of edge and detail structure matching, avoiding depth breaks and artifacts. Combined with post-processing techniques of multi-view consistency verification and spatial filtering, noise and matching errors are suppressed, improving the accuracy and stability of the 3D point cloud. The overall solution, while ensuring a compact system structure and efficient data acquisition, achieves high-precision, robust 3D stereo vision imaging in complex scenes, making it suitable for high-precision industrial inspection and intelligent visual analysis.
[0076] Figure 2 This is a structural block diagram of a four-eye structured light stereo vision imaging system provided in one embodiment of this application. The system includes at least the following modules:
[0077] The image acquisition module is used to arrange four cameras with a fixed viewing angle difference in three-dimensional space and synchronously acquire multi-view image data with structured light stripes under uniform structured light projection conditions.
[0078] The spatial projection module is used to perform image preprocessing and stripe coding recognition on the acquired four-view structured light image data, and to extract the position of the center line of the structured light stripes and its corresponding spatial projection information under each view.
[0079] The point cloud generation module is used to generate high-precision three-dimensional point cloud data of target samples based on the spatial projection information of multi-view structured light stripes, using parallax calculation and three-dimensional reconstruction algorithms, and to complete the spatial registration and filtering optimization of the point cloud.
[0080] The data fusion module is used to fuse point cloud and light intensity data based on the obtained 3D point cloud data and combine multi-view structured light stripe reflection intensity information to generate a composite dataset containing spatial and spectral information.
[0081] The 3D imaging module is used to construct a dense parallax field based on a composite dataset and perform stereo consistency verification to generate a high-fidelity 3D reconstruction model with complete topological relationships and sparse occlusion compensation capabilities.
[0082] For relevant details, please refer to the above method implementation examples.
[0083] Figure 3 This is a block diagram of an electronic device provided in one embodiment of this application. The device includes at least a processor 401 and a memory 402.
[0084] Processor 401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0085] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 are used to store at least one instruction, which is executed by the processor 401 to implement the four-eye structured light stereo vision imaging method provided in the method embodiments of this application.
[0086] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 401, memory 402, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, peripheral devices include, but are not limited to: radio frequency circuits, touch displays, audio circuits, and power supplies.
[0087] Of course, electronic devices may also include fewer or more components, and this embodiment does not limit this.
[0088] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the four-eye structured light stereo vision imaging method of the above method embodiments.
[0089] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the four-eye structured light stereo vision imaging method of the above method embodiments.
[0090] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0091] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A four-lens structured light stereo vision imaging method, characterized in that, The method includes: Four cameras with a fixed viewing angle difference are arranged in three-dimensional space, and multi-view image data with structured light stripes are acquired synchronously under the condition of uniform structured light projection. Image preprocessing and stripe coding recognition are performed on the acquired four-view structured light image data to extract the position of the center line of the structured light stripes and its corresponding spatial projection information under each view. Based on the spatial projection information of multi-view structured light stripes, high-precision three-dimensional point cloud data of target samples are generated using parallax calculation and three-dimensional reconstruction algorithms, and spatial registration and filtering optimization of point clouds are completed. Based on the obtained 3D point cloud data, combined with the multi-view structured light stripe reflection intensity information, the point cloud and light intensity data are fused to generate a composite dataset containing spatial and spectral information. A dense disparity field is constructed based on a composite dataset, and stereo consistency verification is performed to generate a high-fidelity 3D reconstruction model with complete topological relationships and sparse occlusion compensation capabilities.
2. The tetranocular structured light stereo vision imaging method according to claim 1, characterized in that, The arrangement of four cameras with a fixed viewing angle difference in three-dimensional space includes: The four cameras are configured in a non-coplanar manner and are symmetrically distributed around the target area. Their optical axes are oriented to form an observation cone with a certain angle between them around the target. Each camera needs to undergo individual intrinsic parameter calibration, and extrinsic parameter calibration consistent with the global coordinate system is completed by combining the measured target.
3. The tetranocular structured light stereo vision imaging method according to claim 1, characterized in that, The simultaneous acquisition of multi-view image data with structured light stripes under uniform structured light projection conditions includes: The image acquisition timing of the four cameras is controlled by a hardware synchronization mechanism, and a structured light coded pattern is projected onto the sample surface through a preset structured light projection device. After image acquisition is completed, grayscale equalization and gamma correction are performed on the original image in sequence, and preliminary distortion correction is performed on the image according to the previously completed calibration parameters.
4. The tetranocular structured light stereo vision imaging method according to claim 1, characterized in that, The process of image preprocessing and stripe coding recognition of the acquired four-view structured light image data, and extracting the position of the center line of the structured light stripes and its corresponding spatial projection information from each viewpoint includes: Perform normalization processing on each viewpoint image separately; Based on the structured light coding mode adopted, the image is stripe decoding: if it is phase-shift stripe coding, the phase difference of multiple frames of different phase images needs to be calculated and fitted to obtain a continuous phase map; if it is gray-scale modulation coding or binary coding structured light, gray-scale threshold segmentation or pattern template matching method is used to identify each stripe region. The center extraction algorithm is used to extract the pixel sequence of the center line of the continuous stripes and record their precise positions in their respective image coordinate systems. The extracted stripe center lines are projected back to their respective spatial lines of view relative to the world coordinate system by calling the extrinsic parameters of each of the four perspectives.
5. The tetranocular structured light stereo vision imaging method according to claim 1, characterized in that, The spatial projection information based on multi-view structured light stripes, using parallax calculation and 3D reconstruction algorithms to generate high-precision 3D point cloud data of the target sample, and completing the spatial registration and filtering optimization of the point cloud includes: By constructing a matching model between image pairs and combining camera calibration parameters, the center lines of structured light stripes corresponding to the positions in multiple viewpoints are reconstructed by triangulation. For pixel pairs between different viewpoints, the coordinates of their 3D intersection point are estimated using the principle of minimum distance between two rays. A rigid registration algorithm based on feature alignment is used to unify all point clouds into a global coordinate system.
6. The tetranocular structured light stereo vision imaging method according to claim 1, characterized in that, The process of fusing point cloud data with light intensity data based on the obtained 3D point cloud data and combining it with multi-view structured light stripe reflection intensity information to generate a composite dataset containing spatial and spectral information includes: Based on the generated 3D point cloud coordinates, combined with the intrinsic and extrinsic parameter matrices and imaging models of cameras from various viewpoints, the 3D spatial points are back-projected onto the structured light image coordinate system of the corresponding viewpoint to obtain their position index in the image. By extracting the grayscale value or reflection intensity value of the corresponding pixel in the original structured light stripe image through the position index, and mapping it back to the point cloud space, a one-to-one correspondence between three-dimensional geometric points and two-dimensional light intensity is achieved.
7. The tetranocular structured light stereo vision imaging method according to claim 1, characterized in that, The process of constructing a dense disparity field based on a composite dataset and performing stereo consistency verification to generate a high-fidelity 3D reconstruction model with complete topological relationships and sparse occlusion compensation capabilities includes: Based on multiple subsets of point clouds generated from four perspectives, dense stereo matching technology is used to calculate the disparity field, estimate the disparity value pixel by pixel, and map it back to three-dimensional space. By using multi-view consistency verification, the dense point set of the preliminary reconstruction is screened for global consistency, with a focus on handling overlapping, redundant, or drifting points caused by baseline differences or occlusion reconstruction between different cameras.
8. A four-lens structured light stereo vision imaging system, characterized in that, include: The image acquisition module is used to arrange four cameras with a fixed viewing angle difference in three-dimensional space and synchronously acquire multi-view image data with structured light stripes under uniform structured light projection conditions. The spatial projection module is used to perform image preprocessing and stripe coding recognition on the acquired four-view structured light image data, and to extract the position of the center line of the structured light stripes and its corresponding spatial projection information under each view. The point cloud generation module is used to generate high-precision three-dimensional point cloud data of target samples based on the spatial projection information of multi-view structured light stripes, using parallax calculation and three-dimensional reconstruction algorithms, and to complete the spatial registration and filtering optimization of the point cloud. The data fusion module is used to fuse point cloud and light intensity data based on the obtained 3D point cloud data and combine multi-view structured light stripe reflection intensity information to generate a composite dataset containing spatial and spectral information. The 3D imaging module is used to construct a dense parallax field based on a composite dataset and perform stereo consistency verification to generate a high-fidelity 3D reconstruction model with complete topological relationships and sparse occlusion compensation capabilities.
9. An electronic device, characterized in that, The device includes a processor and a memory; the memory stores a program, which is loaded and executed by the processor to implement a tetracular structured light stereo vision imaging method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when executed by a processor, is used to implement a four-eye structured light stereo vision imaging method as described in any one of claims 1 to 7.
Citation Information
Cited By
Near-surface ranging method based on virtual large-baseline four-eye vision, medium and equipment
CN121383872A
A near-ground ranging method, medium and device based on virtual large baseline four-eye vision
CN121383872B
Empty container detection method and system based on multi-modal fusion
CN121482501A
Grain transporting vehicle state monitoring method fusing laser and visual image data
CN121527708A
Multi-frame superposition point cloud super-resolution system for three-dimensional reconstruction
CN121767571A