Spectral polarization three-dimensional 5D imaging system and method driven by cross-modal fusion

The spectral polarization 3D imaging system driven by cross-modal fusion acquires spectral and polarization information using Hadamard templates and polarizers, and achieves high-precision 3D reconstruction through a multi-scale gradient consistency sign correction model. This solves the problem of limited information dimensions in existing technologies and realizes high-precision spectral and depth reconstruction.

CN121280642BActive Publication Date: 2026-05-15NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF SCI & TECH
Filing Date
2025-12-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing hyperspectral imaging and polarization 3D imaging systems have limited information dimensions and target discrimination capabilities in complex natural environments, making it difficult to achieve high-precision 3D reconstruction.

Method used

A cross-modal fusion-driven spectral polarization 5D imaging system is used. It utilizes a single image sensor combined with a projector, a 4F lens, a functional lens, and a grayscale camera. It acquires spectral information and polarization angles through a Hadamard template and a polarizer, and performs high-precision 3D reconstruction by combining a multi-scale gradient consistency sign correction model.

Benefits of technology

High-precision spectral 3D reconstruction was achieved, with spectral resolution reaching the nanometer level and depth accuracy reaching the micrometer level. This solved the azimuth ambiguity problem in polarization imaging and improved the system's compactness and practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121280642B_ABST
    Figure CN121280642B_ABST
Patent Text Reader

Abstract

The application discloses a kind of spectral polarization three-dimensional 5D imaging system and method driven by cross-modal fusion, obtains image using gray scale camera, uses DMD as encoding device, uses polaroid to collect object polarization information, under the inspiration of encoding spectral imaging on image plane, using object plane encoding method, reconstructs the spectral information and coarse three-dimensional information of target. Utilize the rough depth information and the polarization image of four different vibration directions obtained, with rough depth information as constraint, eliminate the azimuth ambiguity problem of polarization three-dimensional imaging, and reconstruct high-precision three-dimensional information. The present application obtains the spectral information and rough depth estimation of the target by encoding and decoding mechanism, introduces multi-scale gradient consistency sign correction model, uses the rough depth information obtained by encoding imaging as priori, guides the polarization three-dimensional reconstruction process, so as to solve the azimuth ambiguity problem in polarization imaging, and realize high-precision three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multimodal imaging technology and relates to a cross-modal fusion-driven spectral polarization three-dimensional 5D imaging system and method. Background Technology

[0002] The demand for multimodal collaborative analysis of matter is driving the development of optical imaging technology towards high-dimensional data perception. In recent years, optical remote sensing instruments and related technologies for multimodal information detection have emerged. A typical implementation method involves independent imaging by a traditional spectral detection system and a traditional 3D imaging system. Polarization information is obtained by adding a polarizer in front of the camera, and then fused into a multi-dimensional data cube using a registration algorithm to achieve multimodal information detection. A multimodal imaging system can acquire a 5D cube, including two spatial dimensions, one spectral dimension, one polarization dimension, and one depth dimension of the target. Multimodal information has wide applications in remote sensing, industrial defect detection, biomedicine, cultural relic protection, national defense, environmental monitoring, and other fields.

[0003] Hyperspectral imaging technology can be divided into non-coded hyperspectral imaging systems and coded hyperspectral imaging systems, depending on whether a modulation structure is introduced. Non-coded systems primarily acquire hyperspectral data point-by-point or line-by-line based on optical beam-splitting elements (such as prisms, gratings, or filter arrays). The imaging process relies on purely physical mechanisms, offering advantages such as intuitive imaging and no need for data reconstruction. However, these methods generally suffer from drawbacks such as complex system structure, large size, low acquisition efficiency, and difficulty in adapting to dynamic scenes, limiting their development for real-time imaging and portable integrated applications. To overcome these shortcomings, coded hyperspectral imaging systems based on computational imaging theory have been developed in recent years. These systems achieve information multiplexing or compressed acquisition by introducing modulation structures in the optical path, and then recover high-dimensional spectral information using reconstruction algorithms. These systems can be further divided into underdetermined systems based on compressed sensing theory and decodable systems based on multiplexing, depending on whether the information is compressed. Typical compressed sensing systems, such as coded aperture snapshot spectral imaging (CASSI), achieve hyperspectral acquisition in a single frame by compressing spectral information. However, their imaging performance is limited by the coding strategy and reconstruction method, requiring a trade-off between the number of measurements, spatial resolution, and spectral resolution. In contrast, Hadamard-coded imaging modulates spectral information without introducing compression, exhibiting better linear response and mathematical interpretability, providing a reliable benchmark for the accuracy verification and algorithm design of hyperspectral imaging systems.

[0004] Meanwhile, 3D shape acquisition, another important direction in computational imaging, plays a crucial role in complex scene perception. Among these, polarization-based 3D imaging, by incorporating a polarization acquisition structure into the imaging system, acquires polarization information. It extracts surface normal information by collecting reflected light modulated by the object's surface, and then uses 3D reconstruction algorithms to reconstruct the object's spatial shape features. Polarization information mainly includes the polarization angle and degree of polarization, reflecting the variation of polarized light direction and intensity, respectively, and is closely related to the normal distribution and material properties of the object's surface. Compared to traditional stereo vision or structured light methods, polarization-based 3D imaging has advantages such as simple structure, no need for active illumination, and sensitivity to minute geometric changes, making it particularly suitable for perceiving the shape of objects with weak textures, smooth surfaces, or rich details.

[0005] However, whether it is hyperspectral imaging or polarization 3D imaging, single-modal information suffers from limited information dimensions and limited target discrimination ability when facing complex natural environments and multi-source interference.

[0006] Therefore, a new cross-modal fusion-driven spectral polarization 3D 5D imaging system is needed to solve the above problems. Summary of the Invention

[0007] The purpose of this invention is to provide a cross-modal fusion-driven spectral polarization three-dimensional 5D imaging system to overcome the shortcomings of the prior art.

[0008] The technical solution of the present invention is as follows:

[0009] A cross-modal fusion-driven spectral polarization 3D 5D imaging system is a single-image sensor imaging system. The single-image sensor imaging system includes a projector, a 4F lens, a functional lens, and a grayscale camera. Therefore, both the projector and the grayscale camera are facing the target. The functional lens and the 4F lens are arranged sequentially between the grayscale camera and the target. The functional lens is a double Amish dispersion prism or a polarizer.

[0010] Furthermore, the projector includes a light source, a hadamard template, and a spatial light modulator (DMD). The light source is a non-polarized LED light source, and the hadamard template is loaded into the spatial light modulator (DMD), which is positioned between the light source and the target. When using a non-polarized LED light source for illumination and object-plane encoded imaging, the hadamard template is loaded into the DMD and projected onto the target. Light reflected from the target passes through a 4F lens and is dispersed parallel to a double Amish dispersion prism before being received by a grayscale camera.

[0011] Furthermore, the wavelength of the unpolarized LED light source is 380-780nm, and the resolution of the Hadamard template is 512*512.

[0012] Furthermore, the wavelength of the dual Amish dispersion prism is 410-660nm.

[0013] Beneficial effects: The cross-modal fusion-driven spectral polarization 5D imaging system of the present invention acquires the spectral information and coarse coded depth information of the target through dual Amish dispersion prisms, obtains polarization images with different polarization angles through polarizers, and achieves high-precision 3D reconstruction by combining the spectral information, coarse coded depth information and polarization images of the target.

[0014] The present invention also discloses a cross-modal fusion-driven spectral polarization three-dimensional 5D imaging method, which employs the cross-modal fusion-driven spectral polarization three-dimensional 5D imaging system described above;

[0015] The spectral polarization three-dimensional 5D imaging method includes the following steps:

[0016] 1) The functional mirror is set as a double Amish dispersion prism, and object plane coding imaging is used to realize object plane coding 4D imaging to obtain four-dimensional data of the measured target. The four-dimensional data includes spectral information and coarse coding depth information; the functional mirror is set as a polarizer, and polarization images with different polarization angles are obtained through polarization three-dimensional imaging.

[0017] 2) Using the coarse encoded depth information and polarization map obtained in step 1), the azimuth and zenith angles are calculated using the Stokes vector method. The azimuth angle is then corrected using a multi-scale gradient consistency sign correction model to obtain the correct azimuth angle, thereby obtaining high-precision three-dimensional information.

[0018] 3) Fuse the high-precision three-dimensional information obtained in step 2) with the spectral information obtained in step 1) to obtain high-quality spectral three-dimensional information.

[0019] Furthermore, in step 1), the four-dimensional data Through the reconstructed single-row spectral three-dimensional data Add one The dimensions are obtained, among which the reconstructed single-row spectral three-dimensional data are obtained. The following formula is used to calculate:

[0020] ,

[0021] In the formula, It is a single-row spectral three-dimensional data. This represents the three-dimensional spectral data after dispersion. This represents the information acquired through compressed acquisition by a two-dimensional detection device. This refers to all the encoded data obtained after n measurements of a single line of data. The Hadamard inverse matrix corresponding to the encoded data. After dispersion by the dispersive device , Modulated by spectral and three-dimensional information , For encoding information, For unencoded shadow acquisition areas, the grayscale increase due to depth is zero. To encode information that was not captured due to depth in uncaptured areas. and These represent the spatial width, height, depth, and number of bands of 4D information, respectively. Represents spatial coordinates, Indicates spectral channels, Indicates the target to be tested in a certain row. This represents the information in the i-th row of the Hadamard matrix. It is the inverse dispersion operation function. Here are the dispersion operation functions. This is a function for modulating three-dimensional information.

[0022] Furthermore, in step 1), the different polarization angles are 0°, 45°, 90° and 135°.

[0023] Furthermore, in step 2), the azimuth and zenith angles are calculated using the following formula:

[0024] ,

[0025] when hour, ,

[0026] when hour, ,

[0027] ,

[0028] In the formula, The zenith angle is obtained through the degree of polarization. This is the uncorrected deflection angle. and It is a two-dimensional vector field composed of gradient combinations. and For the fusion and gradient of direction, For the fusion of fundamental constants, This is the binary result of the Canny edge map. The threshold is the basic constant. For large-scale convolution kernel size, For small-scale convolution kernel size, and for direction and Small-scale gradient in direction and for direction and Large-scale gradient in direction, This is a dynamic threshold.

[0029] Furthermore, , , and The following formula is used to calculate:

[0030] ,

[0031] In the formula, To coarsely encode depth information, , , and They are Small-scale direction Small-scale direction Large-scale directions and Large-scale Sobel convolution kernels in the direction of the direction It is a one-dimensional directional difference weight vector. For linearly weighted parameters, The position of the center row (column), For odd-numbered scales, For the number of rows, For a single-row unnormalized gradient operator, and for direction and Direction Sobel core, Depend on Small-scale Sobel convolution kernels and Large-scale Sobel convolution kernels composition, Depend on Small-scale Sobel convolution kernels and Large-scale Sobel convolution kernels composition.

[0032] Furthermore, in step 2), the multi-scale gradient consistency sign correction model specifically involves: using the dot product to determine whether the gradient directions of the two depth maps are consistent, generating a correction coefficient matrix consisting of ±1, which is used to flip polarization gradients with opposite directions. The dot product of the two-dimensional vector field is calculated pixel-by-pixel using the following formula. :

[0033] ,

[0034] In the formula, and For the fusion and gradient of direction, The zenith angle is obtained through the degree of polarization. This is the uncorrected deflection angle;

[0035] When dot product When the angle between the two directions is less than 90°, the gradient directions of the two depth maps are consistent, and the dot product... At that time, the gradient directions of the two depth maps are opposite.

[0036] The dot product relationship between the reference gradient and the polarization gradient is used to determine whether their directions are consistent. A correction coefficient matrix consisting of ±1 is generated to flip polarization gradients with opposite directions, thereby improving the accuracy and consistency of gradient direction in subsequent polarization depth reconstruction.

[0037] Beneficial effects: The cross-modal fusion-driven spectral polarization 3D imaging method of the present invention obtains the spectral information and roughness depth estimation of the target through an encoding and decoding mechanism, introduces a multi-scale gradient consistency sign correction model, and uses the roughness depth information obtained by encoded imaging as a priori to guide the polarization 3D reconstruction process, thereby solving the azimuth ambiguity problem in polarization imaging and achieving high-precision 3D reconstruction. Attached Figure Description

[0038] Figure 1 It is a diagram of the Hadamard coding system;

[0039] Figure 2 It is an encoding process similar to Hadamard encoding;

[0040] Figure 3 The polarization diagram and the intensity of a point in the polarization image under different polarization angles are shown to vary with the polarization angle, with the selected angle being uniformly distributed in the range of [0°360°].

[0041] Figure 4 A schematic diagram of the object plane encoding for a cross-modal fusion-driven spectral polarization 3D imaging method;

[0042] Figure 5 A schematic diagram of the single-line single-template encoding process for a cross-modal fusion-driven spectral polarization 3D imaging method;

[0043] Figure 6The process for polarization-based 3D reconstruction with physical priors includes: (a) coarse depth information obtained by encoding; (b) polarization images and Stokes vectors in four directions acquired by polarization acquisition; (c) uncorrected azimuth angle; (d) corrected azimuth angle; (e) degree of polarization and the relationship between degree of polarization and zenith angle; (f) uncorrected depth map; and (g) corrected depth map.

[0044] Figure 7 This is a schematic diagram of the structure of a single image sensor imaging system;

[0045] Figure 8 A 2D RGB image as the target;

[0046] Figure 9 A 4-D RGB image as the target;

[0047] Figure 10 Spectral curves for different color regions of the target;

[0048] Figure 11 The reconstruction result is based on wavelength coloring for the target.

[0049] Figure 12 The images are reconstructions of a standard sphere; where (a) is a polarization image, (b) is an object-encoded depth map, (c) is a DRLPR, (d) is an SFPGI, and (e) is the method of the present invention.

[0050] Figure 13 Comparison images of 3D reconstruction of irregular objects of different materials, including (a) grayscale image, (b) object surface encoded depth map, (c) DRLPR, (d) SFPGI, and (e) the method of the present invention.

[0051] In the diagram, 1 is the projector, 2 is the grayscale camera, 11 is the first lens, 12 is the second lens, 13 is the third lens, 14 is the spatial light modulator, 31 is the polarizer, 32 is the dispersive prism, 41 is the fourth lens, 42 is the fifth lens, and 43 is the sixth lens. Detailed Implementation

[0052] The theoretical methods involved in this invention include two parts: Hadamard transform spectral imaging and polarization detection.

[0053] I. Hadamard Transform Spectral Imaging: Encoding and dispersion are two crucial processes in Hadamard-encoded spectral imaging systems for achieving spatial and spectral modulation. Specifically, encoding modulates the light beam containing scene and target information in the spatial dimension within the system's optical path, while dispersion uses prisms or gratings to shift the light beam along different wavelengths in a specific direction. A typical Hadamard-encoded spectral imaging system first encodes the light beam using a template loaded in a DMD, then disperses it, and finally acquires the modulated beam using a camera sensor. The imaging system is as follows: Figure 1 As shown, 2 is a grayscale camera, 11 is the first lens, 12 is the second lens, 13 is the third lens, and 14 is the spatial light modulator.

[0054] Hadamard transform spectral imaging is a typical multiplexed spectral imaging method based on matrix multiplication. The encoding process of this method is as follows: Figure 2 As shown, an m*m Hadamard matrix is ​​decomposed row by row into m coding factors. Each row of coding factors is vertically copied to form an m*m coding template to encode the object. After dispersion modulation, the encoded data is received. One acquisition requires m encoding acquisitions to obtain a cube of encoded data.

[0055] For an m-order Hadamard matrix It can be written as:

[0056] (1)

[0057] If each line is represented as Right now ,but It can be represented as ,

[0058] For the signal to be measured S:

[0059] (2)

[0060] If each column is represented as Right now Then S can be expressed as The process of decomposing the Hadamard matrix into m coding factors by row and performing m measurements on the signal under test can be represented as:

[0061] (3)

[0062] The dispersion process that occurs after each encoding is completed is represented as follows.

[0063] (4)

[0064] in, For encoded data, This is a physical dispersion process. These are measured values. After m encoding processes, m coded dispersive data sheets are obtained, forming a cube of detection data. .

[0065] The specific process of decoding and reconstructing encoded modulation information is as follows: obtaining the cube of probe data. Then, the spectral information is decoded line by line, using the line-by-line information. With the inverse matrix of Hadama Multiplying them together yields the spectral information for each line after dispersion. Each line of the signal to be measured can be obtained by inverse dispersion. The spectral information, after The signal to be tested can be obtained after the second decoding. .

[0066] II. Polarization Detection: Light waves are transverse waves. The vibration intensity of natural light is consistent at every angle on a plane perpendicular to the propagation direction. When light vibrates asymmetrically on a plane perpendicular to the propagation direction, it becomes polarized light. For an object illuminated by natural light, the reflected light becomes polarized due to the different reflectivities of the vibrating light perpendicular and parallel to the incident plane. Using the degree of polarization (DoP) and polarization angle (AoP) of the polarized light, the azimuth and zenith angles of the surface normal can be calculated. Furthermore, the gradients of the surface normal in the x and y directions can be calculated, and the three-dimensional information of the surface can be reconstructed through gradient integration.

[0067] The Stokes vector is commonly used to quantitatively describe polarized light. Since only a few targets affect the circularly polarized light component, imaging detection primarily detects the first three parameters of the Stokes vector. For linearly polarized light, the Stokes vector... , , Represented as:

[0068] (5)

[0069] in, , , , Images representing different polarization phase angles.

[0070] In polarization imaging detection, the linear polarization degree and polarization angle information of the target are the most widely used. The linear polarization degree DoLP represents the proportion of polarized light intensity and can be expressed as...

[0071] (6)

[0072] The polarization angle AoP represents the angle between the polarization direction of linearly polarized light and the reference direction, and is also the azimuth angle of the normal to the object's surface. It can be written as...

[0073] ,

[0074] Detection intensity Angle with linear polarizer The relationship is shown in the following formula.

[0075] ,

[0076] in, and This represents the maximum and minimum values ​​observed by the rotating linear polarizer. This is the azimuth angle of the normal to the object's surface. Therefore, when the polarizer angle... equal to azimuth angle From time to time The above formula and experimental verification yielded... Figure 3 It can be seen that during the 360° rotation of the polarizer, there exists a phase angle. The two maximum values ​​differ by 180°, so the azimuth angle of the surface normal is... or .

[0077] According to Fresnel's equations, the relationship between the zenith angle and the degree of polarization of the normal to the specular reflecting surface can be obtained as follows:

[0078] ,

[0079] in For refractive index, Let be the zenith angle. For a diffuse surface, the above equation can be written as:

[0080] ,

[0081] The zenith angle of the surface normal obtained and azimuth The gradient of the three-dimensional surface normal in the x and y directions can be calculated, and the three-dimensional surface can be obtained by integration using the Frankot-Chellappa algorithm.

[0082] (7)

[0083] (8)

[0084] in and It refers to the Fourier transform and the inverse Fourier transform. and These represent the number of columns and rows of the image, respectively.

[0085] like Figure 1As shown, this invention proposes a cross-modal fusion-driven spectral polarization three-dimensional 5D imaging method, comprising the following steps:

[0086] A. Multiplexed Object Surface Encoding 4D Imaging: Unlike traditional Hadamard transform spectral systems, this invention proposes multiplexed object surface encoding imaging, which transfers the encoding position in the optical system from the image plane to the object plane. The encoding template is projected by a projector to encode the object, which is then dispersed by a prism, and finally the object surface encoding dispersion information is received by a grayscale camera.

[0087] Figures 4-5 The information acquisition and reconstruction process of the proposed method is demonstrated. For example... Figure 4 As shown, the encoding template is obtained by copying each row of the Hadamard matrix row by row. Here, 1 represents the projector, 2 the grayscale camera, 32 the dispersive prism, 41 the fourth lens, 42 the fifth lens, and 43 the sixth lens. Since the information in each row of the encoding template is the same, the analysis of the single-row, single-template encoding acquisition process is as follows: Figure 5 As shown, the encoding template is modulated by the depth information of the object, and then the information is dispersively modulated in the spectral dimension by a prism. Finally, the encoded dispersive information is acquired by a grayscale camera.

[0088] The spectral 3D 4D data is represented cubically as follows: , These represent the spatial width, height, depth, and number of bands of 4D information, respectively. Representing spatial coordinates and spectral channels In a single encoding process, the encoding template is identical for each row of the data cube. A detailed modeling and analysis of the encoding and decoding for single-row information is then performed. The encoding and dispersion process for single-row multi-template encoding is described in detail below.

[0089] For a certain row of target objects Using the information from each row of the Hadamard matrix Encoding it, without considering the modulation of the encoded information by 3D information, can be represented analogously to traditional image plane coding as follows:

[0090] (9)

[0091] Because there is an angle between the projected light and the collected light, the encoded information is modulated by the 3D information. Figure 5It is known that the encoded data space before dispersion contains three types of information: the uncoded shadow acquisition area added due to 3D occlusion, the encoded but unacquired area, and the coded acquisition area. This modulation of the encoded information by depth is called the Modulated by Three-Dimensional Information (MBTI) operation. This allows us to obtain the data after modulation by spectral and 3D information. for:

[0092] (10)

[0093] in For unencoded shadow acquisition areas, the grayscale increase due to depth is zero. To encode information that was not captured due to depth in uncaptured areas. This is a three-dimensional information modulation operation function. Encoded information. It will pass through a dispersive device, undergo linear dispersion, and obtain :

[0094] (11)

[0095] In the formula, Here are the dispersion operation functions;

[0096] After encoding, modulation, and dispersion are completed, the information is compressed and acquired by a two-dimensional detection device to obtain the acquired information. :

[0097] (12)

[0098] Because differential measurement is used, most of the noise in the coded spectral imaging process can be eliminated, and the measured data can be obtained by matrix inversion. The complete encoded data can be obtained after n measurements of a single line of data. The dispersed spectral three-dimensional data can be obtained by multiplying the acquired encoded information by the corresponding Hadamard inverse matrix.

[0099] (13)

[0100] In the formula, This refers to all the encoded data obtained after n measurements of a single line of data. This is the Hadamard inverse matrix corresponding to the encoded data;

[0101] After undergoing an inverse dispersion process, the single-line spectral three-dimensional data can be reconstructed. :

[0102] (14)

[0103] In the formula, It is the inverse dispersion operation function;

[0104] The above explains the encoding and decoding process for single-line data. Because the encoding template is the same for each line, the spectral distribution in the x-direction and the modulation of the template in three dimensions are identical. Therefore, only an x-dimensional dimension needs to be added during data processing to handle four-dimensional data. Encoding and decoding processes are performed to obtain reconstructed 4D information.

[0105] This method can obtain information from both spectral and three-dimensional modes simultaneously through encoding, and has broad prospects for development and application. However, since the three-dimensional information is an adjunct to spectral imaging, the reconstructed three-dimensional depth is coarse, resulting in problems such as low accuracy and incorrect scale.

[0106] B. Polarization 3D Imaging: As mentioned above, spectral information and roughness depth information of the target can be obtained through object surface encoding. Although 3D reconstruction via polarization can capture 3D details, it suffers from inaccurate depth information due to azimuth ambiguity. To address this, this invention proposes a Multi-Scale Gradient Consistency Sign Correction Model to solve the azimuth ambiguity problem in polarization imaging and achieve high-precision 3D reconstruction.

[0107] Figure 6 The system demonstrates a high-precision 3D reconstruction process. The input information includes rough 3D information obtained from object surface encoding, such as... Figure 6 (a) and polarization diagrams at different polarization angles are shown in Figure 1. Figure 6 In (b), the azimuth angle AOP is calculated using the Stokes vector, as shown in (b). Figure 6 (c) and zenith angle DOP, as shown Figure 6 (e) in the image, but due to an incorrect azimuth angle, only the image reconstructed is shown below. Figure 6 The depth information in (f) is biased. To address the depth bias problem caused by azimuth ambiguity, the coarse depth information obtained from coded imaging is used as prior information, such as... Figure 6 In (a), the proposed multi-scale gradient consistency sign correction model is used to correct the azimuth angle, resulting in the correct azimuth angle as shown in (a). Figure 6 (d) in the middle, thus reconstructing as Figure 6 The correct depth information for (g) in the middle.

[0108] To effectively extract depth gradient details from images at different spatial scales, this invention proposes an adjustable-scale two-dimensional Sobel gradient operator construction method. Based on the traditional Sobel operator, this method introduces a scale-aware weighting and normalization strategy, making it more robust and efficient in multi-scale depth map gradient extraction.

[0109] For a given odd scale Construct a pair Two-dimensional convolution kernels of different sizes and These are used to extract depth maps in and Gradient response in direction. Let the position of the center row (column) be:

[0110] (15)

[0111] The one-dimensional directional difference weight vector is defined as:

[0112] (16)

[0113] For each row Introduce linear weighting:

[0114] (17)

[0115] The unnormalized gradient operator is constructed as follows:

[0116] (18)

[0117] To ensure the consistency and stability of responses at different scales, a global normalization operation is introduced, and the final Sobel kernel is defined as follows:

[0118] (19)

[0119] For coarse-coded depth data, the gradient operator is used to calculate its small-scale gradient and large-scale gradient respectively:

[0120] (20)

[0121] in To coarsely encode depth data, , , and They are , Directional small scale and , Large-scale Sobel convolution kernels in the direction of the direction This represents the corresponding gradient.

[0122] Blending scales obtained from convolution kernels of different sizes:

[0123] (twenty one)

[0124] in, and For the fusion and Gradient of direction, For the fusion weight Use edge detection results for spatial adaptive adjustment:

[0125] (twenty two)

[0126] In the formula, For the fusion of fundamental constants, This is a binary result of the Canny edge map. This approach tends to preserve smaller scales (details) at the edges and larger scales (contour structure) in smooth areas.

[0127] To address the impact of extreme high-frequency errors on gradient determination, a dynamic threshold is defined:

[0128] (twenty three)

[0129] The threshold is the basic constant. For large-scale convolution kernel size, This refers to the small-scale convolution kernel size.

[0130] (twenty four)

[0131] All gradients exceeding the threshold are set to 0 to eliminate the impact of extreme high-frequency errors on direction determination.

[0132] Combine the gradients into a two-dimensional vector field:

[0133] (25)

[0134] in , The zenith angle is obtained through the degree of polarization. This is the uncorrected deflection angle.

[0135] Calculate the dot product of the two-dimensional vector field pixel by pixel:

[0136] (26)

[0137] In the formula, and For the fusion and Gradient of direction;

[0138] When dot product When the angle between the two directions is less than 90°, it indicates that the gradient directions of the two depth maps are consistent. When the dot product... When the gradient directions are opposite, the dot product relationship between the reference gradient and the polarization gradient is used to determine whether their directions are consistent. A correction coefficient matrix consisting of ±1 is generated to flip the polarization gradients with opposite directions, thereby improving the accuracy and consistency of the gradient direction in subsequent polarization depth reconstruction.

[0139] After obtaining high-precision three-dimensional information, the three-dimensional information and the spectral information obtained through encoding are fused to obtain high-quality spectral three-dimensional information.

[0140] III. Experiments and Results: To verify the feasibility of the method of the present invention, the following experiments were conducted: Figure 7 The experimental system shown is used to verify the performance of the proposed method. Figure 7 In the diagram: 1 is the projector, 2 is the grayscale camera, 31 is the polarizer, 32 is the dispersive prism, 41 is the fourth lens, 42 is the fifth lens, and 43 is the sixth lens, forming a 4F lens. Unpolarized LED light (380-780nm) is used for illumination. During object surface encoding imaging, a 512*512 resolution Hadamard template is loaded into the DMD and projected onto the target. Light reflected from the target passes through the 4F lens and is dispersed parallel to the double Amish dispersive prism (410-660nm), and is finally received by the grayscale camera. For polarization information acquisition, simply replace the dispersive prism with a polarizer (Hengyang Optics GSP-50, extinction ratio >1000:1), load a 512*512 full-white image into the DMD to ensure it matches the encoding acquisition illumination area, and convert the polarization angle to 0°, 45°, 90°, and 135° to acquire polarized images. The object's refractive index range is [1.3 1.7].

[0141] A 4D image of a complex-shaped, colorful plastic doll was created, and the final image result is as follows. Figures 8-11 As shown. Among them. Figure 8 It is a 2-D RGB image. Figure 9 The 4-D RGB image synthesized using three single-band 4-D information was reconstructed with colors consistent with 2-D RGB, indicating that the acquired spectral information was accurate. Figure 10 The system used a dual-peak LED light source to illuminate the four points on the plastic doll, and the obvious differences in the spectral curves of different color regions confirmed the effectiveness of the reconstructed spectral information. Figure 11The reconstructed image accurately reconstructs spectral information and achieves high depth precision for 4-D information across different spectral bands. A total of 80 4-D data points in the 410-650nm band are reconstructed, with a spectral resolution of 3nm. The method used in this invention is based on a mathematically interpretable reconstruction approach; using a prism with stronger dispersive power would further improve the spectral resolution. The reconstructed image shows that data loss occurs in some areas due to projection occlusion, a common problem in object surface encoding.

[0142] To demonstrate the 3D reconstruction efficiency of the method of the present invention, it was compared with some classic polarization 3D reconstruction methods, including diffuse reflection polarization reconstruction method DRLPR and monocular polarization 3D reconstruction method (SFPGI) which corrects blur by separating specular reflection and diffuse reflection and converting diffuse reflection grayscale into a depth map as a priori condition.

[0143] To quantitatively evaluate the 3D reconstruction accuracy of the proposed method, a standard sphere (50.798 mm in diameter) was selected as the reconstruction target. The reconstruction results are as follows: Figure 12 As shown, Figure 12 (a) in the image shows the captured grayscale image. Figure 12 (b) shows the roughness depth map obtained by surface coding, which clearly shows that there are problems such as incorrect depth ratio and uneven surface. Figure 12 (c) in the figure shows the reconstructed depth map of the DRLPR method, where azimuth ambiguity has a significant impact on the 3D reconstruction results. Figure 12 (d) in the figure shows the depth information reconstructed using the SFPGI method. It can be seen that the grayscale information cannot be converted into effective depth information because of the black spots on the spherical surface. The reconstructed surface has unevenness and low accuracy. Figure 12 (e) in the figure represents the depth map reconstructed by the method proposed in this invention. It can be seen that the method of this invention achieves excellent results, mainly due to the multi-scale gradient consistency sign correction model proposed in this invention. This model fully leverages the advantages of object plane encoding and polarization imaging to achieve high-quality 3D reconstruction. To quantify the 3D accuracy, the reconstructed depth map is fitted to a standard sphere, and the absolute error is calculated by comparing the reconstructed sphere and the fitted standard sphere. The calculated fitting radius of the SFPGI method is 142.51 pixels, and the mean absolute error (MAE) is 1.101. The fitting radius of the method of this invention is 141.21 pixels, and the MAE is 0.504. With a standard sphere diameter of 50.798 mm, the 3D depth accuracy of the method of this invention can be calculated to reach 0.09 mm.

[0144] To verify the effectiveness of 3D reconstruction for different materials, irregular objects made of three different materials were reconstructed, and the reconstruction details were analyzed. The reconstruction results are as follows: Figure 13As shown. The first object is a colorful PVC plastic doll with a smooth surface. The reconstruction result from the DRLPR method shows that the reconstructed shape is distorted because the azimuth blur is not corrected. Compared to DRLPR, SFPGI corrects the azimuth error to some extent, and surface details are also reconstructed; however, the smoothness of the reconstructed smooth surface is not as good as the method proposed in this invention. The second object is a hand-painted colored plaster model with slight unevenness on its surface due to uneven application of paint. Although the surface reconstructed by the DRLPR method is relatively smooth, deformation problems still exist. Grayscale cameras have different quantum efficiencies for different colors, and the grayscale values ​​of different colors in grayscale images differ significantly. This prevents grayscale from being converted into effective depth information as a priori to correct the azimuth error, resulting in a large height difference at the boundaries between different colors in the reconstructed 3D image. Compared with other methods, the method of this invention can generate results that are closest to the real model. The third object is a flat white paper plate with a uniform uneven texture on its surface. While the DRLPR method reconstructs a smooth surface, it loses details of the uneven texture. The SFPGI method can reconstruct uneven texture details, but because the grayscale information of a flat surface has two maximum grayscale values ​​at concave and convex points, the reconstructed 3D information renders concave stripes as convex stripes. Compared with other methods, the method proposed in this invention reconstructs uneven texture information most accurately.

[0145] This invention proposes a cross-modal fusion-driven spectral polarization 3D 5D imaging method. First, a spectral-3D collaborative imaging method based on object surface encoding is constructed to acquire the target's hyperspectral information and roughness depth estimation. Then, a multi-scale gradient consistency sign correction model is introduced, using the roughness depth information acquired through encoded imaging as prior knowledge to guide the polarization 3D reconstruction process, thereby solving the azimuth ambiguity problem in polarization imaging and achieving high-precision 3D reconstruction. Experimental results show that the spectral resolution of this invention reaches the nanometer level, and the depth accuracy reaches the micrometer level, exhibiting good stability and adaptability. The system of this invention adopts a single-camera synchronous acquisition scheme, avoiding the common image registration problem in multimodal imaging from a hardware structure perspective, improving the system's compactness and practicality. The system has a simple structure, low cost, high accuracy, and multiple imaging modalities, and can be applied on a large scale in industrial inspection.

Claims

1. A cross-modal fusion-driven spectral polarization three-dimensional 5D imaging system, characterized in that, The spectral polarization 3D 5D imaging system is a single-image sensor imaging system, which includes a projector, a 4F lens, a functional lens, and a grayscale camera. Therefore, both the projector and the grayscale camera are directly facing the target. The functional lens and the 4F lens are sequentially arranged between the grayscale camera and the target. The functional lens is a double Amicis dispersion prism or a polarizer. The spectral information and coarse-coded depth information of the target are obtained through the double Amicis dispersion prism, and polarized images with different polarization angles are obtained through the polarizer. By combining the spectral information, coarse-coded depth information, and polarized images of the target, high-precision 3D reconstruction is achieved.

2. The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging system as described in claim 1, characterized in that, The projector includes a light source, a hadamard template, and a spatial light modulator (DMD). The light source is a non-polarized LED light source. The hadamard template is loaded in the spatial light modulator (DMD), and the spatial light modulator (DMD) is disposed between the light source and the target.

3. The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging system as described in claim 2, characterized in that, The wavelength of the unpolarized LED light source is 380-780nm, and the resolution of the Hadamard template is 512*512.

4. The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging system as described in claim 1, characterized in that, The wavelength of the dual Amish dispersion prism is 410-660nm.

5. A cross-modal fusion-driven spectral polarization three-dimensional 5D imaging method, characterized in that, The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging system as described in any one of claims 1-4 is employed. The spectral polarization three-dimensional 5D imaging method includes the following steps: 1) The functional mirror is set as a double Amish dispersion prism, and object plane coding imaging is used to realize object plane coding 4D imaging to obtain four-dimensional data of the measured target. The four-dimensional data includes spectral information and coarse coding depth information; the functional mirror is set as a polarizer, and polarization images with different polarization angles are obtained through polarization three-dimensional imaging. 2) Using the coarse encoded depth information and polarization map obtained in step 1), the azimuth and zenith angles are calculated using the Stokes vector method. The azimuth angle is then corrected using a multi-scale gradient consistency sign correction model to obtain the correct azimuth angle, thereby obtaining high-precision three-dimensional information. 3) Fuse the high-precision three-dimensional information obtained in step 2) with the spectral information obtained in step 1) to obtain high-quality spectral three-dimensional information.

6. The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging method as described in claim 5, characterized in that, Step 1) Four-dimensional data Through the reconstructed single-row spectral three-dimensional data Add one The dimensions are obtained, among which the reconstructed single-row spectral three-dimensional data are obtained. The following formula is used to calculate: , In the formula, It is a single-row spectral three-dimensional data. This represents the three-dimensional spectral data after dispersion. This represents the information acquired through compressed acquisition by a two-dimensional detection device. This refers to all the encoded data obtained after n measurements of a single line of data. This is the Hadamard inverse matrix corresponding to the encoded data. for After dispersion by a dispersive device, it is obtained. for Obtained after modulation by spectral and three-dimensional information. For encoding information, For unencoded shadow acquisition areas, the grayscale increase due to depth is zero. To encode information that was not captured due to depth in uncaptured areas. and These represent the spatial width, height, depth, and number of bands of 4D information, respectively. Represents spatial coordinates, Indicates spectral channels, Indicates the target to be tested in a certain row. This represents the information in the i-th row of the Hadamard matrix. It is the inverse dispersion operation function. Here are the dispersion operation functions. This is a function for modulating three-dimensional information.

7. The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging method as described in claim 5, characterized in that, In step 1), the different polarization angles are 0°, 45°, 90° and 135°.

8. The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging method as described in claim 5, characterized in that, Step 2) Central azimuth and zenith By solving the gradient and We obtain, where, gradient and The following formula is used to calculate: , when hour, , when hour, , , In the formula, The zenith angle is obtained through the degree of polarization. This is the uncorrected azimuth angle. and It is a two-dimensional vector field composed of gradient combinations. and For the fusion and Gradient of direction, For the fusion of fundamental constants, This is the binary result of the Canny edge map. The threshold is the basic constant. For large-scale convolution kernel size, For small-scale convolution kernel size, and for direction and Small-scale gradient in direction and for direction and Large-scale gradient in direction, This is a dynamic threshold.

9. The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging method as described in claim 8, characterized in that, , , and The following formula is used to calculate: , In the formula, To coarsely encode depth information, , , and They are Small-scale direction Small-scale direction Large-scale directions and Large-scale Sobel convolution kernels in the direction of the direction It is a one-dimensional directional difference weight vector. For linearly weighted parameters, The position of the center row or center column. For odd-numbered scales, For the number of rows, For a single-row unnormalized gradient operator, and for direction and Direction Sobel core, Depend on Small-scale Sobel convolution kernels and Large-scale Sobel convolution kernels composition, Depend on Small-scale Sobel convolution kernels and Large-scale Sobel convolution kernels composition.

10. The cross-modal fusion-driven spectral polarization three-dimensional 5D imaging method as described in claim 5, characterized in that, Step 2) of the multi-scale gradient consistency sign correction model specifically involves: using the dot product to determine whether the gradient directions of the two depth maps are consistent, generating a correction coefficient matrix consisting of ±1 values ​​to flip polarization gradients with opposite directions. The dot product of the two-dimensional vector field is calculated pixel-by-pixel using the following formula. : , In the formula, and For the fusion and Gradient of direction, The zenith angle is obtained through the degree of polarization. This is the uncorrected azimuth angle; When dot product When the angle between the two directions is less than 90°, the gradient directions of the two depth maps are consistent, and the dot product... At that time, the gradient directions of the two depth maps are opposite.