A computer vision-based stereoscopic image color optimization method
Patent Information
- Application Number
- CN202610926751.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-08-28
AI Technical Summary
单纯基于全局色彩分布的方法忽略了立体图像中隐含的复杂双目几何视差结构(如高视差风险区域、局部深度突变等)以及深层的空间感知关联,导致在执行大幅色彩拉伸时发生高视差区域的空间失真,从而加剧了跨区域、高反差立体图像观看时的视觉疲劳与眩晕感
[0052](1) This invention achieves accurate quantification of global geometric disparity load and generation of spatial adaptive constraints in stereo images by constructing a statistical fitting and adaptive spatial weight mapping mechanism for absolute disparity values. The mean scalar and global standard deviation scalar are calculated based on the absolute disparity matrix of the stereo image, and a weighted summation is performed on the discrete variance values to generate a dizziness threshold. The global proportion of high-risk pixels is statistically determined based on the dizziness threshold, and a logarithmic domain nonlinear compression mapping is performed in conjunction with the global fluctuation benchmark value to generate a global geometric disparity load index. A local mean matrix is calculated using a sliding window and divided element-wise with the global geometric disparity load index. Element-wise multiplication and normalization operations are then performed in conjunction with the local disparity gradient matrix in the same domain to generate an adaptive spatial weight mask. This mechanism converts discrete disparity pixel values into a continuous spatial load distribution field, accurately extracting the spatial weight constraint benchmark for high disparity risk areas, and providing a dynamically adaptive spatial modulation prior for subsequent color enhancement and pixel shifting.
Smart Images

Figure CN122656941A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing and color science technology, and in particular to a method for optimizing the color of stereoscopic images based on computer vision. Background Technology
[0002] With the increasing popularity of stereoscopic display devices, stereoscopic image content is experiencing explosive growth. Classic stereoscopic image rendering pipelines face severe challenges in visual comfort regarding color enhancement and parallax control for massive high-dimensional pixel arrays. Existing stereoscopic image color optimization algorithms, such as global histogram equalization and unified gamma correction, while improving overall image brightness and contrast by utilizing full-image statistical features, primarily rely on pixel-level mapping based on the mean and variance of the global color distribution. Methods based solely on global color distribution ignore the complex binocular geometric parallax structure implicit in stereoscopic images (such as high parallax risk areas and local depth abrupt changes) and deep spatial perception correlations. This leads to spatial distortion in high parallax areas when performing significant color stretching, thus exacerbating visual fatigue and dizziness when viewing cross-regional, high-contrast stereoscopic images. Furthermore, classic color optimization methods often struggle to fully utilize the parallax gradient distribution features in stereoscopic images to constrain the magnitude of color enhancement and the scale of spatial offset. This often results in severe color crosstalk and retinal conflict anomalies when processing scenes with dense depth layers, significantly increasing the neural computational overhead of visual fusion in the brain.
[0003] Therefore, how to provide a computer vision-based method for optimizing the color of stereo images is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] This invention proposes a computer vision-based method for color optimization of stereo images. Through a spatial constraint mapping mechanism based on a global geometric disparity load index and an adaptive spatial weight mask, local suppression modulation and bilinear interpolation inverse remapping are performed on the color enhancement amplitude and independent red and blue color channels of the stereo image. High-dimensional spatial features, including a depth overload coefficient matrix and four-neighbor interpolation weights, are extracted. A Huygens-based inverse wavefront filtering mechanism is introduced based on an improved GCNet model to generate sub-pixel-level offset matrices for the red and blue channels. The adaptive spatial weight mask is embedded as a spatial constraint condition into the color suppression coefficient matrix and the mask constraint offset matrix. By calculating the spatial ratio of local mean to local variance in the local brightness statistical feature matrix, nonlinear gamma curve fitting mapping and residual addition fusion operations are used to iteratively correct the pixel offsets and tone mapping coefficients until the final color-optimized stereo image is output. This mechanism effectively eliminates low-frequency smoothing deviations in the color enhancement process of high-parallax-risk areas by establishing a closed-loop feedback path from "parallax load measurement" to "color space offset." It ensures that the generated spatially constrained tone mapping coefficients dynamically maintain the high-frequency edge abruptness characteristics of the original stereoscopic image, achieving the technical effect of reducing visual dizziness and fusion conflicts while maintaining high-fidelity color reproduction. This invention overcomes the limitations of traditional methods that ignore geometric parallax constraints, global color distortion, and red-blue channel crosstalk, providing an efficient solution for stereoscopic image display and visual comfort optimization.
[0005] A stereo image color optimization method based on computer vision according to an embodiment of the present invention specifically includes:
[0006] S1. Acquire stereo images and calculate disparity maps. Adaptively fit the dizziness threshold using the sum of the mean and standard deviation of the absolute values of disparity. Based on the dizziness threshold, statistically determine the global proportion of high-risk pixels and generate a global geometric disparity load index.
[0007] S2. Based on the dizziness threshold, high parallax risk areas are extracted. The local load index matrix is calculated using a sliding window and compared with the global geometric parallax load index. The adaptive spatial weight mask is generated by combining the local parallax gradient matrix in the same domain.
[0008] S3. Call the adaptive spatial weight mask to spatially suppress the color enhancement amplitude of the stereo image and output the intermediate stereo image;
[0009] S4. Separate the intermediate stereo image into independent color channels, and generate a depth overload coefficient matrix based on the absolute value of disparity in high disparity risk areas and the global geometric disparity load index.
[0010] S5. By improving the GCNet model, spatial orientation priors are extracted based on high parallax risk areas, guiding the mapping of the depth overload coefficient matrix to incident wavefront features. A reverse wavefront filtering mechanism based on Huygens' principle is introduced to cancel low frequencies to generate an anti-smoothing interference coefficient matrix. The element-wise multiplication of the spatial modulation field is combined to output the sub-pixel offset matrix of the red and blue channels.
[0011] S6. Under the constraint of adaptive spatial weight mask, perform bilinear interpolation inverse remapping operation on the independent red and blue color channels according to the sub-pixel offset matrix of red and blue channels to perform pixel offset, and stitch the unoffset green channel to output a synthesized stereo image.
[0012] S7. Use the adaptive spatial weight mask as a constraint to perform local tone mapping fine-tuning on the synthesized stereo image and output the final color-optimized stereo image.
[0013] Optionally, S1 specifically includes:
[0014] S11. Calculate the initial disparity map based on the stereo image by performing stereo matching, and generate the absolute disparity matrix by performing an absolute value operation on the initial disparity map; calculate the mean scalar of the entire image spatial dimension based on the absolute disparity matrix, calculate the corresponding global standard deviation scalar based on the mean scalar, and perform scalar addition on the mean scalar and the global standard deviation scalar to generate the global fluctuation benchmark value.
[0015] S12. Based on the absolute disparity matrix, calculate the sum of squares of the differences between the absolute disparity of each pixel and the mean scalar; divide the sum of squares of the differences by the total number of pixels to generate the discrete variance value; perform a weighted summation of the discrete variance value and the global standard deviation scalar to generate the dizziness threshold.
[0016] S13. Perform a pixel-by-pixel comparison between the dizziness threshold and the absolute disparity matrix, and output a high-risk pixel binarization mask based on the comparison result; perform full-image summation on the high-risk pixel binarization mask to generate a high-risk pixel count value; divide the high-risk pixel count value by the total number of pixels in the stereo image to generate a global proportion coefficient.
[0017] S14. Multiply the global proportion coefficient with the global fluctuation benchmark value to generate the initial load index; perform logarithmic domain nonlinear compression mapping on the initial load index to generate the global geometric disparity load index.
[0018] Optionally, S2 specifically includes:
[0019] S21. Perform a pixel-by-pixel comparison between the adaptive dizziness threshold and the absolute disparity matrix, and output a binarized mask for high disparity risk areas based on the comparison results; use a sliding window to perform a sliding traversal on the absolute disparity matrix, calculate the local mean scalar within each window, and concatenate the local mean scalars according to their spatial positions to generate a local mean matrix; perform an element-by-element division operation between the local mean matrix and the global geometric disparity load index to generate a local load ratio matrix;
[0020] S22. Based on the absolute disparity matrix, perform adjacent pixel difference operations along the horizontal and vertical directions respectively. Perform summation and square root operations on the horizontal and vertical difference results to generate the local disparity gradient matrix in the same domain.
[0021] S23. Perform element-wise multiplication of the local load ratio matrix and the local disparity gradient matrix in the same domain to generate the gradient modulation load matrix; perform maximum and minimum value normalization operation on the gradient modulation load matrix to generate the global spatial weight basis matrix; perform element-wise dot product of the global spatial weight basis matrix and the binarized mask of the high disparity risk region to generate the adaptive spatial weight mask.
[0022] Optionally, S3 specifically includes:
[0023] S31. Convert the stereo image from the red-green-blue color space to the luminance-chrominance color space, and separate the luminance channel matrix and the chrominance channel matrix; calculate the global luminance mean scalar based on the luminance channel matrix;
[0024] S32. Calculate the spatial discrete variance scalar of the luminance channel matrix based on the global luminance mean scalar; perform an addition operation on the global luminance mean scalar and the luminance discrete variance scalar to generate an adaptive color reference value;
[0025] S33. Divide the chroma channel matrix by the adaptive color reference value and perform an element-wise division operation to generate the initial color enhancement matrix;
[0026] S34. Perform a range reversal operation on the adaptive spatial weight mask input to the linear mapping function to generate a spatial suppression coefficient matrix; perform an element-wise multiplication operation between the spatial suppression coefficient matrix and the initial color enhancement matrix to generate a suppressed chroma channel matrix.
[0027] S35. Perform a channel dimension splicing operation on the suppressed chroma channel matrix and the luminance channel matrix, and convert the splicing result from the luminance chroma color space to the red-green-blue color space to generate an intermediate stereo image.
[0028] Optionally, S4 specifically includes:
[0029] S41. Perform channel dimension separation operation on the intermediate stereo image and output three independent color channel matrices: red, green and blue.
[0030] S42. Perform element-wise multiplication of the binarized mask of the high disparity risk area with the disparity absolute value matrix to generate the disparity absolute value matrix of the risk area.
[0031] S43. Perform element-wise division between the absolute value matrix of disparity in the risk area and the global geometric disparity load index to generate the depth overload coefficient matrix.
[0032] Optionally, the improved GCNet model includes a direction prior construction layer, a global context pooling layer, a spatial similarity weighting layer, a Huygens meta-surface interference layer, and a sub-pixel offset fusion layer:
[0033] The orientation prior construction layer is used to extract the inverted disparity gradient of high disparity risk regions as a spatial orientation prior.
[0034] The global context pooling layer is used to extract the depth risk distribution state at the whole graph level by performing a global spatial pooling operation on the depth overload coefficient matrix, and generate a global context feature vector.
[0035] The spatial similarity weight layer is used to extract local spatial features of the depth overload coefficient matrix, calculate the spatial correlation response between the local spatial features and the global context feature vector, and generate a spatial modulation field matrix through probabilistic normalization mapping.
[0036] The Huygens element surface interference layer is used to introduce a reverse wavefront filtering mechanism based on Huygens' principle, specifically including:
[0037] The deep overload coefficient matrix is guided to be mapped to the incident wavefront features of the red and blue channels. The spatial direction prior is transformed into the sampling offset of deformable convolution through channel dimension expansion mapping. On the incident wavefront features of the red and blue channels, the multi-scale ring convolution kernel is guided to deform according to the sampling offset to generate a wavelet excitation matrix along the anisotropic direction indicated by the normal. The wavelet excitation matrix is subjected to nonlinear mapping for phase encoding to generate a phase-encoded feature matrix. The incident wavefront features of the red and blue channels are subjected to phase reversal operation to generate the inverted wavefront features of the red and blue channels. The phase-encoded feature matrix and the inverted wavefront features of the red and blue channels are added element by element. The low frequency is filtered out by difference cancellation and the high frequency abrupt change is retained to generate the anti-smoothing interference coefficient matrix of the red and blue channels.
[0038] The subpixel offset fusion layer is used to perform element-wise multiplication of the red and blue channel anti-smoothing interference coefficient matrix and the spatial modulation field matrix to generate and output the red and blue channel subpixel-level offset matrix.
[0039] Optionally, S6 specifically includes:
[0040] S61. Extract the spatial dimension coefficient matrix of the adaptive spatial weight mask, perform element-wise multiplication of the spatial dimension coefficient matrix with the sub-pixel offset matrix of the red and blue channels, and output the mask constraint offset matrix.
[0041] S62. Extract the grid coordinates of the red and blue independent color channels based on the mask constraint offset matrix, and superimpose the mask constraint offset matrix on the grid coordinates to generate the target sampling coordinate matrix;
[0042] S63. Perform bilinear interpolation inverse remapping on the independent red and blue color channels based on the target sampling coordinate matrix, and calculate the four-neighbor interpolation weights mapped to the discrete pixel coordinate space.
[0043] S64. Based on the four-neighbor interpolation weights, perform a weighted summation on the pixel values corresponding to the discrete pixel coordinates of the independent red and blue color channels, and output the offset red and blue channel matrix.
[0044] S65. Extract the green channel matrix after separating the intermediate stereo image, and perform a stitching operation on the offset red and blue channel matrices and the green channel matrix along the channel dimension to output the synthesized stereo image.
[0045] Optionally, S7 specifically includes:
[0046] S71. Extract the brightness channel matrix of the synthesized stereo image, perform block pooling operation on the brightness channel matrix based on the adaptive spatial weight mask, calculate the local mean and local variance of the block pooling result, and output the local brightness statistical feature matrix.
[0047] S72. Calculate the spatial ratio of the local mean to the local variance in the local brightness statistical feature matrix, and calculate the local contrast enhancement factor matrix based on the spatial ratio.
[0048] S73. Perform an element-wise multiplication operation between the local contrast enhancement factor matrix and the adaptive spatial weight mask to generate a spatially constrained tone mapping coefficient matrix.
[0049] S74. Based on the spatially constrained tone mapping coefficient matrix, perform a nonlinear gamma curve fitting mapping operation on the red, green and blue three-channel matrix of the synthesized stereo image and output a local tone fine-tuning matrix.
[0050] S75. Perform residual addition and fusion operation on the local tone fine-tuning matrix and the synthesized stereo image along the channel dimension to output the final color-optimized stereo image.
[0051] The beneficial effects of this invention are:
[0052] (1) This invention achieves accurate quantification of global geometric disparity load and generation of spatial adaptive constraints in stereo images by constructing a statistical fitting and adaptive spatial weight mapping mechanism for absolute disparity values. The mean scalar and global standard deviation scalar are calculated based on the absolute disparity matrix of the stereo image, and a weighted summation is performed on the discrete variance values to generate a dizziness threshold. The global proportion of high-risk pixels is statistically determined based on the dizziness threshold, and a logarithmic domain nonlinear compression mapping is performed in conjunction with the global fluctuation benchmark value to generate a global geometric disparity load index. A local mean matrix is calculated using a sliding window and divided element-wise with the global geometric disparity load index. Element-wise multiplication and normalization operations are then performed in conjunction with the local disparity gradient matrix in the same domain to generate an adaptive spatial weight mask. This mechanism converts discrete disparity pixel values into a continuous spatial load distribution field, accurately extracting the spatial weight constraint benchmark for high disparity risk areas, and providing a dynamically adaptive spatial modulation prior for subsequent color enhancement and pixel shifting.
[0053] (2) This invention establishes a spatial orientation prior-guided wavefront interferometry and sub-pixel-level offset fusion system by employing an improved GCNet model and a Huygens-based inverse wavefront filtering mechanism. The orientation prior construction layer extracts the inverse disparity gradient of high disparity risk regions as the spatial orientation prior. The Huygens metasurface interferometry layer guides the depth overload coefficient matrix to be mapped to the incident wavefront features of the red and blue channels. The spatial orientation prior is transformed into a deformable convolutional sampling offset through channel dimension expansion mapping. Based on the sampling offset, the multi-scale ring convolution kernel is deformed to generate a wavelet excitation matrix along the anisotropic direction indicated by the normal. The wavelet excitation matrix is phase-encoded by performing a nonlinear mapping, and the incident wavefront features of the red and blue channels are phase-reversed. The phase-encoded feature matrix and the inverted wavefront features of the red and blue channels are added element-wise. The sub-pixel offset fusion layer performs element-wise multiplication of the generated red and blue channel anti-smoothing interferometry coefficient matrix with the spatial modulation field matrix, and outputs the sub-pixel-level offset matrix of the red and blue channels. This system achieves high-frequency edge feature excitation from the depth overload coefficient space to the wavefront interference physical space through spatial direction a priori deformation guidance, wavelet phase encoding and reverse wavefront difference cancellation, ensuring that the output red and blue channel sub-pixel offset matrix has extremely high spatial reshaping capability for the depth abrupt edge of the stereo image. Attached Figure Description
[0054] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0055] Figure 1 This is an overall flowchart of a computer vision-based stereo image color optimization method proposed in this invention;
[0056] Figure 2This is a flowchart illustrating the working principle of the improved GCNet model, which is a computer vision-based method for optimizing stereo image color. Detailed Implementation
[0057] The invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0058] refer to Figure 1 and Figure 2 A computer vision-based method for optimizing the color of stereo images, specifically including:
[0059] S1. Acquire stereo images and calculate disparity maps. Adaptively fit the dizziness threshold using the sum of the mean and standard deviation of the absolute values of disparity. Based on the dizziness threshold, statistically determine the global proportion of high-risk pixels and generate a global geometric disparity load index.
[0060] S2. Based on the dizziness threshold, high parallax risk areas are extracted. The local load index matrix is calculated using a sliding window and compared with the global geometric parallax load index. The adaptive spatial weight mask is generated by combining the local parallax gradient matrix in the same domain.
[0061] S3. Call the adaptive spatial weight mask to spatially suppress the color enhancement amplitude of the stereo image and output the intermediate stereo image;
[0062] S4. Separate the intermediate stereo image into independent color channels, and generate a depth overload coefficient matrix based on the absolute value of disparity in high disparity risk areas and the global geometric disparity load index.
[0063] S5. By improving the GCNet model, spatial orientation priors are extracted based on high parallax risk areas, guiding the mapping of the depth overload coefficient matrix to incident wavefront features. A reverse wavefront filtering mechanism based on Huygens' principle is introduced to cancel low frequencies to generate an anti-smoothing interference coefficient matrix. The element-wise multiplication of the spatial modulation field is combined to output the sub-pixel offset matrix of the red and blue channels.
[0064] S6. Under the constraint of adaptive spatial weight mask, perform bilinear interpolation inverse remapping operation on the independent red and blue color channels according to the sub-pixel offset matrix of red and blue channels to perform pixel offset, and stitch the unoffset green channel to output a synthesized stereo image.
[0065] S7. Use the adaptive spatial weight mask as a constraint to perform local tone mapping fine-tuning on the synthesized stereo image and output the final color-optimized stereo image.
[0066] In this embodiment, S1 specifically includes:
[0067] S11. Calculate the initial disparity map based on the stereo image using stereo matching. Perform an absolute value operation on the initial disparity map to generate a disparity absolute value matrix. Calculate the sum of all pixel values in the entire image in the disparity absolute value matrix and divide it by the total number of pixels to generate a mean scalar. Calculate the difference between each pixel value in the disparity absolute value matrix and the mean scalar. Sum the squares of all differences and divide by the total number of pixels to generate a variance scalar. Perform a square root operation on the variance scalar to generate a global standard deviation scalar. Add the mean scalar and the global standard deviation scalar to generate a global fluctuation baseline value.
[0068] S12. Arrange all pixel values in the absolute disparity matrix in ascending order, extract the value at the 75th percentile as the upper quartile value, and extract the value at the 25th percentile as the lower quartile value. Calculate the difference between the upper and lower quartile values to generate the discrete quartile distance. Set the weight coefficient of the discrete quartile distance to 0.6 and the weight coefficient of the global standard deviation scalar to 0.4. Add the mean scalar to the result of the discrete quartile distance calculation multiplied by the weight coefficient 0.6, and add the result of the global standard deviation scalar multiplied by the weight coefficient 0.4 to generate the dizziness threshold.
[0069] S13. Perform a pixel-by-pixel comparison between each pixel value in the absolute disparity matrix and the dizziness threshold. If the pixel value is greater than the dizziness threshold, mark the pixel position as value 1. If the pixel value is less than or equal to the dizziness threshold, mark the pixel position as value 0. Output a high-risk pixel binarization mask based on the comparison results. Perform full-image summation on all pixels with a value of 1 in the high-risk pixel binarization mask to generate a high-risk pixel count value. Divide the high-risk pixel count value by the total number of pixels in the stereo image to generate a global proportion coefficient.
[0070] S14. Multiply the global proportion coefficient with the global fluctuation benchmark value to generate the initial load index; call the logarithmic function with the natural constant e as the base to perform a logarithmic domain nonlinear compression mapping on the initial load index. Specifically, calculate the result of 1 plus the maximum value between the initial load index and the value 0 plus the value 0.001, and take the logarithm of the result with the natural constant e as the base to generate the global geometric disparity load index.
[0071] In this embodiment, S2 specifically includes:
[0072] S21. Perform a pixel-by-pixel comparison between the adaptive dizziness threshold and the absolute disparity matrix. If the pixel value in the absolute disparity matrix is greater than the adaptive dizziness threshold, mark the pixel position as value 1. If the pixel value is less than or equal to the adaptive dizziness threshold, mark the pixel position as value 0. Output a binary mask for high disparity risk areas based on the comparison results. Set the sliding window size to 3 rows and 3 columns. Use the 3 rows and 3 columns sliding window to perform a sliding traversal on the absolute disparity matrix in the order from left to right and from top to bottom. Calculate the sum of the 9 pixel values contained in each window and divide it by the value 9 to generate a local mean scalar for each window. Concatenate the local mean scalars according to their corresponding spatial positions to generate a local mean matrix with the same size as the absolute disparity matrix. Divide the value of each element in the local mean matrix by the sum of the global geometric disparity load index and the value 0.001 to generate a local load ratio matrix.
[0073] S22. Perform zero-padding on the right and bottom edges of the disparity absolute value matrix. Extract the values of two adjacent pixels in each row of the padded matrix and calculate the difference between the right and left values to generate a horizontal difference result. Extract the values of two adjacent pixels in each column of the padded matrix and calculate the difference between the bottom and top values to generate a vertical difference result. Square the values at the same spatial position in the horizontal and vertical difference results respectively. Add the two squared results and perform a square root operation to generate a local disparity gradient matrix with the same size as the disparity absolute value matrix.
[0074] S23. Perform element-wise multiplication of the local load ratio matrix with the values at the same spatial location in the local disparity gradient matrix of the same domain to generate a gradient modulation load matrix; find the global maximum and global minimum values in the gradient modulation load matrix, subtract the global minimum value from each value in the gradient modulation load matrix and divide by the difference between the global maximum and global minimum values to generate a global spatial weight base matrix; perform element-wise dot product of the global spatial weight base matrix with the values at the same spatial location in the binarized mask of the high disparity risk region to generate an adaptive spatial weight mask.
[0075] The adaptive spatial weight mask generation process proposed in this step is similar to the traditional stereo image salient region extraction mechanism in that it is based on the comparison of local and global features and the spatial gradient perception theory. That is, it captures the depth distribution fluctuation in the spatial neighborhood by performing local window sliding statistics on the disparity matrix, extracts the abrupt response of geometric edges by calculating the difference between adjacent pixels, and adopts a normalized mapping mechanism to transform the feature matrix into a weight mask for subsequent spatial adaptive modulation of the image.
[0076] The difference lies in that this invention breaks through the limitation of traditional salient region extraction, which relies solely on the underlying visual contrast and ignores the distribution of stereo geometric overload. It adds a local load ratio construction step to quantify relative depth pressure and replaces the traditional single threshold segmentation with a gradient modulation mechanism. It performs element-wise division between the local mean matrix and the global geometric disparity load index to generate a local load ratio matrix, performs element-wise multiplication with the local disparity gradient matrix in the same domain, and finally performs element-wise dot product output of an adaptive spatial weight mask in combination with the binarization mask of high disparity risk region, rather than a single gradient magnitude binarization.
[0077] The beneficial effects of the improvements are that, by calculating the local load ratio and using gradient cross-modulation, the physical constraints of the global geometric parallax load are rigidly embedded into the mask spatial distribution. This breaks the limitations of traditional methods, which are prone to color distortion and ghosting due to over-enhancement in areas with drastic changes in depth. It achieves a precise conversion from pure visual saliency to a depth load perception weight field. This design significantly enhances the spatial suppression accuracy of high parallax risk edges, and can accurately separate safe and overloaded areas in the gradient modulation space. Combined with the region locking of the binarized mask, it effectively improves the dynamic fit of the mask to the depth boundary of the stereo image and the absolute safety of subsequent color shift operations.
[0078] In this embodiment, S3 specifically includes:
[0079] S31. Set the conversion coefficients from the red-green-blue color space to the luminance-chrominance color space. Multiply the red channel value of each pixel in the stereo image by 0.299, the green channel value by 0.587, and the blue channel value by 0.114, and sum them to generate a luminance channel matrix. Subtract the corresponding value in the luminance channel matrix from the red channel value of each pixel in the stereo image, divide by 2, and add 128 to generate a first chrominance channel matrix. Subtract the corresponding value in the luminance channel matrix from the blue channel value of each pixel in the stereo image, divide by 2, and add 128 to generate a second chrominance channel matrix. Use the first and second chrominance channel matrices as the chrominance channel matrices. Calculate the sum of all pixel values in the luminance channel matrix and divide by the total number of pixels to generate a global luminance mean scalar.
[0080] S32. Calculate the difference between the value of each pixel in the luminance channel matrix and the global luminance mean scalar. Sum the squares of all differences and divide by the total number of pixels to generate the luminance discrete variance scalar. Perform a scalar addition operation on the global luminance mean scalar and the luminance discrete variance scalar to generate the adaptive color reference value.
[0081] S33. Perform element-wise division by dividing each pixel value in the chroma channel matrix by the sum of the adaptive color reference value and the value 0.001 to generate the initial color enhancement matrix.
[0082] S34. Construct a linear mapping function, setting the slope parameter to -1 and the intercept parameter to 1. Multiply each element value in the adaptive spatial weight mask by -1 and add 1. Perform a maximum value operation on the calculated result and the value 0, and then perform a domain inversion operation to generate a spatial suppression coefficient matrix. Perform an element-wise dot product operation on the spatial suppression coefficient matrix and the values at the corresponding spatial positions in the initial color enhancement matrix to generate a suppressed chroma channel matrix.
[0083] S35. Perform a stitching operation on the first and second chroma channel matrices in the suppressed chroma channel matrix along with the luminance channel matrix in the channel dimension. Set the inverse conversion coefficient from the luminance chroma color space to the red-green-blue color space. Add 2 to the luminance value of each pixel in the stitching result, multiply by the corresponding value of the first chroma channel matrix, and subtract 128 to generate the red channel value. Subtract 0.714 from the luminance value of each pixel in the stitching result, multiply by the corresponding value of the first chroma channel matrix, and subtract 128, and then subtract 0.344 from the corresponding value of the second chroma channel matrix, and subtract 128 to generate the green channel value. Add 2 to the luminance value of each pixel in the stitching result, multiply by the corresponding value of the second chroma channel matrix, and subtract 128 to generate the blue channel value. Stitch the generated red, green, and blue channel values together to generate the intermediate stereoscopic image.
[0084] In this embodiment, S4 specifically includes:
[0085] S41. Extract all pixel values of the intermediate stereo image in the first channel dimension to construct a red channel matrix, extract all pixel values of the intermediate stereo image in the second channel dimension to construct a green channel matrix, extract all pixel values of the intermediate stereo image in the third channel dimension to construct a blue channel matrix, and use the red channel matrix, green channel matrix and blue channel matrix as three independent color channel matrices for output.
[0086] S42. Perform element-wise multiplication of the binarized mask of the high disparity risk area with the values of the corresponding spatial positions in the disparity absolute value matrix. Force the safe area pixels with a value of 0 in the binarized mask to a value of 0, and retain the original disparity absolute values of the high-risk area pixels with a value of 1 in the binarized mask to generate the risk area disparity absolute value matrix.
[0087] S43. Perform an element-wise division operation by dividing each element value in the absolute value matrix of the disparity of the risk area by the sum of the global geometric disparity load index and the value 0.001 to generate the depth overload coefficient matrix.
[0088] In this embodiment, the improved GCNet model includes a direction prior construction layer, a global context pooling layer, a spatial similarity weighting layer, a Huygens metasurface interference layer, and a sub-pixel offset fusion layer:
[0089] The orientation prior construction layer is used to invert the value of each pixel in the binarized mask of the high disparity risk region by subtracting the value of 1, and generate an inverted mask. The inverted mask is then multiplied element-wise with the value of the corresponding spatial position in the local disparity gradient matrix of the same domain. The gradient of the safe region is filtered out and the gradient values of the edge of the high-risk region are retained to generate the spatial orientation prior matrix.
[0090] The global context pooling layer sets the pooling kernel size to the same global size as the depth overload coefficient matrix and the stride to the global size. It performs global spatial pooling on the depth overload coefficient matrix, calculates the sum of the values of all elements in the entire image in the depth overload coefficient matrix and divides it by the total number of pixels to generate the mean scalar of the entire image. The mean scalar of the entire image is broadcast and copied to the same spatial size as the depth overload coefficient matrix to generate the global context feature vector.
[0091] The spatial similarity weight layer sets the size of the local convolutional kernel to 3 rows and 3 columns with a stride of 1. It uses the 3 rows and 3 columns local convolutional kernel to perform a sliding traversal on the depth overload coefficient matrix to extract the values inside each window and generate a local spatial feature matrix. It calculates the absolute value of the difference between the local spatial feature matrix and the values at the same spatial location in the global context feature vector. The absolute value of the difference is multiplied by -1 and used as the exponent. An exponential function mapping with the natural constant e as the base is performed to generate the initial response matrix. The initial response matrix is subjected to a softmax probabilistic normalization mapping. The result of subtracting the global minimum value from each value in the initial response matrix is calculated. All results are divided by the sum of the values of all elements in the initial response matrix to generate the spatial modulation field matrix.
[0092] The Huygens-based surface interference layer is used to introduce a reverse wavefront filtering mechanism based on Huygens' principle, specifically including:
[0093] The depth overload coefficient matrix is input into a 1x1 convolutional kernel containing 32 channels to perform a pointwise convolution operation. The output feature map is divided into the first 16 channels and the last 16 channels in the channel dimension, which are used as the red channel incident wavefront features and the blue channel incident wavefront features, respectively. The spatial orientation prior matrix is input into a 3x3 convolutional kernel containing 64 channels to perform a convolution operation, and then into a 1x1 convolutional kernel containing 128 channels to perform a pointwise convolution operation. The number of output channels is adjusted to 16 channels, the same as the number of red channel incident wavefront feature channels, and multiplied by an expansion coefficient of 3 to generate the sampling offset.
[0094] A multi-scale circular convolutional kernel with two sizes, 3x3 and 5x5, is set up. Based on the values at corresponding positions in the sampling offset, the 3x3 and 5x5 circular convolutional kernels undergo spatial deformation on the red and blue channel incident wavefront features. Convolution is performed along the anisotropic direction indicated by the gradient normal. The convolution results of the two sizes are concatenated along the channel dimension to generate a wavelet excitation matrix. The wavelet excitation matrix is input into a 1x1 convolutional layer with ReLU activation function and 16 channels to perform nonlinear mapping for phase encoding, generating a phase-encoded feature matrix. Each element in the red and blue channel incident wavefront features is multiplied by -1 to perform a phase inversion operation, generating red-blue channel inverted wavefront features. The phase-encoded feature matrix and the red-blue channel inverted wavefront features are added element-wise to the corresponding spatial positions. Low-frequency background is filtered out by difference cancellation while high-frequency abrupt details are preserved, generating a red-blue channel anti-smoothing interference coefficient matrix.
[0095] The subpixel offset fusion layer is used to perform element-wise multiplication of the anti-smoothing interference coefficient matrix of the red and blue channels with the values at the same spatial position in the spatial modulation field matrix, generating and outputting the subpixel offset matrix of the red and blue channels.
[0096] The improved subpixel-level offset prediction process of the GCNet model proposed in this step is similar to the global context attention mechanism of the traditional GCNet model in that it is based on the theory of local-global feature association and spatial weight recalibration. That is, by projecting the input feature matrix to a high-dimensional space to extract local spatial features, using global pooling operation to capture the context distribution state at the whole image level, calculating the association response between local features and global context feature vectors, and using a normalized mapping mechanism to transform the association result into a spatial modulation weight matrix for subsequent adaptive feature fusion.
[0097] The difference lies in that this invention breaks the limitation of the traditional GCNet model, which relies solely on single attention channel fusion and ignores the high-frequency phase fidelity mechanism at the edges of abrupt changes in depth in stereo images. It adds a Huygens metasurface interference layer to replace the traditional element-wise addition feature fusion path, mapping the depth overload coefficient matrix to the incident wavefront features of the red and blue channels. It uses spatial direction priors to transform deformable convolution sampling offsets to guide the deformation of multi-scale ring convolution kernels, generating wavelet excitation matrices along anisotropic directions. Finally, it performs phase reversal on the incident wavefront and performs element-wise addition with the phase-encoded feature matrix. Low frequencies are filtered out by difference cancellation to generate an anti-smoothing interference coefficient matrix, rather than the traditional direct weighted fusion of global context features.
[0098] The beneficial effects of the improvements are that this invention, through wavefront splitting and reverse interference cancellation, forcibly embeds the Huygens physical wavefront diffraction constraint into the forward propagation of the network, breaking the limitation of the traditional GCNet model that is prone to loss of high-frequency details due to smoothing of the receptive field when dealing with depth edge offsets. It realizes the transformation from pure data-driven spatial attention aggregation to high-frequency edge excitation strongly constrained by wave physics mechanism. This design significantly enhances the anti-smoothing defense capability during the sub-pixel offset process of red and blue channels, can accurately cancel low-frequency blurring in the difference interference space and retain the high-frequency components of depth abrupt changes, and combined with the element-wise multiplication of the spatial modulation field, effectively improves the absolute accuracy and visual fit of depth edge pixel relocalization in stereo images.
[0099] In this embodiment, S6 specifically includes:
[0100] S61. Extract the spatial dimension coefficient matrix of the adaptive spatial weight mask, perform element-wise multiplication of the spatial dimension coefficient matrix with the values at the same spatial position in the red and blue channel sub-pixel level offset matrix, filter out the offset of the safe area and retain the offset values of the edge of the high-risk area, and output the mask constraint offset matrix.
[0101] S62. Extract the number of rows and columns of the independent red and blue color channel matrix to generate a grid coordinate point matrix containing the horizontal and vertical coordinates of each pixel. Extract the horizontal coordinate two-dimensional offset matrix and vertical coordinate two-dimensional offset matrix corresponding to the red and blue channels respectively from the mask constraint offset matrix. Add the horizontal coordinate value of each coordinate point in the grid coordinate point matrix to the value of the corresponding position in the horizontal coordinate two-dimensional offset matrix, and add the vertical coordinate value of each coordinate point in the grid coordinate point matrix to the value of the corresponding position in the vertical coordinate two-dimensional offset matrix to generate the target sampling coordinate matrix.
[0102] S63. Extract the x-coordinate and y-coordinate values of each coordinate point in the target sampling coordinate matrix. Round down the x-coordinate and y-coordinate values to obtain the top-left neighbor coordinate point. Round down the x-coordinate value and add 1, keeping the y-coordinate value rounded down to obtain the top-right neighbor coordinate point. Keep the x-coordinate value rounded down and add 1 to the y-coordinate value to obtain the bottom-left neighbor coordinate point. Round down the x-coordinate and y-coordinate values and add 1 to both to obtain the bottom-right neighbor coordinate point. Calculate the difference between the x-coordinate value and the y-coordinate value of the target sampling coordinate matrix. The difference between the x-coordinates of the top-left neighboring coordinates is used to generate the horizontal interpolation ratio. The difference between the y-coordinates of the top-left neighboring coordinates and the y-coordinates of the target sampling coordinates is subtracted to generate the vertical interpolation ratio. The value 1 is subtracted from the horizontal interpolation ratio to generate the left weight. The value 1 is subtracted from the vertical interpolation ratio to generate the upward weight. The left weight is multiplied by the upward weight to generate the top-left interpolation weight. The horizontal interpolation ratio is multiplied by the upward weight to generate the top-right interpolation weight. The left weight is multiplied by the vertical interpolation ratio to generate the bottom-left interpolation weight. The horizontal interpolation ratio is multiplied by the vertical interpolation ratio to generate the bottom-right interpolation weight. The four neighboring interpolation weights are combined to generate the four-neighboring interpolation weights.
[0103] S64. Based on the four-neighbor interpolation weights, extract the pixel values of the corresponding upper-left neighbor coordinates in the red and blue independent color channel matrix and multiply them by the upper-left interpolation weights. Extract the pixel values of the corresponding upper-right neighbor coordinates and multiply them by the upper-right interpolation weights. Extract the pixel values of the corresponding lower-left neighbor coordinates and multiply them by the lower-left interpolation weights. Extract the pixel values of the corresponding lower-right neighbor coordinates and multiply them by the lower-right interpolation weights. Perform an accumulation and summation operation on the four product results and output the offset red and blue channel matrix.
[0104] S65. Extract the green channel matrix after separating the intermediate stereo image, and perform a stitching operation on the offset red and blue channel matrices and the green channel matrix along the channel dimension to output the synthesized stereo image.
[0105] In this embodiment, S7 specifically includes:
[0106] S71. Set the conversion coefficients from red-green-blue color space to luminance-chrominance color space. Multiply the red channel value of each pixel in the synthesized stereo image by 0.299, the green channel value by 0.587, and the blue channel value by 0.114, and sum them up to generate the luminance channel matrix of the synthesized stereo image. Set the block pooling kernel size to 3 rows and 3 columns with a step size of 1. Force the luminance channel matrix pixels corresponding to the safe area with a value of 0 in the adaptive spatial weight mask to a value of 0. Use the 3 rows and 3 columns block pooling kernel to perform sliding traversal on the processed luminance channel matrix. Calculate the sum of the values of 9 pixels in each window and divide it by 9 to generate the local mean. Calculate the squared difference between the values of 9 pixels in each window and the local mean, sum them up, and divide them by 9 to generate the local variance. Concatenate the local mean and local variance as two independent channels to output the local luminance statistical feature matrix.
[0107] S72. Extract the local mean and local variance at the same spatial location from the local brightness statistical feature matrix. Add 0.001 to the local variance and use it as the denominator. Divide the local mean by the denominator to generate the spatial ratio. Construct a piecewise mapping function. When the spatial ratio is greater than 5, set the local contrast enhancement factor to 1.2. When the spatial ratio is less than 1, set the local contrast enhancement factor to 0.8. When the spatial ratio is greater than or equal to 1 and less than or equal to 5, calculate the spatial ratio minus 1, multiply the result by 0.1, and add 0.8 to generate the local contrast enhancement factor. Combine all the local contrast enhancement factors to generate the local contrast enhancement factor matrix.
[0108] S73. Perform element-wise dot product operation between the local contrast enhancement factor matrix and the values at the same spatial position in the adaptive spatial weight mask, filter out the contrast enhancement values in the safe area and retain the contrast enhancement values at the edge of the high-risk area, and generate a spatially constrained tone mapping coefficient matrix.
[0109] S74. Extract the value of each position in the spatially constrained tone mapping coefficient matrix, multiply the value by 0.8 and add 0.2 to generate the adaptive gamma exponent; construct a nonlinear gamma curve fitting mapping function, set the activation function of the nonlinear gamma curve fitting mapping function as a power function, divide the value of each pixel in the red, green and blue three-channel matrix of the synthesized stereo image by 255 to normalize it to the value range of 0 to 1, use the normalized value as the base and the adaptive gamma exponent as the exponent to perform power function mapping, multiply the mapping result by 255 to restore it to the original value range, and output the local tone fine-tuning matrix;
[0110] S75. Perform element-wise addition of the local tone fine-tuning matrix and the values of the same spatial position and the same channel dimension in the synthesized stereo image, perform residual addition and fusion, and output the final color-optimized stereo image.
[0111] Example 1: To verify the feasibility of this invention in color control and visual comfort improvement of stereoscopic display terminals, the method of this invention was applied to the naked-eye 3D display color rendering system of a provincial virtual reality technology company (hereinafter referred to as "Company V"). In traditional naked-eye 3D display systems, global histogram equalization or unified gamma correction algorithms are usually used for color enhancement. These methods not only struggle to accurately quantify the geometric load caused by parallax in complex scenes with dense depth levels, but also fail to perform adaptive color suppression and spatial pixel shift for high parallax risk areas, easily leading to red-blue channel crosstalk caused by excessive color stretching and severe visual dizziness. To solve the above problems, Company V decided to adopt a stereoscopic image color optimization method based on computer vision proposed in this invention.
[0112] During implementation, Company V first uses a front-facing binocular camera to acquire a stereo image video stream. After epipolar correction and synchronization alignment, a stereo image input frame containing left and right eye views is constructed. Simultaneously, Company V's display algorithm engine performs stereo matching on the acquired stereo images to generate a dense initial disparity map, which serves as the benchmark for subsequent disparity load quantization and spatial weight mask generation.
[0113] Company V uses a statistical fitting mechanism based on the absolute disparity value to extract the mean scalar and global standard deviation scalar of the entire image spatial dimension, and performs a weighted summation of the discrete variance values to generate a dizziness threshold. Next, a sliding window is used to perform a sliding traversal on the absolute disparity matrix, and fusion and normalization operations are performed on the local disparity gradient matrix in the same domain to generate an adaptive spatial weight mask. Subsequently, the adaptive spatial weight mask is input into a linear mapping function to perform domain inversion, and then element-wise multiplication is performed with the initial color enhancement matrix to accurately suppress the color enhancement amplitude in high disparity risk areas, outputting an intermediate stereoscopic image.
[0114] In the core offset reshaping and tone fine-tuning stages, this invention improves the GCNet model by extracting the inverted disparity gradient of high disparity risk regions as a spatial orientation prior, guiding the depth overload coefficient matrix to be mapped to the incident wavefront features of the red and blue channels. A Huygens principle-based inverse wavefront filtering mechanism is introduced through a Huygens metasurface interference layer. Low frequencies are canceled by element-wise addition of the phase-encoded feature matrix and the inverted wavefront features of the red and blue channels, generating an anti-smoothing interference coefficient matrix for the red and blue channels. Element-wise dot product is then performed on the spatial modulation field matrix to output the sub-pixel-level offset matrix for the red and blue channels. Finally, under adaptive spatial weight mask constraints, bilinear interpolation is performed on the red and blue channels for inverse remapping, stitching together the un-offset green channel, and combining this with a local tone mapping fine-tuning mechanism to output the final color-optimized stereo image.
[0115] During implementation, the technical team at Company V discovered that, compared to traditional color enhancement and conventional offset correction methods, the method of this invention significantly improves the visual comfort and edge sharpness of stereoscopic images. Traditional methods cannot adaptively constrain color stretching in high parallax regions and are prone to producing low-frequency smoothing blurring for pixel shifts at abrupt depth edges. In contrast, the method of this invention effectively achieves color fidelity and precise depth edge reshaping in stereoscopic images through global geometric parallax load quantization, adaptive spatial weight mask suppression, and high-frequency excitation removal by inverse wavefront filtering based on Huygens' principle.
[0116] To further verify the actual performance of the method of the present invention, Company V conducted a detailed comparative test between the method of the present invention and the traditional method. The specific performance data is shown in Table 1:
[0117] Table 1 Performance Comparison of Company V Naked-Eye 3D Monitor Color Rendering Systems
[0118] Color distortion rate (%) in high parallax regions 18.5 2.1 -88.6% Crosstalk index of red and blue channels 0.32 0.04 -87.5% Deep abrupt edge blur (pixels) 4.8 0.6 -87.5% Subjective rating of visual vertigo (1-10 points) 4.2 8.9 +111.9% Global color enhancement mean deviation 15.3 2.7 -82.4% Single-frame color optimization time (milliseconds) 45 32 -28.9% Success rate of binocular parallax fusion (%) 75.5 98.2 +30.1% High-frequency texture detail retention rate (%) 68.0 95.5 +40.4% End-user visual satisfaction (%) 72.0 96.5 +34.0%
[0119] As shown in Table 1, the performance of the naked-eye 3D display color rendering system was comprehensively improved after applying the method of this invention. The color distortion rate in high parallax regions decreased from 18.5% using traditional methods to 2.1%, and the red-blue channel crosstalk index decreased from 0.32 to 0.04, significantly improving the accuracy and spatial fidelity of color mapping and providing a reliable guarantee for eliminating retinal conflicts. The edge blurring of depth abrupt changes decreased from 4.8 pixels to 0.6 pixels, effectively avoiding the edge smoothing effect during the offset process. The subjective score for visual dizziness increased significantly from 4.2 points to 8.9 points, and the success rate of binocular parallax fusion increased from 75.5% to 98.2%, significantly enhancing system comfort. Furthermore, the high-frequency texture detail retention rate increased from 68.0% to 95.5%, and the end-user visual satisfaction increased from 72.0% to 96.5%.
[0120] Through the method of this invention, Company V has successfully achieved adaptive and precise suppression of stereoscopic image color enhancement and high-frequency anti-smoothing offset of depth edges, effectively reducing visual fatigue and dizziness risks in high parallax scenes, ensuring high-fidelity image quality of naked-eye 3D display, significantly improving the intelligence and physiological adaptation level of color rendering of stereoscopic display terminals, significantly reducing the brain fusion burden when users view stereoscopic content, enhancing the stability and robustness of display systems in complex depth scenes, and providing strong technical support for the research and development of next-generation stereoscopic vision terminals.
[0121] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for optimizing the color of stereo images based on computer vision, characterized in that, Includes the following steps: S1. Acquire stereo images and calculate disparity maps. Adaptively fit the dizziness threshold using the sum of the mean and standard deviation of the absolute values of disparity. Based on the dizziness threshold, statistically determine the global proportion of high-risk pixels and generate a global geometric disparity load index. S2. Based on the dizziness threshold, high parallax risk areas are extracted. The local load index matrix is calculated using a sliding window and compared with the global geometric parallax load index. The adaptive spatial weight mask is generated by combining the local parallax gradient matrix in the same domain. S3. Call the adaptive spatial weight mask to spatially suppress the color enhancement amplitude of the stereo image and output the intermediate stereo image; S4. Separate the intermediate stereo image into independent color channels, and generate a depth overload coefficient matrix based on the absolute value of disparity in high disparity risk areas and the global geometric disparity load index. S5. By improving the GCNet model, spatial orientation priors are extracted based on high parallax risk areas, guiding the mapping of the depth overload coefficient matrix to incident wavefront features. A reverse wavefront filtering mechanism based on Huygens' principle is introduced to cancel low frequencies to generate an anti-smoothing interference coefficient matrix. The element-wise multiplication of the spatial modulation field is combined to output the sub-pixel offset matrix of the red and blue channels. S6. Under the constraint of adaptive spatial weight mask, perform bilinear interpolation inverse remapping operation on the independent red and blue color channels according to the sub-pixel offset matrix of red and blue channels to perform pixel offset, and stitch the unoffset green channel to output a synthesized stereo image. S7. Use the adaptive spatial weight mask as a constraint to perform local tone mapping fine-tuning on the synthesized stereo image and output the final color-optimized stereo image.
2. The method for optimizing stereoscopic image color based on computer vision according to claim 1, characterized in that, S1 specifically includes: S11. Calculate the initial disparity map based on the stereo image by performing stereo matching, and generate the absolute disparity matrix by performing an absolute value operation on the initial disparity map; calculate the mean scalar of the entire image spatial dimension based on the absolute disparity matrix, calculate the corresponding global standard deviation scalar based on the mean scalar, and perform scalar addition on the mean scalar and the global standard deviation scalar to generate the global fluctuation benchmark value. S12. Based on the absolute disparity matrix, calculate the sum of squares of the differences between the absolute disparity of each pixel and the mean scalar; divide the sum of squares of the differences by the total number of pixels to generate the discrete variance value; perform a weighted summation of the discrete variance value and the global standard deviation scalar to generate the dizziness threshold. S13. Perform a pixel-by-pixel comparison between the dizziness threshold and the absolute disparity matrix, and output a high-risk pixel binarization mask based on the comparison result; perform full-image summation on the high-risk pixel binarization mask to generate a high-risk pixel count value; divide the high-risk pixel count value by the total number of pixels in the stereo image to generate a global proportion coefficient. S14. Multiply the global proportion coefficient with the global fluctuation benchmark value to generate the initial load index; perform logarithmic domain nonlinear compression mapping on the initial load index to generate the global geometric disparity load index.
3. The method for optimizing stereoscopic image color based on computer vision according to claim 1, characterized in that, S2 specifically includes: S21. Perform a pixel-by-pixel comparison between the adaptive dizziness threshold and the absolute disparity matrix, and output a binarized mask for high disparity risk areas based on the comparison results; use a sliding window to perform a sliding traversal on the absolute disparity matrix, calculate the local mean scalar within each window, and concatenate the local mean scalars according to their spatial positions to generate a local mean matrix; perform an element-by-element division operation between the local mean matrix and the global geometric disparity load index to generate a local load ratio matrix; S22. Based on the absolute disparity matrix, perform adjacent pixel difference operations along the horizontal and vertical directions respectively. Perform summation and square root operations on the horizontal and vertical difference results to generate the local disparity gradient matrix in the same domain. S23. Perform element-wise multiplication of the local load ratio matrix and the local disparity gradient matrix in the same domain to generate the gradient modulation load matrix; perform maximum and minimum value normalization operation on the gradient modulation load matrix to generate the global spatial weight basis matrix; perform element-wise dot product of the global spatial weight basis matrix and the binarized mask of the high disparity risk region to generate the adaptive spatial weight mask.
4. The method for optimizing stereoscopic image color based on computer vision according to claim 1, characterized in that, S3 specifically includes: S31. Convert the stereo image from the red-green-blue color space to the luminance-chrominance color space, and separate the luminance channel matrix and the chrominance channel matrix; calculate the global luminance mean scalar based on the luminance channel matrix. S32. Calculate the spatial discrete variance scalar of the luminance channel matrix based on the global luminance mean scalar; perform an addition operation on the global luminance mean scalar and the luminance discrete variance scalar to generate an adaptive color reference value; S33. Divide the chroma channel matrix by the adaptive color reference value and perform an element-wise division operation to generate the initial color enhancement matrix; S34. Perform a range reversal operation on the adaptive spatial weight mask input to the linear mapping function to generate a spatial suppression coefficient matrix; perform an element-wise multiplication operation between the spatial suppression coefficient matrix and the initial color enhancement matrix to generate a suppressed chroma channel matrix. S35. Perform a channel dimension splicing operation on the suppressed chroma channel matrix and the luminance channel matrix, and convert the splicing result from the luminance chroma color space to the red-green-blue color space to generate an intermediate stereo image.
5. The method for optimizing stereoscopic image color based on computer vision according to claim 1, characterized in that, S4 specifically includes: S41. Perform channel dimension separation operation on the intermediate stereo image and output three independent color channel matrices: red, green and blue. S42. Perform element-wise multiplication of the binarized mask of the high disparity risk area with the disparity absolute value matrix to generate the disparity absolute value matrix of the risk area. S43. Perform element-wise division between the absolute value matrix of disparity in the risk area and the global geometric disparity load index to generate the depth overload coefficient matrix.
6. The method for optimizing stereoscopic image color based on computer vision according to claim 1, characterized in that, The improved GCNet model includes a direction prior construction layer, a global context pooling layer, a spatial similarity weighting layer, a Huygens metasurface interference layer, and a sub-pixel offset fusion layer: The orientation prior construction layer is used to extract the inverted disparity gradient of high disparity risk regions as a spatial orientation prior. The global context pooling layer is used to extract the depth risk distribution state at the whole graph level by performing a global spatial pooling operation on the depth overload coefficient matrix, and generate a global context feature vector. The spatial similarity weight layer is used to extract local spatial features of the depth overload coefficient matrix, calculate the spatial correlation response between the local spatial features and the global context feature vector, and generate a spatial modulation field matrix through probabilistic normalization mapping. The Huygens element surface interference layer is used to introduce a reverse wavefront filtering mechanism based on Huygens' principle, specifically including: The deep overload coefficient matrix is guided to be mapped to the incident wavefront features of the red and blue channels. The spatial direction prior is transformed into the sampling offset of deformable convolution through channel dimension expansion mapping. On the incident wavefront features of the red and blue channels, the multi-scale ring convolution kernel is guided to deform according to the sampling offset to generate a wavelet excitation matrix along the anisotropic direction indicated by the normal. The wavelet excitation matrix is subjected to nonlinear mapping for phase encoding to generate a phase-encoded feature matrix. The incident wavefront features of the red and blue channels are subjected to phase reversal operation to generate the inverted wavefront features of the red and blue channels. The phase-encoded feature matrix and the inverted wavefront features of the red and blue channels are added element by element. The low frequency is filtered out by difference cancellation and the high frequency abrupt change is retained to generate the anti-smoothing interference coefficient matrix of the red and blue channels. The subpixel offset fusion layer is used to perform element-wise multiplication of the red and blue channel anti-smoothing interference coefficient matrix and the spatial modulation field matrix to generate and output the red and blue channel subpixel-level offset matrix.
7. The method for optimizing stereoscopic image color based on computer vision according to claim 1, characterized in that, S6 specifically includes: S61. Extract the spatial dimension coefficient matrix of the adaptive spatial weight mask, perform element-wise multiplication of the spatial dimension coefficient matrix with the sub-pixel offset matrix of the red and blue channels, and output the mask constraint offset matrix. S62. Extract the grid coordinates of the red and blue independent color channels based on the mask constraint offset matrix, and superimpose the mask constraint offset matrix on the grid coordinates to generate the target sampling coordinate matrix; S63. Perform bilinear interpolation inverse remapping on the red and blue independent color channels based on the target sampling coordinate matrix, and calculate the four-neighbor interpolation weights mapped to the discrete pixel coordinate space. S64. Based on the four-neighbor interpolation weights, perform a weighted summation on the pixel values corresponding to the discrete pixel coordinates of the independent red and blue color channels, and output the offset red and blue channel matrix. S65. Extract the green channel matrix after separating the intermediate stereo image, and perform a stitching operation on the offset red and blue channel matrices and the green channel matrix along the channel dimension to output the synthesized stereo image.
8. The method for optimizing stereoscopic image color based on computer vision according to claim 1, characterized in that, Specifically, S7 includes: S71. Extract the brightness channel matrix of the synthesized stereo image, perform block pooling operation on the brightness channel matrix based on the adaptive spatial weight mask, calculate the local mean and local variance of the block pooling result, and output the local brightness statistical feature matrix. S72. Calculate the spatial ratio of the local mean to the local variance in the local brightness statistical feature matrix, and calculate the local contrast enhancement factor matrix based on the spatial ratio. S73. Perform an element-wise multiplication operation between the local contrast enhancement factor matrix and the adaptive spatial weight mask to generate a spatially constrained tone mapping coefficient matrix. S74. Based on the spatially constrained tone mapping coefficient matrix, perform a nonlinear gamma curve fitting mapping operation on the red, green and blue three-channel matrix of the synthesized stereo image and output a local tone fine-tuning matrix. S75. Perform residual addition and fusion operation on the local tone fine-tuning matrix and the synthesized stereo image along the channel dimension to output the final color-optimized stereo image.